text

Document to Markdown

Convert DOCX, PPTX, XLSX, PDF, TXT, HTML, CSV, and TSV to Markdown locally in your browser.

Local processing
About your filesFiles and content selected in this tool are processed in your browser and are not uploaded to a Browser Kitty backend.
Operation demo See the basic workflow in a short video.

About this tool

Document to Markdown converts Word, PowerPoint, Excel, PDF, text, HTML, CSV, and TSV content to Markdown entirely in the browser. Source documents are not uploaded to a conversion server, and reusable structure such as headings, paragraphs, lists, tables, and links is preserved where possible.

Alongside the Markdown result, Document information reports content as Converted, Check recommended, Not converted, or Partial failure. Embedded DOCX / PPTX images can be saved in a ZIP using relative paths or embedded directly into Markdown as Base64.

Good for

  • Reuse Word or PowerPoint content as Markdown for README files, wikis, notes, or text-based workflows
  • Convert PDF text layers or Excel tables into Markdown that is easier to copy and reuse
  • Batch-convert documents to Markdown without uploading them to an external conversion service

What it can do

Key capabilities available in this tool.

How to use

  1. Choose or drag DOCX, PPTX, XLSX, PDF, TXT, HTML, CSV, or TSV files into the app.
  2. When several files are added, they are queued and converted sequentially.
  3. Select a completed file and review Markdown, Preview, and Document information.
  4. When Document information contains Check recommended, Not converted, or Partial failure items, compare the result with the source document.
  5. Change the output filename if needed, then copy or save the Markdown.
  6. For DOCX / PPTX images, choose a ZIP with images or Base64 embedding; use Save all as ZIP for multiple successful results.

Supported

  • Input: DOCX / PPTX / XLSX / PDF / TXT / HTML / CSV / TSV; 100 MB per file, up to 20 files, 300 MB total.
  • DOCX: converts headings, paragraphs, lists, tables, links, embedded images, and other supported structure where possible.
  • PPTX: handles slide text, tables, links, speaker notes, embedded images, and other supported content.
  • XLSX: converts worksheet structure and cell values to tables; formulas are not executed and cached results or formula text are used.
  • PDF: uses PDF.js to extract the text layer and infer headings, paragraphs, and reading order from layout; OCR is not performed.
  • Export: .md, ZIP with images, Markdown with Base64 images, or ZIP of multiple results; Japanese / English and desktop / mobile.

Limitations and notes

  • Version 1.0.0 does not perform OCR on scanned or image-only PDFs.
  • Complex PDF columns, tables, vertical text, or rotated text may require review of inferred reading order and paragraph structure.
  • Encrypted or password-protected DOCX / PPTX / XLSX files are not supported.
  • PPTX charts, SmartArt, video/audio, animations, and some embedded objects are not fully reconstructed as Markdown.
  • XLSX formulas are not executed; cached results are used when available, otherwise the formula text is retained.
  • Legacy .doc / .ppt / .xls, EPUB, OpenDocument formats, and conversion from remote URLs are not supported.
  • The output does not reproduce the original document's visual layout; review layout-sensitive content with the quality report and preview.

Frequently asked questions

Can it convert Word and PowerPoint to Markdown?

Yes. DOCX and PPTX are supported, including headings, paragraphs, lists, tables, links, PowerPoint speaker notes, embedded images, and other supported content where possible.

Can it convert PDF to Markdown?

Yes. It reads the PDF text layer with PDF.js and infers headings, paragraphs, and reading order from layout. It does not perform OCR on scanned or image-only PDFs.

Can it keep images from the document?

Embedded DOCX / PPTX images can be saved in a ZIP referenced by relative Markdown paths, or embedded directly into one Markdown file as Base64 data.

Can I convert several files in one batch?

Yes. Queue up to 20 files with a 300 MB total limit, then save successful results together with Save all as ZIP.

Are document files uploaded to a server?

No. The hosted page first delivers the app HTML, but parsing, Markdown generation, preview, and export run in the browser and selected document content is not sent to an external conversion server.

Offline version and source code

document-to-markdown.html is a standalone HTML file containing PDF.js, its worker, Japanese CMaps, and the required runtime assets. Open the saved file directly in a supported browser to parse documents and generate Markdown without runtime external network access.

View source on GitHub