Get the text out of a PDF — plain text or Markdown, in reading order
Updated 22 August 2026
Extract Text exports a PDF's text as a plain .txt file or as Markdown with headings, in reading order, with optional page markers — on your device, built on the same page analysis the table tools use. Headings come from the document's tagged structure when it has one and from font sizes when it does not, and the result says which was used. Scanned pages have no text until they are recognised: run OCR PDF first.
Export a PDF's text as .txt or Markdown in reading order — headings kept, page markers optional — in your browser. When to OCR first, and what columns and tables do.
UnboundPDF is a free suite of 54 PDF and image tools that run entirely in your browser — merge, split, compress, edit text, OCR in 126 languages, redact, sign, convert and archive to PDF/A. Your document is read and written by the page on your own device; there is no document-upload endpoint in the core tools, no account, no watermark and no daily cap. Every result can be checked — with the Network tab, or with the Document Passport the Workspace writes for a chain of steps.
- Open Extract Text and add the PDF (run OCR PDF first if it is a scan).
- Choose plain text or Markdown, and whether to insert page markers.
- Download the file; the result notes whether headings came from structure or from font sizes.
Open Extract Text to do this in your browser — the steps above are the whole procedure; the rest of this page is what to expect at each one.
Plain text or Markdown
Plain text is for search, scripts and pasting into a note. Markdown keeps the document's headings — from its own tagged structure when present, from font sizes when not — so a long report becomes an outline you can navigate. The result states which method was used, so you know how much to trust the hierarchy.
Reading order
The exporter reads the page geometry the way the table tools do: lines grouped into blocks, blocks ordered top-to-bottom and left-to-right within columns. Clean reports come out clean; magazine layouts with pull quotes may need a pass by hand.
Mac, Windows, iPhone, Android
Same page; on a phone, share the .txt or .md to your notes app.
Troubleshooting
Empty output. The PDF is a scan — OCR first. Headings missing. The document has no structure and uniform font sizes; plain text is the honest result. Odd characters. Some PDFs map glyphs to private codes; the exporter writes what the file provides.
Related
OCR PDF · PDF to Word · PDF to Excel · Workspace
Frequently asked questions
Why is the text in the wrong order?
Reading order is inferred from the layout. Multi-column pages and text placed in boxes can read out of order; Markdown keeps headings so the structure is still visible. Use the page markers to find where a passage came from.
Does it handle tables?
It exports table text as lines. For a real table with cells, use PDF to Excel, which detects ruled and whitespace tables and writes typed cells.
My scanned PDF produced nothing.
A scan is a picture until it is recognised. Run OCR PDF, then extract.