Extract Text
Plain text or Markdown out of a PDF, in reading order.
Extract Text exports a PDF's text as a plain .txt file or as Markdown with headings, in reading order, with optional page markers — built on the same page analysis the table tools use, on your device. Scanned pages have no text until they are recognised: run OCR PDF first.
Extract Text exports the words a PDF already carries as a .txt or Markdown file, in the order the page lays them out rather than the order they happen to be stored in. Headings become Markdown headings, detected tables become pipe tables, and page markers can be added where each page begins.
How to extract text — steps & common uses
What people use Extract Text for
Anyone who needs the words out of a PDF and into a text editor, a note-taking app or a repository.
- Get a plain-text copy of a report
- Turn a PDF into Markdown for a wiki or a repo
- Pull quotes out of a long document
- Feed a document's text into another program
- Check what text a PDF actually contains
How to extract text
- Add your PDF; it is read on your device.
- Choose plain text or Markdown, and whether to include page markers and detected tables.
- Set the page range if you only want part of the document.
- Click Extract text and download the .txt or .md file.
Good to know: Reading order is inferred from the layout; headings come from the document's structure, or from font sizes when it has none; scans need OCR first.
Questions about Extract Text 7
Are my files uploaded to a server?
No. UnboundPDF processes PDFs locally in the browser — documents are not uploaded to a server. We have no backend that receives, stores, or scans your files.
Why is my scanned PDF empty?
Because a scan is a picture of text and carries no text to export. Run OCR PDF on it first to add a text layer, then come back and extract it.
How are headings decided?
From the document's own tagged structure when it has one, and from relative font sizes when it does not. The result tells you which of the two was used.
Does it keep the reading order of a two-column page?
It uses the same page-structure engine as PDF to Word, which detects column bands and reads them in order rather than straight across the page.
What happens to a word broken across a line?
The hyphen is left exactly as it is. Joining the halves would mean guessing whether the word was hyphenated in the first place, and a wrong join changes the word.
Are tables included?
If you leave tables switched on. Markdown gets a pipe table; plain text gets tab-separated rows you can paste into a spreadsheet.
Does the file leave my device?
No. The PDF is read in the browser tab and the .txt or .md file is built there too.
Read more
How your files are handled
The work runs on your own device, inside the browser tab — so there is no server to send your document to.
- Nothing is uploaded. We have no backend that receives or stores documents, so we cannot see your files during processing or after.
- Your device does the work. Large documents are handled in stages, pages are rendered as you reach them, and each job is checked against your device before it starts — so a long run takes longer on a phone than on a laptop, nothing is queued behind other people's jobs, and you can stop a run at any point without touching your original file.
- No account and no watermark. The output is a clean file you can use anywhere, with no sign-up and no hourly cap.
More legal tools
Other free PDF utilities that work the same way — in your browser, no sign-up.