How to OCR a PDF in more than 100 languages

Updated 19 August 2026

To OCR a PDF in a language other than English, open UnboundPDF’s OCR tool, search the language picker (126 languages, grouped by script) or let auto-detect suggest the script, and run recognition. The language pack is fetched from our site the first time and cached; your document is processed on your device and never uploaded. You get a searchable text layer, a reading-order text file, hOCR, and a confidence figure per page.

Pick from 126 languages or let auto-detect find the right one. OCR runs on your device — the scan is never uploaded.

UnboundPDF is a free suite of 43 PDF and image tools that run entirely in your browser — merge, split, compress, edit text, OCR in 126 languages, redact, sign, convert and archive to PDF/A. Your document is read and written by the page on your own device; there is no document-upload endpoint in the core tools, no account, no watermark and no daily cap. Every result can be checked — with the Network tab, or with the Document Passport the Workspace writes for a chain of steps.

  1. Open OCR PDF and drop in the scanned PDF.
  2. Type in the language search box (or click Auto-detect and accept a suggestion). Tick more than one language if the document mixes them.
  3. Run recognition. Watch the per-page confidence as pages finish.
  4. Download the searchable PDF. The reading-order text and the hOCR file are separate downloads if you need them.

The picker: 126 languages, grouped by script

The list is grouped under script headings — Latin, Cyrillic, Devanagari, Arabic, Han and so on — with a search box, so “viet” finds Vietnamese and “tur” finds Turkish without scrolling. Each entry shows whether its language pack is already on our site or is fetched on first use. Tick several languages for a bilingual document; recognition runs with all of them.

OCR PDF is the tool with this picker, running recognition on your device.

Auto-detect before you start

Auto-detect runs orientation and script detection on a page and suggests candidate languages — useful when you are handed a scan and are not sure whether it is Serbian in Cyrillic or Croatian in Latin. It runs in its own recognition engine, separate from the main pass, so it cannot change how the main recognition behaves. You can always override its suggestion.

“On your device” — what that means for OCR specifically

The recognition engine is a build of Tesseract we host ourselves and load into your browser. The only thing fetched during OCR is the language pack you chose — one file per language, from our own site, the first time you use it. Your scan is never sent anywhere; you can watch the network panel and see only that one pack download. This matters because scans that need OCR are often the most personal documents you have — medical letters, certificates, old official papers.

Reading order, text layer, hOCR

The tool keeps the structure Tesseract finds — blocks, paragraphs, lines — rather than a flat bag of words. The searchable text layer is placed word by word over the scan so searching and selecting work in any reader. The plain-text download is in reading order, with column breaks on two-column pages where a gutter is detected. The hOCR download carries every word’s position for anyone doing further processing.

Per-page confidence

Each page reports Tesseract’s mean confidence — “p1 94%, p2 93%, p3 91%” — so you can see which pages deserve a second look. Low confidence usually means a poor scan rather than a wrong language: resolution under 300 dpi, skew, faint print. Straighten with Deskew PDF and re-run; the OCR pass also upscales and cleans the image it reads internally without touching the scan you keep.

Measured accuracy on our own test fixtures

We measure character accuracy on fixtures we built from public-domain text rendered and scanned at 300 dpi. On those fixtures, Polish and Turkish recognised at 100.00%, Vietnamese at 99.88%, and Simplified Chinese at 96.46% on a representative 13-point page — and 89.95% on a small, dense page, which we quote alongside because it shows where the engine struggles. These are our fixtures, not your scan; a faint fax will score lower in any language.

What is not measured yet

Right-to-left scripts (Arabic, Hebrew) have not been measured — we do not know their accuracy and will not imply it. Dense CJK is measured and disclosed above. Handwriting is not something this engine reads reliably in any language.

Related

The step-by-step for three specific languages: OCR a PDF in Vietnamese, Polish or Turkish. Then edit the recognised words in place with the PDF Editor — see edit text in a scanned PDF. Background on scan quality: OCR a scanned PDF.

Frequently asked questions

Which languages can the OCR tool recognise?

126 languages, listed under 31 script groups with a search box. Tick several for a bilingual document. The language pack is fetched from our own site the first time and cached.

Is my scan uploaded for OCR?

No. The recognition engine runs in your browser; the only download is the language pack you chose, from our site. You can watch the network panel while it runs.

How accurate is it in Polish, Turkish, Vietnamese or Chinese?

On our own 300-dpi test fixtures: Polish 100.00%, Turkish 100.00%, Vietnamese 99.88%, Simplified Chinese 96.46% on a representative page and 89.95% on a small, dense page. Your scan’s quality decides your result.

Does it work for Arabic or Hebrew?

Both are in the picker, but we have not measured their accuracy yet and will not state a number. Treat them as untested.

What does the confidence percentage mean?

It is the engine’s mean word confidence for that page. Low numbers usually point at a poor scan — low resolution, skew, faint print — rather than the wrong language.