How to OCR a PDF in Vietnamese, Polish or Turkish

To OCR a Vietnamese, Polish or Turkish scan, open UnboundPDF’s OCR tool, type the language into the picker (or let auto-detect suggest it), and run recognition. The language pack is fetched once from our site; the scan itself is processed on your device and never uploaded. On our own 300-dpi test fixtures, Polish and Turkish recognised at 100.00% and Vietnamese at 99.88% character accuracy — your scan’s quality decides yours.

Updated 19 August 2026

UnboundPDF is a free suite of 43 PDF and image tools that run entirely in your browser — merge, split, compress, edit text, OCR in 126 languages, redact, sign, convert and archive to PDF/A. Your document is read and written by the page on your own device; there is no document-upload endpoint in the core tools, no account, no watermark and no daily cap. Every result can be checked — with the Network tab, or with the Document Passport the Workspace writes for a chain of steps.

Steps

  1. Open OCR PDF and drop in the scanned PDF.
  2. In the language box type “viet”, “pol” or “tur” and tick the language (tick two for a bilingual document), or click Auto-detect and accept the suggestion.
  3. Run recognition. Each page reports its confidence as it finishes.
  4. Download the searchable PDF; the reading-order text file and hOCR are separate downloads.
  5. Fix any misread word in place with the PDF Editor if you need to.

Why the language matters for these three

Vietnamese stacks tone marks on vowels; Polish has ą, ę, ł, ń, ś, ź, ż; Turkish has dotless ı, ğ, ş and a capital İ. An English model reads those as the nearest plain letter or as noise. Selecting the right pack gives the engine the character set and the word shapes it needs — the same scan goes from guesswork to clean text.

OCR PDF is the tool with this language picker, running on your device.

What we measured, and on what

We built fixtures from public-domain text in each language, rendered, then scanned at 300 dpi, and compared the recognised text character by character. Polish: 100.00%. Turkish: 100.00%. Vietnamese: 99.88%. Those are our fixtures — clean print, straight pages. A faint photocopy or a phone photo at an angle will score lower, in any language.

Get the scan right first

300 dpi or better. Straight — if the page is tilted, Deskew PDF first. Dark text on plain background. The OCR pass also upscales and cleans the image it reads internally, but it cannot invent detail a 72-dpi scan never captured.

Reading the confidence figure

Each page shows the engine’s mean confidence. A page in the low 80s usually means the scan, not the language: look at that page, and re-scan or deskew if you can. Confidence is a guide to where to look, not a promise.

Fixing a word afterwards

Open the result in the PDF Editor, click the word, retype it. The correction blends into the scan’s own paper rather than a white box — see editing text in a scanned PDF.

Related

All 126 languages, auto-detect and the limits: OCR a PDF in more than 100 languages. Background on scans: OCR a scanned PDF.

Frequently asked questions

Is Vietnamese OCR accurate?

On our 300-dpi test fixture, 99.88% character accuracy. Real scans vary with resolution, skew and print quality.

Do I need to download anything?

The tool fetches the language pack from our own site the first time you use that language; nothing else is downloaded and your scan is not uploaded.

Can I OCR a document that mixes Polish and English?

Yes — tick both languages before running.

The page confidence is low. What should I do?

Check that page’s scan quality: deskew it, re-scan at 300 dpi if you can, and make sure the right language is ticked. Then run again.