Turn a scanned PDF into searchable text
A scan is only a picture of a page, so you cannot search it or copy a word from it. OCR reads the picture, and you download a searchable PDF, a plain-text file, or both.
- Files never leave your browser.
- No account needed
- No watermark added
- Works on any device
How it works
Step 1: Choose or drop a scanned PDF, or a JPG or PNG image.
Step 2: Pick one or more languages that match the text, and the output: searchable PDF, text file, or both.
Step 3: Choose whether to skip pages that already have text, and select 200 or 300 DPI.
Step 4: Select OCR PDF, wait for each page to be processed, then download the result.
What OCR does to a scanned document
OCR stands for optical character recognition. Each page is drawn as an image, the recognition engine (Tesseract.js, running as WebAssembly in your browser) reads the letters in it, and the text is recorded together with where each word sits on the page. For a searchable PDF, that text is laid invisibly behind the original page image, so the page looks unchanged but you can search it, select it and copy from it.
You can also download the recognized text as a plain .txt file, which is handy for pasting into an editor or feeding another tool. You choose the languages first, then whether to skip pages that already contain text, and the resolution used to render pages (200 or 300 DPI). Besides PDFs, the tool accepts JPG and PNG images as input.
Documents are never uploaded. The one thing that is fetched over the network is the OCR engine and the language data, which are downloaded from a public CDN (jsDelivr) the first time you use the tool, a few megabytes in total. The CDN can see your IP address and standard request headers, but not your documents.
Why use this tool
Documents are not uploaded
Recognition runs in your browser. Only the engine and language files are downloaded, once, from a public CDN; your pages never leave your device.
Searchable and selectable
Latin-script languages give you a PDF where you can press Ctrl+F, highlight a paragraph and copy it, while the page keeps its original look.
Ten languages, mixed if needed
English, French, Spanish, Arabic, German, Italian, Portuguese, Dutch, Turkish and Russian are available, and you can select several for bilingual pages.
Skips what is already text
In a file that mixes scans with digital pages, you can leave the pages that already have text alone, which saves time.
Common uses
Searching an archive of old scanned letters
A family archive or a small office has hundreds of scans named by date only. After OCR, you can search for a surname across the results instead of opening each file.
Copying a passage from a scanned book chapter
A student photographs a library chapter and needs a quotation. OCR turns it into text that can be pasted into a paper, ready to be proofread.
Extracting text from an Arabic or Russian document
For these scripts the tool provides the recognized text as a .txt download, so a scanned contract or notice can be read into an editor or translated.
Preparing a scan for another tool
PDF to Word, PDF to Excel, PDF to Text and Compare PDF read text and find nothing in a scan. Run OCR PDF first to give them something to work with.
Digitizing receipts and invoices
A freelancer with a phone photo of a receipt uses the JPG as input and gets text to copy amounts and dates from.
Scanned pages versus pages that already have text
Not every PDF needs OCR. A file exported from a word processor already contains real text, and running OCR over it would only add work and risk small errors. A scan, on the other hand, holds images only. A quick test: open the file and try to select a word with the cursor. If nothing highlights, or Ctrl+F finds nothing, the pages are pictures.
Mixed files are common, for example a digital contract with a few scanned signature pages. The option to skip pages that already have text handles them: the digital pages are left untouched and only the scanned ones are processed. Once done, the searchable PDF supports every text tool, so you can go on to Compress PDF, PDF to Word or PDF to Text with useful results.
- Test first: if you can select text, the page is not a scan.
- Use the skip option for mixed documents.
- Search for a distinctive word afterwards to confirm the result.
Choosing languages and understanding scripts
The engine reads better when it knows what to expect, so select the language of the text and not just the one you happen to speak. If a page mixes French and English, select both; recognition takes slightly longer but it stops mistaking one language's accents for noise. Selecting many languages you do not need slows things down and can cause stray misreadings, so pick only the ones that are on the page.
There is one important difference between scripts. For Latin-script languages (English, French, Spanish, German, Italian, Portuguese, Dutch and Turkish), the tool can build a searchable PDF whose invisible text layer sits behind the page image. For Arabic and Russian, which use other scripts, that searchable layer is not available in this tool. It still recognizes them, but delivers the result as a .txt file. If you need a searchable Arabic PDF, this tool will not produce it.
- Bilingual page: select both languages.
- Arabic or Russian: use the .txt download.
- Numbers, dates and Latin names inside an Arabic page are best proofread by eye.
Getting the most accurate result
Image quality matters more than any setting. The most reliable input is a page scanned at 300 DPI, straight on the glass, in even light, with dark text on a light background. Skewed photos, shadows, creases, faded ink and very small print all reduce accuracy. If you are capturing a page with a phone, hold it parallel to the paper, avoid your own shadow and fill the frame with the page.
Use the 300 DPI option for small print, footnotes and dense pages, and 200 DPI when text is large and you want faster processing. Higher resolution also means more memory and more time per page. Tables and multi-column layouts are read line by line, so the order of the text may not match the visual layout; check those areas by eye.
- Scan at 300 DPI or higher where possible.
- Straighten and crop pages before OCR if they are tilted.
- Do not expect reliable handwriting recognition.
- Proofread names, figures and dates against the original page.
Supported formats
- Input
- PDF files (.pdf) and images (.jpg, .jpeg, .png).
- Output
- A searchable PDF (page image plus invisible text) and/or a plain-text .txt file.
The searchable PDF text layer covers Latin-script languages. For Arabic and Russian, the recognized text is provided as a .txt download.
Good to know
- The invisible text layer of the searchable PDF supports Latin-script languages only. For Arabic and Russian you get the recognized text as a .txt file, not a searchable PDF.
- Accuracy depends on the scan: 300 DPI, straight, evenly lit, high-contrast pages give the best results. Handwriting is not reliably recognized.
- OCR is slow on large documents: seconds per page on a laptop, longer on a phone.
- The first use downloads the engine and language data from a public CDN (jsDelivr), so an internet connection is needed at that point.
- Recognition can make mistakes, especially with small print, tables and unusual fonts, so proofread anything important.
Your files never leave your browser
This tool runs entirely on your device using your browser. Your files are not uploaded to our servers and are not stored anywhere. Closing the tab discards everything.
The OCR engine and language data are downloaded once from a public CDN the first time you run OCR. Your documents themselves are never uploaded.
Frequently asked questions
Are my documents uploaded to be recognized?
No. Recognition runs in your browser, and your pages are never uploaded. The OCR engine and language files are downloaded once from a public CDN (jsDelivr), which sees your IP address and standard request headers but not your documents.
Which languages does OCR PDF support?
English, French, Spanish, Arabic, German, Italian, Portuguese, Dutch, Turkish and Russian. You can choose one or several, for example when a page mixes two languages.
Can I get a searchable PDF in Arabic?
No. The searchable PDF text layer works for Latin-script languages. For Arabic and Russian the tool recognizes the text but provides it only as a .txt file.
How long does OCR take?
Expect a few seconds per page on a laptop and longer on a phone, so a long document can take several minutes. The first run also spends time downloading the language data.
Does it recognize handwriting?
Not reliably. OCR works best on printed text. Handwriting usually comes out with many errors or is missed.
Why is the result wrong in some places?
Usually because of image quality: low resolution, skew, shadows, faded ink or a font the engine does not know. Rescan at 300 DPI, select the right language and try again.
Related tools
Related guides
- How to make a scanned PDF searchableTurn a scanned PDF into a searchable one with OCR: how to check if your file needs it, language and DPI settings, what to expect, and how to verify the result.
- What is OCR and how does it work?OCR turns images of text into real, searchable text. Learn how it works step by step, what affects accuracy, its limits with handwriting, and what it outputs.
- How to scan documents to PDFTurn paper into a clean PDF with a phone camera or scanner: lighting and framing tips, page order, the document filter, file size, and adding searchable text.