OCR PDF
Recognize the text in a scanned PDF so you can search, select and copy it — or export it as plain text. Runs in your browser — your files are never uploaded.
or drop a file here
Up to 50 MB- 1Choose
Choose a scanned PDF.
- 2Adjust
Pick the language of the document and the output: a searchable PDF or plain text.
- 3Download
Click “Recognize text”. Expect a few seconds per page.
Turn a scan into a PDF you can search
A scanned PDF is a stack of pictures: you cannot search it, select a sentence, or copy a number out of it. OCR (optical character recognition) reads the text in those pictures. ShyPDF places the recognized words as an invisible layer exactly on top of the scanned ones, so Ctrl+F, text selection and copy-and-paste work while the page looks exactly as it did.
The original pages are not re-compressed or redrawn, so there is no loss of quality and the file only grows by the size of the text. If you just want the words, choose “Plain text (.txt)” instead. Pages that already contain selectable text are skipped by default, which makes mixed documents faster.
Getting good results
Pick the language the document is written in — it is the single biggest factor in accuracy. English, Spanish, Portuguese, French, German, Italian, Japanese and Simplified Chinese are available, and “Also recognize English” helps with documents that mix English terms into another language.
Recognition runs in your browser using Tesseract, a long-established open-source OCR engine, compiled to WebAssembly. Expect a few seconds per page, depending on your device. Scans of bank statements, IDs, medical records and signed contracts are exactly the kind of file that should not be uploaded to an OCR service; here they never leave your device.
Questions
Are my files uploaded to a server?
No. Everything runs inside your browser, so your files never leave your device. Close the tab and nothing is left behind.
Does the PDF look different afterwards?
No. The pages are left exactly as they are; ShyPDF only adds an invisible text layer on top of them, so the scan keeps its quality and the file grows very little.
How accurate is it?
Clean, straight scans at 200 dpi or more are recognized very well. Blurry phone photos, handwriting, stamps and unusual fonts are much less reliable. Choosing the right language matters more than anything else.
Why does the first run take longer?
The recognition engine and the language data (roughly 5–10 MB) are downloaded from this site the first time and then kept in your browser. The PDF itself is never uploaded.
Can I run OCR on a photo or a JPG?
Yes, in two steps: turn the images into a PDF with JPG to PDF, then run that PDF through OCR PDF.