Remove a Known PDF Password from an Authorized Copy
Learn how to create an unprotected working copy of a PDF when you know the password and are authorized to modify the document.
Read guide →Choose embedded PDF text extraction, OCR for scanned pages, or image OCR based on what the source actually contains.
In this guide
Text extraction reads characters already stored inside a digital PDF. Optical character recognition (OCR) estimates characters from page pixels. Use extraction when a selectable text layer exists; use OCR for scans, photographs, or pages whose embedded text is empty or unusable.
FiloTool's Smart extraction checks pages individually, so a mixed PDF can use its embedded layer where useful and OCR only the pages that need it.
Complete this task using FiloTool's PDF to Text directly from your browser.
Open PDF to Text →A digital PDF stores characters as data. A scan stores a photograph of characters and nothing else, even though both look identical on screen. The quickest check needs no tool: try selecting a line of text in your PDF viewer. If a text cursor appears and the words highlight, a text layer exists. If your selection draws a rectangle over the page instead, you are looking at an image.
A third case is the awkward one. Some PDFs carry a text layer that exists but is unusable — a handful of stray characters from a header, or a stream of substituted glyphs that select as nonsense. These files look extractable and are not, which is why deciding per document rather than per page tends to produce disappointing output.
Smart mode makes that decision one page at a time, using a deliberately simple test. It strips all whitespace from whatever embedded text the page yields, then requires two things: at least 10 characters remaining, and at least half of those characters being letters or digits rather than punctuation and symbols.
A page that passes is extracted directly. A page that fails is rendered and sent to OCR. That is why a single document can come back with different pages processed by different methods — a scanned appendix inside an otherwise digital contract gets OCR while the rest does not, without you having to identify which pages those were.
The threshold also explains a specific failure. A page carrying only a running header of six characters fails the ten-character test and goes to OCR even though it technically had a text layer, which is usually the right outcome — that header was not the content you wanted.
Extraction is fast, exact and free of guesswork: the characters it returns are the characters stored in the file, so there is no accuracy question to worry about. It also asks nothing of your device beyond reading the document.
Its weakness is arrangement rather than accuracy. A PDF stores text as positioned fragments, and the order those fragments appear in the file is the order the document was drawn, which need not match the order a human reads them. A two-column layout can therefore extract as interleaved lines, a table can lose its column structure, and a sidebar can land in the middle of a paragraph. Every character is correct and the sequence is not.
This is worth knowing before you blame the tool. Reading order problems are a property of how the PDF was built, and no extractor can reliably infer an intent the file never recorded.
OCR is an estimate. It renders each page to an image and recognises characters from pixels, which means it can be wrong, and it returns a confidence score so you can see where it was least sure. Extraction returns no confidence score, because there is nothing to be uncertain about.
Because recognition works from pixels, resolution matters. The quality setting controls how large each page is rendered before recognition: roughly 1.5x for Fast, 2x for Balanced and 2.5x for High. Higher settings give the recogniser more pixels per character and take longer, which is the entire trade.
Preprocessing helps more than people expect on poor scans. Converting to greyscale is on by default and a modest contrast increase is applied; for a faint or unevenly lit scan you can also apply a hard black-and-white threshold, invert a light-on-dark page, rotate a sideways one, or upscale before recognition. A skewed, low-contrast photograph of a page will produce worse output than a flat 300 dpi scan no matter which settings you choose.
In most cases you do not need to decide at all — leave Smart mode on and let the per-page test choose. Reach for a forced mode when you know something the heuristic cannot.
Try selecting text in your viewer to see whether a text layer exists at all
Leave Smart mode on unless you have a specific reason to override it
Choose Embedded when you want speed and exactness and would rather see empty output than guessed output
Choose OCR when the text layer exists but you have reason to distrust it
Check the confidence figures and spot-check names, numbers and dates against the original
Fix reading-order problems in the output text rather than by re-running with different settings
Recognition degrades predictably, and knowing the pattern saves time. Handwriting is not the target of this kind of recognition and should not be expected to work. Low-resolution scans, heavy JPEG artifacts, skew, shadows and show-through from the reverse side of a page all reduce accuracy. Dense tables and multi-column layouts frequently lose their structure even when individual words are read correctly, because a recogniser reports text rather than layout.
Decorative or unusual typefaces, very small print, and pages mixing several languages are also harder. Language selection matters: the recogniser is told which language to expect, and a document in a language you did not select will be read poorly.
Extraction accepts PDFs up to 100 MB and 300 pages. OCR is capped lower at 100 pages in one run, because recognising a page is far more expensive than reading one, and each page is rendered within a 25-megapixel ceiling to stay inside browser memory.
Seven recognition languages are available through PDF to Text: English, Hindi, Spanish, French, German, Italian and Portuguese. One detail is worth stating plainly rather than leaving implied — the English recognition model is served from this site, while the other languages are fetched from the tesseract.js project's own distribution when you select them. Your document is not part of either request; recognition runs in your browser and the file is not uploaded.
If the source is an image rather than a PDF, Image to Text applies the same recognition to a picture directly, which avoids the intermediate step of assembling a PDF just to read it.
Open the tool, upload your file and complete the task in a few simple steps.
Try It Now →Common questions
Clear answers to common questions about this topic and the related FiloTool tool.
No. Extraction returns the characters actually stored in the file, so it is exact. OCR estimates characters from pixels and can be wrong. Extraction is only unavailable when there is no usable text layer to read.
It strips whitespace from the page's embedded text and requires at least 10 characters, of which at least half must be letters or digits. Pages that pass are extracted; pages that fail are rendered and sent to OCR. The test runs per page, so one document can use both methods.
A PDF stores text as positioned fragments in the order the page was drawn, which need not match reading order. Columns can interleave and sidebars can land mid-paragraph. The characters are still correct; only the sequence is wrong, and it is a property of the file rather than the extractor.
Not reliably. Recognition reports text, not layout, so rows and columns commonly collapse even when the individual words are read correctly. Expect to rebuild table structure by hand.
Balanced renders pages at about twice their nominal size and suits most scans. Fast uses roughly 1.5x and is quicker on long documents; High uses about 2.5x and helps with small or faint print, at the cost of time.
Extraction handles up to 300 pages while OCR is limited to 100 in one run. Recognising a page means rendering it to an image and analysing pixels, which costs far more time and memory than reading stored characters.
No. Recognition runs in your browser. The English model is served from this site and other language models are fetched from the tesseract.js distribution when selected, but your document is not part of those requests.
Useful Tools
These workflows are selected for the next steps most closely related to this guide.
About the author
Software engineer, developer of FiloTool
Urvish Parmar is a software engineer and the developer of FiloTool. He builds and maintains the browser-local PDF and image tools published on this site, and writes the guides that document how they behave.
Because the same person implements a tool and documents it, these guides describe what the code actually does — including the formats it rejects, the limits it hits and the results it cannot promise.
Read how FiloTool tests and updates its content on the About page, or report an error through Contact.
An updated date is shown only when the article's instructions, evidence or material guidance changed.
Continue reading
Explore more practical PDF and image guides from FiloTool.
Learn how to create an unprotected working copy of a PDF when you know the password and are authorized to modify the document.
Read guide →Learn how PDF page rendering differs from recovering original embedded images, and what JPG export preserves or loses.
Read guide →Choose the right order of merge, split and page organization steps while preserving page appearance and checking document-level features.
Read guide →