OCR vs Text Extraction: Which PDF-to-Text Method Should You Use?
Choose embedded PDF text extraction, OCR for scanned pages, or image OCR based on what the source actually contains.
Read guide →Distinguish removable visual covering from flattened PDF redaction, then verify text, metadata and every affected page before sharing.
In this guide
Placing a black rectangle over text is not automatically redaction. If the rectangle is an annotation or drawing object, the original text or image may remain selectable, searchable, removable, or extractable underneath. The document looks redacted on screen and is not redacted in the file.
A safer workflow replaces the affected page content with rendered pixels and bakes opaque redaction areas into that new page image. This page explains what that rebuild actually does to a PDF, what the export checks on your behalf before it hands you a file, and the three checks you should still run yourself.
Complete this task using FiloTool's Redact PDF directly from your browser.
Open Redact PDF →Blur and pixelation are visual effects, not reliable removal methods. A solid overlay can also be unsafe while it remains editable. FiloTool's Redact PDF Secure Flatten workflow rasterizes each affected page, draws solid redactions into the pixels and builds a replacement page; unaffected pages are preserved where possible.
Blur and pixelation transform the pixels you can see; they do not remove the information those pixels encode. Pixelation replaces a region with a grid of averaged blocks, which is a deterministic function of the original. When the hidden value comes from a small, predictable set—a six-digit account number, a date, a name from a known list—an attacker can render candidate values through the same transformation and compare the output until one matches. Blur is a convolution, and approximate inverses for it are well studied.
There is a second, simpler failure. If the blur or pixelation is applied as an overlay object rather than baked into replacement pixels, the untouched content is still sitting underneath it in the file. FiloTool's Pixelate Image tool exists for creative and preview purposes on images, and is deliberately not offered as a document redaction method.
Text search locates exact matches inside the embedded text layer, with case-sensitivity and whole-word controls. Pattern suggestions cover four shapes: email addresses, digit runs that look like phone numbers, IPv4-style addresses, and http or https URLs. These are heuristics, and they are deliberately loose—the phone pattern matches most separated digit runs of nine or more characters, so it will also flag order numbers and reference codes, and the IPv4 pattern accepts impossible values such as 999.999.999.999.
Matched values are masked in the review list, showing the first and last two characters with the middle replaced by dots, so that reviewing a long list of hits does not put the sensitive values back on your screen. Scanned pages have no embedded text layer for any of this to search, so mark those regions by drawing areas over them instead.
Names, identifiers and account numbers
Faces, signatures, barcodes and QR codes
Headers, footers and repeated page regions
Comments, form values and metadata that need separate review
Only pages carrying at least one redaction area are rebuilt. Each of those pages is rendered to a canvas at 1.75x scale on Standard quality or 2.5x on High, your redaction rectangles are filled into those pixels as solid colour, and any replacement label you added is drawn and clipped to the rectangle. The finished canvas is encoded as a JPEG—quality 0.88 on Standard, 0.94 on High—and embedded as a single full-page image at the original page dimensions.
The consequence is that a rebuilt page has no text layer, no annotations, no form fields and no links, because it is now one flat image. Pages you did not mark are copied across untouched and keep their original text, vectors and searchability. That preserves quality where nothing needed removing, and it creates the limitation described further down this page.
Work from a copy and define what must be removed before drawing. Include enough surrounding area to cover every glyph, shadow or edge. Secure Flatten invalidates existing digital signatures and removes searchable text, links, forms and annotations from affected pages.
Keep the unedited original in a controlled location
Search embedded text and inspect every suggested match
Draw areas over scanned or non-text content
Choose Standard or High quality and optionally remove document metadata
Export, then verify the downloaded file—not the editor preview
Secure Flatten does not simply save and hand over the result. After writing the new document it re-reads its own output and fails loudly rather than returning a file that did not meet its conditions. It confirms the saved bytes begin with a valid %PDF- signature, reopens the document and confirms the page count still matches the source, then re-parses every page it redacted and checks two things: that no extractable text remains on that page, and that no annotations remain on it.
If any of those checks fail you receive an error message instead of a download. That is a useful floor, and it is not a guarantee. The check only inspects pages that carried a redaction area, so it can confirm that the pages you marked were flattened correctly—it cannot know about a page you forgot to mark.
Verify the downloaded file, not the editor preview. The preview is a rendering of your intent; the download is the artifact you will actually send. These three checks are independent, which is the point—each one can catch a failure the others miss.
Look: open the output in a different PDF viewer and inspect every affected page at high zoom, watching for glyph edges or descenders peeking out from under a rectangle
Extract: select all and copy, then separately run the file through PDF to Text and search the extracted output for the removed value
Inspect properties: open document properties and check title, author, subject and keywords, which are carried over from the source unless you chose to remove metadata
The realistic risk is not a defeated redaction. It is a value you never marked. Because unmarked pages are copied through with their text layer intact, and because the built-in verification only inspects the pages you redacted, a sensitive string sitting on page 34 of a 40-page contract is fully extractable from the exported file and nothing in the workflow will complain.
So run the search across the whole exported document rather than page by page, and search for the value itself rather than trusting that you found every occurrence visually. Repeated headers, footers, page-level watermarks, appendices and signature blocks are the usual places a second copy hides.
Flatten PDF converts supported form controls into ordinary page content. It is useful when form values must stop being editable, but flattening a form is not the same as securely removing sensitive pixels or text. Use the dedicated redaction workflow for content removal.
Redact PDF accepts one file at a time, up to 100 MB and between 1 and 500 pages. A rendered page is capped at 40 million pixels, which is why a very large page can fail on High quality and succeed on Standard—the lower scale factor produces a smaller canvas.
Rebuilding pages as images has costs worth planning for. Redacted pages are no longer selectable or searchable, which also means they are no longer readable by a screen reader; if the document must remain accessible, redact the smallest possible set of pages. The file can also grow, because a JPEG of a text page is often larger than the text instructions that drew it, so compare the exported size against the original before sending.
Text search depends on the embedded text layer, so unusual encodings and subset-embedded fonts can defeat it, and pattern suggestions can be both over- and under-inclusive. Existing digital signatures are invalidated, since the pages they signed no longer exist in their original form. And no browser workflow can reach copies that already left your machine.
Open the tool, upload your file and complete the task in a few simple steps.
Try It Now →Common questions
Clear answers to common questions about this topic and the related FiloTool tool.
Only when the exported page has replaced the underlying content. An editable rectangle can leave the original material intact.
Do not rely on blur or pixelation for sensitive information. Use a fully opaque redaction that is flattened into the exported page pixels.
Run three independent checks on the downloaded file: inspect every affected page at high zoom in a separate viewer, extract the text with PDF to Text and search it for the removed value, and review document properties for metadata that still contains it.
No. Rebuilding an affected page means the content the signature covered no longer exists in its original form, so existing digital signatures are invalidated.
No. Only pages carrying at least one redaction area are rebuilt as images. Every other page is copied across untouched and keeps its original text, vectors and searchability—which also means any sensitive text on those pages is still extractable.
Each redacted page becomes a full-page JPEG, and an image of a text page is often larger than the instructions that drew that text. Redact the smallest set of pages you can, and compress the result if file size matters.
Not on pages that were redacted. Those pages are flattened to a single image, so they contain no text layer, annotations, form fields or links. This also removes screen-reader access to those pages.
You get an error rather than a file. Before offering the download, the export re-reads its output and fails if the PDF signature is wrong, the page count changed, or any redacted page still contains extractable text or annotations.
Useful Tools
These workflows are selected for the next steps most closely related to this guide.
About the author
Software engineer, developer of FiloTool
Urvish Parmar is a software engineer and the developer of FiloTool. He builds and maintains the browser-local PDF and image tools published on this site, and writes the guides that document how they behave.
Because the same person implements a tool and documents it, these guides describe what the code actually does — including the formats it rejects, the limits it hits and the results it cannot promise.
Read how FiloTool tests and updates its content on the About page, or report an error through Contact.
An updated date is shown only when the article's instructions, evidence or material guidance changed.
Continue reading
Explore more practical PDF and image guides from FiloTool.
Choose embedded PDF text extraction, OCR for scanned pages, or image OCR based on what the source actually contains.
Read guide →Learn how to create an unprotected working copy of a PDF when you know the password and are authorized to modify the document.
Read guide →Learn how PDF page rendering differs from recovering original embedded images, and what JPG export preserves or loses.
Read guide →