How to extract selected pages
Upload a PDF, enter a range such as 1-5, 7, 9, choose Smart extraction, select the printed language and review the page statuses. Results remain in document order even if the input range is not ordered.
FOCUSED BROWSER-BASED TOOL
Read embedded text and OCR scanned pages privately in your browser.
PDF Tool
PDF to Text
Extract embedded or scanned text from a PDF and export an editable copy.
Drag and drop a PDF or browse your device
Processed locally
PDF text extraction guide
A PDF-to-text converter turns document content into editable plain text. Digital PDFs already contain character data, while scans contain page images. Smart extraction checks each selected page, reads a meaningful embedded layer directly, and renders only pages that need OCR. This is generally faster and more accurate than OCRing everything.
Upload a PDF, enter a range such as 1-5, 7, 9, choose Smart extraction, select the printed language and review the page statuses. Results remain in document order even if the input range is not ordered.
Use the correct language and orientation. Try Balanced quality first; High accuracy renders more pixels and uses more memory. Grayscale, threshold, inversion and enlargement can help difficult scans but can also remove faint detail.
Supported encrypted PDFs can be opened with a password held only in page state. Parsing, rendering and recognition happen locally, results are not retained after reset, and embedded PDF JavaScript evaluation is disabled.
Plain text cannot exactly retain fonts, tables, columns or visual layout. OCR is an estimate, particularly for handwriting, receipts, forms and damaged scans. Proofread important output against the original.
Process one PDF up to 100 MiB and 300 pages. A job that may use OCR is limited to 100 selected pages, and each rendered page must stay within the 25-megapixel safety limit.
Edit the combined result before copying it or downloading UTF-8 TXT. JSON output includes page-level method and status information; DOCX and structured spreadsheet export are not part of this tool.
Create editable drafts from scanned reports, invoices, receipts, academic papers, forms, manuals, notes and archived documents. The page-by-page report identifies embedded, OCR and mixed extraction and isolates failures so one difficult page does not discard earlier results. For an important document, compare names, dates, amounts and headings against the page image before using the text elsewhere.
A digital page returns little text: try Smart extraction so pages without a meaningful embedded layer can fall back to OCR.
OCR confuses characters: confirm the printed language, correct page rotation and compare Balanced with High accuracy. Thresholding can help high-contrast scans but remove faint marks.
Columns read in the wrong order: plain-text extraction cannot reconstruct every layout. Process a smaller page range and rearrange the editable result while checking the original.
The PDF is protected: enter the password when prompted if supported, or use Unlock PDF only for a document you are authorized to open.
Use OCR PDF when English recognition of scanned pages is the whole task, Image to Text for a photograph or screenshot, and the OCR versus text extraction guide when you are unsure whether a PDF has selectable text. The PDF password guide covers protected-file limitations.
Add one PDF, choose all pages or a valid page range, select an extraction mode and language, then choose Extract text. Review the editable result before copying or downloading it.
Yes. Smart extraction OCRs pages without a meaningful text layer, while OCR Every Page recognises every selected page.
Normal extraction reads characters already stored in a digital PDF. OCR estimates characters from rendered page pixels and can contain recognition errors.
Yes. Process every page or enter values such as 1-5, 7, 9. Duplicate, empty, reversed and out-of-range selections are rejected.
Yes when PDF.js supports its encryption. Enter the password when prompted; it stays in component memory and is cleared on reset.
It preserves basic text and line breaks where practical, but does not reproduce fonts, exact layout, columns or page geometry.
It may recover table words, but rows and columns are not guaranteed to remain structured. Use the PDF to Excel tool for supported digital tables.
English, Hindi, Spanish, French, German, Italian and Portuguese OCR data are available. Choose the language manually.
No. PDF parsing, page rendering, preprocessing and OCR run locally in your browser. The selected OCR language data is downloaded by the OCR library, but your document is not sent to it.
Yes. The combined output is an editable text area, and TXT and JSON exports use your current edited text.
Yes. TXT export is UTF-8 and uses a sanitized version of the PDF filename.
Not currently. This version offers TXT and JSON because no lightweight, reliable browser DOCX exporter is installed.
Small text, blur, handwriting, unusual fonts, rotation, poor contrast, complex columns and an incorrect language choice can reduce accuracy.
Digital text extraction is usually quick. OCR runs one page at a time and may take several minutes for long or high-quality jobs, depending on the device.
Yes. PDFs are limited to 100 MB and 300 pages. Jobs that may use OCR are limited to 100 selected pages, and individual renders have a pixel-area safety limit.
Explore More
Extract text from scanned and mixed PDFs privately in your browser.
Extract editable text from images with browser-based OCR.
Remove password protection from PDF files online.
Extract selected pages from a PDF file.
Convert PDF pages into JPG or PNG images.
Reduce PDF file size directly in your browser.
Related guide: Compare OCR with embedded PDF text extraction.