PDF to Text

Extract the text from a PDF, page by page, into a box you can copy. The PDF is read in your browser.

Updated

PDF Tools● Free No upload Instant

Loading PDF to Text…

Your browser is preparing the tool. It runs 100% locally.

Quick answer

Upload a PDF and the tool reads its text layer page by page and puts the result in a box you can copy. It works on PDFs that contain real, selectable text. A scanned PDF is just an image of the page with no text layer, so it returns nothing and needs OCR instead. Your browser opens and reads the file, so it never goes to a server.

What the PDF to Text does

This tool pulls the words out of a PDF so you can reuse them. It reads the document's text page by page and gives you plain text, ready to copy into a note, an editor or a search box.

Use it on PDFs that already have selectable text, such as reports, exported documents and articles, to get the text out without retyping it.

How it works

The tool opens the PDF with a rendering engine and, for each page, reads the text items the PDF stores (the actual characters placed on the page). It joins them, puts a blank line between pages, and shows the result in a copyable box.

This works because a normal PDF carries a text layer: the words are stored as text, even though the page looks like a fixed image. A scan or photograph of a page has no such layer, so there is nothing to read and the result is empty.

Processing pipeline

  1. Open the PDF. Load the document in the browser with a PDF engine.
  2. Read each page. For every page, read the text items the PDF stores.
  3. Join the text. Join the items into text, separating pages with a blank line.
  4. Show it to copy. Put the extracted text in a box ready to copy.

What it reads

for each page: read the stored text items → join them with spaces pages separated by a blank line · a scanned page with no text layer → nothing
Worked example
a 3-page text PDF → its text, page by page, in a copyable box

It reads the existing text layer only. A scanned PDF (an image of text) has no text layer, so it returns nothing and needs OCR. Columns and tables run together, because items are joined in reading order.

Privacy

The PDF stays on your device. Your browser opens it and reads the text, with no upload and no copy kept anywhere.

There is no account, and the page saves nothing. Refresh or close the tab and both the PDF and the extracted text are cleared from memory, so it is fine for private documents.

Standards and references

  • Text layer extraction: A PDF stores its words as text items positioned on each page. The tool reads those items, which is why a PDF with selectable text extracts cleanly and quickly.
  • Scanned PDFs need OCR: A scanned or photographed page is an image with no text layer, so there is nothing to extract and the result is empty. Getting that text back takes OCR (optical character recognition), which reads the image.
  • Layout is lost: Text items are joined in the order the PDF stores them, which is close to reading order. Columns, tables and complex layouts collapse into a single stream of text.

Accuracy and limits

For a PDF with a real text layer the extraction is fast and faithful: the words come out as stored, page by page. That covers most exported and generated PDFs.

A scanned PDF returns nothing. If the page is an image of text, there is no text layer to read and the box stays empty. Those documents need an OCR tool, which reads the picture of the words.

The layout isn't kept. Because the text items are joined into one stream, multi-column pages, tables and side-by-side text can come out interleaved or run together, even though the words themselves are right.

Spacing can vary. The tool joins text items with spaces, which usually reads well, but some PDFs split or space words oddly, which can add or drop a space compared with the original.

Real-world uses

Reusing PDF text

Copy text out of a report or document.

Quoting a passage

Pull a section of text to quote elsewhere.

Searching content

Get the text out so you can search or analyse it.

Moving to an editor

Take PDF text into a word processor or your notes.

When it fits, and when it doesn't

Good for

  • PDFs with selectable text
  • Copying text out of a document
  • Reusing or quoting passages
  • Getting text for search or analysis

Not the best choice for

  • Scanned or photographed PDFs
  • Keeping columns and tables intact
  • Keeping the original formatting
  • Password-protected PDFs

For a scanned PDF, use an OCR tool, which reads the image and produces text. To keep the layout, export to a format that preserves it. This tool reads only the text layer the PDF already has.

Frequently asked questions

Why did it return no text?
Most likely the PDF is a scan: an image of the page with no text layer, so there is nothing to extract. To get text from a scan, use an OCR tool, which reads the image of the words.
What is a text layer?
It is the actual text a PDF stores behind the page image, the characters you can select and search. The tool reads this layer. A scan has only a picture, so nothing can be extracted.
How is a scanned PDF different?
A scanned PDF is a photograph or image of a page, so its words are pixels, not text. PDFs with selectable text store the words as text. Only those can be extracted here; scans need OCR first.
Does it keep the original layout?
No. Text items are joined into a single stream in reading order, so columns, tables and complex layouts run together. The words are correct, but their arrangement on the page is lost.
Are the pages separated in the output?
Yes. A blank line separates each page's text, so you can see where one page ends and the next begins.
Can it extract text from a password-protected PDF?
A protected PDF has to be unlocked first. If the document is encrypted, the tool can't read its text until the password is supplied or the protection is removed.
Why are there extra or missing spaces?
The tool joins stored text items with spaces, and some PDFs split or space words in unusual ways. That can add or drop a space compared with the original, though the words themselves are intact.
Is my PDF uploaded to extract the text?
No. Your browser opens and reads the PDF, so it is never uploaded and no copy is kept.
Will it get text from a form or a table?
Form field text and table cells that are stored as text can be extracted, but the structure is lost, so a table comes out as a run of values instead of rows and columns.
Can I extract just one page?
The tool extracts the whole document, page by page, into one box. To get a single page, pull that page out first with a page-extraction tool, then run it through this one.
Does it work on every language?
It reads whatever text the PDF stores, including many languages and scripts. If the characters are in the text layer, they come through.
What can I do with the extracted text?
Copy it into a document, your notes, a search box or an editor. It is plain text, so once it is out of the PDF you can reuse, quote, search or analyse it.

References

PDF.js reads the text layer the PDF already has. A scanned PDF is only a picture of text with no text layer, so it comes back empty and needs OCR instead.

PopularHot

Merge PDF

Combine several PDFs into one file in the order you choose, without leaving your browser.

PDFOpen Tool

Text to PDF

Turn plain text into a multi-page PDF set in Helvetica on US Letter pages.

PDFOpen Tool
Popular

Split PDF

Copy the pages you name, like 1-3,5, out of a PDF into a new file, in your browser.

PDFOpen Tool
Popular

PDF to JPG

Turn every page of a PDF into its own JPG image, rendered in your browser with PDF.js.

PDFOpen Tool
Popular

JPG to PDF

Turn JPGs or PNGs into one PDF, one image per page, each page sized to its image, with no re-compression.

PDFOpen Tool
Hot

Compress PDF

Re-save a PDF with object-stream compression. Text-heavy files get smaller and nothing on the page changes.

PDFOpen Tool

Back to the PDF to Text

Your file is processed on this device and never uploaded, and there's no account to create.