PDF to Text

Extract the text from a PDF in your browser: choose pages, keep line breaks or reflow into paragraphs, add page separators, then copy or download a .txt file.

About this tool

PDF to Text pulls the words out of a PDF and gives you plain text you can copy or save as a .txt file, entirely in your browser. It reads the PDF's text layer with pdf.js, groups the words into lines using their position on the page, and lets you choose which pages to read, whether to keep the PDF's line breaks or reflow them into paragraphs, and how pages are separated in the output. It works on PDFs that were created digitally (exported from Word, a browser, a report generator). A scanned PDF has no text layer, so for those run OCR PDF first.

How to use it

  1. Upload a PDF. Thumbnails of every page appear, all selected by default; click a thumbnail to leave that page out.
  2. Choose Keep line breaks (one output line per PDF line, good for code, addresses and lists) or Reflow paragraphs (lines inside a paragraph are joined, good for prose you will paste elsewhere).
  3. Choose what goes between pages: a "--- Page N ---" marker, a blank line, or nothing.
  4. Click Extract Text. The text appears in the box with a word and character count. Use Copy to put it on the clipboard or Download .txt to save it.

FAQ

Why is the extracted text empty or missing pages?

The PDF is probably a scan or a photo of a document. A scanned PDF contains only a picture of each page and no text for the tool to read, so those pages come out empty. The result lists how many selected pages had no selectable text. Run the file through OCR PDF to add a text layer, then extract the text from the result.

What is the difference between Keep line breaks and Reflow paragraphs?

A PDF stores text as lines positioned on the page, not as paragraphs. Keep line breaks writes one output line per PDF line, which preserves the layout of code, addresses, poems and tables of short lines. Reflow joins the lines inside each paragraph with spaces, and rejoins a word that was hyphenated at a line end (exam- / ple becomes example), which reads better when you paste prose into an email or document. In both modes a larger vertical gap between lines is kept as a blank line.

Why do multi-column pages come out jumbled?

Text is read line by line across the page, so when two columns share the same baseline their lines are merged into one, and the columns can interleave. If the PDF has columns, extract it, then fix the order by hand, or split the pages first. Tables are not reconstructed either: the cells of a row appear on one line separated by spaces.

How is this different from PDF to Markdown?

PDF to Markdown also guesses headings from font size and turns bullets and numbered lines into Markdown, so it is meant for content you will publish or edit as Markdown. PDF to Text gives plain text with no formatting, page selection, optional reflow and page markers, which suits search, copying quotes, feeding text to another tool or word counting.

Can I copy text from one page only?

Yes. Deselect the pages you do not want in the thumbnail grid, or use Clear and click only the pages you need, then extract. You can also just select text in a PDF reader, but this tool is faster when you need several pages or want the page markers.

Does it keep bold, fonts or images?

No. The output is plain text only: no fonts, bold, colours or images. If you need the pages as pictures use PDF to PNG or PDF to JPG, and for slides with editable text use PDF to PowerPoint.

Is my PDF uploaded to a server?

No. The PDF is read by pdf.js in your browser tab and the text is built in memory, so nothing is transmitted. That makes it safe for contracts, statements and other private documents.

My PDF is password-protected. Can I extract its text?

Not directly, because an encrypted PDF cannot be read without its password. If you know the password, remove the protection with Unlock PDF first and extract the text from the unlocked copy.

Related tools

Guides & comparisons

← All tools · PDF Tools · Blog · Cheat sheets · FAQ