Extract the text from a PDF in your browser: choose pages, keep line breaks or reflow into paragraphs, add page separators, then copy or download a .txt file.
PDF to Text pulls the words out of a PDF and gives you plain text you can copy or save as a .txt file, entirely in your browser. It reads the PDF's text layer with pdf.js, groups the words into lines using their position on the page, and lets you choose which pages to read, whether to keep the PDF's line breaks or reflow them into paragraphs, and how pages are separated in the output. It works on PDFs that were created digitally (exported from Word, a browser, a report generator). A scanned PDF has no text layer, so for those run OCR PDF first.
The PDF is probably a scan or a photo of a document. A scanned PDF contains only a picture of each page and no text for the tool to read, so those pages come out empty. The result lists how many selected pages had no selectable text. Run the file through OCR PDF to add a text layer, then extract the text from the result.
A PDF stores text as lines positioned on the page, not as paragraphs. Keep line breaks writes one output line per PDF line, which preserves the layout of code, addresses, poems and tables of short lines. Reflow joins the lines inside each paragraph with spaces, and rejoins a word that was hyphenated at a line end (exam- / ple becomes example), which reads better when you paste prose into an email or document. In both modes a larger vertical gap between lines is kept as a blank line.
Text is read line by line across the page, so when two columns share the same baseline their lines are merged into one, and the columns can interleave. If the PDF has columns, extract it, then fix the order by hand, or split the pages first. Tables are not reconstructed either: the cells of a row appear on one line separated by spaces.
PDF to Markdown also guesses headings from font size and turns bullets and numbered lines into Markdown, so it is meant for content you will publish or edit as Markdown. PDF to Text gives plain text with no formatting, page selection, optional reflow and page markers, which suits search, copying quotes, feeding text to another tool or word counting.
Yes. Deselect the pages you do not want in the thumbnail grid, or use Clear and click only the pages you need, then extract. You can also just select text in a PDF reader, but this tool is faster when you need several pages or want the page markers.
No. The output is plain text only: no fonts, bold, colours or images. If you need the pages as pictures use PDF to PNG or PDF to JPG, and for slides with editable text use PDF to PowerPoint.
No. The PDF is read by pdf.js in your browser tab and the text is built in memory, so nothing is transmitted. That makes it safe for contracts, statements and other private documents.
Not directly, because an encrypted PDF cannot be read without its password. If you know the password, remove the protection with Unlock PDF first and extract the text from the unlocked copy.
← All tools · PDF Tools · Blog · Cheat sheets · FAQ