PDF to text
Get the words out of a PDF as a plain text file. Paragraphs are rebuilt, and everything else is left behind.
How the files you choose are handled is set out in the privacy policy.
How to convert a PDF to text
- Tap Choose a PDF.
- Tick Mark where each page starts if you want a line such as “Page 3” between pages.
- Press Convert to text, then Save text file.
What a text file gives you
A .txt file holds letters and line breaks and nothing else. That is its strength. It opens in every editor on every device, pastes cleanly into an email or a form, and can be searched, compared or fed to another program without formatting getting in the way.
| Kept | Left behind |
|---|---|
| Every word, in reading order | Fonts, sizes, bold and colour |
| Paragraphs, with wrapped lines rejoined | Pictures, charts and logos |
| Headings, each on its own line | The layout of tables: cells arrive as lines of text |
| Accented letters and non-Latin scripts | Columns, headers placed in margins, exact positions |
Why use text, and not Word?
- Quoting and pasting. Copying from a PDF viewer often breaks every line. The text file has whole paragraphs.
- Searching many documents. Text files can be searched together by any file manager or code editor.
- Feeding other tools. Translation, summarising, word counts and text-to-speech all work best on plain text.
- Small size. A 5 MB report is often 50 KB of text.
If you want to keep headings as headings and go on editing in a word processor, PDF to Word is the better choice.
Scans have no text
A PDF can be a picture of a page. Scans and photographed documents look like text to a person, but hold no letters for a program to read. If your PDF is one of those, the tool reports that no text was found.
A quick test: open the PDF and try to select a word. If the selection takes a word, there is text. If it takes nothing, or draws a box around the whole page, it is a picture. Getting text from a picture needs text recognition (OCR), which this site does not offer. Many phone scanning apps can save a scan with recognised text, and such a file converts normally.
Odd results and what causes them
- Two columns run together. A PDF records positions, not reading order. Most multi-column pages come out correctly, but dense layouts can interleave.
- Repeated headers. A running title printed on every page appears once per page in the text. Page markers make these easy to find and delete.
- Strange symbols. A few PDFs use fonts that do not record which letter each shape stands for. Their text cannot be recovered correctly by any converter.
The file is saved as UTF-8, which every current editor reads correctly.
Questions
How do I copy all the text from a PDF?
Convert it here and open the text file. Select all, copy, and paste where you need it.
Can it read a scanned PDF?
No. A scan is a picture of the page and contains no text. It needs text recognition (OCR), which this tool does not do.
Does it keep tables?
The words in a table are kept, but its grid is not. Each row or cell arrives as a line of text.
What encoding is the text file?
UTF-8, with Windows-style line endings so it opens correctly in Notepad as well as on phones and Macs.
Can I convert a password-protected PDF?
Yes, if you know the password. You are asked for it when you choose the file.