PDF to Text Converter

Convert PDF to text you can copy, search, and edit. Lines are rebuilt from where the words sit on the page, broken lines are joined back into paragraphs, and each page is marked, all in your browser.

Updated
Your files are processed in your browser and never uploaded.

Your PDF

Choose a PDF or drop it here. One PDF, up to 100 MB.

    Settings

    Only the text is taken. Images, tables, and formatting are left out; table cells come out one per line.

    Text

    The text of your file will appear here
    1. Choose or drop a PDF on the left
    2. The text is read right away
    3. Copy it or download a .txt file

    How to extract text from a PDF

    1. Choose a PDF or drop it on the box. The text appears right away.
    2. Keep Join lines into paragraphs on for flowing text, or turn it off to keep the PDF's line breaks, for example for poems or addresses.
    3. Keep Mark where pages start on to see which page each part came from.
    4. Press Copy, or Download .txt to save a text file.

    How it works

    A PDF does not store sentences. It stores small pieces of text with an exact position on the page, often a word or a few letters at a time. pdf.js reads those pieces, and the tool puts them back together:

    • Pieces whose baselines are within 40% of the font size of each other form one line, sorted left to right.
    • A gap wider than about a sixth of the font size between two pieces becomes a space.
    • A gap between lines more than 1.4 times the usual line spacing, or a change in font size, starts a new paragraph.
    • With joining on, the lines of a paragraph become one line, and a word split by a hyphen at the end of a line is put back together.

    Examples

    The first page of the sample report comes out as:

    --- Page 1 ---
    Utilza sample document
    
    Quarterly Field Report
    
    This sample document has three pages. Use it to try merging, splitting, ...

    A line that ends in inspec- followed by a line starting with tions is joined as inspections.

    Limitations

    • Scanned PDFs are pictures of text with no text layer. The tool says "No text found"; getting text from them needs OCR, which is not offered here.
    • Text in several columns may come out with the columns side by side on one line.
    • Tables come out as text, one row per line, without cell borders.
    • Password-protected PDFs can't be opened.

    Frequently asked questions

    Why does my PDF give no text?

    It is most likely a scan or a photo saved as PDF. The pages are images, so there are no letters to read. You need an OCR program to recognize the text.

    Can I keep the line breaks exactly as in the PDF?

    Yes. Turn off Join lines into paragraphs, and every line of the PDF stays on its own line.

    Does it work with PDFs in other languages?

    Yes. Text in any language is read, including Arabic, Chinese, and Cyrillic, as long as the PDF has a text layer.

    Is my PDF uploaded to extract the text?

    No. The text is read in your browser, and nothing is sent to a server.

    Often used together with the PDF to Text.

    • DOCX to Text

      Takes the plain text out of a Word .docx file, paragraph by paragraph.

    • Document Word Counter

      Counts words, characters, sentences, and pages in DOCX, PDF, ODT, and text files.

    • PDF to Images

      Saves each PDF page as a JPG or PNG image at 72 to 300 DPI.