DOCX to Text
Takes the plain text out of a Word .docx file, paragraph by paragraph.
Convert PDF to text you can copy, search, and edit. Lines are rebuilt from where the words sit on the page, broken lines are joined back into paragraphs, and each page is marked, all in your browser.
Choose a PDF or drop it here. One PDF, up to 100 MB.
Only the text is taken. Images, tables, and formatting are left out; table cells come out one per line.
A PDF does not store sentences. It stores small pieces of text with an exact position on the page, often a word or a few letters at a time. pdf.js reads those pieces, and the tool puts them back together:
The first page of the sample report comes out as:
--- Page 1 --- Utilza sample document Quarterly Field Report This sample document has three pages. Use it to try merging, splitting, ...
A line that ends in inspec- followed by a line starting with tions is joined as inspections.
It is most likely a scan or a photo saved as PDF. The pages are images, so there are no letters to read. You need an OCR program to recognize the text.
Yes. Turn off Join lines into paragraphs, and every line of the PDF stays on its own line.
Yes. Text in any language is read, including Arabic, Chinese, and Cyrillic, as long as the PDF has a text layer.
No. The text is read in your browser, and nothing is sent to a server.
Often used together with the PDF to Text.
Takes the plain text out of a Word .docx file, paragraph by paragraph.
Counts words, characters, sentences, and pages in DOCX, PDF, ODT, and text files.
Saves each PDF page as a JPG or PNG image at 72 to 300 DPI.