PDF to Text
Pulls the text out of a PDF with lines, paragraphs, and page marks rebuilt.
Convert DOCX to text without opening Word: every paragraph of the document comes out as plain text you can copy, paste anywhere, or save as a .txt file. The document is read in your browser.
Choose a DOCX file or drop it here. Word documents (.docx), up to 20 MB.
Only the text is taken. Images, tables, and formatting are left out; table cells come out one per line.
A .docx file is a ZIP archive of XML files. The text lives in word/document.xml, split into paragraphs (w:p) and runs of text (w:t). The mammoth library reads that XML in your browser and writes out each paragraph's text in order, dropping fonts, colors, and images.
List items and table cells are paragraphs in Word too, so each comes out on its own line. Page breaks and section breaks leave no mark.
The sample document, a set of meeting notes, gives:
Project Kickoff Notes This sample Word document shows what the document tools can read. ... Decisions The team agreed to ship the first version in six weeks. ...
The headings become ordinary lines. A three-column table with two rows gives six lines, one per cell.
No. The document is read directly from the .docx file in your browser.
Not directly. The old .doc format is a different binary format. Save the file as .docx in Word, Google Docs, or LibreOffice, then open it here.
No, only the words. Bold, fonts, colors, and images are left out, which is what makes the text easy to paste anywhere.
No. It is opened and read in your browser, and nothing is sent to a server.
Often used together with the DOCX to Text.
Pulls the text out of a PDF with lines, paragraphs, and page marks rebuilt.
Makes a Word .docx file from plain text, with headings and your font.
Counts words, characters, sentences, and pages in DOCX, PDF, ODT, and text files.