PDF to Text
Pulls the text out of a PDF with lines, paragraphs, and page marks rebuilt.
Count the words in a DOCX, PDF, ODT, TXT, Markdown, or HTML file without opening it. The word count document summary also shows characters, sentences, paragraphs, pages, reading time, and the words you use most. Paste text if you prefer.
Choose a document or drop it here. DOCX, PDF, ODT, TXT, MD, or HTML. Up to 20 MB (PDF 100 MB).
| Word | Times | Share |
|---|
First the text is taken out of the file: DOCX with mammoth, PDF with pdf.js, ODT from its content.xml, HTML without its tags, scripts, and styles, and TXT or Markdown as it is. Then it is counted with the browser's own text segmenter (Intl.Segmenter), which knows where words and sentences end in every language, including ones written without spaces.
The quick brown fox jumps over the lazy dog counts as 9 words, 43 characters, 35 characters without spaces, and 1 sentence.Word, Google Docs, and this tool can disagree by a few words. The usual reasons: hyphenated words (well-known is one word here, as in Word), numbers and symbols (a lone dash is not a word), and text Word counts but a raw read skips, such as footnotes and text boxes. For school or publishing limits, the count in the app you submit from is the one that matters.
Drop the PDF on the box. The words are read from its text layer, and the page count is shown with the other numbers.
Yes. It uses the browser's word segmenter, which finds word boundaries even in languages that don't put spaces between words.
The word count divided by 238, the average silent reading speed of adults. Speaking time uses 150 words per minute.
No. The file is opened and counted in your browser.
Often used together with the Document Word Counter.
Pulls the text out of a PDF with lines, paragraphs, and page marks rebuilt.
Takes the plain text out of a Word .docx file, paragraph by paragraph.
Text in 12 cases at once, from Title Case to snake_case.