Convert a .docx into a clean text PDF.
Add a .docx file and download a PDF. The document is read in your browser and rebuilt as a PDF with selectable text — headings, bold and italic, and lists all carry over. Because the rebuild happens here rather than in Word, spacing and page breaks will not match the original exactly.
Choose or drop a .docx file. It is read in the tab and never uploaded, which matters for the contracts and letters people usually convert.
The panel states plainly what carries over and what does not. This is worth a moment before you convert, because the answer determines whether this tool is the right one for your document.
The document is parsed, then rebuilt as a PDF with selectable text. Headings, bold and italic, bullet and numbered lists all survive, and the output is searchable and copyable rather than a picture of a page.
Open both side by side. Line breaks and page breaks will fall differently. If the document is a letter or a contract, that is usually fine. If it is a designed layout, it will not be.
A .docx is a zip archive containing XML. Rename one to .zip and open it and you will find the document text in one file, styles in another, images in a folder. That open structure is why reading one in a browser is possible at all, and why the older binary .doc format is not — it is an undocumented format with decades of accumulated quirks.
The XML describes content and style intent: this paragraph uses Heading 1, this run is bold. It does not describe where lines break or where pages end. Those are decided at layout time by an engine that knows the exact font metrics, the hyphenation dictionary, the justification algorithm and dozens of compatibility flags accumulated since the nineties. Word has that engine. Reproducing it faithfully is the reason commercial converters run Word or LibreOffice on a server rather than parsing the file directly.
It reads the semantic structure — headings, paragraphs, emphasis, list nesting — and lays it out fresh with its own spacing rules on Letter pages. The result is a clean, consistent document that is not a facsimile of the original. For text documents this is often an improvement on a Word file assembled with manual spacing. For a designed document it is a poor substitute.
Unzip a .docx and the body text sits in word/document.xml. Every paragraph is a w:p element containing one or more w:r runs, and a run is a span of text sharing identical formatting. Type a sentence and italicise one word in the middle and Word splits that paragraph into three runs. Styles are referenced by name and defined separately in styles.xml, so a heading is not marked up as a heading in any structural sense — it is a paragraph whose style reference happens to point at the Heading 1 definition. Reading a .docx faithfully therefore means resolving style references, walking nested run properties, and reassembling runs that Word split for reasons of its own. That is what this converter does, and it is why the semantic structure survives even though the visual layout does not.
Word does not store a list as a list. It stores a sequence of ordinary paragraphs, each carrying a numbering reference that points into numbering.xml, where the actual list definitions live with their levels, formats and restart rules. Nesting is expressed as a level number on each paragraph rather than by containment. Reconstructing a readable list means following those references, tracking which list each paragraph belongs to, and restarting counters correctly when a new list begins. Get it slightly wrong and numbering runs continuously through the whole document, which is the classic symptom of a naive converter.
The output uses the standard PDF font family, which is guaranteed present in every reader and needs no embedding. That keeps the file small and universally readable, at the cost of not matching whatever typeface the original used. Typographic characters that Word inserts automatically — curly quotes, en and em dashes, ellipses — are mapped to their nearest supported equivalent rather than dropped, so the text reads correctly.
The file is in the older .doc format. Open it in Word, Google Docs, or LibreOffice and save as .docx, then convert that.
Table cells are flattened to a single line of text separated by dots rather than dropped entirely, so the content survives but the grid does not. For a document where the table is the point, export to PDF from Word itself.
Expected, and explained above. The layout is rebuilt rather than reproduced. If page breaks matter, use Word’s own export.
Images are not embedded. This converter produces a text document. A file where the images matter needs a real converter.
Each list gets its own counter, and nested lists count independently. If a list in Word was manually numbered as plain text rather than using list formatting, it comes through as text and keeps whatever numbers were typed.
The file may be corrupt or password-protected. Open it in Word and save a fresh copy.
If you have Word, use it. Nothing else will match its layout, because nothing else has its layout engine. This tool exists for the case where you do not have Word open, do not have it installed, or do not want the document leaving your machine.
Uploading to Google Docs and exporting as PDF gives good fidelity for straightforward documents and is free. It also means the document is on Google’s servers, which for a confidential file is the thing you were trying to avoid.
LibreOffice is free, runs locally, and its Word compatibility is good though not perfect. For regular conversion of complex documents it is the right answer. It is a large installation for one conversion.
They typically run Word or LibreOffice server-side, so fidelity is genuinely high. The document goes to their server to get it. For a contract or anything under NDA, that is a real disclosure to weigh against a better-looking result.
Not exactly. Matching Word page for page needs Word’s own layout engine running on a server. This rebuilds the content faithfully — text, headings, emphasis, lists — with clean spacing of its own. For a contract or a letter that is usually what you want; for a heavily designed brochure it is not.
Headings, paragraphs, bold, italic, bullet and numbered lists all carry over, and the text stays selectable and searchable. Images, tables, columns, headers and footers are dropped.
Only .docx. The older .doc format is a different, undocumented binary format. Open it in Word or Google Docs and save as .docx first.