Word to PDF

Convert a .docx into a clean text PDF.

How it works

Add a .docx file and download a PDF. The document is read in your browser and rebuilt as a PDF with selectable text — headings, bold and italic, and lists all carry over. Because the rebuild happens here rather than in Word, spacing and page breaks will not match the original exactly.

Step by step

Add the file

Choose or drop a .docx file. It is read in the tab and never uploaded, which matters for the contracts and letters people usually convert.

Read the note about fidelity

The panel states plainly what carries over and what does not. This is worth a moment before you convert, because the answer determines whether this tool is the right one for your document.

Convert and download

The document is parsed, then rebuilt as a PDF with selectable text. Headings, bold and italic, bullet and numbered lists all survive, and the output is searchable and copyable rather than a picture of a page.

Compare against the original

Open both side by side. Line breaks and page breaks will fall differently. If the document is a letter or a contract, that is usually fine. If it is a designed layout, it will not be.

Good for

Use something else for

How it actually works

What a .docx file actually is

A .docx is a zip archive containing XML. Rename one to .zip and open it and you will find the document text in one file, styles in another, images in a folder. That open structure is why reading one in a browser is possible at all, and why the older binary .doc format is not — it is an undocumented format with decades of accumulated quirks.

Why exact fidelity needs Word itself

The XML describes content and style intent: this paragraph uses Heading 1, this run is bold. It does not describe where lines break or where pages end. Those are decided at layout time by an engine that knows the exact font metrics, the hyphenation dictionary, the justification algorithm and dozens of compatibility flags accumulated since the nineties. Word has that engine. Reproducing it faithfully is the reason commercial converters run Word or LibreOffice on a server rather than parsing the file directly.

What this converter does instead

It reads the semantic structure — headings, paragraphs, emphasis, list nesting — and lays it out fresh with its own spacing rules on Letter pages. The result is a clean, consistent document that is not a facsimile of the original. For text documents this is often an improvement on a Word file assembled with manual spacing. For a designed document it is a poor substitute.

What the XML actually looks like inside

Unzip a .docx and the body text sits in word/document.xml. Every paragraph is a w:p element containing one or more w:r runs, and a run is a span of text sharing identical formatting. Type a sentence and italicise one word in the middle and Word splits that paragraph into three runs. Styles are referenced by name and defined separately in styles.xml, so a heading is not marked up as a heading in any structural sense — it is a paragraph whose style reference happens to point at the Heading 1 definition. Reading a .docx faithfully therefore means resolving style references, walking nested run properties, and reassembling runs that Word split for reasons of its own. That is what this converter does, and it is why the semantic structure survives even though the visual layout does not.

Lists, and why they are the hardest part

Word does not store a list as a list. It stores a sequence of ordinary paragraphs, each carrying a numbering reference that points into numbering.xml, where the actual list definitions live with their levels, formats and restart rules. Nesting is expressed as a level number on each paragraph rather than by containment. Reconstructing a readable list means following those references, tracking which list each paragraph belongs to, and restarting counters correctly when a new list begins. Get it slightly wrong and numbering runs continuously through the whole document, which is the classic symptom of a naive converter.

Fonts and characters

The output uses the standard PDF font family, which is guaranteed present in every reader and needs no embedding. That keeps the file small and universally readable, at the cost of not matching whatever typeface the original used. Typographic characters that Word inserts automatically — curly quotes, en and em dashes, ellipses — are mapped to their nearest supported equivalent rather than dropped, so the text reads correctly.

When something goes wrong

It says only .docx works

The file is in the older .doc format. Open it in Word, Google Docs, or LibreOffice and save as .docx, then convert that.

My tables disappeared

Table cells are flattened to a single line of text separated by dots rather than dropped entirely, so the content survives but the grid does not. For a document where the table is the point, export to PDF from Word itself.

The page breaks are in different places

Expected, and explained above. The layout is rebuilt rather than reproduced. If page breaks matter, use Word’s own export.

Images are missing

Images are not embedded. This converter produces a text document. A file where the images matter needs a real converter.

Numbered lists restart oddly

Each list gets its own counter, and nested lists count independently. If a list in Word was manually numbered as plain text rather than using list formatting, it comes through as text and keeps whatever numbers were typed.

It could not read the file

The file may be corrupt or password-protected. Open it in Word and save a fresh copy.

Compared with the alternatives

Against Word’s own Save as PDF

If you have Word, use it. Nothing else will match its layout, because nothing else has its layout engine. This tool exists for the case where you do not have Word open, do not have it installed, or do not want the document leaving your machine.

Against Google Docs

Uploading to Google Docs and exporting as PDF gives good fidelity for straightforward documents and is free. It also means the document is on Google’s servers, which for a confidential file is the thing you were trying to avoid.

Against LibreOffice

LibreOffice is free, runs locally, and its Word compatibility is good though not perfect. For regular conversion of complex documents it is the right answer. It is a large installation for one conversion.

Against the upload converters

They typically run Word or LibreOffice server-side, so fidelity is genuinely high. The document goes to their server to get it. For a contract or anything under NDA, that is a real disclosure to weigh against a better-looking result.

Common questions

Will it look identical to my Word document?

Not exactly. Matching Word page for page needs Word’s own layout engine running on a server. This rebuilds the content faithfully — text, headings, emphasis, lists — with clean spacing of its own. For a contract or a letter that is usually what you want; for a heavily designed brochure it is not.

What carries over and what does not?

Headings, paragraphs, bold, italic, bullet and numbered lists all carry over, and the text stays selectable and searchable. Images, tables, columns, headers and footers are dropped.

Does .doc work, or only .docx?

Only .docx. The older .doc format is a different, undocumented binary format. Open it in Word or Google Docs and save as .docx first.

Guides


All tools · Guides · About · Privacy · Contact

Unpacking...