PDF TOOLS
How to Convert PDF to Word: What Survives, and What Does Not
8 min read · ToolsBay editorial · Published · Updated
Just want to do it now?
Extract the text of a PDF into an editable Word (DOCX) document.
"Convert PDF to Word without losing formatting" is a promise nobody can keep, including us. Some formatting always goes, and which parts go is predictable once you know what a PDF file actually holds. This post is the list: what comes through a conversion intact, what comes through mangled, what never comes through at all, and how to tell in ten seconds which kind of document you are holding.
A PDF page is a set of drawing instructions
Open a PDF in a text editor and, past the binary compression, a page is a stream of operators. Here is what a single line of a heading looks like in that stream:
BT
/F1 18 Tf
72 708 Td
[(W) -80 (aterfall) 20 (Repor) -15 (t)] TJ
ETThat is: begin text, use font F1 at 18 points, move to a point 72 units from the left and 708 units up from the bottom of the page, then show these glyphs with these adjustments between them. The numbers inside the array are kerning nudges, expressed in thousandths of an em and subtracted from the horizontal position, which is why the word "Waterfall" arrives split into two pieces with a number wedged in the middle.
Nothing in that stream says "heading". Nothing says "paragraph", "column", "table cell", "footnote" or "list item". A PDF records where ink goes. The structure you see on the page is something your eye assembles from spacing and size, and a converter has to assemble it again from the same evidence — coordinates and font metrics — with no access to what the author meant.
That is the whole story. Everything below follows from it.
What survives, more or less reliably
The characters themselves. If the PDF has a text layer, the words come out. This is the part that works.
Reading order within a single column of prose. PDF to Word does not trust the order the glyphs appear in the file — that is paint order, which generators emit in whatever sequence suits them. It groups glyphs into lines by vertical position and sorts each line left to right. For a report, a letter, a contract or a chapter of a book, this reproduces the text in the order a person reads it.
Bold and italic, when the font name admits it. pdf.js reports a font name per text item, and the converter marks a run bold if that name contains "bold" and italic if it contains "italic". This is a heuristic, and its edges are worth knowing. A font called Helvetica-Semibold matches "bold" and comes out fully bold. A font called Avenir-Heavy or Futura-Black matches nothing and comes out regular, despite sitting heavier on the page than the semibold did. An oblique face named Helvetica-Oblique comes out upright, because "oblique" is not "italic".
Relative text size. Each run keeps a size taken from the measured glyph height, so headings stay visibly larger than body text in the DOCX. It is a measurement, not the author's declared point size, so expect 11.5 where the original said 12.
A file Word will actually open. Text lifted out of PDFs routinely carries bytes that are illegal in XML — stray control characters, half of a surrogate pair left behind by a broken encoding. One of those inside document.xml makes the whole DOCX unparseable, and Word responds by offering to repair it. Those characters are stripped before the document is written. Unglamorous, and the difference between a file that opens and one that does not.
What does not survive
Multi-column layouts. Two columns side by side share baselines. Grouping by vertical position therefore puts the first line of the left column and the first line of the right column into the same line, sorts them left to right and joins them. An academic paper comes out as alternating half-sentences. There is no repair for this short of a layout analysis pass that works out where the column boundaries are, which this converter does not attempt.
Tables. The same cause, worse consequences: a row of cells is just text at similar heights, so it becomes one long line, and the borders you can see are unrelated vector drawings that are not carried across at all. Nothing in a DOCX produced this way is a real Word table.
Paragraph shape. Each visual line in the PDF becomes its own paragraph in the Word file. That keeps the page recognisable, but the text does not reflow — delete a sentence and the following lines stay where they are instead of sliding up. If you intend to edit heavily, expect to join lines back together first. This is the trade-off nobody mentions, and it is the one that costs the most time.
Hyphenation. A word broken across a line break was broken with a real hyphen character, drawn on the page. You get docu- at the end of one paragraph and ment at the start of the next.
Lists. A bullet is a glyph. It arrives as a literal • at the start of the line, not as Word list formatting, and the indent that made it look like a list is gone.
Images, colour, alignment, margins, and headers and footers as page furniture. The converter writes text runs and nothing else. Every image is dropped. A running header repeats as ordinary text at the top of each page's worth of content, because that is exactly what it is in the file.
Ligatures, and fonts with no `ToUnicode` map. A subset font needs a table telling the extractor which Unicode character each glyph stands for. When that table is missing or wrong — not rare in older exports and in some LaTeX output — the text extracts as plausible-looking nonsense, and no converter can fix it, because the information needed to fix it was never written down. Ligatures are a milder version of the same thing: fi may come out as the single character U+FB01 rather than the letters f and i, which then fails a search for "file".
Filled form fields. A form's values commonly live in annotation objects rather than in the page's content stream, so text extraction walks straight past them.
Word boundaries, occasionally. Because kerning splits a word into separate items, the converter has to choose between joining adjacent items and separating them. It separates them, on the grounds that a stray space inside "Waterfall" is easier to spot and fix than two words silently fused into one.
Why the output has page banners in it
The DOCX opens with --- Page 1 --- as a Heading 2, and carries one before every page after that. People assume it is a bug. It is there because a PDF page break carries information that a reflowable document has no way to express, and dropping it silently makes a fifty-page report harder to check against the original. Remove them with a find-and-replace once you have confirmed the conversion landed correctly.
Scanned PDFs are not a formatting problem
If your PDF came from a scanner, a phone camera, or one of the signature workflows that flattens everything on the way out, each page is a photograph. There are no glyphs, no font names, no coordinates — one image, drawn to fill the page. Text extraction returns nothing, and the converter says so rather than handing back a document containing only page banners.
Recovering words from a picture is optical character recognition: a trained model reading shapes and guessing letters, with an error rate attached. It is a genuinely different problem, and not one a browser-only tool solves without shipping a model to your machine first.
The fastest way to find out which kind of document you have is to run it through PDF to Text before anything else. Either text comes back or nothing does, and that answer tells you whether to convert or to go looking for OCR.
The cheapest fix is not to convert at all
If the PDF was exported from a Word document, ask whoever sent it for that document. Every conversion is reconstruction; the original is not. When the flow only needs to run the other way, Word to PDF is close to lossless by comparison, because a DOCX genuinely knows what its own headings, lists and paragraphs are — it is discarding structure rather than inventing it.
And if what you want is the words rather than the layout — pasting into a wiki, feeding a script, quoting a section — take the text and drop the pretence of preserved formatting entirely. PDF to Markdown uses relative font size to mark headings, and gives you structure you can trust precisely because it claims much less.
One last note on where this runs. All of it happens in the page you are on: the PDF is read into your browser, parsed there, and the DOCX is assembled there. Nothing is uploaded, which matters more for this tool than for most, because the documents people need to make editable are contracts, statements, medical letters and things under NDA. There is no retention policy to read, because there is nothing holding the file to retain it.