PDF TOOLS
PDF Table Extraction: Why It Fails, and What Works Instead
8 min read · ToolsBay editorial · Published · Updated
Just want to do it now?
Render an Excel spreadsheet as a paginated PDF table.
Copy a table out of a PDF, paste it into a spreadsheet, and you get one long column of run-together text. Every value is there. The grid is gone. That is not your PDF reader being lazy, and it is not a setting you failed to find. It follows from what the file actually holds.
What is inside the file
I built a small sales table as a PDF by hand, uncompressed, so the inside is readable. The whole file is 884 bytes. The part that paints the table is 340 of them:
BT
/F1 10 Tf
1 0 0 1 72 700 Tm (Region) Tj
1 0 0 1 220 700 Tm (Units) Tj
1 0 0 1 320 700 Tm (Revenue) Tj
1 0 0 1 72 684 Tm (North) Tj
1 0 0 1 220 684 Tm (1,204) Tj
1 0 0 1 320 684 Tm (48,160.00) Tj
ET
0.5 w
64 692 m 420 692 l STm sets a position and Tj paints a string there. m, l and S move to a point, line to another point, and stroke the line between them. That is the whole vocabulary: put these glyphs at this coordinate, draw this rule between those two.
Search the file for "Table" and you get zero hits. Same for "Row" and "Cell". The header Revenue and the value 48,160.00 below it are two unrelated drawing instructions that happen to share an x coordinate. The rule between the header and the body is a third instruction, and nothing anywhere says it belongs to either. A person looking at the page infers a table. The file does not contain one.
A real generator compresses that stream and subsets the font, so you will not see it in a text editor. The operators are identical.
The same file, two different answers
Here is poppler's pdftotext on that file with no flags:
Region North South
Units 1,204 987
Revenue 48,160.00 39,480.00Same file, same binary, with -layout:
Region Units Revenue
North 1,204 48,160.00
South 987 39,480.00The first result is not a bug. Given three blocks of text separated by wide horizontal gaps, poppler's default reading-order analysis decided the page had three columns and read each one top to bottom. That is right for a newspaper and wrong for a table, and nothing in the file distinguishes the two cases. Every extractor has to make that call, and it makes it with a heuristic.
Ruled and unruled tables are different problems
Extraction splits into two families, and the good libraries let you pick.
Follow the lines. Find the strokes, find where they intersect, treat each enclosed box as a cell. Camelot and Tabula both call this lattice mode; pdfplumber calls it the lines strategy. When a table is fully ruled it is close to exact, because the generator has told you where the cell boundaries are. It falls over when the "lines" are not strokes at all — plenty of HTML-to-PDF engines draw a border as a very short filled rectangle — or when the rules live in a background image, or when only the header row is ruled and the body is bare.
Follow the whitespace. With no rules to follow, group text runs into rows by their y position, then look for vertical gaps that appear in every row and call those the column boundaries. This is the mode most people end up in, and it is fragile in specific ways: one long description that spills into the next gutter destroys that boundary for the entire table; a right-aligned number column closes its own gutter as soon as one value is wider than the rest; a header cell spanning two columns merges them; and a row that wraps onto two lines gets read as two rows.
Camelot will hand back an accuracy figure and a whitespace figure per table in its parsing report. Those are worth reading as a smell test, but they describe how tidily text landed inside the cells the algorithm chose. They are not a check that the numbers are the right numbers.
The values may already be damaged
This is the part that gets missed, and it is the reason the best fix is not a better extractor.
A PDF holds what was printed, not what was computed. If the source cell contained 48159.9962 and the column was formatted to two decimals, the file holds 48,160.00 and the remaining digits were destroyed at export time. Nothing downstream recovers them. Sum the extracted column and your total will disagree with the total printed at the bottom of the page — not because extraction failed, but because you are adding rounded numbers and the page added the unrounded ones.
The same applies to everything else formatting does. Negatives may be printed as (1,234) or with a trailing minus. Thousands separators in European exports are often a narrow or non-breaking space, which whitespace clustering can read as a column boundary. A percentage displayed as 12% was probably 0.12. A numeric column that was too narrow in the source spreadsheet can print as a row of hashes, and hashes are what the PDF contains.
You are not extracting the data. You are extracting a picture of a formatted view of the data.
Multi-page tables are many tables
A twenty-page table is not one table to an extractor. It is twenty detected regions across twenty pages, each returned separately.
The header is repeated at the top of every page, so nineteen header rows land in the middle of your data. Page numbers, footers and "continued" notes sit inside or beside the detected region and arrive as rows of their own. Subtotal rows look exactly like data rows. A row whose text wraps can straddle the page break and come back as two half-rows.
Stitching is the step that gets skipped, and it corrupts results quietly rather than loudly: a stray header row parked in a column of numbers is text, and a spreadsheet's SUM skips text without complaining. The total looks plausible and is short by a page.
Going the other way is easy for exactly the reason this direction is hard. Excel to PDF repeats the header on every page deliberately, because a spreadsheet does know where its cells are and can say so.
Scans, and what OCR actually buys
If PDF to Text comes back with nothing, the pages are images. The tool says so rather than handing you an empty file and calling it a success, and that answer is worth having early — it rules out every other approach on this page.
OCR converts pixels into positioned words with a per-word confidence score. It does not convert them into a table. You land back at the identical geometry problem, except now every value is itself a guess: 0 against O, 1 against l, 5 against S, 8 against B, and a decimal point a few pixels wide. Rules on a page scanned a degree off true stop being straight lines, so line detection needs the image deskewed first. OCR moves the starting line back; it does not move the finish line closer. Every figure from an OCR'd table needs checking against the page.
The exception worth checking first
A PDF can carry a structure tree: a /StructTreeRoot holding real Table, TR, TH and TD elements alongside the drawing instructions. Tagging is required by PDF/UA and by the a conformance levels of PDF/A, and Word produces it when you export with document structure tags for accessibility enabled. Where those tags exist the cell boundaries were recorded rather than inferred, and extraction becomes a read instead of a guess.
Most PDFs in the wild are untagged — invoices out of an ERP, bank statements, regulatory filings. My hand-built file has no structure tree either. But checking costs a minute: qpdf --qdf or mutool will show you whether one is present. A plain grep for StructTreeRoot finds it in many files, though it will miss the ones whose catalog sits inside a compressed object stream.
Why this site has no PDF to Excel
ToolsBay does not ship one. Ask for /tools/pdf-to-excel and you are redirected to the text extractor instead.
The reason is everything above. The converter we could put on that URL would be the whitespace heuristic, running on a document we cannot see, with no dial to turn and no way to tell you when it guessed. It would produce a spreadsheet-shaped file most of the time and a wrong spreadsheet-shaped file some of the time, and you would have to check every value regardless. A tool that is usually right about numbers is worse than no tool, because it moves the checking to a step people skip.
What to do instead, cheapest first
Go upstream. Almost every PDF table is the print view of something else: a query result, a report, an export. Ask whoever sent it for the CSV or the workbook. If what they can give you is an API response, JSON to CSV flattens it into columns without any of this. This path works far more often than people expect, and it is the only one that gives you the unrounded values.
One table, once. Extract the text and reshape it by hand. Tedious, correct, and usually finished in less time than it takes to tune an extractor.
A repeating pipeline. Camelot, Tabula or pdfplumber, pinned to that specific layout, with a check that fails loudly — compare the row count and at least one column total against the figures printed on the page. Tabula's interface lets you draw the table region yourself, which beats automatic detection on a document you process once a month.
A scan. OCR first, then treat every figure as unverified until you have read it.
The honest answer to "how do I convert this PDF to Excel" is usually that you should not convert it at all. You should go one step back and ask for the thing the PDF was made from.