How to Extract a Table from a PDF to Excel
Jun 21, 2026 • 7 min read
How to extract a table from a PDF to Excel three ways (copy-paste, Excel Power Query, and AI extraction) and which one keeps your rows and columns intact.
// Try it now, no signup required
PDF, JPG, PNG, BMP, HEIC, TIFF
Upload a document to extract
Drop files here or click to upload
Up to 50 files
Free plan extracts the first 5, rest can be unlocked after
Uploading...
Free on your own files. No credit card, no signup to test.
A table inside a PDF looks like data, but it is really just lines and text positioned on a page. The moment you try to move it into a spreadsheet, that structure falls apart. You select the rows, paste them into Excel, and everything lands in a single column, or the numbers shift a row out of place and stop lining up with their labels.
This is one of the most common document chores there is, and it has the same fix whether the table is a list of invoice line items, a supplier price list, a row of figures in a report, or a scanned data table. This guide covers how to extract a table from a PDF to Excel three ways, by hand, with Excel itself, and with AI, and how to keep your rows and columns intact through the process.
Why copy-paste and basic converters break tables
Copy and paste is the first thing everyone tries, and it almost always disappoints. A PDF stores a table as text fragments placed at coordinates, not as cells. When you copy a block and paste it into Excel, the column boundaries are gone, so several columns collapse into one cell and you are left untangling them by hand. Rows can also drift, especially where a label wraps onto two lines, which pushes the figures out of alignment with the wrong account or product.
Basic converters help a little but hit the same walls. Multi-page tables get split, so a header on page one and its continuation on page three become two disconnected blocks. Merged cells and nested headers confuse the column detection. And if the table is a scan or a photo, there is no selectable text at all, so copy-paste and most simple converters return nothing usable. You need a tool that recognizes the table as a structure, not one that lifts raw text.
Three ways to extract a table from a PDF to Excel
There are three practical methods, and the right one depends on how many PDFs you process and how clean they are. Here is an honest look at each.
1. Manual copy-paste or retyping
For a single, short table you can copy it across, fix the columns by hand, or just retype the values. This is fine when you have one document and a few rows, and it needs no software beyond Excel. The trouble is that it does not scale and it introduces errors. Every transposed digit and misaligned row is a mistake you have to catch later, and keying a long table is slow and tedious. Once you are dealing with more than a handful of tables, manual work becomes the bottleneck.
2. Excel's Data > From PDF (Power Query)
Recent versions of Excel can import tables directly. Go to Data, then Get Data, then From File, then From PDF, point it at the document, and Power Query previews the tables it found so you can load the ones you want. For a native, digital PDF with a clean, simple table this works well and keeps the columns separated. It is built in, so there is nothing extra to install.
The limits show up quickly on harder documents. Power Query reads the PDF's underlying text layer, so a scanned or photographed table returns nothing. It also struggles with complex layouts, multi-page tables, merged cells, and unusual headers, and it can detect the wrong table boundaries and need manual cleanup. It is a good first choice for clean digital PDFs and a poor fit for scans and messy reports.
3. AI extraction
An AI extraction tool combines OCR with a model that understands table structure. It reads both digital and scanned tables, keeps each row and column where it belongs, follows a table across multiple pages, and exports a clean .xlsx or .csv. Because it recognizes the table as a structure, it handles the layouts that break copy-paste and Power Query. The trade-off is that it is a separate tool rather than a feature already in Excel, and like any automated method it benefits from a quick review on lower-quality scans. For volume and for mixed document quality, this is the approach that holds up.
How to extract a scanned PDF table
A scanned PDF is an image, so there is no text to select and no table for Excel to find. To get anything out of it you need optical character recognition. OCR turns the image into text, and on its own that is only half the job, because raw recognized text still has no column structure. The step that matters is the layer on top of OCR that maps the recognized text back into rows and columns, so the data lands in a spreadsheet shaped the way it looked on the page.
Scan quality drives the result. A sharp, high-contrast, straight scan reads far better than a faint, skewed, or low-resolution photo. When a value is hard to read, a good tool flags it rather than guessing silently, so you can confirm it before the number is used.
Step-by-step with DocuOCR
DocuOCR is built to pull tables out of PDFs and return them as structured spreadsheets, including from scans. The flow is short:
- Upload the PDF. Drop in the file, digital or scanned. There is nothing to configure to get started.
- The AI detects the table and its columns. It locates the table on the page, identifies the header and each column, and reconstructs the rows, following the table across pages where it continues.
- Review any low-confidence cells. Each cell carries a confidence score, so instead of checking everything you confirm only the few values the model was unsure about, which is where poor scans usually need a human eye.
- Download as Excel or CSV. Export the finished table to .xlsx or .csv with the columns and rows intact, ready to sort, filter, or load into another system.
Accuracy depends on the document. On clean tables, extraction is typically in the 95 to 99 percent range, and on faded or skewed scans it is lower, which is exactly why the confidence scores and review step exist. The honest way to judge any tool, including this one, is to run it on your own worst documents rather than a clean sample, and to keep a person in the loop for the figures that feed a model or a report.
Tips for cleaner results
- Start with a good scan. Higher resolution, straight pages, and clean contrast all improve recognition. A better input is the cheapest accuracy gain there is.
- Keep layouts consistent. When the same kind of table always arrives in the same format, detection is more reliable and your review is faster.
- Check the totals. Reconcile any subtotals or totals against the rows above them after export, and confirm that figures in parentheses came through as negative values.
- Use the API for volume. If you are processing tables in bulk, send them through the API so extraction runs automatically and you only handle the exceptions.
Getting a table out of a PDF and into Excel does not have to mean fighting collapsed columns and misaligned rows. If you have a single clean PDF, Excel's own importer may be enough; if you are working through scans, complex layouts, or volume, an AI tool that understands table structure will keep your data intact. You can try it on your own file with the PDF to Excel converter, and the same engine powers broader PDF data extraction software for the rest of your documents. For one of the most common tabular documents of all, a row of transactions, a dedicated bank statement OCR tool handles the per-bank layouts straight to a spreadsheet.
Extract your documents with DocuOCR
DocuOCR's AI OCR software turns any document into clean, structured data in seconds. No template setup required.
Start free