ScanPilot ← All Articles

How to Get Data from PDF to Excel: 4 Ways That Actually Work

August 20, 2026 · 4 min read · By ScanPilot Team
Use AI to summarize this article and ask questions

Getting data from a PDF into Excel is one of the most common office tasks there is, and one of the most quietly frustrating. The data is right there on the page; it just refuses to arrive in your spreadsheet as usable rows and columns. There are four practical ways to do it, and each fits a different kind of document. This guide covers all four honestly, including when the free built-in options are all you need.

If your document is a scan or the table is complex, you can skip straight to the AI PDF to Excel converter: the first page converts free with no signup.

Method 1: Copy-Paste (Fast, Fragile)

Select the table in your PDF viewer, copy, and paste into Excel. For a small, simple, digital table this sometimes just works, and it's always worth ten seconds to try.

Why it usually fails: PDF stores characters and positions, not tables. On paste, multi-line descriptions split into extra rows, columns merge, and numbers arrive as text that breaks SUM. On a scanned PDF there is nothing to select at all.

Use it when: the table is small, digital, and single-page, and you don't mind fixing a few cells.

Method 2: Excel's Built-In "Get Data from PDF" (Free, Digital PDFs Only)

Excel can import PDF data directly, and surprisingly few people know it:

  1. Open Excel (Microsoft 365 on Windows).
  2. Go to Data → Get Data → From File → From PDF.
  3. Pick your PDF. The Navigator shows each table and page Excel detected.
  4. Select a table, click Load (or Transform Data to clean it in Power Query first).

This is genuinely good for clean, digital PDFs with simple tables, and it's already on your machine. Its limits are structural: it cannot read scanned PDFs (no text layer, nothing detected), it frequently mis-splits tables with merged cells or wrapped text, and each page's table arrives separately, so multi-page statements need manual stitching.

Use it when: the PDF is digital, the table is well-behaved, and you're comfortable with Power Query for cleanup.

Method 3: Adobe Acrobat's Export (Decent, Subscription)

Acrobat Pro's "Export a PDF → Spreadsheet" produces an .xlsx and handles simple documents reasonably. It has limited OCR for scans, struggles with complex and multi-page tables, and requires the subscription. If you already pay for Acrobat, try it; if you don't, it's not worth subscribing for this task alone.

Use it when: you already have Acrobat Pro and the document is straightforward.

Method 4: AI Extraction (Scans, Complex Tables, Handwriting)

AI-based extraction reads the rendered page the way a person does: it finds the table, understands which values belong together, and rebuilds the structure as real rows and columns. That difference is what handles the documents the first three methods can't:

With ScanPilot's PDF to Excel converter, you upload the PDF, preview the first page as a table in seconds (free, no signup), and download the full document as XLSX, CSV, or JSON. For specific document types, there are tuned versions: bank statements, scanned PDFs, and handwritten documents.

Use it when: the document is scanned, the table is complex, the file is long, or you've already watched methods 1-3 scramble it.

Which Method Should You Use?

Your situation Best method
Small digital table, one page Copy-paste, then fix a few cells
Clean digital PDF, simple tables Excel's Get Data from PDF (free)
You already pay for Acrobat Pro Adobe export, with expectations managed
Scanned PDF or photo AI extraction (OCR + table detection)
Bank statement, invoice, register AI extraction tuned for the document type
Complex, multi-page, or handwritten AI extraction

Key Takeaways

Frequently Asked Questions

How do I get data from a PDF into Excel?

There are four practical ways: copy-paste (fast but breaks tables), Excel's built-in Get Data from PDF (good for clean digital PDFs), Adobe's export to spreadsheet (decent, subscription required), and AI extraction with a tool like ScanPilot (works on scans, complex tables, and handwriting). Which one fits depends on your document.

Can Excel import data from a PDF directly?

Yes. In Excel for Windows (Microsoft 365), use Data → Get Data → From File → From PDF. Excel detects tables in the PDF and lets you load them via Power Query. It works well on clean, digital PDFs, but it cannot read scanned PDFs and often mis-detects complex or multi-line tables.

How do I extract a table from a PDF to Excel?

If the PDF is digital and the table is simple, Excel's Get Data from PDF can pull it in directly. For scanned PDFs, complex layouts, or tables that continue across pages, an AI extraction tool reads the table structurally and returns clean rows and columns.

How do I get data from a scanned PDF to Excel?

Scanned PDFs have no text layer, so copy-paste and Excel's Get Data find nothing. You need OCR combined with table detection: upload the scan to an AI converter like ScanPilot, and the recognized table comes back as an editable spreadsheet. The first page converts free.

Why does my PDF data come out misaligned in Excel?

Most extraction methods rebuild tables from character positions. Multi-line cells, empty cells, and merged headers shift that reconstruction, so values land in the wrong columns. AI extraction reads the visual layout instead, which is why it keeps columns aligned.