space ocr
GuidesArticlesPricingDocs
documents

How to Extract Table Data from an Image to CSV

Turn a photo of a table, order sheet, or delivery slip into a clean CSV file. See how space-ocr reads line items and how row-level verification surfaces the values worth checking.

4 min read· 2026-08-31

Getting data from a scanned table into a spreadsheet is a classic chore. You have a crisp image of a delivery slip or a purchase order, full of line items. But it's just pixels. The next step is usually tedious, manual data entry, copying each item, quantity, and price into a new row, one by one. The process is slow and a single typo can throw off your entire dataset.

A table-style order/delivery document
A line-item table — many rows, one consistent shape out.

A better approach is to define the table's structure as a schema. Instead of pulling out one block of text, you declare the columns you need. The repeating line-item section becomes an array field with its own children, and the columns that always hold a number get a declared type.

1
2
3
4
5
6
7
8
9
10
{
  "name": "items",
  "type": "array",
  "children": [
    { "name": "name",   "type": "string" },
    { "name": "qty",    "type": "integer" },
    { "name": "price",  "type": "number" },
    { "name": "amount", "type": "number" }
  ]
}

You never state how many rows the page has — that is up to the document. Each row comes back at an indexed path (values.items[0], values.items[1], and so on), and the same path is the key into the per-value coordinate and verification map, cells["items[0].price"]. A declared type does not change what is extracted, because the type is never shown to the model; it adds a second, deterministic layer beside the values, data.normalized.

Define an array field for line items, then upload an image to extract the table into a structured grid.

This works even for dense tables with repeating values, which can be a challenge. The system uses a large language model to propose the initial extracted text, but it doesn't stop there. For each value, like an item name of "刻みたくあん" or a price of "580", it cross-validates the result. The engine checks the language model's reading against the document's column structure and performs a character-by-character match against the symbols originally detected on the page. When a value slides into the neighbouring row, the characters at those coordinates usually stop agreeing, and the field is flagged for review instead of passing quietly.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
"data": {
  "values": {
    "items": [
      { "name": "刻みたくあん", "qty": "3",
        "price": "580", "amount": "1,740" }
    ]
  },
  "cells": {
    "items[0]": {
      "box": { "xmin": 263, "ymin": 460,
               "xmax": 738, "ymax": 523 },
      "quad": [ { "x": 263, "y": 460 }, { "x": 738, "y": 460 },
                { "x": 738, "y": 523 }, { "x": 263, "y": 523 } ],
      "verified": null, "review": null
    },
    "items[0].name": {
      "box": {…}, "quad": […],
      "verified": true, "review": null,
      "evidence": { "text_match": true, "match_ratio": 1.0 }
    },
    "items[0].qty": { "box": {…}, "quad": […],
                      "verified": true, "review": null },
    "items[0].price": {
      "box": { "xmin": 693, "ymin": 460,
               "xmax": 738, "ymax": 488 },
      "quad": […],
      "verified": false,
      "review": { "reasons": ["text_mismatch"] },
      "evidence": { "text_match": false, "match_ratio": 0.62 }
    }
  },
  "review": {
    "unit": "field",
    "flagged": [
      { "path": "items[0].price", "reasons": ["text_mismatch"] }
    ]
  },
  "normalized": {
    "items": [ { "qty": 3, "price": 580, "amount": 1740 } ]
  }
}

The flagged list is the work queue: data.review.flagged names the exact path and the reasons behind it, and cells["items[0].price"] holds the coordinates that were checked. It is not a catch-all. A neighbouring row printing the same number reads as consistent wherever the box landed, and when both engines make the same misread there is nothing left to disagree with. So the row and column errors character matching cannot see are worth handing over as declared rules: pattern for a code column, enum for values your master data already holds, min and max for a plausible range. Every violation arrives in that same review.flagged list.

The last layer is your own arithmetic, and it belongs downstream. Because qty was declared integer and the money columns number, data.normalized carries them parsed — "1,740" is 1740 — so a line check costs nothing beyond the multiplication.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
const { values, normalized, review } = data;

const rows = values.items.map((item, i) => {
  const n = normalized.items[i];
  const flagged = review.flagged.some((f) =>
    f.path.startsWith(`items[${i}]`)
  );
  return {
    name: item.name,
    qty: n.qty,
    price: n.price,
    amount: n.amount,
    // look at this row by hand: flagged, or the line does not tie
    check:
      flagged || n.qty == null || n.qty * n.price !== n.amount,
  };
});

A leaf that will not parse comes back null in normalized, with the kind of failure in cells[path].normalized.error and type_mismatch in the review list. Compute from normalized, and keep values for what you print — that is the side carrying the coordinates and the verification.

Once data is extracted, a single click exports the entire sheet to a clean CSV file.
✓ Verified

Each value is held against the page it came from. The model's reading is matched character by character against the OCR symbols actually detected there, and evidence.match_ratio records how much of it lined up; 0.85 or higher is a confident match. The verdict itself is verified, the mirror of review: false when anything was flagged, true when a check ran and nothing was, null when there was nothing to check — a row's union box, for instance. Coordinates come from the matched symbols as box and quad, normalized to a 0–1000 scale against data.image. That is evidence of where a value came from rather than proof that it is the value you wanted, and the values that did not check out are the ones listed in review.flagged.

The cost is based on usage, at $0.05 per image processed. Your account includes 100 free scans every month. If an extraction fails for any reason, there is no charge.

  1. Define a Sheet Schema
    Create a new Sheet and define your columns. For line items, use the 'array' type and add child columns for name, quantity, price, etc.
  2. Upload Your Image
    Drag and drop or use the API to upload an image of the table to the Sheet.
  3. Review the Extracted Data
    The image will be processed against your schema. Each line item from the table appears as a structured row in the Sheet.
  4. Correct if Needed
    Click on any cell to see the corresponding area on the image. Manually correct any values directly in the grid.
  5. Export to CSV
    Click the 'Export' button and choose CSV. Your table data, including all line items, is downloaded as a clean, structured file.
What if my table has merged cells or a complex layout?
The system is designed for standard row and column tables. For highly complex layouts, you can define multiple schemas or manually adjust the data in the sheet after the initial extraction.
How does CSV export handle the line items?
If your array column is named 'items' and has children 'name' and 'price', the CSV will have headers 'items.name' and 'items.price'. Each line item from the image becomes a separate row in the CSV file.
Can I process PDF files with tables?
Yes, in the web app. You can drop a PDF file, and it will automatically render each page as an image for processing. The API itself accepts raster image formats like JPEG and PNG.
How are the coordinates for each cell determined?
For each extracted value, the system matches its characters against the OCR symbols detected on the page. The matched symbols give the value its `box` and `quad`, normalized to a 0–1000 scale against `data.image`, the page as it was read. Those coordinates are the evidence for where the value came from, so you can draw them back onto the image and check the value yourself.
Is there a limit to the number of rows in a table?
There is no row count to configure — how many rows come back is up to the page. What does have a ceiling is a synchronous call: processing is capped at 180 seconds, and going over returns 504 `ocr_engine_timeout`, which is not charged. Density rather than pixel count is the usual cause, so for a very dense table send one page per image, or use the asynchronous `POST /upload`.

Turn Your Image Tables into Data

Get 100 free scans every month. No credit card required to start.

Related