space ocr
GuidesArticlesPricingDocs
AI OCR

AI OCR that you don't have to take on faith

space-ocr structures documents with a model, then cross-checks every value against the page: data.cells[path] returns box, quad, verified and review, and data.review.flagged is the review queue.

AI OCR sounds like the answer to messy documents: hand a receipt or an invoice to a model and get clean, structured fields back. The trouble is what happens when the model is wrong. A language model returns a confident, well-formatted value whether or not it actually read it off the page, and most tools pass that value on with no way to tell the difference.

space-ocr splits the job instead. A multimodal model does the structuring and never produces coordinates; a separate OCR pass reads the page and is the only source of geometry. The extracted value is then matched character by character against the symbols that pass detected. The response is separated the same way: data.values holds the business data in the schema you declared, and data.cells[path] holds the box, the quad, the verified verdict, the review reasons and the supporting evidence for that same path. Which OCR and model implementations run underneath is an implementation detail that can change — the response contract is what stays stable.

See the AI's output, checked

Hover any field below — the box on the receipt is where that value was actually located on the page, not where the model claimed it was. The values, boxes and verification flags here are read from a real parsed result, not a mockup.

Receipts with extracted-field bounding boxes
Verified fields
KINSHO · 合計 2,045
ライフ · 合計 4,286

Each value with a box carries a verified on-page location — in data.cells[path], that is box + 4-point quad + evidence.match_ratio — on a 0–1000 normalized grid (0,0 top-left → 1000,1000 bottom-right), the same shape the live API returns. Hover a field to trace it back to the pixels it came from.

Three shapes, one contract
Take the page as declared fields (`POST /ocr/fields`), as layout-preserving Markdown (`POST /ocr/markdown`), or as plain text in reading order (`POST /ocr/text`). All three return the same envelope — `data.values`, `data.cells`, `data.review`, `data.image` — so one review screen can read every endpoint. Markdown includes its elements by default; on `/ocr/text`, set `includeBlocks: true` to get per-block cells.
The model never returns coordinates
Values come from the model; geometry comes from the OCR pass and from matching each value against it. `evidence.source` records which mechanism produced the box, and on `/ocr/fields` `evidence.printed_text` returns the raw glyphs at those coordinates, so you can compare them with the value yourself.
Every value addressable by path
The keys of `data.cells` — `total`, `items[0].amount` — use the same path grammar as `data.review.flagged`. `box` is axis-aligned on a 0–1000 normalized grid, `quad` is four points that follow the tilt of the page, and `data.image` gives the width and height that convert either into pixels.
Declare the fields, or let the model propose them
Pass `fields` with `name`, `type` and `children`, or set `autoFields` and let the model suggest the structure. Declarations such as `required`, `pattern`, `min`/`max`, `enum` and `near` are never shown to the model: they are checked after extraction, so a violation becomes a review reason rather than a rewritten value. Declaring a scalar type adds `data.normalized`, a deterministic parse kept in its own layer.
Audit trail: original beside edited
Extraction runs through a model and is not deterministic run to run, so the response JSON is the record worth storing. In the app, correcting a cell saves your value beside the original OCR value instead of overwriting it — what the model read and what a person changed both stay visible.
Line items checked row by row
In an `array` field each row gets its own paths (`items[0].amount`) and the row itself carries a union box. A column of repeated values is where a model's own token hints are least reliable, so the engine leans on column and row consistency and raises reasons such as `ambiguous_occurrence` instead of taking the row on trust.
No language setting
Japanese, Korean, Chinese and English are handled by one engine, mixed scripts included. There is no language parameter in the public API and nothing to configure per document.

How AI OCR works in space-ocr

Send an image to POST /ocr/fields with imageType set to url or base64. An OCR pass reads the page first, and that pass is the only source of coordinates. A multimodal model then reads the document into the schema you declared and returns values only. Matching each value, character by character, against the detected symbols is what produces the box, the quad and the evidence in data.cells[path].

verified is a verdict, not a character score. It is false whenever the cell carries review reasons of any kind, true when a check ran and nothing was flagged, and null when there was nothing to check. The character comparison itself is evidence.text_match, which is why verified: false together with text_match: true is a normal combination: the glyphs agreed, and a rule you declared caught the value anyway.

Silent mismatches get surfaced this way, but nothing here promises to catch every error. The model and the OCR pass are independent and can still agree on the same misread. Coordinates are evidence of where a value came from, not proof that it is correct, so keep your own business rules downstream.

You do not have to write a schema. Declare fields, or set autoFields and let the model propose the structure. The web app rasterizes PDF pages before reading them; the public API takes raster images directly.

declare the fields — every value comes back located and reviewed
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
curl -s https://api.space-ocr.com/ocr/fields \
  -H "Authorization: Bearer $SPACE_OCR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "image": "https://example.com/receipt.jpg",
    "imageType": "url",
    "fields": [
      { "name": "store_name", "type": "string", "required": true },
      { "name": "date", "type": "date", "required": true },
      { "name": "total", "type": "number", "required": true, "min": 0 },
      {
        "name": "items",
        "type": "array",
        "children": [
          { "name": "name", "type": "string" },
          { "name": "amount", "type": "number" }
        ]
      }
    ]
  }'

How to run AI OCR you can verify

  1. Send a document
    Post an image to /ocr/fields with imageType set to url or base64. In the app you can drop a PDF and each page is rasterized first; the public API takes raster images.
  2. Declare the schema
    Pass fields with name, type and children, or set autoFields and let the model propose the structure. Add required, pattern, min, max, enum or near wherever a rule matters.
  3. Read the checked result
    data.values holds the business data, data.cells[path] the box, quad, verified verdict, review reasons and evidence, data.normalized the parsed values for declared scalar types, and data.image the frame that converts coordinates to pixels.
  4. Work the review queue
    Iterate data.review.flagged, take reasons[0] as the primary reason, and draw that cell's box or quad over the page so a reviewer sees the region the value came from.
  5. Store and query
    Keep results in a sheet with POST /create and POST /upload, then read them back with GET /view using where, sort, select and limit. Those reads are not charged and do not re-run OCR.

Simple, predictable pricing

One credit is one page processed, at $0.05 including tax, with 100 credits free every month and no credit card. Failed scans are never charged. Reading stored data back with GET /space, GET /view or GET /jobs is free. Flat plans add monthly credits, more sheets and storage.

Free
$0
  • 100 credits / month
  • 3 sheets
  • 1 GB storage
Free — no card
Starter
$19/mo
  • 500 credits / month
  • 15 sheets
  • 10 GB storage
Start free
Most popular
Pro
$39/mo
  • 1,100 credits / month
  • Unlimited sheets
  • 100 GB storage
Start free
What makes this AI OCR different from a model that just returns JSON?
The model structures the document, but it does not have the last word. Business data comes back in data.values, while data.cells[path] carries the box and quad where that value was located, a verified verdict, the review reasons and the supporting evidence. data.review.flagged lists the paths worth a human look, so you review the model's output instead of accepting it.
Does the AI return the coordinates?
No. The model returns values only. Coordinates come from the OCR pass over the page and from matching each value, character by character, against the symbols that pass detected. evidence.source records which mechanism produced the box, and on /ocr/fields evidence.printed_text returns the raw glyphs at those coordinates.
How do I know whether to trust a given value?
Read data.review.flagged rather than a fixed score. Each item has a path and a reasons array ranked with the primary reason first, and the number of items to review is flagged.length. Open data.cells[path] for that path to see the verdict, the coordinates and the evidence, including evidence.match_ratio, which describes character coverage as supporting evidence rather than an acceptance rule.
Can the AI propose the fields for me?
Yes. Set autoFields and the model suggests a schema for the document, or declare fields yourself — including an array field with children for line items. Declarations such as required, pattern, min, max, enum and near never reach the model; they are checked after extraction, so they add review reasons and coordinate anchors rather than changing the extracted value.
What happens to the original value when I fix the AI's output?
In the app your edit is stored beside the original OCR value, not on top of it, so the model's reading and the human correction both stay on record. Because extraction runs through a model and is not deterministic run to run, the response JSON is the record worth keeping for audit; the coordinate cross-check and the normalized parse are the deterministic parts.
How much does it cost?
One credit is one page processed, at $0.05 including tax, with 100 credits free every month and no credit card. Failed scans are never charged, and reading stored data back with GET /space, GET /view or GET /jobs is free. Starter and Pro add monthly credits, more sheets and storage — see the plans above.

Use AI on your documents without trusting it blindly

Free tier — 100 credits a month, no credit card. Every value comes back with its coordinates, a review verdict and the evidence behind it.

Related