space ocr
GuidesArticlesPricingDocs

How space-ocr is different from LLM OCR: verifiable, structured extraction

How space-ocr differs from prompting a raw LLM for OCR: fields come back with source coordinates on the page, a verified verdict, and a review list of what to check.

9 min read· 2026-08-31

You can hand a receipt or an invoice to GPT-4o, Gemini, or Claude and ask for the total, the vendor, and the line items. Most of the time you get sensible-looking JSON back. The trouble starts when you try to trust it at volume: the model returns a string, and a string has no address. If the total comes back as 48,200, which pixels on the page did it read? Was that number actually printed on the document, or did the model fill in something plausible? With a raw LLM call you cannot answer that without re-reading the page yourself.

That gap is the whole difference between using a general-purpose LLM as an OCR tool and using space-ocr. space-ocr is not anti-LLM. Under the hood it currently pairs an OCR engine (Google Cloud Vision) with Gemini for structuring — an implementation detail rather than part of the API contract. What it adds is the layer built around the model: each value it returns is checked against what the OCR pass actually saw on the page, given a verdict, and, once you upload into a sheet, stored as a row you can query. Those are the two things a raw LLM call leaves to you — per-value provenance you can verify, and structured output you can query without standing up a database.

Raw LLM OCR vs space-ocr

Raw LLM OCR (GPT-4o / Gemini / Claude)space-ocr
Per-value locationa JSON-extraction call gives you text; source coordinates cross-checked against a separate OCR pass are not part of the contracta box (xmin, ymin, xmax, ymax on a 0–1000 grid) and a four-point oriented quad for each value that can be anchored, and the ones that cannot come back in review.flagged as nobox
Per-value verificationnone, you trust the stringcells[path].verified is the verdict and review.reasons says why; the cross-check itself rides in evidence — text_match, source, and match_ratio, the share of the value's characters found among the page's detected symbols
Declared rulesyou write your own validatorsrequired, pattern, enum, min/max, near and not_near travel with the request and are checked server-side; a violation lands in review.flagged
Check a value in contextre-read the document yourselfclick a cell in the app and its exact region lights up on the original image
Output shapeprompt-dependent JSON that varies run to runa fixed schema: declare fields once, or let autoFields propose them, and data.values comes back in that shape
Storage and queryyou build itPOST /ocr/fields answers in the response and stores no image; upload into a sheet instead and each page lands as a row you query with GET /view (where, sort, select, limit, offset), no re-OCR and no charge
Scriptsdepends on the model and promptJapanese, Korean, Chinese, English and more, auto-detected, with no language parameter
Setupyour own prompt, retry, parse, and validate pipelineone HTTPS call with a Bearer key
✓ Verified

Here is what "verified" actually means, because it is easy to overclaim. The language model does not emit coordinates. It returns each value's text plus word-token hints, and the engine then matches that text character by character against the symbols Google Cloud Vision detected on the page. It lands the box on those real symbols, and the outcome of that comparison rides in cells[path].evidence: text_match for the character cross-check itself, match_ratio for the share of the value's characters that were located, source for how the box was found. verified sits on top as the verdict — false when the cell carries any reason at all, true when a check ran and nothing was raised, null when there was nothing to check. A weak match raises a reason such as low_ratio, and that path then appears in review.flagged, which is the list you route to a person. values is the model's reading rather than a byte-for-byte copy, so when you need an exact string to match against master data, evidence.printed_text holds the glyphs the OCR pass read at those coordinates. This is not a promise that the model can never be wrong: the token hints can still drift, and both engines can agree on the same misreading. It means each value is checked against the page and reported on, instead of taken on faith.

POST /ocr/fields — one field in the response
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
{
  "status": "success",
  "data": {
    "values": {
      "vendor": "ACME Trading Co."
    },
    "cells": {
      "vendor": {
        "box": { "xmin": 120, "ymin": 84, "xmax": 512, "ymax": 118 },
        "quad": [
          { "x": 120, "y": 84 }, { "x": 512, "y": 84 },
          { "x": 512, "y": 118 }, { "x": 120, "y": 118 }
        ],
        "verified": true,
        "review": null,
        "evidence": {
          "text_match": true,
          "source": "vision_symbol_match",
          "match_ratio": 1.0,
          "ocr_confidence": 0.97
        }
      }
    },
    "review": {
      "unit": "field",
      "declared": 1,
      "returned": 1,
      "boxed": 1,
      "verified": 1,
      "flagged": []
    },
    "image": { "width": 1654, "height": 2339 }
  }
}

The box is returned on a 0–1000 grid that is independent of the image size, so you scale it to pixels when you draw it: pixel_x = xmin / 1000 * image_width. Take the width and height from data.image in the same response — that is the page as it was actually read, with EXIF rotation baked in and a large photo already downscaled, not the file you sent. The four-point quad follows a tilted or rotated scan, ordered top-left, top-right, bottom-right, bottom-left. When something does not check out, the same path turns up in data.review.flagged as { "path": "total", "reasons": ["text_mismatch"] }, so what you hand a reviewer is a list you iterate, not a score you have to pick a threshold for. A JSON-extraction call to a general model does not come with any of this, so any audit trail you want, you assemble by hand.

From values to a queryable table

A raw LLM call ends at the JSON. You still have to persist it, and the moment you want to ask "which invoices this quarter are over 40,000?" you are building a database and a query layer first. POST /ocr/fields ends at the JSON too — it answers in the response and keeps no copy of the image. What differs is that the same extraction runs through a storage layer when you want one: create a sheet with your columns (POST /create), upload pages into it (POST /upload), and each page lands as a row carrying the same values, cells and review. Querying is then an API call: GET /view with where=total>=40000, sort=-invoice_date, select=vendor,total, plus limit and offset for paging. It runs server-side, does not re-run OCR, and is not charged. Export the sheet to CSV (UTF-8 with a BOM, so Japanese, Korean, and Chinese text and currency open correctly in Excel, and line-item arrays expand into their own rows).

When a raw LLM is the better tool

A general-purpose LLM is the right choice when you want a one-off read, a loose summary, or reasoning about what a document means. "What is this contract about?" is a model question, not an OCR-with-coordinates question. Reach for space-ocr when you are processing documents at volume and need each value to be verifiable, consistently structured, and queryable: accounts-payable automation, expense reconciliation, importing business cards into a CRM, or digitizing a backlog of receipts. The honest framing is that space-ocr is an LLM-backed OCR with a verification and storage layer, not a rival to the models themselves.

Billing is pay-as-you-go per scan with optional monthly plans and a monthly allowance of free scans, and failed scans are never charged. The pricing page has the current figures.

Does space-ocr replace GPT-4o or Gemini for OCR?
No. space-ocr uses an LLM under the hood — the current stack pairs Google Cloud Vision for OCR with Gemini for structuring, which is an implementation detail rather than part of the API contract. The difference is the layer around the model: each value is matched back to the page's symbols and given a verdict, and once you upload into a sheet the result is stored as a queryable row. It is an LLM-backed OCR, not a replacement for the models.
How do I know a value was not hallucinated?
Read three things. cells[path].verified is the verdict: false when the cell carries any reason, true when a check ran and nothing was raised, null when there was nothing to check. data.review.flagged is the list of paths worth a look, each with its reasons. And cells[path].evidence carries the cross-check itself — text_match, match_ratio, source — on the values where a check could run; those keys are absent when it could not, which means undetermined rather than fine. It is verification against the real page, not a guarantee that the model can never err: both engines can agree on the same misreading.
What coordinate format does space-ocr return?
A box as four integers — xmin, ymin, xmax, ymax — on a normalized 0–1000 grid, plus a four-point oriented quad for tilted or rotated pages, ordered top-left, top-right, bottom-right, bottom-left. Convert to pixels against data.image from the same response, for example pixel_x = xmin / 1000 * image_width.
Can I query the extracted data without my own database?
Yes, once the pages are in a sheet. POST /ocr/fields returns the JSON and stores nothing; create a sheet (POST /create) and upload pages into it (POST /upload) and each page becomes a row. GET /view then filters them server-side with where, sort, select, limit, and offset (for example where=total>=40000). It does not re-run OCR and is not charged, and you can export the sheet to CSV.
Do I have to tell space-ocr the document's language?
No. Language is auto-detected across Japanese, Korean, Chinese, English and more. There is no language parameter to set.
Can it read a PDF?
The web app rasterizes each PDF page to an image and then OCRs it. The API itself takes raster images (JPEG, PNG, GIF, BMP, TIFF, WebP), one image per call, so convert PDF pages to images before sending them.
How do I call it?
Send one HTTPS request to POST /ocr/fields with a Bearer spocr_ key, passing the image plus a fields schema — or autoFields: true to have one proposed. The response keeps the two apart: data.values holds your schema and nothing else, data.cells[path] holds the box, quad, verified, review and evidence for each value, and data.review summarizes what to check.