space ocr
GuidesArticlesPricingDocs
Image OCR

Image OCR that returns structured fields, not a wall of text

Run OCR on JPEG, PNG, and other images with space-ocr: declare the fields you need and every value comes back with its box and quad coordinates, plus a data.review.flagged list of what to check.

Most image OCR hands you a wall of plain text and stops there. You snap a receipt, run it, and get back a blob of lines you still have to read, split, and retype into the right columns. The structure that was obvious to your eye on the page is gone.

space-ocr reads an image into structured fields instead — store name here, date there, total over there, line items as rows. Every value also carries the exact spot on the image it was read from: an axis-aligned box and a four-point quad under data.cells[path]. And when a value does not hold up against the page, it shows up in data.review.flagged with its reasons — so what you get back is a work list, not a number you have to take on faith.

See a real extraction you can check

This is one image — a photo of two receipts — read into fields. Hover any value below and the box on the image is exactly where it was read. The values, boxes, and character-match figures shown here come straight from a real parsed result, not a mockup. The match figure is supporting evidence about how much of the value was located on the page, not a pass/fail score.

Receipts with extracted-field bounding boxes
Verified fields
KINSHO · 合計 2,045
ライフ · 合計 4,286

Each value with a box carries a verified on-page location — in data.cells[path], that is box + 4-point quad + evidence.match_ratio — on a 0–1000 normalized grid (0,0 top-left → 1000,1000 bottom-right), the same shape the live API returns. Hover a field to trace it back to the pixels it came from.

Three shapes, one contract
Take the page as named fields, as layout-preserving Markdown (POST /ocr/markdown), or as plain text in reading order (POST /ocr/text). All three answer with the same envelope — values, cells, review, image — so coordinates and the verified verdict work the same way whichever shape you pick.
Structured fields, not a text dump
An image comes back under data.values as named values and rows — store, date, total, line items — in the exact shape you asked for, instead of one long string you have to split yourself.
Every value located
data.cells[path] holds a box (xmin/ymin/xmax/ymax) and a quad of four points, both on a 0–1000 normalized grid. data.image reports the width and height to convert them to pixels: x = box.xmin / 1000 × width.
Phone photos welcome
EXIF orientation is baked in before the page is read, so coordinates line up with the photo as displayed. There is no deskew step, so the quad follows the tilt of a handheld shot, and a very large photo may be downscaled first — which is why data.image, not the file you sent, is the frame for every coordinate.
Declare the fields you need
Send a fields array — name and type per field, plus optional required, pattern, min/max, enum, label, and near — or set autoFields to true and let the reader come back with field names taken from the document itself.
Line items, not just totals
An array field with children returns repeating rows, and every cell keeps its own path and position: items[0].amount addresses one cell, items[0] the row it belongs to.
Clean exports
CSV with a UTF-8 BOM (Excel- and CJK-safe, line items unfolded) from the app, and JSON over a REST API with async jobs (POST /upload → GET /jobs/{jobId}) and HMAC-signed webhooks.

How image OCR works in space-ocr

Send an image to /ocr/fields as a URL or as plain base64 — JPEG, PNG, GIF, BMP, TIFF, and WebP are all read directly. EXIF orientation is applied before reading, and a very large photo may be downscaled to 4000px on the longest side, so data.image describes the page the coordinates actually belong to.

You describe the result you want with a fields array: a name and a type (string, number, integer, date, array, object) per field, an array field with children for a line-item table, and optional declarations such as required, pattern, min/max, or enum. If you would rather start from the document, send autoFields: true and work from the field names that come back. Declarations are not shown to the model, so they do not steer the reading — they decide what lands in data.review.flagged and what the deterministic data.normalized layer parses. (PDFs go through the web app, which renders each page to an image first; the API itself reads images.)

extract fields from an image
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
curl -s https://api.space-ocr.com/ocr/fields \
  -H "Authorization: Bearer $SPACE_OCR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "image": "https://example.com/receipt-photo.jpg",
    "imageType": "url",
    "fields": [
      { "name": "store_name", "type": "string", "required": true },
      { "name": "date", "type": "date", "required": true },
      { "name": "total", "type": "number", "required": true, "min": 0 },
      {
        "name": "items",
        "type": "array",
        "children": [
          { "name": "name", "type": "string" },
          { "name": "price", "type": "number" }
        ]
      }
    ]
  }'

How to OCR an image

  1. Send your image
    Post a JPEG, PNG, GIF, BMP, TIFF, or WebP to /ocr/fields as a URL or plain base64, or drop it into the app. EXIF orientation is applied before the page is read.
  2. Declare your fields
    Send a fields array with a name and type per value — an array field with children for line-item tables — or set autoFields to true and start from the field names that come back.
  3. Read the structured result
    Business data stays in data.values. data.cells maps each path to its box and quad, a verified verdict and evidence, and data.image is the frame that converts those coordinates to pixels.
  4. Work the review list
    Iterate data.review.flagged: each entry names a path and its reasons. Open cells[path], draw the box on the image, and compare the value with what is printed there. Edits are stored beside the original OCR value.
  5. Export or query
    Download CSV (UTF-8 BOM, line items unfolded), or query a stored sheet with GET /view using where, sort, and select — reading stored rows does not re-run OCR and is not charged.

Simple, predictable pricing

$0.05 per image (tax included), with 100 credits free every month and no credit card. Failed scans are never charged. Flat plans add monthly credits, more sheets, and storage.

Free
$0
  • 100 credits / month
  • 3 sheets
  • 1 GB storage
Free — no card
Starter
$19/mo
  • 500 credits / month
  • 15 sheets
  • 10 GB storage
Start free
Most popular
Pro
$39/mo
  • 1,100 credits / month
  • Unlimited sheets
  • 100 GB storage
Start free
What image formats can space-ocr OCR?
The public API reads raster images directly — JPEG, PNG, GIF, BMP, TIFF, and WebP. Images are converted to RGB automatically. PDFs go through the web app, which renders each page to an image before OCR.
Does image OCR give me structured fields or just text?
Structured fields. An image is read into named values and rows under data.values — store, date, total, line items — in the shape you declared, rather than one long block of plain text you have to parse yourself.
Can I OCR a photo taken on my phone?
Yes. EXIF orientation is baked in before the page is read, so the returned coordinates match the photo as displayed. There is no deskew step, so the four-point quad follows the tilt of a handheld shot, and data.image reports the page dimensions those coordinates belong to.
Does image OCR keep the location of each value?
Yes. Every path in data.cells carries a box (xmin/ymin/xmax/ymax) and a quad of four points on a 0–1000 normalized grid, and data.image gives the width and height needed to convert them to pixels. cells[path].evidence holds the supporting comparison, including match_ratio and, on this endpoint, printed_text — the raw glyphs read at those coordinates.
How do I know which values need checking?
Read data.review.flagged. Each entry has a path and a rank-ordered reasons array — text_mismatch, missing, pattern_mismatch, out_of_range and the other documented codes — and the count of items to review is simply flagged.length. Use the path to open cells[path] and draw its box beside the value. Two independent readings can still agree on the same misread, so keep your own business rules downstream.
How do I send the image to the API?
Send it to POST /ocr/fields as a URL (imageType 'url') or as plain base64 (imageType 'base64', no data-URI prefix). Authenticate with a Bearer token; keys are prefixed spocr_. Pass a fields array describing what you need, or set autoFields to true.
How much does image OCR cost?
$0.05 per image (tax included), with 100 credits free every month and no credit card, and failed scans are never charged. Flat plans (Starter and Pro) add monthly credits, more sheets, and storage — see the plans above.

Turn your own images into checkable data

Free tier — 100 credits a month, no credit card. Every value comes back with its on-image location.

Related