space ocr
GuidesArticlesPricingDocs
PDF OCR

PDF OCR that turns documents into data you can check

Extract structured data from PDF pages with space-ocr: declare the fields you need, get every value back with its source coordinates in data.cells, and read data.review.flagged for what to check.

PDFs are where data goes to hide. An invoice, a stack of receipts, a delivery note — the numbers are right there on the page, but getting them into a spreadsheet usually means retyping. PDF OCR promises to fix that: read the document, get structured fields back. The catch is that most tools stop at a plausible guess and leave you to trust it.

space-ocr answers a stricter question. You declare the fields you want, and each value comes back with the region of the page it was read from — a box and a quad under data.cells — alongside data.review.flagged, the list of paths that did not check out. Instead of one score to interpret, you get a work list.

See a real extraction you can check

Hover any field below — the box on the receipt is where that value was read. The values, boxes, and match ratios shown here come from a real parsed result, not a mockup.

Receipts with extracted-field bounding boxes
Verified fields
KINSHO · 合計 2,045
ライフ · 合計 4,286

Each value with a box carries a verified on-page location — in data.cells[path], that is box + 4-point quad + evidence.match_ratio — on a 0–1000 normalized grid (0,0 top-left → 1000,1000 bottom-right), the same shape the live API returns. Hover a field to trace it back to the pixels it came from.

Three shapes, one contract
Take the page as named fields (POST /ocr/fields), as layout-preserving Markdown (POST /ocr/markdown), or as plain text in reading order (POST /ocr/text). All three answer with the same envelope — values, cells, review, image — so one review path handles all of them. Markdown returns per-element cells by default; for text, ask for them with includeBlocks: true.
Every value located
Each path in data.cells carries a box (xmin/ymin/xmax/ymax on a 0–1000 normalized grid) and a quad of four points that follows the page's tilt. data.image gives the width and height those coordinates are measured against, so converting to pixels is one multiplication.
Line items, not just totals
Declare a field of type array with children and every row lands at an indexed path — items[0].amount — with a cell of its own. The row itself gets a union box, so a wrapped or merged line is still traceable.
Your fields, not a fixed schema
Send the fields you actually store: name and type, plus required, pattern, min/max, enum, label, or near where they apply. Sampling an unfamiliar layout? Send autoFields: true and a schema is proposed for you. Declarations never reach the model — they decide what gets flagged, not how the page is read.
Clean exports
CSV with a UTF-8 BOM (Excel- and CJK-safe, line items unfolded) and JSON over a REST API, with async jobs at GET /jobs/{jobId} and HMAC-signed webhooks for ocr.completed.
Languages on autopilot
Japanese, Korean, Chinese, and English in one engine — there is no language setting to pick, and mixed scripts are handled.
Phone photos and skewed scans
EXIF orientation is baked into the pixels before reading, and nothing is deskewed — the quad follows the document's tilt, so the outline you draw sits on the text as printed.

How PDF OCR works in space-ocr

Drop a PDF into the web app and each page is rendered to a PNG in your browser, then read and turned into structured fields — a multi-page PDF becomes a set of rows you can sort, filter, and export.

The public API is stricter: it reads raster images, never PDF bytes. POST /upload rejects a PDF outright and tells you to render the pages first, so rasterize them yourself and send the page images. One page is one credit either way.

What a /ocr/fields call returns:

  • data.values — the business data, in exactly the schema you declared.
  • data.cells[path] — box, quad, verified, review, and evidence for that path.
  • data.review — declared, returned, boxed, verified, and flagged, the work list.
  • data.normalized — parsed values for declared number, integer, and date fields.
  • data.image — the width and height every coordinate is measured against.

verified is a verdict rather than a character-match result: false whenever the cell carries review reasons, true when a check ran and nothing was flagged, null when there was nothing to check. The character comparison itself sits at evidence.text_match, with evidence.match_ratio beside it as supporting detail.

rasterize the pages, then extract fields
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
# the API reads images, not PDF bytes — render the pages first
pdftoppm -r 200 -jpeg invoice.pdf page

curl -s https://api.space-ocr.com/ocr/fields \
  -H "Authorization: Bearer $SPACE_OCR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "image": "https://example.com/page-1.jpg",
    "imageType": "url",
    "fields": [
      { "name": "invoice_no", "type": "string", "required": true },
      { "name": "issue_date", "type": "date", "required": true },
      { "name": "total", "type": "number", "required": true, "min": 0 },
      { "name": "items", "type": "array", "children": [
        { "name": "description", "type": "string" },
        { "name": "quantity", "type": "number" },
        { "name": "amount", "type": "number" }
      ] }
    ]
  }'

How to OCR a PDF

  1. Rasterize the pages
    In the web app, drop the PDF and each page is rendered to a PNG in the browser. Calling the API directly, render the pages yourself and send them as url or base64 to POST /ocr/fields, or upload the images to a sheet with POST /upload.
  2. Declare your fields
    List the fields you want: name and type, plus required, pattern, min/max, enum, label, or near where they apply. Use an array field with children for line-item tables, or send autoFields: true to have a schema proposed.
  3. Read the structured result
    data.values holds the business data in your schema, data.cells[path] holds box, quad, verified, review, and evidence for each path, and data.normalized holds parsed values for declared number, integer, and date fields.
  4. Work through review.flagged
    Iterate data.review.flagged. Each entry is a path and a rank-ordered reasons array, with index 0 as the primary reason. Open data.cells[path], draw its box or quad against data.image, and correct the value beside the region it was read from.
  5. Export or query
    Download CSV (UTF-8 BOM, line items unfolded), or query a stored sheet with GET /view using where, sort, select, and boxes — that read is free and does not re-run OCR.

Simple, predictable pricing

One credit covers one page, at $0.05 including tax. Every account gets 100 credits a month without a card, and failed scans are never charged. Flat plans add monthly credits, more sheets, and storage.

Free
$0
  • 100 credits / month
  • 3 sheets
  • 1 GB storage
Free — no card
Starter
$19/mo
  • 500 credits / month
  • 15 sheets
  • 10 GB storage
Start free
Most popular
Pro
$39/mo
  • 1,100 credits / month
  • Unlimited sheets
  • 100 GB storage
Start free
Can I OCR a PDF with space-ocr?
Yes, with one step in between. The web app takes PDFs directly: it renders each page to a PNG in the browser and runs OCR on those, so a multi-page PDF becomes structured rows. The public API reads raster images only — POST /upload rejects a PDF and asks you to render the pages first, then send those images.
Does PDF OCR keep the location of each value?
Yes. Each path in data.cells returns a box (xmin/ymin/xmax/ymax on a 0–1000 normalized grid) and a quad of four points that follows the page's tilt, both measured against data.image. The cell's verified field states whether it was flagged, review.reasons says why, and supporting signals such as evidence.match_ratio and evidence.printed_text sit under evidence.
Can it extract tables and line items from a PDF?
Yes. Request line items as a field of type 'array' whose children describe one row (description, quantity, amount, and so on). Each cell keeps its own indexed path such as items[2].amount, and the row itself gets a union box, so a wrapped or merged line item is still traceable to its position.
How do I know which values to check?
Read data.review.flagged. Each entry is a path plus a rank-ordered reasons array — missing, text_mismatch, type_mismatch, out_of_range, pattern_mismatch, nobox and the rest — and the count of items to review is simply its length. Use the path to open data.cells[path] and show the region the value came from. There is no fixed score threshold to implement.
What can I export PDF OCR results to?
CSV with a UTF-8 BOM so Excel opens Japanese, Korean, and Chinese text correctly (line items unfold into sub-rows), and JSON over the REST API. A stored sheet can also be queried server-side with GET /view using where, sort, select, and boxes — that read is free and does not re-run OCR.
How much does PDF OCR cost?
$0.05 including tax per credit, and one credit covers one page. Every account gets 100 credits a month with no card, and failed scans are never charged. Starter and Pro add monthly credits, more sheets, and storage — see the plans above.
Which languages does it handle?
Detection is automatic — Japanese, Korean, Chinese, and English in one engine, including mixed scripts and full-width or half-width characters. There is no language setting to pick.

Turn your own PDFs into checkable data

Free tier — 100 credits a month, no credit card. Every value comes back with its on-page location.

Related