Skip to content

OCR Models

EdgeParse extracts born-digital PDF text without any ML stack. Optional OCR recovers text from image-embedded / raster tables (and related image regions) when you turn it on.

SurfaceFlag
Web SDKParseOptions.enableOcr (default true); create with ocr: 'off' to disable
WASM / configrasterTableOcr
Pythonraster_table_ocr
NoderasterTableOcr
CLI--raster-table-ocr

When OCR is off or models are declined, born-digital text still extracts; image tables degrade to PDF text with quality: "degraded".

Live demo: edgeparse.com/demo/ — toolbar OCR toggle → enableOcr without recreating the client.

Terminal window
npm install @edgeparse/web edgeparse-wasm
import { EdgeParse } from '@edgeparse/web';
const ep = await EdgeParse.create({
models: 'lazy', // or 'preload' | 'manual' | 'off'
ocr: 'small', // 'tiny' | 'small' | 'medium' | 'off'
onBeforeDownload: async (model) =>
confirm(`Download ${model.id} (~${(model.bytes / 1e6).toFixed(1)} MB)?`),
});
const job = ep.parse(file, { format: 'markdown', enableOcr: true });
const { markdown, quality } = await job.result;

Architecture:

  1. Parse worker — ParseSession (Rust/WASM): plan candidates, finish with OCR words
  2. OCR worker — onnxruntime-web (WebGPU → WASM+SIMD), PP-OCRv6
  3. ModelManager — pinned models.json, OPFS then Cache API (not IndexedDB), sha256 verify, Web Locks

Default tier is small. Approximate download sizes (det + rec + dict from the shipped manifest):

TierApprox. sizeminCapability
tiny~6.4 MBsimd
small~31 MBsimd
medium~139 MBwebgpu

Weights are Apache-2.0 PP-OCRv6 artifacts (Hugging Face snowfluke/ppu-paddle-ocr-models). Attribute the license in redistributions.

  • npm ships the pinned manifest only (models/models.json + sha256). Weight blobs are fetched at runtime.
  • Primary host: Hugging Face; dictionary failover includes jsDelivr.
  • After a successful download, artifacts stay in OPFS (preferred) or the Cache API.
  • models: 'preload' (or ep.models.preload()) once online → that tier works offline later.
  • If the browser is offline and the model is not cached → error code OFFLINE (parse degrades).
  • With navigator.connection.saveData, downloads ask consent again or skip (ABORTED).

onBeforeDownload runs before each uncached artifact. Return false to skip OCR for that model; the parse continues with quality: "degraded".

Pass a custom manifest whose urls point at your CDN (keep the same sha256 / bytes):

const ep = await EdgeParse.create({
ocr: 'small',
manifest: myManifest, // same schema as models/models.json
onBeforeDownload: () => true,
});

CSP connect-src must allow those hosts. See CSP.md.

Release CI runs npm run check-models and fails closed on placeholder sha256s. Published @edgeparse/web includes the pinned manifest, not the multi‑MB weight files.

PathBehavior
@edgeparse/web/nodeNullOcrBackend by default; optional NodePpocrBackend (onnxruntime-node)
Native ocrsOptional .rten models via EDGEPARSE_OCRS_MODEL_DIR — not the web PP-OCR ORT packages