pdf_ocr_ondevice library
On-device, downloadable OCR for dart_pdf_editor.
Implements PdfOcrEngine so PdfEditor.applyOcr can add a selectable,
searchable, invisible text layer over scanned PDF pages - running entirely
on the device, with no per-page network call. The (small, ~21 MB) PP-OCR
model is downloaded once via PdfOcrModelManager and then runs locally on
ONNX Runtime.
Supported on the native platforms (Android, iOS, macOS, Windows, Linux);
on the web use the HTTP-backed pdf_ocr_vlm engine instead.
final manager = PdfOcrModelManager();
final model = PdfOcrModels.ppOcrV5Mobile;
if (!await manager.isDownloaded(model)) {
await manager.download(model, onProgress: (p) => print(p.fraction));
}
final engine = await OnDeviceOcrEngine.fromDownloadedModel(manager, model);
final editor = PdfEditor(PdfDocument.open(bytes));
await editor.applyOcr(0, engine);
await engine.dispose();
Classes
- CtcDecoder
- Greedy CTC decoder over a fixed character dictionary.
- CtcResult
- The text decoded from a recognizer's per-timestep scores.
- DetectedBox
- A text region found by detection, as an axis-aligned box in the original image's pixel space (top-left origin, y down) with the mean detector probability over its pixels.
- DetectionResize
- The target size and the per-axis scale factors of a detection resize.
- IsolateOcrModelRunner
- Runs an unloaded, isolate-sendable OcrModelRunner on a long-lived worker isolate.
- OcrImage
- A decoded RGBA raster in plain Dart memory - the form the OCR pipeline works in, so all of its geometry/cropping/resizing is pure Dart and unit testable off a GPU.
- OcrModelRunner
- The inference backend behind OnDeviceOcrEngine: turns a rasterized page into recognized text lines.
- OnDeviceOcrEngine
-
A
PdfOcrEnginethat recognizes pages on device, with no network call at recognition time - the model is downloaded once (see PdfOcrModelManager) and then runs locally. - OnnxOcrModelRunner
- An OcrModelRunner that runs a PP-OCR-style detect-then-recognize pipeline on ONNX Runtime.
- PdfOcrDownloadCancelToken
- Cooperative cancellation handle for PdfOcrModelManager.download.
- PdfOcrDownloadProgress
- Progress of a PdfOcrModelManager.download, emitted on the download's stream as bytes arrive.
- PdfOcrModel
- A complete on-device OCR model: a text-detection network, a text-recognition network, the recognizer's character dictionary, and an optional orientation classifier - the four pieces of a classic detect-then-recognize OCR pipeline (PP-OCR family).
- PdfOcrModelFile
- One downloadable file that makes up an on-device OCR model - typically an ONNX network or a character dictionary.
- PdfOcrModelManager
- Downloads, caches, verifies, and removes on-device OCR models.
- PdfOcrModels
- Built-in model descriptors.
- RecognizedTextLine
- One recognized text line: the characters and where they sit in the page raster's pixel space (top-left origin, y down - the same space the OCR model worked in). OnDeviceOcrEngine maps pixelBounds back to PDF user space.
Functions
-
detectionResize(
int srcWidth, int srcHeight, {int sideLimit = 960, int multiple = 32}) → DetectionResize -
Computes the detection input size for an
srcWidthxsrcHeightpage: longest side scaled down to at mostsideLimit, each side rounded to a multiple ofmultiple(and never below it). Pages already within the limit are only rounded, never upscaled. -
extractDetectionBoxes(
Float32List probMap, int width, int height, {double threshold = 0.3, double boxScoreThreshold = 0.5, double unclipRatio = 1.6, int minSize = 3, double scaleX = 1.0, double scaleY = 1.0}) → List< DetectedBox> - Extracts text boxes from a DB-style detection probability map.
-
parseDictionary(
String contents, {bool addSpace = true}) → List< String> -
Parses a PP-OCR character dictionary file: one token per line, blank lines
preserved as a single space (some dicts encode the space as an empty
line). A trailing space token is appended when the file does not already
end with one, matching PP-OCR's
use_space_chardefault. -
recognitionInput(
OcrImage crop, {int targetHeight = 48, int maxWidth = 320}) → ({Float32List tensor, int width}) -
Recognition preprocessing for one cropped text line: resize to the
network's fixed
targetHeightkeeping aspect (width clamped tomaxWidth), then normalize to NCHW float32 in[-1, 1](PP-OCR rec convention:(pixel/255 - 0.5) / 0.5). -
toNchwFloat32(
OcrImage image, {List< double> mean = const [0.485, 0.456, 0.406], List<double> std = const [0.229, 0.224, 0.225]}) → Float32List -
Normalizes
imageinto an NCHW float32 tensor ([1, 3, height, width], channel order R, G, B) using per-channelmeanandstdon the 0..1 range - the standard ImageNet-style detection input.
Exceptions / Errors
- PdfOcrModelDownloadCanceled
- Thrown when a model download is canceled by PdfOcrDownloadCancelToken.cancel.
- PdfOcrModelException
- Thrown when a model file cannot be downloaded, fails its integrity check, or is requested on an unsupported platform.