Authoritative documentation · English
AI, OCR, and privacy
Harnesser is provider-optional. Retrieval and Evidence Intelligence run locally. Book discovery is still being connected to the shipped desktop interface and is not yet available; when released, it will also run locally. External AI is a bounded, explicitly selected accelerant for generation, answering, verification, and translation. OCR runs locally on the device that holds the Book and is never a cloud service.
OCR scope and privacy
Local OCR supports printed English and formal Arabic, including meaning-bearing Arabic marks when the source and engine preserve them. Scanned and mixed PDFs, image inputs (PNG, JPEG, WebP, TIFF), exact text, a separately derived Arabic search skeleton, reading order, confidence, warnings, and page or image-region evidence locations are all produced locally.
OCR processing keeps the following on the device: original page and image objects, OCR text and Arabic search forms, confidence and warning metadata, page or region evidence locators and content digests, and checkpoints for resumable processing. No image, text, or metadata leaves the device for OCR.
The product does not claim handwriting recognition, cloud OCR, or reliable recognition of arbitrary dialects. Unsupported input is reported as unsupported; it is not silently sent to an external service or replaced with guessed text.
Formal-Arabic fidelity
Arabic letters, harakat, tashkil, and Quranic marks are meaning-bearing. The OCR output preserves them as exact text in the evidence channel. A separate Arabic search skeleton may widen recall by stripping selected marks, but the skeleton alone never validates semantic equivalence, merges candidates, or answers a question. Exact vocalized evidence is always prioritized.
When two admissible items with different vocalizations produce different meanings and the surrounding question does not resolve the ambiguity, the answer is exactly I don't know.
Engine benchmark gates
Launch OCR is blocked unless one supported engine configuration reaches the published benchmark gates: at least 98% character and 95% word accuracy on clean printed fixtures, 92% character and 85% word accuracy on degraded fixtures, 95% harakat and Quranic-mark F1 on clean vocalized Arabic, and 95% correct reading order and evidence location on the versioned acceptance corpus.
Provider privacy posture
Every external model operation sends a bounded, explicitly selected slice of plaintext — never the Book. The rest of the Book, other Notes, other Sources, provider credentials, application logs, release bundles, and unrelated content are never sent.
A default provider configuration uses a documented no-training and zero-data-retention posture when the provider offers it, disables provider-side request logging, pins one endpoint and model identity, requires structured outputs, and disables silent fallback. A non-conforming advanced configuration requires an explicit warning and remains visibly identified.
Certified configurations
A signed compatibility manifest records model configurations certified separately for generation, answering, verification, and translation roles. A model must pass structured-output, grounding, exact-abstention, citation, English, formal-Arabic, latency, and usage-reporting acceptance gates before default selection.
Cost and usage reporting
Every completed AI operation reports normalized provider, model, token, and cost usage. A failed or uncertain request retains its conservative reservation until reconciled. Set token and cost budgets before a run; the operator always sees what was spent and what remains.