Authoritative documentation · English

Importer

The Importer brings external file sources into a Book as Documents. Harnesser preserves source references so generated knowledge can point back to original evidence, and it processes imports incrementally so repeat or unchanged content is not re-billed.

Supported sources

  • Text-bearing files — exact text files, office files, and text-layer PDFs.
  • Scanned and mixed PDFs — processed by local OCR on the device that holds the Book; OCR is never a cloud service.
  • Images — PNG, JPEG, WebP, and TIFF, with text recognized locally and image-region evidence locations preserved.

Imported Documents become file sources alongside your committed Notes. Both feed the same extraction and validation workflow.

What the Importer preserves

  • Original page and image objects, kept on the device.
  • OCR text and Arabic search forms, with confidence and warning metadata.
  • Page or region evidence locators and content digests, so answers cite exact provenance.
  • Checkpoints for resumable processing on large or slow imports.

Incremental and exact imports

Changed sources process only changed blocks; unchanged blocks are not re-extracted. Exact duplicate imports create no repeated billable generation. Large imports resume from the last checkpoint after interruption rather than restarting from zero.

What the Importer does not do

  • It does not handwrite-recognize, cloud-OCR, or guess unsupported input. Unsupported input is reported as unsupported.
  • It does not publish or validate content. An imported Document becomes a source; extraction, validation, and publication remain explicit stages.
  • It does not send original content wholesale to a provider. External generation receives only a bounded, explicitly selected slice.