Analyze exported documents
Examines exported documents through Markdown, assets and source references.
Goal
Answer document questions from the exported source artefact, not from a hidden
retrieval index. The runtime context supplies the absolute Dokument-Markdown
path and, where available, Asset-Verzeichnis path for each attached document.
Workflow
- Start with the supplied absolute Markdown path; never derive a path from a file ID or filename.
- Read the heading structure with
rg '^#{1,6} ' <document.md>. - Search literal terms with
rg -n -i '<terms>' <document.md>. - Read the surrounding range with
sed -n '<start>,<end>p' <document.md>. - For broad or full-coverage work, progress through
document.mdin bounded line windows and keep the covered ranges in your reasoning. - Follow
assets/...references for tables or pictures whose visual content can change the answer. Resolve them relative to the export directory.
Artefact layout
Refined Docling artefacts are structured as:
<source-file-stem>__<file-key>/docling-<source-hash>/
├── document.md
└── assets/
An AnyDoc fast artefact is instead typically
.../fast-<source-hash>/document.md. It can be read immediately but may have
no assets and no page/BBox provenance. A refined artefact preserves Docling's
reading order and can include visual assets. Never infer page or bounding-box
provenance from Markdown markup. UI renderers obtain that information from the
document-element graph data.
Rules
- Do not look for former document retrieval tools, Chroma collections,
embedding IDs, or
FileChunknodes. They no longer exist; the exported Markdown file is the document index. - Prefer the supplied absolute paths; do not guess storage paths or use file IDs as retrieval keys.
- Quote or cite Markdown line ranges and asset names when this helps a user verify the source.
- Do not claim visual facts from an asset filename or a placeholder alone.
- If a document artefact is missing, say so instead of inventing content from metadata.
