A changing corpus, limited compute.
Much of the work at Immigrant Invest depends on documents. The corpus contains more than 100,000 confidential files and keeps changing as documents are added, edited, and replaced. Formats, filenames, and folder structures are inconsistent, so downstream systems cannot rely on the files arriving ready to use.
I designed and built a shared document intelligence platform to classify documents, extract their contents, and maintain a structured representation for other systems. Confidentiality requirements made local OCR and LLM processing central to the design. Limited local compute made routing and throughput equally important.
Route the work to the right processing configuration.
The ingestion process captures source metadata, document IDs, and versions. Agents route documents to processing configurations suited to their format. Local OCR extracts text, and a locally hosted Qwen3.6-35B-A3B model supports document classification and information extraction.
I used batching and format-aware routing to make better use of the available hardware. Preparing documents in advance also separates the cost of extraction from the latency of the downstream agents that need the results.
100K+ confidential documents
New files, edits, and revisions
-
Track source versions
Metadata and document IDs
-
Route, classify, extract
Format-specific configurations · local OCR · Qwen3.6
Batched local processing
Assess extraction confidence
Confidence ≥ 0.90
Continue to normalization.
Lower confidence: retry
Retry and reassess, with selective Azure LLM fallback.
Still below 0.90
Escalate to legal reviewers.
Normalized, versioned records
Accepted fields indexed by document ID and version
- Operational tables
- Azure embeddings / RAG
- Downstream agents
Source changes update the shared document representation.
Confidence determines the next step.
Extraction confidence guides repeat processing and escalation. Lower-confidence results go through retries before reaching a person. Azure-hosted LLMs provide a selective fallback, used for less than 0.5% of the reported processing volume.
Values that remain below the 0.90 confidence threshold are escalated to legal reviewers. This gives the workflow an explicit route for uncertain extractions and focuses human attention on unresolved cases.
0.90 is a workflow threshold, not a measured accuracy rate. The Azure fallback share refers to LLM processing; it does not include Azure embeddings or retrieval services.
Structured records that stay connected to their sources.
Extracted information is normalized into structured records and tables, indexed by document ID and version. Synchronization updates the derived data as the original files change, keeping the shared representation connected to its source material.
These outputs support downstream automation and RAG, with Azure used for embeddings and retrieval preparation. Other systems consume the prepared representation instead of independently reconstructing the same information from files.
A reusable foundation for other agents.
The platform made it easier to connect additional systems to the document corpus. Classification, extraction, normalization, and version tracking were already handled in one place, giving each new integration a consistent starting point.
Downstream agents can work with prepared data without waiting for OCR and extraction during each request. Reusing that work also makes better use of constrained compute, while source synchronization gives connected systems a common view of the documents they depend on.
Its value extends beyond processing individual files: it provides a maintained document data foundation that many agents and business workflows can share.