DocAgent is a production-grade, autonomous document processing agent designed to streamline accounts payable and enterprise invoice operations. It extracts structured data from multi-modal documents (PDF, Scanned Images, TXT), validates compliance against 9 deterministic business rules, performs statistical fraud & anomaly screening, detects SHA-256 duplicates, and manages human-in-the-loop approval workflows with full database persistence and an immutable audit trail.
View Interactive Mermaid Architecture Source
graph TB
classDef clientStyle fill:#EBF5FB,stroke:#2980B9,stroke-width:2px,color:#1B4F72;
classDef gatewayStyle fill:#E8F8F5,stroke:#16A085,stroke-width:2px,color:#0E6251;
classDef agentStyle fill:#FEF9E7,stroke:#F39C12,stroke-width:2px,color:#7D6608;
classDef toolStyle fill:#FDEDEC,stroke:#E74C3C,stroke-width:2px,color:#78281F;
classDef dataStyle fill:#F4ECF7,stroke:#8E44AD,stroke-width:2px,color:#512E5F;
subgraph L1["PRESENTATION LAYER"]
UI["Streamlit Web UI<br/>(Live Ingestion, Review Queue, Analytics)"]:::clientStyle
API_CLIENT["External REST Clients / ERP Systems"]:::clientStyle
end
subgraph L2["API GATEWAY & CONTROLLER"]
FASTAPI["FastAPI Application<br/>(File Parsing, Async Endpoints, Auth & CORS)"]:::gatewayStyle
end
subgraph L3["AGENTIC ORCHESTRATION ENGINE (LangGraph)"]
GRAPH["StateGraph Workflow Runner<br/>• Self-Correction Extraction Loop<br/>• Deterministic Business Rule Engine (9 Rules)<br/>• Composite Risk Scoring & Threshold Routing<br/>• Human-in-the-Loop Checkpoint Pausing"]:::agentStyle
end
subgraph L4["AI & INTELLIGENCE SERVICES"]
direction LR
LLM_PRIMARY["Gemini Multimodal<br/>(Vision & Structured JSON)"]:::toolStyle
LLM_FALLBACK["GitHub Models / GPT-4o<br/>(Automatic Fallback)"]:::toolStyle
ANOMALY_ENGINE["Statistical & Forensic Engine<br/>(Z-Score Outliers, Round Numbers, Structuring)"]:::toolStyle
end
subgraph L5["PERSISTENCE & CACHE LAYER"]
direction LR
DB[("PostgreSQL / SQLite<br/>(Invoices, Line Items, Audit Logs, Validations)")]:::dataStyle
CACHE[("Redis / In-Memory Cache<br/>(SHA-256 Deduplication Registry)")]:::dataStyle
end
UI -->|Upload Document / Human Approval| FASTAPI
API_CLIENT -->|POST /process & POST /resume| FASTAPI
FASTAPI -->|Initialize & Run State| GRAPH
GRAPH <-->|Extract Structured Data| LLM_PRIMARY
LLM_PRIMARY -.->|On Failure / Failover| LLM_FALLBACK
GRAPH <-->|Detect Fraud & Pattern Outliers| ANOMALY_ENGINE
GRAPH <-->|Deduplication Check| CACHE
GRAPH <-->|Query Spend History & Persist Records| DB
FASTAPI -->|Stream Analytics & Status| UI
Raw Invoice (PDF/Image/TXT)
│
├─► 1. Pre-LLM Smart File Hash Gate (Zero-model-cost SHA-256 intercept with appeal/reprocess support)
│
├─► 2. Text Prep & Concrete Tampering Inspection (White text #FFFFFF, microscopic fonts < 3pt, covered overlays)
│
├─► 3. Primary Reader A (Multimodal Gemini 3.8 Flash vision extraction with auto-retry)
│
├─► 4. Verifier Reader B (Independent text-only GitHub Models / GPT-4o-mini anchor field extraction)
│
├─► 5. Dual-Reader Verification & Grounding (Evidence snippet verification in raw text -> Full/Partial/None)
│
├─► 6. Legibility & OCR Quality Gate (Readability score calculation and tampering alert flagging)
│
├─► 7. Compliance Validation Node (9 Business Rules: Approved Vendors, Tax Math, Spends, Dates)
│
├─► 8. Semantic & Near-Duplicate Detection (Levenshtein distance checks excluding invoice's own case ID)
│
├─► 9. 3-Tier Vendor Anomaly Detection (Tier 0 Cold-Start, Tier 1 Small-Sample, Tier 2 Empirical Z-Scores)
│
├─► 10. Risk Scoring & Decision Routing Node (Calibrated composite risk score; enforces Auto-Approve Invariant)
│ ├── Risk < 0.20 & Verification == 'FULL' ──► Auto-Approve
│ ├── Risk ≥ 0.70 or Tampering/Duplicates ──► Auto-Reject
│ └── Risk ≥ 0.20 or Partial Verification ──► Flag for Human Review (Manager / Director)
│
├─► 11. Human Review Queue & Correction Loop (Max 2 re-entries limit, separation of duties, case reopening)
│
└─► 12. Tamper-Evident Hash-Chained Audit Trail (Cryptographic SHA-256 running hash chain per step)
| Rule Name | Severity | Description |
|---|---|---|
max_amount |
ERROR |
Invoice amount must not exceed departmental budget limit (₹500,000 / $500,000). |
approved_vendor |
ERROR |
Vendor must be in the approved master vendor registry (24+ verified enterprise vendors + real database vendors). |
date_validity |
ERROR |
Invoice date must be a valid ISO format and not post-dated in the future. |
visual_text_consistency |
ERROR |
Embedded visual scan amount must match digital line items (tampering alert). |
line_item_math |
WARNING |
Individual line items (qty × rate) must sum to subtotal within ±1% tolerance. |
total_math |
WARNING |
Declared Subtotal + Tax Amount must match Total Amount within ±1% tolerance. |
vendor_id_present |
WARNING |
Vendor must supply a valid tax registration identifier (GSTIN, VAT ID, EIN). |
currency_consistency |
WARNING |
Currency must be a recognized ISO 4217 standard currency code (INR, USD, EUR, GBP). |
payment_terms_check |
INFO |
Invoice must explicitly declare payment terms (e.g., Net 30, Due on Receipt). |
DocAgent has been benchmarked directly on 16 authentic enterprise PDF invoices (including Amazon Web Services, Flipkart B2B, Azure, QualityHosting, Coolblue, and Microsoft Document Intelligence samples) evaluated against a historical baseline of 500 real U.S. federal contract awards ($146.1B volume across 178 contractors via USASpending.gov):
# Run the real-world enterprise benchmark suite
python eval/run_real_benchmark.py| Real-World Benchmark Metric | Result | Target | Status |
|---|---|---|---|
| Document Extraction Rate | 100.0% (16/16 invoices) |
≥ 90.0% |
PASSED |
| Vendor Recognition Accuracy | 100.0% (16/16 vendors) |
≥ 90.0% |
PASSED |
| Total Amount Accuracy | 100.0% (16/16 totals) |
≥ 95.0% |
PASSED |
| Invoice Date Accuracy | 100.0% (16/16 dates) |
≥ 90.0% |
PASSED |
| Line Items Extraction Rate | 100.0% (16/16 items) |
≥ 85.0% |
PASSED |
| Average Extraction Confidence | 95.9% |
≥ 85.0% |
PASSED |
| Average Processing Latency | 11.23s / document |
≤ 15.0s |
PASSED |
Full real-world reports: eval/real_benchmark_report.md and eval/real_benchmark_report.json.
DocAgent also includes a rule-verification test harness evaluated against 20 curated test scenarios across 5 distinct compliance categories:
# Run offline compliance rule tests
python eval/run_eval.py --mode offline| Benchmark Metric | Result | Target | Status |
|---|---|---|---|
| Rule Compliance Accuracy | 100.0% (180/180 checks) |
≥ 95.0% |
PASSED |
| Anomaly Detection Precision | 1.00 |
≥ 0.90 |
PASSED |
| Anomaly Detection Recall | 1.00 |
≥ 0.90 |
PASSED |
| Anomaly Detection F1 Score | 1.00 |
≥ 0.90 |
PASSED |
| Decision Routing Accuracy | 100.0% (20/20 invoices) |
≥ 95.0% |
PASSED |
| Approval Level Routing | 100.0% (20/20 invoices) |
≥ 95.0% |
PASSED |
Full offline evaluation reports: eval/eval_report.md and eval/eval_report.json.
Upload PDF, Image, or TXT invoices to initiate multi-stage extraction, compliance verification, and routing.
Live visual feedback with extracted table breakdown, composite risk gauge, and compliance rule results.
Transparent observability into every step executed by the LangGraph state machine.
| Decision Area | Technology / Pattern | Engineering Rationale |
|---|---|---|
| Agent Framework | LangGraph (StateGraph) | Deterministic state machine, checkpointing/resume capabilities, conditional routing, and granular step observability. |
| Multi-Modal LLMs | Google Gemini + GitHub Models (GPT-4o) | High-speed multi-modal vision with automatic failover for high availability and zero vendor lock-in. |
| Persistence Layer | SQLAlchemy 2.0 ORM | Enterprise relational data layer with auto-fallback to SQLite when PostgreSQL is offline for zero-friction local development. |
| Duplicate Prevention | Redis + SHA-256 Fingerprinting | O(1) idempotent hash lookup over key invoice attributes (vendor:number:total) with configurable TTL. |
| Human-in-the-Loop | LangGraph Interrupts | Pauses workflow execution for manager/director sign-off without dropping state; resumes via /resume/{thread_id}. |
| Backend API | FastAPI | Async I/O, automatic OpenAPI Swagger documentation, dependency injection, and Pydantic validation. |
| Frontend UI | Streamlit | Rapid, reactive UI rendering with interactive audit logs, pending review queues, and live analytics. |
- Python 3.12+
- Gemini API Key (or GitHub Personal Access Token for GitHub Models)
- Docker & Docker Compose (optional, for containerized execution)
-
Clone the repository:
git clone https://github.com/ayushcodes27/doc-agent.git cd doc-agent -
Configure environment variables:
cp .env.example .env # Edit .env and supply your GEMINI_API_KEY (or GITHUB_TOKEN) -
Launch the entire stack:
docker-compose up --build
-
Access the services:
- Streamlit UI: http://localhost:8501
- FastAPI Docs (Swagger): http://localhost:8000/docs
- PostgreSQL:
localhost:5432 - Redis:
localhost:6379
-
Create and activate a virtual environment:
python -m venv .venv # Windows (PowerShell): .venv\Scripts\Activate.ps1 # macOS/Linux: source .venv/bin/activate
-
Install dependencies:
pip install -r requirements.txt
-
Configure environment:
cp .env.example .env # Ensure your GEMINI_API_KEY / GITHUB_TOKEN are set in .env -
Start the FastAPI backend server:
uvicorn api.main:app --host 127.0.0.1 --port 8000 --reload
-
In a separate terminal, launch the Streamlit frontend:
streamlit run ui/app.py
The test suite covers data models, agent nodes, validation rules, anomaly detection, deduplication, database persistence, and the complete LangGraph workflow.
# Run all 53 unit and integration tests
pytest -v
# Run real-world benchmark suite
python eval/run_real_benchmark.py
# Run offline compliance rule tests
python eval/run_eval.py --mode offlinedoc-agent/
├── agent/ # LangGraph StateGraph workflow, nodes, & human-in-the-loop logic
├── api/ # FastAPI REST service (/process, /resume, /invoices, /analytics)
├── data/ # Real procurement datasets (USASpending.gov) & benchmark invoices
├── db/ # SQLAlchemy 2.0 persistence layer, ORM models, & repository CRUD
├── eval/ # Benchmark runners, ground truth datasets, and evaluation reports
├── tools/ # Multi-modal extractors, 9-rule validator, anomaly & dedup engines
├── ui/ # Streamlit frontend (Document ingestion, review queue, analytics)
├── scripts/ # Ingestion, seeding, and downloading utilities
├── tests/ # Pytest unit & integration test suite (53 tests)
├── config.py # Environment configuration & LLM provider settings
└── docker-compose.yml # Container orchestration (API, UI, PostgreSQL, Redis)
Distributed under the MIT License. See LICENSE for more information.



