Agentic PDF Document Extraction for Infra Admins (Deprecated) — Unique AI Documentation

Agentic PDF Document Extraction for Infra Admins (Deprecated)

Target Audience

Infrastructure Admins who will deploy and run the Agentic Ingestion service

Description

Agentic PDF Document Extraction is an advanced document processing service that leverages AI-powered extraction techniques to convert PDF documents into structured, searchable content. The service uses multiple extraction methods including MDI (Microsoft Document Intelligence), Vision-based extraction, and hybrid approaches to provide superior text extraction accuracy compared to traditional OCR methods.

Our platform previously relied on basic OCR and manual document processing workflows. While functional, this approach had several limitations for enterprise requirements, particularly in handling complex financial documents, tables, and multi-format content.

Trigger

Activated when pdfReadMode = CUSTOM_SINGLE_PAGE_API with the Agentic Ingestion API identifier configured in CUSTOM_API_DEFINITIONS.

Processing Flow

Step-by-step:

  1. node-ingestion-worker splits the PDF into individual pages
  2. For each page, the Custom API Definition Parser sends the full page as base64-encoded PDF to POST /agentic-ingestion/extractions
  3. agentic-ingestion enqueues a job in Redis (taskiq:pdf-content-extraction) and returns a job_id
  4. The worker picks up the job and selects the extraction method based on the extractionMethod parameter:
    • MDI — Structured extraction via Azure Document Intelligence only
    • VISION — Image-based extraction via Azure OpenAI vision model only
    • MDI_VISION — Hybrid approach: MDI for structured content + Vision for image/chart content
  5. The extraction runs against external services (Azure MDI and/or Azure OpenAI)
  6. Optionally, the Page Content Optimizer post-processes the output (evaluator/generator loop to correct layout errors and improve readability)
  7. The result (extracted markdown) is stored in Redis
  8. node-ingestion-worker polls GET /agentic-ingestion/extractions/{job_id} until the job completes, then receives the extracted markdown

Code Path (node-ingestion-worker)

PDF Ingestor Service
  └─ pdfReadMode === CUSTOM_SINGLE_PAGE_API
      └─ CustomApiDefinitionParser.runCustomApi()
          ├─ POST {definition.url}/extractions  (full page PDF base64)
          └─ Poll GET {definition.url}/extractions/{job_id}

Code Path (agentic-ingestion)

POST /agentic-ingestion/extractions
  └─ Enqueue in Redis (taskiq:pdf-content-extraction)
      └─ Worker: process_extraction_job()
          ├─ MDI path → Azure Document Intelligence
          ├─ VISION path → render page → Vision LLM
          ├─ MDI_VISION path → both in parallel
          └─ Optional: PageContentOptimizer (evaluator/generator loops)

Environment Variables and Secrets

Name Description Example Value Default Required
API_BASE Base URL for the Unique AI API <http://node-chat.finance-gpt.svc.cluster.local:8092/public> "" Yes
FEATURE_FLAG_ENABLE_PDF_CONTENT_EXTRACTION false No, must be set to true to enable the feature
MDI_LOCATION MDI service location identifier "switzerland-north" "" Yes
MDI_ENDPOINT_DEFINITIONS JSON array of MDI endpoint definitions `[{