Agentic Image Content Extraction for Infra Admins — Unique AI Documentation
Agentic Image Content Extraction for Infra Admins
This feature is currently in BETA. It may change before general availability, but it is targeted to be high quality and stable. Documentation may lag behind feature updates. Use in production environments at your own discretion. Please refer to our Upgrade and Release Process for more information.
Overview
| Field | Value |
| Feature name | Agentic Image Content Extraction |
| What it is | Vision AI figure extraction within the MDI pipeline (supersedes deprecated Agentic PDF Extraction) |
| What it does | Extracts textual content from figures, charts, diagrams, and other visual elements detected within PDF pages during ingestion. Uses vision-capable AI models (Azure OpenAI) to interpret cropped figure images and merges the extracted text back into the page markdown at the correct position. Enhances the standard MDI (Microsoft Document Intelligence) pipeline — does not replace it. |
| Who it's for | Knowledge base admins who need higher-quality ingestion of image-heavy documents (financial reports, research papers, regulatory filings). End users benefit from charts and graphs becoming searchable in AI chat. |
| When to use | Pilot tenants with image-heavy document sets (financial reports, research papers, regulatory filings) where chart/graph content is currently lost during ingestion. Recommended to test on a sample folder before enabling space-wide. |
| Deprecation | Agentic PDF Document Extraction (CUSTOM_SINGLE_PAGE_API with the Agentic Ingestion API identifier) is deprecated and should not be enabled for new tenants. Image Content Extraction is the recommended replacement. |
Processing Flow
Step-by-step:
- node-ingestion-worker splits the PDF into individual pages
- For each page, the MDI Client calls Azure Document Intelligence with
extractFigures: true - MDI returns the page text/layout plus figure bounding polygons for each detected figure
- The MDI Page Composer renders the PDF page as an image at the configured DPI
- Each detected figure is cropped from the rendered page image using the polygon coordinates
- Cropped figure images are sent (including the detected language by MDI) to the Image Extraction Adapter (up to 5 in parallel)
- The adapter creates async jobs via
POST /image-content-extraction/extractionsand polls for results - Inside agentic-ingestion, the worker picks up each job and applies the configured strategy:
- ONE_STEP — Single vision LLM call to directly extract content from the figure
- TWO_STEP — First classify the image (chart, table, diagram, icon, etc.), then extract with a category-specific prompt; non-informational categories (icons, decorative images) are skipped
- The vision LLM is called via
API_BASE(node-chat gateway). If the primary model fails, an automatic fallback to a secondary model is attempted - Results are returned to the composer, which merges figure texts into the page markdown at the correct positions, preserving reading order and captions
Fallback Behavior
- Primary → Fallback model: If primary LLM fails (except content filter), retries with fallback model
- Per-figure resilience: If extraction fails for an individual figure → empty text for that figure; rest of page composes normally
- Pipeline fallback: If entire figure extraction pipeline fails → falls back to standard MDI output (no figure text)
Code Path (node-ingestion-worker)
PDF Ingestor Service
└─ pdfReadMode === DOC_INTELLIGENCE_DEFAULT
└─ imageContentExtraction.enabled === true
└─ applyDocIntelligenceOnPageWithImageContentExtraction()
└─ MSDocumentIntelligence.analyzeDocument(page, { extractFigures: true })
└─ MdiPageComposer.composePageWithFigures()
├─ Render PDF page as image (at configured DPI)
├─ Crop each figure using MDI polygon coordinates
└─ AgenticIngestionImageExtractionAdapter.extractImageContent()
├─ POST /image-content-extraction/extractions
└─ Poll GET /image-content-extraction/extractions/{job_id}
Code Path (agentic-ingestion)
POST /image-content-extraction/extractions
└─ Enqueue in Redis (taskiq:image-content-extraction)
└─ Worker: process_image_extraction_job()
└─ run_image_extraction()
├─ ONE_STEP → ClassifierAndExtractor → single vision LLM call
└─ TWO_STEP → Classifier (categorize image)
├─ Non-informational category → skip (return empty)
└─ Informational category → Extractor (category-specific prompt)
How to enable it
Enabling this feature requires changes in four places: deploying the agentic-ingestion service, configuring agentic-ingestion, configuring node-ingestion-worker, enabling the UI feature flag on knowledge-upload, and activating it per scope / space.
A. Deploy the agentic-ingestion service
The agentic-ingestion service must be deployed and running before the feature can be used. Image content extraction is a module within this service — no separate deployment is needed, but the service itself is a prerequisite.
- Helm chart: The agentic-ingestion service is deployed via its own Helm chart as part of the Agentic Ingestion bundle
Environment variables required for setup:
| Env var | Description | Example | Default | Required |
| API_BASE | Base URL for the Unique AI API (node-chat). Used for all LLM vision completions. | http://node-chat.finance-gpt.svc.cluster.local:8092/public | none (mandatory) | Yes |
| FEATURE_FLAG_ENABLE_IMAGE_CONTENT_EXTRACTION_UN_17223 | Enable/disable the image content extraction module. When false, the /image-content-extraction endpoint is not registered. | `