Agentic Metadata Extraction for Infra Admins — Unique AI Documentation
Agentic Metadata Extraction for Infra Admins
This feature is currently in BETA.
The Agentic Metadata Extraction Service automatically extracts structured metadata from document content using Large Language Models (LLMs). Unlike traditional metadata extraction relying on document properties, this service analyzes actual text content to intelligently extract business-relevant metadata fields.
Processing Flow
- User uploads document → Content is ingested and chunked
- Platform sends webhook to /metadata-extraction/webhook
- Service fetches document chunks and configuration
- Chunks merged up to maxInputTokens limit
- LLM extracts metadata based on configured schema
- Extracted metadata merged with existing content metadata
- Returns 200 OK when complete synchronous processing
The service processes already-ingested content chunks, not raw documents. Processing is synchronous - the webhook completes when extraction finishes.
Code Path (agentic-ingestion)
POST /metadata-extraction
└─ Receive content + metadata schema
└─ LLM completion with structured output
└─ Return extracted metadata fields per schema
Provisioning
Prerequisites
Service-Specific Requirements
- Feature flag: FEATURE_FLAG_ENABLE_AGENTIC_METADATA_EXTRACTION_UN_15619 =true
- Unique AI API access for content retrieval and updates
- LLM endpoints (Azure OpenAI or compatible) supporting structured output
- Network access to Unique AI API and LLM endpoints
Deployment
Environment Variables and Secrets
| Change | Environment Variable Name | Application Default Value (if env variable unset) |
Example | Required | Applications | Short Description |
|---|---|---|---|---|---|---|
| Added | FEATURE_FLAG_ENABLE_AGENTIC_METADATA_EXTRACTION_UN_15619 |
"false" |
"true" |
No | - web-app-knowledge-upload- backend-service-agentic-ingestion |
Enable metadata extraction |
| Added | DEFAULT_ENCODER_NAME |
"o200k_base" |
"o200k_base" |
No | - backend-service-agentic-ingestion |
Token encoder |
| Added | METADATA_EXTRACTION_LANGUAGE_MODELS |
"" |
"AZURE_GPT_4o_2024_0806:GPT-4o (2024-08-06),AZURE_GPT_4o_2024_1120:GPT-4o (2024-11-20)" |
Yes | - web-app-knowledge-upload |
Comma-separated list of available LLM models for UI selection |
| Added | DEFAULT_MAX_TOKENS_USED_FOR_METADATA_EXTRACTION |
10000 |
10000 |
No | - web-app-knowledge-upload |
Sizing
Performance
| Scenario | Processing time |
|---|---|
| Simple (3-5 fields) | 2-5 seconds |
| Complex (10+ field) | 5-15 seconds |
Resource recommendations
We recommend a deployment with 2-4 replicas for High Availability.
Cost Estimation
LLM costs (GPT-4o example):
- ~$0.05-0.15 per document
- 10,000 docs/month: ~$500-1500/month
Configuration
Webhook Setup
To trigger the Agentic Metadata Extraction app, we need to send a POST request to the/metadata-extraction/webhook endpoint.
Metadata Schema
Schemas are configured per-folder in the Unique AI platform UI (Knowledge Upload → Folder Settings → Configure file ingestion).
Field Properties
| Key | Description | Value |
|---|---|---|
| type | Data type | string, number, boolean |
| description | Natural language description to guide LLM | Text |
| required | Whether field must be extracted or not | true/false |
Example Metadata Schema
{
"summary": {
"type": "string",
"required": true,
"description": "A brief one-sentence summary of the document"
},
"documentType": {
"type": "string",
"required": true,
"description": "What type of document is this? (e.g., invoice, report, email)"
}
}
Operating & Troubleshooting
Authentication Methods
The service uses:
- API Keys: For Unique AI API access
API Endpoints
Webhook Endpoint: POST /metadata-extraction/webhook Payload:
{
"event": "unique.content.ingestion.finished",
"companyId": "company-123",
"userId": "user-456",
"payload": {
"contentId": "content-789"
}
}
Troubleshooting
The metadata extraction feature is not visible on the UI, Why?
- Verify
FEATURE_FLAG_ENABLE_AGENTIC_METADATA_EXTRACTION_UN_15619=trueis set in Frontend (knowledge-upload app)
Metadata’s are not visible after Ingestion has finished, Why?
- Verify
FEATURE_FLAG_ENABLE_AGENTIC_METADATA_EXTRACTION_UN_15619=trueis set in Backend (agentic-ingestion app) - Check if module is loaded
Key Metrics to Monitor:
- Webhook processing lag/queue length
- Processing time per document
- Success/failure rates
- Resource utilization (CPU, Memory)
- LLM API response times
- Content API health & authentication success rate
- Token consumption & costs
- Extraction quality (fields extracted, empty extraction rate)