Agentic Metadata Extraction for Infra Admins — Unique AI Documentation

Agentic Metadata Extraction for Infra Admins

This feature is currently in BETA.

The Agentic Metadata Extraction Service automatically extracts structured metadata from document content using Large Language Models (LLMs). Unlike traditional metadata extraction relying on document properties, this service analyzes actual text content to intelligently extract business-relevant metadata fields.

Processing Flow

  1. User uploads document → Content is ingested and chunked
  2. Platform sends webhook to /metadata-extraction/webhook
  3. Service fetches document chunks and configuration
  4. Chunks merged up to maxInputTokens limit
  5. LLM extracts metadata based on configured schema
  6. Extracted metadata merged with existing content metadata
  7. Returns 200 OK when complete synchronous processing

The service processes already-ingested content chunks, not raw documents. Processing is synchronous - the webhook completes when extraction finishes.

Code Path (agentic-ingestion)

POST /metadata-extraction
  └─ Receive content + metadata schema
      └─ LLM completion with structured output
          └─ Return extracted metadata fields per schema

Provisioning

Prerequisites

Service-Specific Requirements

Deployment

Environment Variables and Secrets

Change Environment Variable Name Application Default Value
(if env variable unset)
Example Required Applications Short Description
Added FEATURE_FLAG_ENABLE_AGENTIC_METADATA_EXTRACTION_UN_15619 "false" "true" No - web-app-knowledge-upload
- backend-service-agentic-ingestion
Enable metadata extraction
Added DEFAULT_ENCODER_NAME "o200k_base" "o200k_base" No - backend-service-agentic-ingestion Token encoder
Added METADATA_EXTRACTION_LANGUAGE_MODELS "" "AZURE_GPT_4o_2024_0806:GPT-4o (2024-08-06),AZURE_GPT_4o_2024_1120:GPT-4o (2024-11-20)" Yes - web-app-knowledge-upload Comma-separated list of available LLM models for UI selection
Added DEFAULT_MAX_TOKENS_USED_FOR_METADATA_EXTRACTION 10000 10000 No - web-app-knowledge-upload

Sizing

Performance

Scenario Processing time
Simple (3-5 fields) 2-5 seconds
Complex (10+ field) 5-15 seconds

Resource recommendations

We recommend a deployment with 2-4 replicas for High Availability.

Cost Estimation

LLM costs (GPT-4o example):


Configuration

Webhook Setup

To trigger the Agentic Metadata Extraction app, we need to send a POST request to the/metadata-extraction/webhook endpoint.

Metadata Schema

Schemas are configured per-folder in the Unique AI platform UI (Knowledge Upload → Folder Settings → Configure file ingestion).

Field Properties

Key Description Value
type Data type string, number, boolean
description Natural language description to guide LLM Text
required Whether field must be extracted or not true/false

Example Metadata Schema

{
  "summary": {
    "type": "string",
    "required": true,
    "description": "A brief one-sentence summary of the document"
  },
  "documentType": {
    "type": "string",
    "required": true,
    "description": "What type of document is this? (e.g., invoice, report, email)"
  }
}

Operating & Troubleshooting

Authentication Methods

The service uses:

API Endpoints

Webhook Endpoint: POST /metadata-extraction/webhook Payload:

{
  "event": "unique.content.ingestion.finished",
  "companyId": "company-123",
  "userId": "user-456",
  "payload": {
    "contentId": "content-789"
  }
}

Troubleshooting

The metadata extraction feature is not visible on the UI, Why?

Metadata’s are not visible after Ingestion has finished, Why?


Key Metrics to Monitor:

  1. Webhook processing lag/queue length
  2. Processing time per document
  3. Success/failure rates
  4. Resource utilization (CPU, Memory)
  5. LLM API response times
  6. Content API health & authentication success rate
  7. Token consumption & costs
  8. Extraction quality (fields extracted, empty extraction rate)