Agentic Metadata-Extraction — Unique AI Documentation

Agentic Metadata-Extraction

This feature is currently in BETA. It may change before general availability, due to user and client feedback, but it targeted to be high quality and stable. Documentation may lag behind feature updates. Use in production environments at your own discretion. Please refer to our Upgrade and Release Process for more information.

Overview

Traditional document metadata relies on file properties (author, creation date) or manual tagging, which often fails to capture the semantic richness of document content. In the Financial Services Industry, where documents contain critical information like client names, account numbers, document types, and regulatory classifications, automated metadata extraction is essential for effective document management and retrieval.

AI-Powered Metadata Extraction addresses this challenge by using Large Language Models (LLMs) to intelligently analyze document content and extract structured metadata based on user-defined schemas. This approach ensures that documents are automatically tagged with meaningful, searchable metadata that enhances discoverability and enables sophisticated document workflows.

When you enable Metadata Extraction in your document processing workflow, our platform:

This intelligent approach ensures your knowledge base is enriched with accurate, consistent metadata that improves document discoverability and enables automated workflows.

Who It's For

How It Works

Backbone Components

Example Use Cases

Document Classification

Policy Documents: "Classify documents by policy type and effective date"

Client Documents: "Extract client information from onboarding documents"

Step-by-Step Guide

1. Enable Metadata Extraction

  1. Click Configure File Ingestion in a folder of choice
  2. Locate the AI Metadata Extraction section in the configuration panel
  3. Toggle the feature ON to enable metadata extraction

2. Configure Language Model

In the metadata extraction configuration:

  1. Select your preferred Language Model from the available options, for instance:
    • GPT-4o (2024-08-06)
    • GPT-4o (2024-11-20)
  2. Set the Max Input Tokens (1000-10000):
    • Lower values process faster but may miss content at the end of documents
    • Higher values capture more context but increase processing time
    • Recommended: Start with 5000 tokens and adjust based on your documents

3. Define Metadata Schema

Create a JSON schema defining the metadata fields to extract. Each field requires:

Example Schema:

{
  "author_name": {
    "type": "string",
    "description": "Name of the document author or creator",
    "required": true
  },
  "publication_date": {
    "type": "string",
    "description": "Publication or creation date in YYYY-MM-DD format",
    "required": true
  },
  "document_type": {
    "type": "string",
    "description": "Type of document (e.g., report, memo, policy, analysis)",
    "required": true,
    "enum": ["invoice", "contract", "report", "other"]
  }
}

4. Upload Documents

Upload your documents through the standard Unique AI interface. The system will automatically:

  1. Process the document through standard ingestion (reading, chunking, embedding)
  2. Trigger metadata extraction upon ingestion completion
  3. Analyze document content using the configured LLM
  4. Extract metadata according to your schema
  5. Update the document's metadata fields

5. Verify Results

Review the extracted metadata on your documents:

  1. Navigate to the folder containing your uploaded documents
  2. Select a document to view its details
  3. Check the Metadata section for extracted fields
  4. Verify accuracy and completeness of extracted values

6. Re-run Extraction on Existing Files

Enabling metadata extraction does not automatically process files already in the folder. To backfill existing documents:

  1. Check "Extract metadata on existing files in this folder"
  2. Optionally check "Also extract on files in subfolders" to include the entire folder tree
  3. Click Save

Only fully ingested documents are targeted. New uploads are processed automatically and do not require this step.

Configuration Options

Language Models

Language Models may differ across companies and environments.

Requirements: All models must be Azure OpenAI deployments with Structured Output support (e.g., GPT-4o series). This allows companies to adopt newer models as they become available by simply updating the configuration, without requiring code changes.

Schema Field Types

Type Description Example Values
string Text values "John Smith", "2024-01-15"
number Numeric values 42, 3.14, 1000000
boolean True/false values true, false

Enum Constraints & Field Validation

After the LLM responds, every extracted field is validated against your schema before being saved. Fields that fail are dropped (not saved), never causing an error on the document.

Data type validation — each field's value is checked against its declared type:

If the LLM returns... Outcome
Correct type Field saved
Wrong type (e.g. a string for a number field) Field dropped with a warning
null on an optional field Field skipped silently
A required field missing entirely Field dropped with a warning

Enum constraints — add an optional enum property to restrict the LLM to a fixed set of allowed values:

{
  "document_type": {
    "type": "string",
    "description": "Type of document",
    "required": true,
    "enum": ["invoice", "contract", "report", "other"]
  }
}

Values outside the enum are dropped after extraction. For array fields, only the items within the allowed set are kept; if none remain, the field is dropped entirely.

Token Settings

Setting Description Default Recommended Range
maxInputTokens Maximum document length to process 10000 1000-10000