Question Extractor — Unique AI Documentation
Question Extractor
Functionality
This module is designed to extract all questions from a document, which can for example be a meeting transcription or a request for proposal.
Sequence of events within module
File Handling:
Automatically detects the relevant file name from user input.
- Possibility to upload directly upload the transcript document and process its content.
Company/Project Name Extraction: Knowing the company or the project name improves the model better understand the transcript (Particularly with extended questions).
Question Extraction: Identifies and extracts questions from the transcript.
Question Extending: Question extracted from the transcript usually lack context. Thus, we extend the original questions extracted in the previous step with using the context from the chunk.
Topic Assignment: Assigns topics to the extracted questions for better categorisation.
Chat Output Streaming: After processing each chunk, the message on the chat interface is updated with the new questions that have been extracted.
Excel Export: Exports the collected questions to an Excel file and links it in the chat for downloading.
Word Export: Exports a version of the transcript document with highlights of extracted questions.
Input
The input document is preferably a Word file (the module also works with PDF but no highlighting functionality is available in this case.) uploaded to the designated scope in the Knowledge Center, which is connected to the extractor module or directly to the chat.
A user then uses the assistant linked to this module to extract all questions, e.g.
- Extract questions from meeting_transcript_xyz.pdf.
If the file is directly uploaded to the chat, the user can use the following prompt:
- Extract questions from this uploaded file.
Meeting Transcript Format:
- Each statement in the discussion must be preceded by the name of the speaker. Below you can see an example document:
Output
The output consists of:
- Table in chat interface containing all extracted questions
- Downloadable excel file with extracted questions
- Downloadable word file with extracted questions highlighted in yellow
Reference in Code in AI Module Template
QuestionExtractor
Configuration settings
This section contain all configurable parameters.
Default Configuration
{
"languageModel": "AZURE_GPT_4o_2024_1120",
"companyNameExtraction": {
"systemMessage": "You are given a file name. Your task is to extract the name of the company contained in the file name. If you can not extract a company name, return \"NONE\". \"IC\" is not a company name.",
"userMessage": "$document_name",
"exampleMessages": [
{
"role": "user",
"content": "20240215_annual_report_ubs.pdf"
},
{
"role": "assistant",
"content": "{\n \"company_name\": \"UBS\"\n}"
}
]
},
"questionExtraction": {
"useCompanyName": true,
"additionalInputKeywords": [],
"dummyExtendedQuestionPlaceholder": "Additional Input",
"sheet1Name": "Sheet1",
"sheet2Name": "Sheet2",
"originalQuestion": {
"system": "Your role as an assistant is to analyze a transcript from a meeting about an investment opportunity.",
"trigger": "Transcript:\n''\n${chunk}\n''\n\nJSON output::"
},
"additionalInputs": {
"system": "Your purpose is to convert a text with bullet points into a structured JSON object.",
"trigger": "Input:\n''\n${chunk}\n''\n\nJSON output:"
}
},
"excelGenerator": {
"uploadScopeId": null,
"uploadToChat": true,
"tableDataFormat": {
"bg_color": "#FFFFFF",
"bold": false,
"font_color": "black",
"text_wrap": true,
"border": 1,
"valign": "top"
}
}
}
General parameters
| Parameter | Description | Type | Default |
| languageModel | used LLM model | string | AZURE_GPT_4o_2024_1120 |
| companyNameExtraction | Configuration for the service that attempts to extract a company name from the uploaded document’s filename. | object | |
| questionExtraction | Configuration for question extraction service | object | |
| excelGenerator | See Excel Generator | object | |