AI Verify Foundation — Unique AI Documentation

AI Verify Foundation

4 min read

Executive Summary

We teamed up with QuantPi in the AI Verify pilot to test our RAG-based Investment Research Assistant. We focused on the big risks: making sure answers are grounded in the right documents, the search pulls the right info, and results stay reliable (even with typos or tricky questions), plus staying within policy and protecting data. We ran about 200 realistic tests in a secure setup and measured both the search and the AI’s answers with simple stats and an AI judge checked by domain experts. The results showed where retrieval quality drives answer quality, helped us tweak search and prompts, and gave us a repeatable way to keep the assistant accurate and trustworthy in real use.

Background: Unique at the AI Verify Foundation

Our Pairing and Use Case

Risks We Prioritized

From a broader risk list, we selected the risks most relevant to our financial research context:

How We Measured Reliability

We combined metrics suited to our RAG pipeline with qualitative calibration:

Test Design and Implementation

What Have We Learnt

Challenges We Faced

Key Insights We Took Away