RAG Assessment and Improvement — Unique AI Documentation

RAG Assessment and Improvement

Evaluation of Information Retrieval Systems: A Comprehensive Assessment Framework

This report provides a comprehensive evaluation of information retrieval (IR) systems, with a particular focus on the performance of semantic search and its enhancements through combined methodologies. The scope of this assessment is not limited to research in retrieval-augmented generation (RAG), but extends to evaluating IR setups across a broader range of applications.

We begin by introducing the assessment approach that was implemented to systematically evaluate our current retrieval setup. Two core evaluation metrics were employed: Recall, which measures the completeness of retrieval, and Normalized Discounted Cumulative Gain (NDCG), which assesses both the relevance and ranking of the retrieved results. These metrics offer complementary perspectives—ensuring not only that relevant information is captured, but also that it is prioritized effectively.

To ground our evaluations in realistic usage scenarios, we constructed a novel assessment dataset using large language models (LLMs). These models were leveraged to generate diverse queries and rank document chunks, closely simulating the types of interactions expected in production systems.

The assessment benchmarks the performance of our baseline method—semantic search—against alternative and combined strategies aimed at enhancing retrieval effectiveness. In particular, we examine the integration of reranking techniques, which reorder retrieved results based on relevance signals to improve precision.

Key findings reveal that combining semantic search with reranking and other hybrid retrieval strategies leads to significant improvements in both recall and ranking metrics. These improvements underscore the practical value of layered retrieval techniques, especially in environments demanding high accuracy and coverage.

The report concludes with an analysis of trade-offs observed in performance, such as latency versus precision, and identifies areas for future optimization and experimentation. These insights will inform ongoing enhancements to our IR infrastructure and support the development of more robust and responsive systems.

Evaluation Metrics for information retrieval

There are different metrics that can be used to assess the performance of Information Retrieval (IR) systems. In this report, we will restrict ourselves to two metrics:

Recall

The recall is a metric used to measure the system's ability to retrieve all relevant documents from a given dataset. It quantifies the proportion of relevant documents retrieved by the system out of the total number of relevant documents available. In simpler terms, recall assesses how well the system avoids missing relevant information. A high recall indicates that the system effectively retrieves a large portion of the relevant documents, while a low recall suggests that the system may be overlooking important information. When used carefully, recall is easy to implement and can provide valuable information about the performance of the IR system. However, its main disadvantage is being order-unaware metric. This means that it cannot account for the relevance level of the retrieved items.

NDCG

NDCG (Normalized Discounted Cumulative Gain) is a metric commonly used in Information Retrieval systems to evaluate the quality of ranked search results. It takes into account both relevance and the position of relevant documents in the ranked list. NDCG assigns higher scores to relevant documents appearing higher in the list and discounts the relevance score based on the document's position. The metric is normalized to allow comparison across different search queries and systems, making it a valuable measure for assessing the effectiveness of retrieval algorithms in providing relevant and highly-ranked search results. However, this metric is harder to interpret than recall.

Constructing Evaluation Dataset

Now that we know that what metrics will be used to assess and compare our IR algorithms, we need to have a dataset on which we can run our tests. As you already know, it is difficult to obtain such a dataset, especially when we are using private customer data. To solve this issue, we need to be creative and rely on our best friend, LLMs, particularly GPT4 family from OpenAI. In this section, we will focus on explaining the different steps that were used to construct our assessment dataset.

Question Generation

To start the construction of dataset journey, we first need to generate questions. This can be done by using OpenAI completion model by feeding it a chunk as context and ask it to generate questions about it. To achieve this, we used the following Prompt:

Welcome to your new and specialized position as a Query Generation Wizard. Your role involves crafting questions that are used to extract information from an extensive database of banking documents using semantic search. The documents in the database contain information about various banking products, services, and regulations.
You possess the unique ability to empathize and adopt the perspectives of individuals from different departments (HR, relationship manager,...).
As part of your query generation process, you will receive a header of a document and a piece of text from that document. Your task is to thoroughly contemplate these excerpts and conceive questions or queries that a human with interest in the subject matter might pose, ensuring that the answers to these questions can likely be found within the provided text segment.
You should generate 10 questions: 3 should be open questions, 3 should be specific questions and 4 should be instruction questions. Instruction questions are questions where you ask in first person how to do something, or if you are allowed to do something. These are 2 examples:

  1. "As a banker, am I allowed to… ?"
  2. "What documents do I need to… ?"
    You should not use general references like "What documents do I need to submit to demonstrate compliance with this directive?" or "What are the requirements for compliance with this directive?".
    Instead, you should use specific references like "What documents do I need to submit to demonstrate compliance with the deposit guarantee schemes directive?".

We provide below an example of generate question from a specific chunk:

Example of a chunk

The recovery time objective (RTO) is the time within which an application, system and/or process must be recovered. The recovery point objective (RPO) is the maximum tolerable period during which data is lost.

Operational resilience refers to the institution's ability to restore its critical functions in case of a disruption within the tolerance for disruption. That is to say, the institution's ability to identify threats and possible failures, to protect itself from them and to respond to them, to restore normal business operations in the event of disruptions and to learn from them, so as to minimise the impact of disruptions on the provision of critical functions. An operationally resilient institution has designed its operating model in such a way that it is less exposed to the risk of disruptions in relation to its critical functions. Operational resilience thus reduces not only the residual risks of disruptions, but also the inherent risk of disruptions occurring. Effective operational risk management helps strengthen the institution's operational resilience.

Grouping Chunks

So far, we have a set of question related with each chunk. But, what if a question requires multiple chunks to be answered? We need to find a way to assign to each question the potential set of chunks containing the answer. To achieve this, we start with the following assumption:

Similar chunks should produce similar questions!

Consider the case where we have a set of fund fact sheet documents. The user should more or less ask the same question but only changing the fund name. Thus, in theory, questions asked about funds should be highly similar. Thus, to construct our assessment dataset, we can proceed as follows:

  1. We compute the embeddings of each query
  2. We compute the pairwise similarity of the queries
  3. We only keep questions over a certain level of threshold (e.g. sim>0.9)
  4. We create sets of highly similar questions
  5. We recover the set of chunks that have been used to generate the similar sets questions

Using this approach, we end up 53% of the original generated queries. Now, we have two remaining problems:

Ranking Chunks

As you may have already noticed, when things get hairy, our go-to move is to unleash the mighty force of ChatGPT, our problem-solving pal! Ranking chunks becomes as easy as using a prompt to ask GPT-4 to get the task done for us. Below, you can see the prompt used for this purpose:

You are a helping assistant in a company. You are asked to rank the relevance of a chunk of text with respect to a given question. The relevance score should reflect how well the chunk of text answers the question. For example if the chunk of text doesn't answer the question at all, it should be ranked as 0. If the chunk of text answers the question perfectly, it should be ranked as 4.
You can use the following scale to rank the relevance of the chunks:

Now, you can leave the model label your dataset and go for your lunch break. The histogram below show that most of the chunks are irrelevant. This is probably where our assumption failed. Thus, we simply need to drop irrelevant question-chunk pairs and we are left with 47% of the initial generated questions.

Critics

In this section, we present a workaround solution for acquiring an assessment dataset. Although this approach minimizes human intervention, it places sole responsibility on the completion model, which may fail in some cases to generate human-like queries or accurately rank relevant chunks. Consequently, our current setup falls short of the ideal. An optimal approach would entail human involvement, either through manual labeling of the dataset or thorough review of the generated one to ensure precision and relevance.

Evaluation

In this section, we will start by evaluating the current IR systems that are available and will later propose a reranking method that has the potential to improve the retrieval performance.

Semantic Search vs Combined Search

In this paragraph, we will compare the performance of the two IR method:

  1. Semantic Search (aka Vector)
  2. Combined Search

Semantic search relies on using the context embeddings generated with the model ‘text-embedding-ada-002' from OpenAI to compute the similarity between the documents and users’ queries. On the other hand, the combined search adds Full Text Search to the mix in the computation of query-document similarity. Full Text Search relies on keyword cooccurrence in order to determine relevant chunks.

Reranking

After comparing the existing IR systems, we're diving into reranking using cross encoders to see if we can spice up our current setup. We tested three different models:

  1. cross-encoder/ms-marco-MiniLM-L-12-v2 (Reranker 1 - Monolingual EN)
  2. mixedbread-ai/mxbai-rerank-xsmall-v1 (Reranker 2 - Monolingual EN)
  3. cross-encoder/msmarco-MiniLM-L12-en-de-v1 (Reranker 3 - Bilingual EN-DE)

Results

Our findings are summarized in the table below. Initially, we observe that the combined search followed by reranking with a bilingual model surpasses all other approaches in terms of recall and NDCG. This is particularly crucial given the chatbot model's limited input token size, emphasizing the need to prioritize relevant chunks in ranking.

Metric Method \ Top k 10 20 30 40 50
Recall Semantic (Baseline) 0.65 0.745 0.79 0.815 0.825
Combined 0.705 (+8.46%) 0.8 (+7.38%) 0.835 (+5.70%) 0.86 (+5.52%) 0.865 (+4.85%)
Semantic + Reranker 1 0.665 (+2.31%) 0.705 (-5.37%) 0.725 (-8.23%) 0.755 (-7.36%) -

Discussion

The study we conducted shows that the Combine Search is a very good retrieval approach (Recall@20: 80%). Moreover, when we add a reranking step, we gain 2.5% in Recall@20. But most importantly, the NDCG@20 score jumped from 49.5% to 60.5%. This means that we are 11% better at pushing the relevant chunks in the top of retrieved documents with the combined search.