OpenFest 2026

Evaluation Driven Development for RAG Systems with Coding Agents
Език: English

A RAG pipeline can look great in a demo and still fail when the data, queries, or requirements change. The harder part is not adding another retriever or swapping another LLM. It is knowing what broke, why it broke, and whether the change actually made the system better.

This talk explores Evaluation Driven Development for RAG systems with Coding Agents. We will start by defining what good looks like, create an evaluation dataset, establish a baseline, and use measurable signals to guide development. We will evaluate both retrieval and generated responses, covering metrics such as Hit Rate, Recall, MRR, nDCG, relevance, and faithfulness.

Then we will bring Coding Agents into the loop. Instead of asking an Agent to simply build a RAG pipeline, we will give it evaluations that it can run, inspect, and use to iterate on the implementation. The goal is a development workflow where every change can be tested against evidence rather than judged by whether the latest demo looks better.


We will start with the practical steps: defining the task, creating a representative dataset, choosing the right signals, and establishing a baseline.

From there, we will evaluate the two major parts of a RAG pipeline separately. For retrieval, we will use metrics such as Hit Rate, Recall, MRR and nDCG to understand whether the right context is being found and ranked. For generation, we will examine answer quality, relevance, faithfulness and other useful signals.