Real-Time Evaluation Models for RAG: Who Detects Hallucinations Best?

Abstract

This article surveys Evaluation models to automatically detect hallucinationsin Retrieval-Augmented Generation (RAG), and presents a comprehensive benchmarkof their performance across six RAG applications. Methods included in our studyinclude: LLM-as-a-Judge, Prometheus, Lynx, the Hughes Hallucination EvaluationModel (HHEM), and the Trustworthy Language Model (TLM). These approaches areall reference-free, requiring no ground-truth answers/labels to catch incorrectLLM responses. Our study reveals that, across diverse RAG applications, some ofthese approaches consistently detect incorrect RAG responses with highprecision/recall.