
RAG evaluation case study — hybrid retrieval pipeline benchmarked with an independent LLM judge, 91% faithfulness, full failure-mode diagnosis.
The RAG Reliability Case Study is a production-style Retrieval-Augmented Generation (RAG) HR assistant evaluated against a 20-question golden dataset with an independent judge model. It documents real failure modes, bottleneck distributions, and the engineering decisions behind a faithfulness score of 0.911. The repository is organized into several resources. CASE_STUDY.md provides a full narrative: system overview, evaluation methodology, results, and what this demonstrates. methodology.md covers evaluation framework design — why an independent judge, how the benchmark was constructed. The evaluation/ directory contains a headline metrics summary and stripped results table. The failure-analysis/ directory contains per-failure root-cause write-ups for Q16, Q17, and Q20. The architecture/ directory contains a system architecture diagram — hybrid retrieval + evaluation pipeline. The screenshots/ directory contains UI screenshots of the chat interface and demo assets. A demo video is also referenced as RAG.case.study.mp4.
Teams use the case study to review a full narrative of system overview, evaluation methodology, and results.
Teams use the methodology document to understand why an independent judge was used and how the benchmark was constructed.
Teams use the evaluation directory to review headline metrics and a stripped results table.
Teams use the failure-analysis directory to examine per-failure root-cause write-ups for Q16, Q17, and Q20.
Teams use the architecture directory to view a system architecture diagram of hybrid retrieval and evaluation pipeline.
Teams use the screenshots directory to see UI screenshots of the chat interface and demo assets.
Lynote combines AI-draft rewriting and AI-text detection with note-taking, transcription, video summaries, and study tools.
LMSYS Chatbot Arena is a crowdsourced open platform for LLM evals. Collected over 1,000,000 human pairwise comparisons to rank LLMs with the Bradley-Terry model and display the model ratings in Elo-scale.
Zhipu Qingyan is a GLM-based AI assistant whose official description emphasizes understanding goals, breaking down tasks, and using tools.
字节跳动AI助手