Bỏ qua để đến nội dung
Search & RAG

10. Nguồn tham khảo

⚠️ Danh sách này viết từ trí nhớ. Tên tác giả / năm / venue nên kiểm tra lại trước khi trích dẫn trong tài liệu chính thức.

Sách — đọc trước nếu muốn nền tảng chắc

Phần tiêu đề “Sách — đọc trước nếu muốn nền tảng chắc”
NguồnGhi chú
Manning, Raghavan & Schütze — Introduction to Information Retrieval, Cambridge UP, 2008. Chương 8: Evaluation in information retrievalĐiểm khởi đầu tốt nhất. Miễn phí: https://nlp.stanford.edu/IR-book/
Croft, Metzler & Strohman — Search Engines: Information Retrieval in Practice, 2009. Chương 8Thực dụng hơn, gần với engineering
Baeza-Yates & Ribeiro-Neto — Modern Information Retrieval, 2nd ed., 2011. Chương 4: Retrieval EvaluationBao quát nhất về metrics
Harman — Information Retrieval Evaluation, Synthesis Lectures, Morgan & Claypool, 2011Ngắn, tập trung hoàn toàn vào evaluation
Sakai — Laboratory Experiments in Information Retrieval, Springer, 2018Phần thống kê / topic set size design mạnh nhất
Voorhees & Harman (eds.) — TREC: Experiment and Evaluation in Information Retrieval, MIT Press, 2005Lịch sử và phương pháp luận TREC
MetricNguồn
nDCGJärvelin & Kekäläinen, “Cumulated gain-based evaluation of IR techniques”, ACM TOIS 20(4):422–446, 2002
nDCG — phân tích lý thuyếtWang, Wang, Li, He, Liu, “A Theoretical Analysis of NDCG Type Ranking Measures”, COLT 2013
ERRChapelle, Metlzer, Zhang, Grinspan, “Expected Reciprocal Rank for Graded Relevance”, CIKM 2009
RBPMoffat & Zobel, “Rank-biased precision for measurement of retrieval effectiveness”, ACM TOIS 27(1), 2008
bprefBuckley & Voorhees, “Retrieval evaluation with incomplete information”, SIGIR 2004
infAPYilmaz & Aslam, “Estimating average precision with incomplete and imperfect judgments”, CIKM 2006
Mọi metric = mô hình người dùng (C/W/L)Moffat, Bailey, Scholer, Thomas, “Incorporating User Expectations and Behavior into the Measurement of Search Effectiveness”, ACM TOIS 35(3), 2017
Metric vs hành vi thật của người dùngMoffat, Thomas, Scholer, “Users versus models: what observation tells us about effectiveness metrics”, CIKM 2013
Diversity (α-nDCG)Clarke et al., “Novelty and diversity in information retrieval evaluation”, SIGIR 2008

Phương pháp luận & thống kê — phần quan trọng nhất cho golden set nhỏ

Phần tiêu đề “Phương pháp luận & thống kê — phần quan trọng nhất cho golden set nhỏ”
Chủ đềNguồn
Cần bao nhiêu topicVoorhees & Buckley, “The effect of topic set size on retrieval experiment error”, SIGIR 2002
So sánh các test thống kê cho IRSmucker, Allan, Carterette, “A comparison of statistical significance tests for information retrieval evaluation”, CIKM 2007
Bootstrap trong IRSakai, “Evaluating evaluation metrics based on the bootstrap”, SIGIR 2006
Độ tin cậy của thí nghiệm IR lớnZobel, “How reliable are the results of large-scale information retrieval experiments?”, SIGIR 1998
Phương sai topic > phương sai hệBanks, Over, Zhang, “Blind men and elephants: Six approaches to TREC data”, Information Retrieval 1(1–2), 1999
Multiple testingCarterette, “Multiple testing in statistical analysis of systems-based information retrieval experiments”, ACM TOIS 30(1), 2012
Danh sách lỗi thường gặpFuhr, “Some Common Mistakes In IR Evaluation, And How They Can Be Avoided”, SIGIR Forum 51(3), 2017 — đọc cái này, ngắn và trực tiếp
Phản biện FuhrSakai, “On Fuhr’s Guideline for IR Evaluation”, SIGIR Forum 54(1), 2020 — đọc cùng cái trên để thấy tranh luận thật
Bất đồng giữa người gán nhãnVoorhees, “Variations in relevance judgments and the measurement of retrieval effectiveness”, Information Processing & Management 36(5), 2000
Thang đo / interval scale của metricFerrante, Ferro, Maistro, “Towards a Formal Framework for Utility-oriented Measurements of Retrieval Effectiveness”, ICTIR 2015
Topic set size designSakai, “Topic set size design”, Information Retrieval Journal 19(3), 2016

Benchmark & evaluation cho retrieval hiện đại

Phần tiêu đề “Benchmark & evaluation cho retrieval hiện đại”
NguồnGhi chú
Thakur, Reimers, Rücklé, Srivastava, Gurevych — “BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models”, NeurIPS Datasets & Benchmarks 2021Chuẩn de-facto cho so sánh retrieval zero-shot; dùng nDCG@10
Muennighoff, Tazi, Magne, Reimers — “MTEB: Massive Text Embedding Benchmark”, EACL 2023Leaderboard embedding; đọc để biết cách nó có thể gây nhầm lẫn (overfit leaderboard)
MS MARCO — Nguyen et al., 2016Dataset gốc; qrels rất thưa (~1 doc/query) → giải thích tại sao MRR@10 là metric của nó
NguồnGhi chú
Es, James, Espinosa-Anke, Schockaert — “RAGAS: Automated Evaluation of Retrieval Augmented Generation”, EACL 2024 (System Demonstrations)Nguồn gốc của faithfulness / answer relevance / context precision-recall
Saad-Falcon, Khattab, Potts, Zaharia — “ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems”, NAACL 2024Cách train judge nhẹ thay vì gọi LLM lớn
Rashkin et al. — “Measuring Attribution in Natural Language Generation Models”, Computational Linguistics 2023Khung AIS — định nghĩa chặt chẽ của “attributable to identified sources”
Gao, Yen, Yu, Chen — “Enabling Large Language Models to Generate Text with Citations” (ALCE), EMNLP 2023Đo citation precision/recall một cách hệ thống
Liu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni, Liang — “Lost in the Middle: How Language Models Use Long Contexts”, TACL 2024Bằng chứng thực nghiệm cho việc precision của context thật sự quan trọng, không chỉ recall
Zheng et al. — “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”, NeurIPS 2023Các bias của LLM-judge: độ dài, vị trí, tự thiên vị
TREC RAG Track (từ 2024)Nỗ lực chuẩn hoá evaluation cho RAG theo phương pháp TREC

Công cụ (để tham khảo cách người khác implement)

Phần tiêu đề “Công cụ (để tham khảo cách người khác implement)”
ToolGhi chú
trec_evalImplementation tham chiếu chính thức của TREC. Khi nghi ngờ công thức, so với cái này.
pytrec_eval (Van Gysel & de Rijke, SIGIR 2018)Binding Python của trec_eval
ir_measuresAPI thống nhất cho nhiều metric, dễ dùng hơn trec_eval
ranxCó sẵn cả bootstrap significance test và fusion
ragasLibrary của paper RAGAS

Phần trên viết từ trí nhớ nên chỉ có tên. Dưới đây là các nguồn có link truy cập được ngay — dùng cái này khi cần đọc chứ không chỉ cần trích dẫn.

NguồnLink
IR Book — ch.8 Evaluation (bản HTML)https://nlp.stanford.edu/IR-book/html/htmledition/evaluation-in-information-retrieval-1.html
IR Book — toàn văn PDFhttps://nlp.stanford.edu/IR-book/pdf/irbookonlinereading.pdf
Stanford CS276 — slide/handouthttps://web.stanford.edu/class/cs276/
BEIR (paper)https://arxiv.org/abs/2104.08663
MTEB (paper)https://arxiv.org/abs/2210.07316
MTEB leaderboardhttps://huggingface.co/spaces/mteb/leaderboard
RAGAS (paper)https://arxiv.org/abs/2309.15217
Lost in the Middle (paper)https://arxiv.org/abs/2307.03172
Self-RAG (paper)https://arxiv.org/abs/2310.11511
Searching for Best Practices in RAG (paper)https://arxiv.org/abs/2407.01219
trec_evalhttps://github.com/usnistgov/trec_eval
ir_measureshttps://github.com/terrierteam/ir_measures
ranxhttps://github.com/AmenRa/ranx
ragashttps://github.com/explodinggradients/ragas
Offline evaluation cho retrieval — Pinecone Learnhttps://www.pinecone.io/learn/offline-evaluation/
Interleaving thay A/B test — Airbnb Engineeringhttps://medium.com/airbnb-engineering/beyond-a-b-test-speeding-up-airbnb-search-ranking-experimentation-through-interleaving-7087afa09c8e
Team Draft Interleaving — OpenSource Connectionshttps://opensourceconnections.com/blog/2025/08/06/a-b-testing-with-team-draft-interleaving/
[VI] Đánh giá các mô hình học máy — Viblohttps://viblo.asia/p/danh-gia-cac-mo-hinh-hoc-may-RnB5pp4D5PG
Phần 2 — Metrics