10. Nguồn tham khảo
⚠️ Danh sách này viết từ trí nhớ. Tên tác giả / năm / venue nên kiểm tra lại trước khi trích dẫn trong tài liệu chính thức.
Sách — đọc trước nếu muốn nền tảng chắc
Phần tiêu đề “Sách — đọc trước nếu muốn nền tảng chắc”| Nguồn | Ghi chú |
|---|---|
| Manning, Raghavan & Schütze — Introduction to Information Retrieval, Cambridge UP, 2008. Chương 8: Evaluation in information retrieval | Điểm khởi đầu tốt nhất. Miễn phí: https://nlp.stanford.edu/IR-book/ |
| Croft, Metzler & Strohman — Search Engines: Information Retrieval in Practice, 2009. Chương 8 | Thực dụng hơn, gần với engineering |
| Baeza-Yates & Ribeiro-Neto — Modern Information Retrieval, 2nd ed., 2011. Chương 4: Retrieval Evaluation | Bao quát nhất về metrics |
| Harman — Information Retrieval Evaluation, Synthesis Lectures, Morgan & Claypool, 2011 | Ngắn, tập trung hoàn toàn vào evaluation |
| Sakai — Laboratory Experiments in Information Retrieval, Springer, 2018 | Phần thống kê / topic set size design mạnh nhất |
| Voorhees & Harman (eds.) — TREC: Experiment and Evaluation in Information Retrieval, MIT Press, 2005 | Lịch sử và phương pháp luận TREC |
Paper nền tảng theo metric
Phần tiêu đề “Paper nền tảng theo metric”| Metric | Nguồn |
|---|---|
| nDCG | Järvelin & Kekäläinen, “Cumulated gain-based evaluation of IR techniques”, ACM TOIS 20(4):422–446, 2002 |
| nDCG — phân tích lý thuyết | Wang, Wang, Li, He, Liu, “A Theoretical Analysis of NDCG Type Ranking Measures”, COLT 2013 |
| ERR | Chapelle, Metlzer, Zhang, Grinspan, “Expected Reciprocal Rank for Graded Relevance”, CIKM 2009 |
| RBP | Moffat & Zobel, “Rank-biased precision for measurement of retrieval effectiveness”, ACM TOIS 27(1), 2008 |
| bpref | Buckley & Voorhees, “Retrieval evaluation with incomplete information”, SIGIR 2004 |
| infAP | Yilmaz & Aslam, “Estimating average precision with incomplete and imperfect judgments”, CIKM 2006 |
| Mọi metric = mô hình người dùng (C/W/L) | Moffat, Bailey, Scholer, Thomas, “Incorporating User Expectations and Behavior into the Measurement of Search Effectiveness”, ACM TOIS 35(3), 2017 |
| Metric vs hành vi thật của người dùng | Moffat, Thomas, Scholer, “Users versus models: what observation tells us about effectiveness metrics”, CIKM 2013 |
| Diversity (α-nDCG) | Clarke et al., “Novelty and diversity in information retrieval evaluation”, SIGIR 2008 |
Phương pháp luận & thống kê — phần quan trọng nhất cho golden set nhỏ
Phần tiêu đề “Phương pháp luận & thống kê — phần quan trọng nhất cho golden set nhỏ”| Chủ đề | Nguồn |
|---|---|
| Cần bao nhiêu topic | Voorhees & Buckley, “The effect of topic set size on retrieval experiment error”, SIGIR 2002 |
| So sánh các test thống kê cho IR | Smucker, Allan, Carterette, “A comparison of statistical significance tests for information retrieval evaluation”, CIKM 2007 |
| Bootstrap trong IR | Sakai, “Evaluating evaluation metrics based on the bootstrap”, SIGIR 2006 |
| Độ tin cậy của thí nghiệm IR lớn | Zobel, “How reliable are the results of large-scale information retrieval experiments?”, SIGIR 1998 |
| Phương sai topic > phương sai hệ | Banks, Over, Zhang, “Blind men and elephants: Six approaches to TREC data”, Information Retrieval 1(1–2), 1999 |
| Multiple testing | Carterette, “Multiple testing in statistical analysis of systems-based information retrieval experiments”, ACM TOIS 30(1), 2012 |
| Danh sách lỗi thường gặp | Fuhr, “Some Common Mistakes In IR Evaluation, And How They Can Be Avoided”, SIGIR Forum 51(3), 2017 — đọc cái này, ngắn và trực tiếp |
| Phản biện Fuhr | Sakai, “On Fuhr’s Guideline for IR Evaluation”, SIGIR Forum 54(1), 2020 — đọc cùng cái trên để thấy tranh luận thật |
| Bất đồng giữa người gán nhãn | Voorhees, “Variations in relevance judgments and the measurement of retrieval effectiveness”, Information Processing & Management 36(5), 2000 |
| Thang đo / interval scale của metric | Ferrante, Ferro, Maistro, “Towards a Formal Framework for Utility-oriented Measurements of Retrieval Effectiveness”, ICTIR 2015 |
| Topic set size design | Sakai, “Topic set size design”, Information Retrieval Journal 19(3), 2016 |
Benchmark & evaluation cho retrieval hiện đại
Phần tiêu đề “Benchmark & evaluation cho retrieval hiện đại”| Nguồn | Ghi chú |
|---|---|
| Thakur, Reimers, Rücklé, Srivastava, Gurevych — “BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models”, NeurIPS Datasets & Benchmarks 2021 | Chuẩn de-facto cho so sánh retrieval zero-shot; dùng nDCG@10 |
| Muennighoff, Tazi, Magne, Reimers — “MTEB: Massive Text Embedding Benchmark”, EACL 2023 | Leaderboard embedding; đọc để biết cách nó có thể gây nhầm lẫn (overfit leaderboard) |
| MS MARCO — Nguyen et al., 2016 | Dataset gốc; qrels rất thưa (~1 doc/query) → giải thích tại sao MRR@10 là metric của nó |
RAG evaluation
Phần tiêu đề “RAG evaluation”| Nguồn | Ghi chú |
|---|---|
| Es, James, Espinosa-Anke, Schockaert — “RAGAS: Automated Evaluation of Retrieval Augmented Generation”, EACL 2024 (System Demonstrations) | Nguồn gốc của faithfulness / answer relevance / context precision-recall |
| Saad-Falcon, Khattab, Potts, Zaharia — “ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems”, NAACL 2024 | Cách train judge nhẹ thay vì gọi LLM lớn |
| Rashkin et al. — “Measuring Attribution in Natural Language Generation Models”, Computational Linguistics 2023 | Khung AIS — định nghĩa chặt chẽ của “attributable to identified sources” |
| Gao, Yen, Yu, Chen — “Enabling Large Language Models to Generate Text with Citations” (ALCE), EMNLP 2023 | Đo citation precision/recall một cách hệ thống |
| Liu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni, Liang — “Lost in the Middle: How Language Models Use Long Contexts”, TACL 2024 | Bằng chứng thực nghiệm cho việc precision của context thật sự quan trọng, không chỉ recall |
| Zheng et al. — “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”, NeurIPS 2023 | Các bias của LLM-judge: độ dài, vị trí, tự thiên vị |
| TREC RAG Track (từ 2024) | Nỗ lực chuẩn hoá evaluation cho RAG theo phương pháp TREC |
Công cụ (để tham khảo cách người khác implement)
Phần tiêu đề “Công cụ (để tham khảo cách người khác implement)”| Tool | Ghi chú |
|---|---|
trec_eval | Implementation tham chiếu chính thức của TREC. Khi nghi ngờ công thức, so với cái này. |
pytrec_eval (Van Gysel & de Rijke, SIGIR 2018) | Binding Python của trec_eval |
ir_measures | API thống nhất cho nhiều metric, dễ dùng hơn trec_eval |
ranx | Có sẵn cả bootstrap significance test và fusion |
ragas | Library của paper RAGAS |
Link mở sẵn được (đã kiểm tra)
Phần tiêu đề “Link mở sẵn được (đã kiểm tra)”Phần trên viết từ trí nhớ nên chỉ có tên. Dưới đây là các nguồn có link truy cập được ngay — dùng cái này khi cần đọc chứ không chỉ cần trích dẫn.