用多智能体辩论提升检索评估准确率,仅需3.5%人力标注。
Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks
- 多智能体分正反方辩论,循环互评提升判断可靠性。
- 实现95.2%标注准确率,仅需3.5%人类介入。
- 发现29,824个遗漏相关段落,适用于检索与RAG评估改进。
信息检索(IR)评估因基准数据集存在未标注的相关片段而面临挑战。尽管大语言模型(LLM)及人机协作策略降低了人工成本,但仍易受模型自满和无效人机升级影响。为此,我们提出DREAM框架——基于多轮辩论的归因评估机制,由持对立立场的LLM智能体进行迭代互评。通过基于共识的辩论模式,该方法在特定场景下提升标注准确性,在不确定时生成更可靠的AI到人类升级请求。实验显示,仅需3.5%的人力参与即可实现95.2%的标注准确率。利用DREAM构建了BRIDGE基准数据集,揭示了29,824个此前缺失的相关片段,有效缓解评估偏差,支持更公平的检索器对比。重新评估后发现,未填补的盲区不仅扭曲检索器排名,还导致检索-生成不一致。DREAM框架与BRIDGE数据集均已开源。
原文摘要 · Abstract (English)
Information retrieval (IR) evaluation remains challenging due to incomplete IR benchmark datasets that contain unlabeled relevant chunks. While LLMs and LLM-human hybrid strategies reduce costly human effort, they remain prone to LLM overconfidence and ineffective AI-to-human escalation. To address this, we propose DREAM, a multi-round debate-based relevance assessment framework with LLM agents, built on opposing initial stances and iterative reciprocal critique. Through our agreement-based debate, it yields more accurate labeling for certain cases and more reliable AI-to-human escalation for uncertain ones, achieving 95.2% labeling accuracy with only 3.5% human involvement. Using DREAM, we build BRIDGE, a refined benchmark that mitigates evaluation bias and enables fairer retriever comparison by uncovering 29,824 missing relevant chunks. We then re-benchmark IR systems and extend evaluation to RAG, showing that unaddressed holes not only distort retriever rankings but also drive retrieval-generation misalignment. The relevance assessment framework is available at https: //github.com/DISL-Lab/DREAM-ICLR-26; and the BRIDGE dataset is available at https://github.com/DISL-Lab/BRIDGE-Benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。