arXiv:2506.00314cs.IR2025-06中稿 · SIGIR 2026被引 2

提出细粒度对话评估方法,让系统评价更准更可解释。

FACE: A Fine-Grained Reference-Free Evaluator for Conversational Information Access

  • 用强化学习优化提示词,分项评估对话细节
  • 与人工评分相关性达0.9,显著优于现有方法
  • 输出可解释得分,帮助定位系统问题

对话式信息获取(CIA)系统的系统化、可靠且低成本评估仍是开放挑战。现有基于参考答案的方法难以应对信息获取对话的动态特性,而现有的大模型参考自由方法存在评估偏差且泛化能力有限。本文提出FACE:一种细粒度、基于维度的对话评估方法,可对对话中不同轮次和整体层面进行多维度评分。FACE利用束搜索与博弈优化技术,为每个评估维度生成最优的大模型指令;通过选定指令对原子信息单元(particles)打分,并聚合为最终得分。实验表明,FACE与人工评分具有强相关性(系统相关性达0.9),显著超越当前最优对话评估方法。进一步验证其优化后的指令可在多种大模型和数据集间迁移。此外,不同于现有方法仅提供不可解释的单一分数,FACE能揭示系统表现细节,支持精准定位对话中的缺陷。

原文摘要 · Abstract (English)

A systematic, reliable, and low-cost evaluation of Conversational Information Access (CIA) systems remains an open challenge. Existing reference-based evaluation methods are proven insufficient for evaluating the dynamic nature of information access conversations, while existing LLM-based reference-free methods suffer from evaluation bias and limited generalizability. This work proposes FACE: a Fine-grained, Aspect-based Conversation Evaluation method that provides evaluation scores for diverse turn and dialogue-level aspects of conversations. FACE leverages beam search and bandit optimization to select optimized LLM instructions per evaluation aspect. It assigns scores to atomic information units (particles) using the selected instructions and then aggregates them into a single score. We show that FACE achieves a strong correlation with human judgments, achieving system correlation of 0.9, outperforming state-of-the-art conversation evaluation methods by a large margin. We further demonstrate its optimized instructions are transferable across various LLMs and datasets. Additionally, unlike existing LLM-based methods that provide single uninterpretable scores, FACE provides insights into the system performance and enables identifying and locating problems within conversations.

对话评估大模型评测可解释性信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。