用搜索策略主动发现会议助手的语义失效点
Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants
- 把评估当作自适应搜索,聚焦容易出错的对话结构
- 比随机测试多发现2.5倍失效,失败主要来自语用理解而非事实错误
- 适合评估大模型在真实会议场景中的真实表现
LLM驱动的会议助手已大规模部署,但对其语义准确性的系统性评估仍局限于静态基准,难以捕捉与特定话语结构或推理需求相关的失效模式。本文提出评价即搜索(EaS)方法,将质量评估建模为对会议参与者可能提出的问题空间的自适应搜索。通过迭代反馈学习,EaS集中探测最可能出现失效的认知需求,利用基于UCB的覆盖率图和盲态多维质量评估实现优化。基于此,我们构建了MeetingProbe基准,包含超过3000个标注的问题-答案对,覆盖20段来自三种会议类型的转录文本及三类LLM助手。消融实验表明,自适应搜索发现的失效率(7.1%)是随机探测(2.9%)的2.5倍,其中策略规划模块贡献最大。在三类模型中观察到清晰的能力梯度,并识别出八类常见失效,以话语-语用挑战为主,而非事实记忆错误。进一步验证显示,该基准在多个模型家族和提供商间具有可复现能力,揭示了所有模型均无法处理的一组通用失效。MeetingProbe已公开发布,支持会议助手语义准确性的可复现评估。
原文摘要 · Abstract (English)
LLM-powered meeting assistants are deployed at scale, yet systematic evaluation of their grounding fidelity remains limited to static benchmarks that miss failure modes tied to specific discourse structures or reasoning demands. We propose Evaluation-as-Search (EaS), a feedback-driven methodology that frames quality evaluation as an adaptive search over the space of natural questions a meeting participant might ask. Rather than sampling uniformly, EaS learns from evaluator feedback across iterations to concentrate probing effort on cognitive demands where failures are most likely, guided by a UCB-scored coverage map and blind multi-dimensional quality evaluation. Using EaS, we construct MeetingProbe, a benchmark of over $3{,}000$ annotated question--answer pairs spanning 20 transcripts from three meeting genres and three LLM assistants. In ablations, adaptive search surfaces $2.5\times$ more failures than random probing ($7.1\%$ vs. $2.9\%$ finding rate), with the strategic planner contributing the largest individual effect. Across three models, we observe a clear capability gradient and identify eight recurring failure categories dominated by discourse-pragmatic challenges rather than factual recall errors. We further validate MeetingProbe across multiple model families and providers, finding a clean capability gradient and a curated subset of universal failures that no model handles. MeetingProbe is released publicly to support reproducible evaluation of meeting assistant grounding fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。