为真实企业会议设计多维度评估与双策略智能助手,兼顾速度与准确性。
MeetBench-XL: Calibrated Multi-Dimensional Evaluation and Learned Dual-Policy Agents for Real-Time Meetings
- 基于231场企业会议构建多模态数据集,覆盖金融、医疗等多领域真实场景。
- 提出多维评估体系,量化判断回答的准确性、完整性与响应效率。
- 开发双策略代理,自动选择快慢推理路径,提升任务执行效率和精度。
企业会议环境要求AI助手在低延迟、低成本和高隐私约束下完成多样操作任务,包括实时事实核查和跨会分析。现有基准主要聚焦简化问答,难以反映真实企业工作流中由多方协作自然产生的问题,且常需长时上下文和工具增强推理。为此,我们构建了基于231场企业会议(共140小时)的多语言多模态语料库MeetAll,通过领域专家评审和人类可区分性研究验证问题注入协议。该协议围绕认知负荷、时间跨度、领域专长和可执行任务四维度进行校准。进一步提出MeetBench XL,一个与人类判断对齐的多维度评估框架,涵盖事实准确性、意图一致性、响应效率、结构清晰度和完整性。最后,设计了MeetMaster XL,一种学习型双策略代理,联合优化查询路由至快速/慢速推理路径及工具调用(如检索、跨会聚合、网络搜索)。轻量级分类器实现精准路由,开销极小,优于单一模型基线。实验对比商业系统显示持续提升,经消融实验、鲁棒性测试及真实部署案例验证。
原文摘要 · Abstract (English)
Enterprise meeting environments require AI assistants that handle diverse operational tasks, from rapid fact checking during live discussions to cross meeting analysis for strategic planning, under strict latency, cost, and privacy constraints. Existing meeting benchmarks mainly focus on simplified question answering and fail to reflect real world enterprise workflows, where queries arise organically from multi stakeholder collaboration, span long temporal contexts, and require tool augmented reasoning. We address this gap through a grounded dataset and a learned agent framework. First, we introduce MeetAll, a bilingual and multimodal corpus derived from 231 enterprise meetings totaling 140 hours. Questions are injected using an enterprise informed protocol validated by domain expert review and human discriminability studies. Unlike purely synthetic benchmarks, this protocol is grounded in four enterprise critical dimensions: cognitive load, temporal context span, domain expertise, and actionable task execution, calibrated through interviews with stakeholders across finance, healthcare, and technology sectors. Second, we propose MeetBench XL, a multi dimensional evaluation protocol aligned with human judgment that measures factual fidelity, intent alignment, response efficiency, structural clarity, and completeness. Third, we present MeetMaster XL, a learned dual policy agent that jointly optimizes query routing between fast and slow reasoning paths and tool invocation, including retrieval, cross meeting aggregation, and web search. A lightweight classifier enables accurate routing with minimal overhead, achieving a superior quality latency tradeoff over single model baselines. Experiments against commercial systems show consistent gains, supported by ablations, robustness tests, and a real world deployment case study.Resources: https://github.com/huyuelin/MeetBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。