arXiv:2604.17260cs.CL2026-04ACL

用细粒度时间分析取代单一评分,提升会议效果评估的精度与实用性。

Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation

论文配图:Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation
图 1 · 摘自论文原文
  • 按话题片段逐段评估会议目标达成率,实现时间维度上的精细分析。
  • 构建包含2459段标注的AMI-ME数据集,覆盖130场真实会议。
  • 基于大模型裁判机制,支持从原始语音到自动评分的端到端评估。

会议效果评估对提升组织效率至关重要。现有方法依赖事后调查,仅给出单一时段的粗粒度评分,存在可扩展性差、成本高、结果难复现等问题,且无法捕捉协作讨论的动态特性。本文提出一种新范式:以目标达成速率为核心指标,对会议中各话题片段进行细粒度时间评估。为此,我们构建了AMI Meeting Effectiveness(AMI-ME)数据集,包含来自130场AMI语料库会议的2,459个经人工标注的片段。同时,开发了一个自动评估框架,利用大语言模型(LLM)作为裁判,对每个片段相对于整体会议目标的有效性进行打分。通过大量实验,建立了该任务的全面基准,验证了框架在商业场景至非结构化讨论等不同会议类型中的泛化能力。此外,从原始语音开始评测端到端系统性能。结果表明该框架有效,并为未来会议分析与多方对话研究提供强有力基线。数据集与代码将公开可用。

原文摘要 · Abstract (English)

Evaluating meeting effectiveness is crucial for improving organizational productivity. Current approaches rely on post-hoc surveys that yield a single coarse-grained score for an entire meeting. The reliance on manual assessment is inherently limited in scalability, cost, and reproducibility. Moreover, a single score fails to capture the dynamic nature of collaborative discussions. We propose a new paradigm for evaluating meeting effectiveness centered on novel criteria and temporal fine-grained approach. We define effectiveness as the rate of objective achievement over time and assess it for individual topical segments within a meeting. To support this task, we introduce the AMI Meeting Effectiveness (AMI-ME) dataset, a new meta-evaluation dataset containing 2,459 human-annotated segments from 130 AMI Corpus meetings. We also develop an automatic effectiveness evaluation framework that uses a Large Language Model (LLM) as a judge to score each segment's effectiveness relative to the overall meeting objectives. Through substantial experiments, we establish a comprehensive benchmark for this new task and evaluate the framework's generalizability across distinct meeting types, ranging from business scenarios to unstructured discussions. Furthermore, we benchmark end-to-end performance starting from raw speech to measure the capabilities of a complete system. Our results validate the framework's effectiveness and provide strong baselines to facilitate future research in meeting analysis and multi-party dialogue. Our dataset and code will be publicly available. The AMI-ME dataset and the Automatic Evaluation Framework are available at: this URL.

会议评估大模型应用细粒度分析多说话人对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。