为多人对话生成设计了可量化的评估基准,解决评价标准模糊问题。
MPCEval: A Benchmark for Multi-Party Conversation Generation
- 分三维度评估:说话人建模、内容质量、说话人-内容一致性
- 提出无参考、可复现的指标,区分局部续写与整体对话生成
- 揭示模型在参与平衡、内容进展等方面的系统性差异
多人对话生成(如智能回复、协作助手)是生成式AI的重要能力,但其评估仍是关键瓶颈。相比双人对话,多人场景面临复杂的轮流发言、角色依赖的说话行为、长程对话结构及多个合理延续等挑战。为此,我们提出MPCEval,一个面向多人对话生成的任务感知评估与基准套件。该套件将生成质量分解为说话人建模、内容质量与说话人-内容一致性,并明确区分局部下一轮预测与全局全对话生成。它提供新颖的、量化的、无参考且可复现的指标,可跨数据集和模型扩展。我们在多种公开与真实世界数据集上应用MPCEval,评估现代生成方法与人工撰写对话。结果揭示了模型在参与平衡、内容推进与新颖性、说话人-内容一致性上的系统性特征,表明评估目标显著影响模型评估,而单一得分评估会掩盖多人对话行为的本质差异。MPCEval实现与相关评估代码已开源。
原文摘要 · Abstract (English)
Multi-party conversation generation, such as smart reply and collaborative assistants, is an increasingly important capability of generative AI, yet its evaluation remains a critical bottleneck. Compared to two-party dialogue, multi-party settings introduce distinct challenges, including complex turn-taking, role-dependent speaker behavior, long-range conversational structure, and multiple equally valid continuations. Accordingly, we introduce MPCEval, a task-aware evaluation and benchmarking suite for multi-party conversation generation. MPCEval decomposes generation quality into speaker modeling, content quality, and speaker--content consistency, and explicitly distinguishes local next-turn prediction from global full-conversation generation. It provides novel, quantitative, reference-free, and reproducible metrics that scale across datasets and models. We apply MPCEval to diverse public and real-world datasets and evaluate modern generation methods alongside human-authored conversations. The results reveal systematic, dimension-specific model characteristics in participation balance, content progression and novelty, and speaker--content consistency, demonstrating that evaluation objectives critically shape model assessment and that single-score evaluation obscures fundamental differences in multi-party conversational behavior. The implementation of MPCEval and the associated evaluation code are publicly available at https://github.com/Owen-Yang-18/MPCEval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。