arXiv:2601.05564cs.SDcs.CL2026-01被引 7

ICASSP 2026 挑战赛评测大模型时代的类人对话系统能力

The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era

  • 分情感智能与全双工交互双赛道评估对话系统
  • 基于真实人类对话数据集,测试长期情绪理解与实时响应
  • 适合关注对话系统智能化、人机交互的研究者

随着大语言模型(LLMs),特别是音频大模型和全能模型的快速发展,语音对话系统已显著进步,逐步缩小人机与人际交流之间的差距。实现真正‘类人’沟通需要双重能力:感知并共鸣用户情绪的情感智能,以及应对对话动态自然流的稳健交互机制,如实时轮换发言。因此,我们在 ICASSP 2026 发起首届类人语音对话系统挑战赛(HumDial),用于评测这两项核心能力。该挑战基于大规模真实人类对话数据集,设立两个赛道:(1)情感智能,聚焦长期情绪理解与共情生成;(2)全双工交互,系统性评估在‘边听边说’条件下的实时决策能力。本文总结了数据集构建、赛道设置及最终结果。

原文摘要 · Abstract (English)

Driven by the rapid advancement of Large Language Models (LLMs), particularly Audio-LLMs and Omni-models, spoken dialogue systems have evolved significantly, progressively narrowing the gap between human-machine and human-human interactions. Achieving truly ``human-like'' communication necessitates a dual capability: emotional intelligence to perceive and resonate with users' emotional states, and robust interaction mechanisms to navigate the dynamic, natural flow of conversation, such as real-time turn-taking. Therefore, we launched the first Human-like Spoken Dialogue Systems Challenge (HumDial) at ICASSP 2026 to benchmark these dual capabilities. Anchored by a sizable dataset derived from authentic human conversations, this initiative establishes a fair evaluation platform across two tracks: (1) Emotional Intelligence, targeting long-term emotion understanding and empathetic generation; and (2) Full-Duplex Interaction, systematically evaluating real-time decision-making under `` listening-while-speaking'' conditions. This paper summarizes the dataset, track configurations, and the final results.

对话系统情感智能全双工

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。