大模型在会议对话中预测发言者和轮次转换,表现超越人类。
Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings

- 构建多任务评估框架,比较文本与多模态大模型性能。
- 大模型在发言者预测上优于人类,但视觉音频信息利用仍不足。
- 对话上下文对预测至关重要,尤其在频繁换人时更难判断。
我们利用大语言模型(LLMs)研究多模态多人对话中的轮次转换问题。构建了针对三个任务的评估框架:发言对象识别、轮次切换预测和下一说话者预测。对比了监督模型、基于文本的LLM、多模态大模型(MM-LLMs)以及人类参与者。在AMI语料库上的实验表明,尽管未在目标领域训练且无音频/视觉信息,大模型在下一说话者预测上仍优于监督模型和人类。多模态大模型在发言对象识别和轮次切换预测上优于纯文本大模型,但仍未达到人类水平,表明其难以有效利用原始音视频信号。消融分析显示,对话上下文对预测极为关键,尤其是下一说话者预测。观察发现,人类与大模型的预测模式相似,频繁轮换区间对两者均构成挑战。
原文摘要 · Abstract (English)
We investigate turn-taking in multimodal multi-party conversations using large language models (LLMs). We construct an evaluation framework for three tasks: addressee detection, turn-change prediction, and next speaker prediction. We compare supervised models trained for these tasks, text-based LLMs, multimodal LLMs (MM-LLMs), and human subjects. Experiments on the AMI corpus showed that LLMs outperformed supervised models and humans in next speaker prediction, despite not being trained on the target domain and without access to audio or visual information. An MM-LLM performed better than text-based LLMs on addressee detection and turn-change prediction but remained below human performance, indicating difficulty leveraging raw audio-visual signals. Ablation analyses revealed that conversational context was critical, particularly for next speaker prediction. We observed that human and LLM prediction patterns were similar, and intervals with frequent turn changes were difficult for both.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。