arXiv:2608.27988cs.CLcs.SD2026-08

通过眼神与声音分析,预测多人对话中谁将接话

Predicting Turn-Taking Outcomes in Multi-Party Conversation: Interpretable Modelling of Speech and Gaze Dynamics with Interpersonal Closeness

论文配图:Predicting Turn-Taking Outcomes in Multi-Party Conversation: Interpretable Modelling of Speech and Gaze Dynamics with Interpersonal Closeness
图 1 · 摘自论文原文
  • 用眼神模式和音量等行为特征建模对话接棒
  • 结合眼神与音量可使预测准确率达76%
  • 适合研究人机对话交互或社交行为的学者

流畅的发言切换是有效对话的基础,依赖于对话者对何时介入对话的准确判断。这一能力取决于对言语与非言语线索的解读与表达,这些线索提示说话人是否希望接话或让出发言权。在嘈杂、自然的四人多轮对话中,该过程更为复杂。本研究基于GaMMA语料库,利用对话前提取的可解释行为特征(包括眼神转移模式、行为对比、熵值、注视对象识别、相互凝视、音量等)训练逻辑回归模型,分类发言权交接为沉默间隙或重叠。结果表明,眼神特征具有预测性,结合音量后性能提升(ROC AUC = 0.76 ± 0.04)。音量反映说话主导权,眼神分散与指向则对应听众准备度与竞争性介入。模型在不同噪声条件下表现稳定,说明眼神是语音之外的抗干扰补充线索。

原文摘要 · Abstract (English)

Smooth speaker transitions are fundamental to effective conversation and rely on an interlocutor's ability to predict when to enter the conversation. This ability depends on accurately interpreting and expressing the verbal and non-verbal cues that signal when a speaker wishes to take or relinquish the floor. The process becomes even more complex in noisy, natural, multi-party settings, with multiple interlocutors available. This study models how gaze and speech, together with perceived interpersonal closeness, signal conversational floor changes in free four-person dialogue. Using the GaMMA corpus, we trained logistic regression models using interpretable, behaviourally motivated features extracted before each turn-taking event to classify floor-transfer outcomes as gaps or overlaps. Predictors included gaze features such as transition motifs and behavioural contrasts, entropy, gaze-based addressee identity, and mutual gaze, alongside speech features derived from speaker loudness, as well as perceived interpersonal closeness (IOS) between speakers. Results show that gaze features capture predictive structure, and that combining them with loudness improves performance (ROC AUC = 0.76 +- 0.04). Loudness reflected speaker control, while gaze dispersion and addressing indexed listener readiness and competitive entry. Performance remained robust across noise conditions, indicating that gaze provides a complementary, noise-resilient cue to turn-taking dynamics.

对话分析眼神追踪语音交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。