用声音和肢体特征让机器人判断何时该换话题
Let's move on: Topic Change in Robot-Facilitated Group Discussions
- 用语音和动作特征训练模型预测合适的话题切换时机
- 模型在识别不当换题上表现优于传统规则方法
- 声音特征效果接近全模态,适合轻量部署
机器人主持的群体讨论有望促进人类参与者的高效互动。以往对话代理中的话题管理多关注用户参与度与个性化,且代理通常主动引导话题。尽管已有研究证实机器人在群体中具有价值,但如何让机器人学会适时换题仍需探索。为此,我们研究了机器学习模型与音视频非语言特征在预测恰当话题切换中的适用性。基于机器人主持人与人类参与者的真实互动数据,我们进行了标注,并提取了声学与体态语言特征。通过分析不同特征集下序列与非序列数据的机器学习性能,结果表明模型在分类不当话题切换方面表现优异,超越规则基方法。此外,声学特征的表现与完整多模态特征相当,具备良好鲁棒性。我们公开了标注数据集,地址为 https://github.com/ghadj/topic-change-robot-discussions-data-2024。
原文摘要 · Abstract (English)
Robot-moderated group discussions have the potential to facilitate engaging and productive interactions among human participants. Previous work on topic management in conversational agents has predominantly focused on human engagement and topic personalization, with the agent having an active role in the discussion. Also, studies have shown the usefulness of including robots in groups, yet further exploration is still needed for robots to learn when to change the topic while facilitating discussions. Accordingly, our work investigates the suitability of machine-learning models and audiovisual non-verbal features in predicting appropriate topic changes. We utilized interactions between a robot moderator and human participants, which we annotated and used for extracting acoustic and body language-related features. We provide a detailed analysis of the performance of machine learning approaches using sequential and non-sequential data with different sets of features. The results indicate promising performance in classifying inappropriate topic changes, outperforming rule-based approaches. Additionally, acoustic features exhibited comparable performance and robustness compared to the complete set of multimodal features. Our annotated data is publicly available at https://github.com/ghadj/topic-change-robot-discussions-data-2024.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。