让机器人在多人对话中判断何时何人该回应。
Whom to Respond To? A Transformer-Based Model for Multi-Party Social Robot Interaction
- 用Transformer多任务模型分析对话上下文,决定回应时机与对象。
- 新损失函数提升对说话人和目标对话的识别准确率。
- 适合研发社交机器人或人机交互系统的团队参考。
以往人机交互研究多聚焦单用户场景,机器人无需考虑回应时机与对象。但在商场、医院等多人场景中,社交机器人需理解上下文并决策何时以及向谁回应。本文提出基于Transformer的多任务学习框架,改进机器人在多用户环境中的决策能力。针对人机交互特性,设计两种新损失函数:一是约束活跃说话人以增强场景建模,二是引导响应选择聚焦于直接指向机器人的语句。同时构建了包含真实复杂性(如视线错位)的多用户人机交互数据集。实验表明,该模型在回应决策上达到当前最优性能,显著优于传统启发式与单任务方法。研究推动了具备自然、情境感知能力的社交机器人发展。
原文摘要 · Abstract (English)
Prior human-robot interaction (HRI) research has primarily focused on single-user interactions, where robots do not need to consider the timing or recipient of their responses. However, in multi-party interactions, such as at malls and hospitals, social robots must understand the context and decide both when and to whom they should respond. In this paper, we propose a Transformer-based multi-task learning framework to improve the decision-making process of social robots, particularly in multi-user environments. Considering the characteristics of HRI, we propose two novel loss functions: one that enforces constraints on active speakers to improve scene modeling, and another that guides response selection towards utterances specifically directed at the robot. Additionally, we construct a novel multi-party HRI dataset that captures real-world complexities, such as gaze misalignment. Experimental results demonstrate that our model achieves state-of-the-art performance in respond decisions, outperforming existing heuristic-based and single-task approaches. Our findings contribute to the development of socially intelligent social robots capable of engaging in natural and context-aware multi-party interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。