arXiv:2507.14412cs.RO2025-07被引 3

用端到端语音语言模型提升机器人情感支持的自然互动能力

Personalized Socially Assistive Robots With End-to-End Speech-Language Models For Well-Being Support

  • 采用端到端语音语言模型实现机器人对话系统
  • 用户测试显示能自然回应并共情,但反馈较重复
  • 适合心理支持、人机交互研究者关注

社交陪伴机器人(SARs)在补充心理健康支持方面展现出巨大潜力。然而,现有对话系统在实时延迟、回应性提示和个性化语音对话方面仍存局限。为此,本文提出使用集成的端到端语音-语言模型(SLMs)与SAR结合。研究通过一项小规模被试内用户实验(N=11,大学学生)评估了SLM驱动的对话系统可用性,并基于用户反馈识别出关键改进方向。结果显示,参与者认为该系统能提供共情回应、自然换言、回应性提示及自适应响应。但同时指出,机器人的非语言行为缺乏变化且与对话不同步,语音输出则显得通用且重复。研究强调需实现实时动作同步、优化提示或微调以更契合心理健康实践,以及提升语音表达力与适应性。

原文摘要 · Abstract (English)

Socially assistive robots (SARs) have shown great potential for supplementing well-being support. However, prior studies have found that existing dialogue pipelines for SARs remain limited in real-time latency, back-channeling, and personalized speech dialogue. Toward addressing these limitations, we propose using integrated end-to-end speech-language models (SLMs) with SARs. This work 1) evaluated the usability of an SLM-enabled SAR dialogue system through a small user study, and 2) identified remaining limitations through study user feedback to inform future improvements. We conducted a small within-participant user study with university students (N = 11) whose results showed that participants perceived an SLM-enabled SAR system as capable of providing empathetic feedback, natural turn-taking, back-channeling, and adaptive responses. We also found that participants reported the robot's nonverbal behaviors as lacking variability and synchronization with conversation, and the SLM's verbal feedback as generic and repetitive. These findings highlighted the need for real-time robot movement synchronized with conversation, improved prompting or fine-tuning to generate outputs better aligned with mental health practices, and more expressive, adaptive vocal generation.

社交机器人语音生成情感支持人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。