arXiv:2512.02569cs.HCcs.RO2025-12被引 2

用虚拟机器人+大模型打造更安全、智能、有共情的人机交互。

Reframing Human-Robot Interaction Through Extended Reality: Unlocking Safer, Smarter, and More Empathic Interactions with Virtual Robots and Foundation Models

  • 通过扩展现实技术让虚拟机器人突破硬件限制,实现灵活部署。
  • 大模型使虚拟机器人能理解情境、感知情绪并长期适应用户需求。
  • 适合关注人机共情、未来交互设计与伦理问题的研究者。

本文从扩展现实(XR)视角重构人-机器人交互(HRI),提出由大型基础模型(FMs)驱动的虚拟机器人可作为具认知基础且富有同理心的交互代理。相较于实体机器人,XR原生代理不受硬件约束,可按需实例化、适配与扩展,同时保持身体存在感与共在感。我们整合了XR、HRI与认知人工智能的研究成果,表明此类代理能在高风险场景中提升安全性,在跨领域实现社会与认知共情,并借助XR与AI融合拓展物理能力边界。多模态大基础模型(如大语言模型、视觉模型、视觉-语言模型)使虚拟机器人具备情境感知推理、情绪敏感响应及长期适应能力,使其成为认知与共情中介而非单纯仿真工具。同时,文章指出过度信任、文化与表征偏差、生物特征传感隐私及数据治理透明性等挑战。最后,提出以用户为中心、伦理为基的科研议程,强调多层次评估框架、多用户生态、虚实混合具身化及社会伦理设计实践,展望由大模型赋能的虚拟代理重塑未来人机交互为更高效自适应范式。

原文摘要 · Abstract (English)

This perspective reframes human-robot interaction (HRI) through extended reality (XR), arguing that virtual robots powered by large foundation models (FMs) can serve as cognitively grounded, empathic agents. Unlike physical robots, XR-native agents are unbound by hardware constraints and can be instantiated, adapted, and scaled on demand, while still affording embodiment and co-presence. We synthesize work across XR, HRI, and cognitive AI to show how such agents can support safety-critical scenarios, socially and cognitively empathic interaction across domains, and outreaching physical capabilities with XR and AI integration. We then discuss how multimodal large FMs (e.g., large language model, large vision model, and vision-language model) enable context-aware reasoning, affect-sensitive situations, and long-term adaptation, positioning virtual robots as cognitive and empathic mediators rather than mere simulation assets. At the same time, we highlight challenges and potential risks, including overtrust, cultural and representational bias, privacy concerns around biometric sensing, and data governance and transparency. The paper concludes by outlining a research agenda for human-centered, ethically grounded XR agents - emphasizing multi-layered evaluation frameworks, multi-user ecosystems, mixed virtual-physical embodiment, and societal and ethical design practices to envision XR-based virtual agents powered by FMs as reshaping future HRI into a more efficient and adaptive paradigm.

人机交互虚拟机器人大模型扩展现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。