让3D虚拟角色具备社交智能,实现自然互动。
SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters
- 构建统一的多模态交互框架,根据用户输入生成语音与动作。
- 用合成数据解决社交互动数据稀缺问题,提升响应自然度。
- 支持沉浸式VR交互,适合游戏、元宇宙等场景应用。
人类是社会性动物。如何赋予3D自主角色类似的社会智能,使其能够感知、理解并与人类互动,仍是开放且根本性的问题。本文提出SOLAMI,首个面向沉浸式交互的3D自主角色社会性视觉-语言-动作(VLA)建模框架。SOLAMI从三方面构建3D角色:(1)社会VLA架构:提出统一的社会化多模态框架,基于用户多模态输入生成语音与动作,驱动角色进行社会互动;(2)交互式多模态数据:提出SynMSI,一个通过自动化流水线仅利用现有动作数据集生成的合成多模态社交互动数据集,缓解数据稀缺问题;(3)沉浸式VR界面:开发支持用户沉浸式交互的VR界面,可与不同架构驱动的角色互动。大量定量实验与用户研究验证,该框架能生成更精确、自然的语音与动作响应,符合用户预期,且延迟更低。
原文摘要 · Abstract (English)
Human beings are social animals. How to equip 3D autonomous characters with similar social intelligence that can perceive, understand and interact with humans remains an open yet foundamental problem. In this paper, we introduce SOLAMI, the first end-to-end Social vision-Language-Action (VLA) Modeling framework for Immersive interaction with 3D autonomous characters. Specifically, SOLAMI builds 3D autonomous characters from three aspects: (1) Social VLA Architecture: We propose a unified social VLA framework to generate multimodal response (speech and motion) based on the user's multimodal input to drive the character for social interaction. (2) Interactive Multimodal Data: We present SynMSI, a synthetic multimodal social interaction dataset generated by an automatic pipeline using only existing motion datasets to address the issue of data scarcity. (3) Immersive VR Interface: We develop a VR interface that enables users to immersively interact with these characters driven by various architectures. Extensive quantitative experiments and user studies demonstrate that our framework leads to more precise and natural character responses (in both speech and motion) that align with user expectations with lower latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。