用检索增强生成脚本,让虚拟人物动作更自然多样。
SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation
- 用大模型+检索增强生成连贯多样的交互脚本
- 基于物理的多条件控制器实现风格化动作生成
- 适合做虚拟角色动画、游戏或影视特效的人看
在真实环境中模拟风格化人-场景交互(HSI)是一项挑战性任务。以往方法侧重长期执行,但难以同时实现风格多样性与物理合理性。为此,我们提出一种分层框架SIMS,将高层脚本驱动意图与低层控制策略无缝衔接,提升交互表现力与多样性。具体地,利用具备检索增强生成(RAG)的大语言模型生成连贯且多样化的长篇脚本,为运动规划提供丰富基础;同时设计一个多功能物理驱动控制策略,通过脚本中的文本嵌入编码风格线索,同时感知环境几何结构并完成任务目标。通过结合检索增强脚本生成与多条件控制器,该方法统一解决了风格化HSI运动生成问题。我们还构建了一个由RAG生成的综合规划数据集和包含多样化行走与交互动作的风格化运动数据集。大量实验表明,SIMS在多种任务执行与跨场景泛化方面显著优于现有方法。
原文摘要 · Abstract (English)
Simulating stylized human-scene interactions (HSI) in physical environments is a challenging yet fascinating task. Prior works emphasize long-term execution but fall short in achieving both diverse style and physical plausibility. To tackle this challenge, we introduce a novel hierarchical framework named SIMS that seamlessly bridges highlevel script-driven intent with a low-level control policy, enabling more expressive and diverse human-scene interactions. Specifically, we employ Large Language Models with Retrieval-Augmented Generation (RAG) to generate coherent and diverse long-form scripts, providing a rich foundation for motion planning. A versatile multicondition physics-based control policy is also developed, which leverages text embeddings from the generated scripts to encode stylistic cues, simultaneously perceiving environmental geometries and accomplishing task goals. By integrating the retrieval-augmented script generation with the multi-condition controller, our approach provides a unified solution for generating stylized HSI motions. We further introduce a comprehensive planning dataset produced by RAG and a stylized motion dataset featuring diverse locomotions and interactions. Extensive experiments demonstrate SIMS's effectiveness in executing various tasks and generalizing across different scenarios, significantly outperforming previous methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。