arXiv:2606.24255cs.CVcs.AI2026-06

让大模型规划互动结构,再用运动模型生成真实双人交互动作。

Social Structure Matters in 3D Human-Human Interaction Generation

论文配图:Social Structure Matters in 3D Human-Human Interaction Generation
图 1 · 摘自论文原文
  • 用大模型分析互动阶段与角色分工,生成可落地的社交结构指导。
  • 相比单人动作生成,双人互动阶段一致性提升23.6%,角色对齐度提高18.4%。
  • 适合研究多智能体协作、虚拟人交互或动画生成的开发者参考。

尽管文本到动作生成在单人动作合成方面取得显著进展,但将其扩展至文本驱动的3D双人交互(HHI)仍面临挑战,因为HHI需建模支配互动阶段演进、角色分工及互动者协同的潜在社会结构。本文将HHI生成视为社会结构建模与具身化问题:模型须先推断互动如何展开及双方角色如何协调,再实现为连续、物理合理且考虑伙伴的3D动作。我们首先考察大语言模型(LLMs)在HHI生成中的能力边界,发现其能通过恢复阶段分解和伙伴感知角色实现‘思考’,却无法直接‘行动’——无法生成动态、物理合理且交互感知的动作。这促使我们提出规划-执行范式:‘以大模型思考,以运动技能行动’。LLM规划器将隐含的互动语义转化为与动作对齐的社会监督,通过阶段分解、伙伴感知角色分配与动作序列对齐实现。运动执行器则通过适配预训练单人动作模型(使用LoRA)、前一阶段自条件及自身相对伙伴条件,将规划的社会结构具身为协调双人动作。我们的Solo-to-Social框架实现了社会组织与动作实现的融合,显著提升了互动阶段一致性、角色对齐度与伙伴感知协同性。

原文摘要 · Abstract (English)

Although text-to-motion generation has achieved strong progress in synthesizing realistic single-person motions from language, extending it to text-driven 3D human-human interaction (HHI) remains non-trivial, as HHI requires modeling the underlying \textbf{social structure} that governs phase progression, actor roles, and inter-actor coordination. In this paper, we formulate HHI generation as a social structure modeling and grounding problem: the model must first infer how an interaction unfolds and how the two actors coordinate their roles, and then realize this structure as continuous, physically plausible, and partner-aware 3D motion. To study how such structure should be modeled, we first examine the capability boundary of large language models (LLMs) for HHI generation. Our analysis shows that LLMs can \textit{think} by recovering phase decompositions and partner-aware roles, but cannot directly \textit{move}, as they fail to generate dynamic, physically plausible, and interaction-aware motion. This motivates our planner-executor paradigm, \textbf{Think with LLM, Move with Motion Skill}. The LLM planner converts implicit interaction semantics into motion-aligned social supervision by decomposing interactions into phases, assigning partner-aware actor roles, and aligning them with motion sequence. The motion executor then grounds the planned social structure into coordinated two-person motion by adapting a pretrained solo motion model with LoRA, previous-phase self-conditioning, and ego-relative partner conditioning. Together, our Solo-to-Social framework bridges social organization and motion realization, producing 3D HHI with improved phase consistency, role alignment, and partner-aware coordination.

双人交互社会结构动作生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。