arXiv:2605.05711cs.CVcs.GR2026-05

让3D场景生成与用户互动闭环联动,提升沉浸感与交互体验

Closing the Loop: Unified 3D Scene Generation and Immersive Interaction via LLM-RL Coupling

论文配图:Closing the Loop: Unified 3D Scene Generation and Immersive Interaction via LLM-RL Coupling
图 1 · 摘自论文原文
  • 用大模型生成结构化场景,再通过强化学习优化布局
  • 在虚拟现实中实现人机交互反馈,使场景更贴合真实需求
  • 适合做智能交互系统、元宇宙内容生成的研究者

大语言模型显著提升了语言驱动的3D内容生成能力,但现有方法大多将场景生成与用户交互分开处理,限制了系统的适应性与沉浸感。本文提出一个统一框架,实现语言驱动3D场景生成与沉浸式用户交互的闭环整合。给定自然语言指令,系统首先利用大模型构建结构化场景表示,再在几何与语义约束下通过强化学习优化空间布局。生成环境部署于虚拟现实场景中,支持人机交互闭环,用户行为提供持续反馈以对齐人类感知与可用性。通过紧密耦合生成与交互,该框架实现了更响应迅速、自适应且真实的多媒体体验。在ALFRED基准上的实验表明,其任务导向的场景生成达到当前最优性能。定性结果与用户研究进一步验证了沉浸感、交互质量与任务效率的显著提升,凸显了生成与交互闭环集成对未来多媒体系统的重要性。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have significantly improved language-driven 3D content generation, but most existing approaches still treat scene generation and user interaction as separate processes, limiting the adaptability and immersive potential of interactive multimedia systems. This paper presents a unified framework that closes the loop between language-driven 3D scene generation and immersive user interaction. Given natural language instructions, the system first constructs structured scene representations using LLMs, and then optimizes spatial layouts via reinforcement learning under geometric and semantic constraints. The generated environments are deployed in a virtual reality setting to facilitate HRI-in-the-loop, where user interactions provide continuous feedback to align generated content with human perception and usability. By tightly coupling generation and interaction, the proposed framework enables more responsive, adaptive, and realistic multimedia experiences. Experiments on the ALFRED benchmark demonstrate state-of-the-art performance in task-based scene generation. Furthermore, qualitative results and user studies show consistent improvements in immersion, interaction quality, and task efficiency, highlighting the importance of closed-loop integration of generation and interaction for next-generation multimedia systems. Our project page can be found at https://proj-showcase.github.io/h3ds/.

3D生成人机交互强化学习虚拟现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。