arXiv:2603.12238cs.CV2026-03被引 3

用视觉反馈让AI理解自然语言生成任意3D场景

SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation

论文配图:SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation
图 1 · 摘自论文原文
  • 通过视觉反馈与视觉语言模型协作,实现开放式词汇3D场景生成
  • 支持自然语言指令编辑已有场景,生成质量优于现有方法
  • 适合数字内容创作、交互式设计等需要灵活场景生成的场景

从自然语言生成3D场景在数字内容创作中极具吸引力。然而,现有方法大多受限于特定领域或依赖预定义的空间关系,难以实现无约束的开放词汇3D场景合成。本文提出SceneAssistant,一种基于视觉反馈驱动的智能体,用于开放词汇3D场景生成。框架结合现代3D物体生成模型与视觉语言模型(VLM)的空间推理和规划能力。为支持开放词汇场景组合,我们向VLM提供一组完整的原子操作(如Scale、Rotate、FocusOn)。在每一步交互中,VLM接收渲染的视觉反馈并据此采取行动,迭代优化场景布局,实现更连贯的空间结构与更强的文本对齐。实验结果表明,该方法可生成多样化、开放词汇且高质量的3D场景。定性分析与定量的人类评估均证明其优于现有方法。此外,该方法允许用户通过自然语言指令编辑已有场景。代码已公开于https://github.com/ROUJINN/SceneAssistant。

原文摘要 · Abstract (English)

Text-to-3D scene generation from natural language is highly desirable for digital content creation. However, existing methods are largely domain-restricted or reliant on predefined spatial relationships, limiting their capacity for unconstrained, open-vocabulary 3D scene synthesis. In this paper, we introduce SceneAssistant, a visual-feedback-driven agent designed for open-vocabulary 3D scene generation. Our framework leverages modern 3D object generation model along with the spatial reasoning and planning capabilities of Vision-Language Models (VLMs). To enable open-vocabulary scene composition, we provide the VLMs with a comprehensive set of atomic operations (e.g., Scale, Rotate, FocusOn). At each interaction step, the VLM receives rendered visual feedback and takes actions accordingly, iteratively refining the scene to achieve more coherent spatial arrangements and better alignment with the input text. Experimental results demonstrate that our method can generate diverse, open-vocabulary, and high-quality 3D scenes. Both qualitative analysis and quantitative human evaluations demonstrate the superiority of our approach over existing methods. Furthermore, our method allows users to instruct the agent to edit existing scenes based on natural language commands. Our code is available at https://github.com/ROUJINN/SceneAssistant

3D生成视觉语言模型自然语言控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。