arXiv:2410.16770cs.CVcs.AI2024-10CVPR被引 39

用程序+文字+嵌入表示场景,实现高保真可控生成

The Scene Language: Representing Scenes with Programs, Words, and Embeddings

  • 用程序定义场景层级关系,文字描述语义,嵌入保留视觉身份
  • 无需训练即可从图文输入生成复杂场景,支持3D/4D图像渲染
  • 适合需要精确编辑与高质量生成的视觉创作场景

我们提出场景语言(Scene Language),一种简洁精准的视觉场景表征方式,包含三个核心组件:描述实体间层次与关系的程序、总结每个实体语义类别的自然语言词汇,以及捕捉每个实体视觉身份的嵌入。该表征可通过无需训练的推理技术,从预训练语言模型中获取,输入为文本或图像。生成的场景可使用传统、神经或混合图形渲染器还原为图像。整体构成一个鲁棒、自动的高质量3D与4D场景生成系统。相比现有场景图等表征,场景语言在生成复杂场景时保真度更高,并显式建模结构,支持精确控制与编辑。

原文摘要 · Abstract (English)

We introduce the Scene Language, a visual scene representation that concisely and precisely describes the structure, semantics, and identity of visual scenes. It represents a scene with three key components: a program that specifies the hierarchical and relational structure of entities in the scene, words in natural language that summarize the semantic class of each entity, and embeddings that capture the visual identity of each entity. This representation can be inferred from pre-trained language models via a training-free inference technique, given text or image inputs. The resulting scene can be rendered into images using traditional, neural, or hybrid graphics renderers. Together, this forms a robust, automated system for high-quality 3D and 4D scene generation. Compared with existing representations like scene graphs, our proposed Scene Language generates complex scenes with higher fidelity, while explicitly modeling the scene structures to enable precise control and editing.

场景生成视觉表征神经渲染程序生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。