让文字精准控制全景图中物体位置,解决方向描述错位问题。
Bridging Language and Spherical Space: Object-Centric Control for Text-to-Panorama Generation

- 将文本中的方向描述转为球面空间的物体级结构条件。
- 在扩散模型中引入物体感知注意力与空间残差增强,提升定位精度。
- 适用于虚拟现实、3D内容生成等需精确空间控制的场景。
全景图像生成在虚拟现实、增强现实和3D内容创作等沉浸式应用中日益重要。与透视图像不同,全景图以观察者为中心呈现360°环绕空间,方向性表达如左右前后在空间理解中起核心作用。然而,现有文本到全景图方法多依赖隐式空间推理,难以准确对齐物体级方向描述与球面全景场景。直接引入显式布局又需人工指定空间条件,降低语言交互灵活性,且无法根本解决以自我为中心的方向语言与全景空间的错位问题。为此,我们提出PanoCtrl,一种面向可控文本到全景图生成的物体中心框架。该方法通过将文本描述转化为结构化的物体级球面条件,并将其融入扩散过程,实现自然语言与球面空间的显式桥接。具体而言,我们设计了文本条件解析器PanoParse,用于预测物体语义及球面视场(BFoV)参数;以及PanoControl模块,通过物体感知注意力与空间残差增强,向扩散变换器注入物体级语义与空间引导。为支持该任务,我们构建了包含物体级球面标注和多样化方向描述的PanoGround数据集。大量实验表明,PanoCtrl在空间对齐性和图像质量上均达到当前最优水平。
原文摘要 · Abstract (English)
Panoramic image generation is increasingly important for immersive applications such as virtual reality, augmented reality, and 3D content creation. Unlike perspective images, panoramic images represent a viewer-centered $360^\circ$ surrounding space, where directional expressions such as left, right, front, and behind play a central role in spatial understanding. However, existing text-to-panorama methods largely rely on implicit spatial reasoning and often fail to faithfully ground object-level directional descriptions in spherical panoramic scenes. A straightforward alternative is to introduce explicit layouts, but requiring manually specified spatial conditions reduces the flexibility of language-based interaction and does not directly resolve the misalignment between egocentric directional language and panoramic image space. To address this issue, we propose PanoCtrl, an object-centric framework for controllable text-to-panorama generation. Our method explicitly bridges natural language and spherical panoramic space by converting textual descriptions into structured object-level spherical conditions and integrating them into the diffusion process. Specifically, we introduce PanoParse, a text-conditioned parser that predicts object semantics and spherical bounding field-of-view (BFoV) parameters, and \textbf{PanoControl}, which injects object-level semantic and spatial guidance into the diffusion transformer through object-aware attention and spatial residual enhancement. To support this task, we construct PanoGround, a dataset with object-level spherical annotations and diverse directional descriptions for controllable panoramic generation. Extensive experiments demonstrate that PanoCtrl achieves state-of-the-art performance in both spatial alignment and image quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。