解决遥感图文生成中的空间错位问题,让生成图像更准确反映文本描述的地理布局。
Uni-RS: A Spatially Faithful Unified Understanding and Generation Model for Remote Sensing
- 将文本指令转化为空间布局规划,分离几何设计与图像合成过程
- 通过空间感知查询监督,强化模型对空间关系的理解与生成
- 引入几何一致的空间变换数据增强,提升生成结果的空间保真度
统一的遥感多模态模型存在显著的空间反转缺陷:虽能准确识别和描述图像中物体的位置,但在文本到图像生成中却难以忠实再现这些空间关系,而空间关系在遥感领域是核心语义信息。为此,我们提出Uni-RS,首个专为遥感设计的统一多模态模型,旨在显式解决理解与生成之间的空间不对称性。具体而言,首先引入显式的空间布局规划,将文本指令转换为空间布局计划,实现几何规划与视觉合成的解耦;其次,采用空间感知查询监督,引导可学习查询显式关注指令中指定的空间关系;最后,设计图像-标题空间布局变化机制,使模型暴露于系统性的、几何一致的空间变换中。在多个基准上的大量实验表明,该方法显著提升了文本到图像生成的空间保真度,同时保持了图像描述、视觉定位和VQA等多模态理解任务的强性能。
原文摘要 · Abstract (English)
Unified remote sensing multimodal models exhibit a pronounced spatial reversal curse: Although they can accurately recognize and describe object locations in images, they often fail to faithfully execute the same spatial relations during text-to-image generation, where such relations constitute core semantic information in remote sensing. Motivated by this observation, we propose Uni-RS, the first unified multimodal model tailored for remote sensing, to explicitly address the spatial asymmetry between understanding and generation. Specifically, we first introduce explicit Spatial-Layout Planning to transform textual instructions into spatial layout plans, decoupling geometric planning from visual synthesis. We then impose Spatial-Aware Query Supervision to bias learnable queries toward spatial relations explicitly specified in the instruction. Finally, we develop Image-Caption Spatial Layout Variation to expose the model to systematic geometry-consistent spatial transformations. Extensive experiments across multiple benchmarks show that our approach substantially improves spatial faithfulness in text-to-image generation, while maintaining strong performance on multimodal understanding tasks like image captioning, visual grounding, and VQA tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。