arXiv:2605.25294cs.CV2026-05被引 3

将自然图像建模为球面结构,提升生成质量

Geometry-Aware Image Flow Matching

论文配图:Geometry-Aware Image Flow Matching
图 1 · 摘自论文原文
  • 发现语义信息集中在方向分量,幅值可近似为全局均值
  • 在多个数据集上,新方法比欧氏基线提升1.5~2.3个点
  • 适合追求高质量图像生成的开发者和研究者

最近的生成模型进展表明,几何感知建模在流形约束场景中具有强大潜力。然而,对于自然图像,该领域仍局限于欧氏假设,未能充分利用数据内在几何结构。本文研究自然图像的几何特性,发现语义信息主要编码在方向分量中,而幅值分量可由全局平均近似。这一性质在RGB空间和潜在空间均成立,表明自然图像可有效建模于超球面。基于此,我们提出球面最优传输流匹配(SOT-CFM),利用角度距离;以及球面流匹配(SFM),直接在流形上约束动态过程。实验表明,这些几何感知方法显著优于欧氏基线。本工作为黎曼流形建模与自然图像生成之间建立了新桥梁。

原文摘要 · Abstract (English)

Recent advances in generative models highlight the power of geometry-aware modeling in manifold-constrained settings. Yet, for natural images, the field remains confined to Euclidean assumptions, failing to exploit the potential of intrinsic geometric structures within the data. In this work, we investigate the geometry of natural images and observe that semantic information is predominantly encoded in directional components, while norm components can be approximated by the global average. This property holds across both RGB and latent spaces, suggesting that natural images can be effectively modeled on a hypersphere. Building on this finding, we introduce Spherical Optimal Transport Flow Matching (SOT-CFM), which utilizes angular distance, and Spherical Flow Matching (SFM), which constrains dynamics directly on the manifold. Our experiments demonstrate that these geometry-aware methods achieve superior performance against Euclidean baselines. Ultimately, this work provides a novel perspective that bridges the gap between Riemannian manifold-based modeling and natural image generation.

图像生成几何建模流形学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。