arXiv:2609.04603cs.CV2026-09

提出新评估指标,解决人物在场景中多视角生成的一致性难题

An Evaluation Framework for Generating Multi-View Images of a Person in a Scene

论文配图:An Evaluation Framework for Generating Multi-View Images of a Person in a Scene
图 1 · 摘自论文原文
  • 设计分离相机运动与头部姿态的评估方法
  • 通过实验证明该指标可有效衡量人物场景的空间视差
  • 适合研究图像生成、3D 重建及可控编辑的开发者

近期的生成式图像编辑扩散变换模型(DiTs)虽具备出色的语义编辑能力,但在空间一致的视角变换上仍存在困难。训练基础模型实现自由形式、提示驱动的视角变化主要受限于缺乏专用训练数据。尽管通用三维环境和物体的多视角数据集已存在,但缺少人物在自然场景中固定位置的成对多视角数据,尤其是正面与侧视图。在非受限环境中采集此类多相机数据既困难又难以扩展。本文首先尝试多种先进图像编辑模型合成此类数据,发现输出常出现头部朝向与背景不一致的幻觉。为此,提出头-场景旋转差异(HSRD)度量,将相机运动与局部头部姿态解耦,可定量评估人物在场景中的视角变化。实验表明,该指标为评估人物场景的三维空间视差提供了可靠管道,有助于构建高质量的多视角合成数据集。

原文摘要 · Abstract (English)

Recent generative image-editing Diffusion Transformers (DiTs) demonstrate impressive semantic editing capabilities but still struggle with spatially consistent camera angle changes. A primary bottleneck in training foundation models to execute free-form, promptable camera angle changes is the lack of specialized training data. While multi-view datasets exist for generic 3D environments and objects, there remains an absence of paired, multi-view datasets featuring human subjects at fixed locations in natural scenes, including frontal and side-profile views. Capturing such multi-camera data in unconstrained environments is logistically challenging and unscalable. In this paper, we first experiment with multiple state-of-the-art image editing models to create this data synthetically, but find that the outputs are frequently prone to hallucinations involving how much the subject's head turns relative to the background, often producing inconsistent environments. To address this issue, we propose the Head Scene Rotation Difference (HSRD) metric to quantitatively evaluate camera movements around a person. The proposed metric operates by decoupling camera movement from localized head pose manipulation. As demonstrated by the extensive experimentation, HSRD provides the pipeline necessary to evaluate 3D spatial parallax for a person in a scene, paving the way to reliably construct high-quality multi-view synthetic datasets.

图像生成多视角扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。