arXiv:2506.23711cs.CV2025-06ICCV被引 5

用文字和草图重建人眼感知的场景,不依赖大量训练数据。

Subjective Camera 1.0: Bridging Human Cognition and Visual Reconstruction through Sequence-Aware Sketch-Guided Diffusion

  • 基于扩散模型优化,通过草图和文字逐步重建场景。
  • 在两个数据集上达到图像质量与语义对齐的最新水平。
  • 适合关注主观认知重建与生成式视觉的科研与应用者。

我们提出主观相机的概念,用于重建物理相机无法捕捉的有意义瞬间。本文介绍主观相机1.0,一种从易获取的主观输入(即文本描述和逐笔绘制的粗略草图)中重建真实场景的框架。该方法基于扩散模型的优化对齐,避免了大规模成对训练数据并缓解了泛化问题。为解决现实场景中多抽象概念融合的挑战,我们设计了序列感知草图引导扩散框架,包含三项损失项,按主观输入的自然顺序实现概念级的渐进优化。在两个数据集上的实验表明,该方法在图像质量以及空间和语义对齐方面均达到当前最优表现。40名用户的主观评估进一步验证了该方法的一致偏好性。项目页面:subjective-camera.github.io

原文摘要 · Abstract (English)

We introduce the concept of a subjective camera to reconstruct meaningful moments that physical cameras fail to capture. We propose Subjective Camera 1.0, a framework for reconstructing real-world scenes from readily accessible subjective readouts, i.e., textual descriptions and progressively drawn rough sketches. Built on optimization-based alignment of diffusion models, our approach avoids large-scale paired training data and mitigates generalization issues. To address the challenge of integrating multiple abstract concepts in real-world scenarios, we design a Sequence-Aware Sketch-Guided Diffusion framework with three loss terms for concept-wise sequential optimization, following the natural order of subjective readouts. Experiments on two datasets demonstrate that our method achieves state-of-the-art performance in image quality as well as spatial and semantic alignment with target scenes. User studies with 40 participants further confirm that our approach is consistently preferred. Our project page is at: subjective-camera.github.io

视觉重建草图生成扩散模型主观认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。