arXiv:2604.14302cs.CV2026-04

仅用一张手绘草图生成多视角一致的3D场景,突破几何扭曲限制。

Geometrically Consistent Multi-View Scene Generation from Freehand Sketches

论文配图:Geometrically Consistent Multi-View Scene Generation from Freehand Sketches
图 1 · 摘自论文原文
  • 设计新型注意力模块与损失函数,从扭曲草图中推断几何结构
  • 在9000组数据上训练,实现多视角一致性,FID降低60%以上
  • 无需参考图或迭代优化,推理速度提升3.7倍,适合快速原型设计

本文解决从单张手绘草图生成几何一致的多视角场景这一新问题。手绘草图几何信息贫乏,抽象笔触伴随空间扭曲,难以支持一致的3D重建。现有方法需照片或文本输入,或依赖多视图草图与耗时优化。我们提出三项协同贡献:(i) 构建约9000个草图-多视图样本的标注数据集,通过自动化生成与过滤流程;(ii) 并行相机感知注意力适配器(CA3),向视频变换器注入几何先验;(iii) 基于运动恢复结构(SfM)的稀疏对应监督损失(CSL)。所提框架在单一去噪过程中合成所有视图,无需参考图像、迭代优化或场景级调优。相比最先进两阶段基线,显著提升真实感(FID降低超60%)和几何一致性(相关准确率提升23%),并实现最高3.7倍的推理加速。

原文摘要 · Abstract (English)

We tackle a new problem: generating geometrically consistent multi-view scenes from a single freehand sketch. Freehand sketches are the most geometrically impoverished input one could offer a multi-view generator. They convey scene intent through abstract strokes while introducing spatial distortions that actively conflict with any consistent 3D interpretation. No prior method attempts this; existing multi-view approaches require photographs or text, while sketch-to-3D methods need multiple views or costly per-scene optimisation. We address three compounding challenges; absent training data, the need for geometric reasoning from distorted 2D input, and cross-view consistency, through three mutually reinforcing contributions: (i) a curated dataset of $\sim$9k sketch-to-multiview samples, constructed via an automated generation and filtering pipeline; (ii) Parallel Camera-Aware Attention Adapters (CA3) that inject geometric inductive biases into the video transformer; and (iii) a Sparse Correspondence Supervision Loss (CSL) derived from Structure-from-Motion reconstructions. Our framework synthesizes all views in a single denoising process without requiring reference images, iterative refinement, or per-scene optimization. Our approach significantly outperforms state-of-the-art two-stage baselines, improving realism (FID) by over 60% and geometric consistency (Corr-Acc) by 23%, while providing up to a 3.7$\times$ inference speedup.

草图生成多视图几何一致性扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。