首次量化图像线索对单图3D生成的影响,揭示几何线索比纹理更关键。
Cue3D: Quantifying the Role of Image Cues in Single-Image 3D Generation
- 通过系统扰动七类图像线索,评估其对3D生成质量的影响。
- 发现形状合理性比纹理更重要,阴影等几何线索起决定性作用。
- 适用于想提升3D生成模型可解释性与鲁棒性的研究人员。
人类和传统计算机视觉方法依赖多种单目线索(如阴影、纹理、轮廓等)从单张图像推断三维结构。尽管深度生成模型在单图3D生成上取得显著进展,但尚不清楚这些方法实际利用了哪些图像线索。本文提出Cue3D,首个全面且模型无关的框架,用于量化单图3D生成中各类图像线索的影响。该统一基准评估了七种前沿方法,涵盖回归式、多视角及原生3D生成范式。通过系统扰动阴影、纹理、轮廓、透视、边缘和局部连续性等线索,我们测量其对3D输出质量的影响。分析表明,形状合理性而非纹理主导泛化性能;几何线索(尤其是阴影)对3D生成至关重要。此外,发现模型过度依赖给定轮廓,且对透视和局部连续性的敏感度在不同模型族间差异显著。通过剖析这些依赖关系,Cue3D深化了对现代3D网络如何利用经典视觉线索的理解,并为开发更透明、稳健、可控的单图3D生成模型提供了方向。
原文摘要 · Abstract (English)
Humans and traditional computer vision methods rely on a diverse set of monocular cues to infer 3D structure from a single image, such as shading, texture, silhouette, etc. While recent deep generative models have dramatically advanced single-image 3D generation, it remains unclear which image cues these methods actually exploit. We introduce Cue3D, the first comprehensive, model-agnostic framework for quantifying the influence of individual image cues in single-image 3D generation. Our unified benchmark evaluates seven state-of-the-art methods, spanning regression-based, multi-view, and native 3D generative paradigms. By systematically perturbing cues such as shading, texture, silhouette, perspective, edges, and local continuity, we measure their impact on 3D output quality. Our analysis reveals that shape meaningfulness, not texture, dictates generalization. Geometric cues, particularly shading, are crucial for 3D generation. We further identify over-reliance on provided silhouettes and diverse sensitivities to cues such as perspective and local continuity across model families. By dissecting these dependencies, Cue3D advances our understanding of how modern 3D networks leverage classical vision cues, and offers directions for developing more transparent, robust, and controllable single-image 3D generation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。