arXiv:2506.02473cs.CV2025-06NeurIPS被引 1

用运动视频生成物体形状与材质的多种可能解,捕捉视觉模糊性。

Generative Perception of Shape and Material from Differential Motion

  • 基于微小运动视频,用扩散模型生成形状与材质的联合分布
  • 静态时输出多模态猜测,运动时预测更准确并收敛
  • 适合需要应对视觉不确定性的机器人感知系统

从单张图像感知物体形状与材质本质上存在歧义,尤其在光照未知且自由的情况下。尽管如此,人类常通过轻微移动头部或旋转物体来缓解这种不确定性。受此启发,我们提出一种新型条件去噪扩散模型,能够从一段物体发生微小运动的短视频中生成形状与材质图的样本。该参数高效架构直接在像素空间训练,可同时生成多个解耦属性。在少量带标注的合成物体-运动视频上训练后,模型展现出显著的涌现行为:对于静态观察,生成多样且多模态的合理形状-材质组合,体现固有的模糊性;当物体运动时,预测分布趋于收敛,解释更准确。此外,模型对真实世界中较不模糊的物体也生成高质量的形状-材质估计。通过从单视角转向连续运动观测,并利用生成式感知捕捉视觉模糊性,本工作为物理具身系统中的视觉推理提供了新思路。

原文摘要 · Abstract (English)

Perceiving the shape and material of an object from a single image is inherently ambiguous, especially when lighting is unknown and unconstrained. Despite this, humans can often disentangle shape and material, and when they are uncertain, they often move their head slightly or rotate the object to help resolve the ambiguities. Inspired by this behavior, we introduce a novel conditional denoising-diffusion model that generates samples of shape-and-material maps from a short video of an object undergoing differential motions. Our parameter-efficient architecture allows training directly in pixel-space, and it generates many disentangled attributes of an object simultaneously. Trained on a modest number of synthetic object-motion videos with supervision on shape and material, the model exhibits compelling emergent behavior: For static observations, it produces diverse, multimodal predictions of plausible shape-and-material maps that capture the inherent ambiguities; and when objects move, the distributions converge to more accurate explanations. The model also produces high-quality shape-and-material estimates for less ambiguous, real-world objects. By moving beyond single-view to continuous motion observations, and by using generative perception to capture visual ambiguities, our work suggests ways to improve visual reasoning in physically-embodied systems.

生成感知形状推断材质识别扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。