点云输入微小扰动会导致3D生成模型崩溃,输出分裂成数百碎片。
Meltdown: Circuits and Bifurcations in Point-Cloud-Conditioned 3D Diffusion Transformers
- 发现扩散模型在点云输入下存在'熔毁'故障,由早期注意力层写操作引发。
- 实测89.9%-100%的形状在真实数据集上触发熔毁,且两种主流架构均受影响。
- 提出测试时控制方法PowerRemap,可恢复98.3%的生成结果,提升鲁棒性。
稀疏点云是3D表面重建的常见输入,尤其在手术导航和自动驾驶等关键场景中。近期基于点云条件的3D扩散变换器利用学习先验取得了最先进性能。我们发现这些模型在现实输入变化下可能灾难性失效,并揭示其机制。识别出一种称为‘熔毁’的故障模式:点云表面极微小扰动即可导致重建输出分裂为数百个不连通部分。对抗搜索在两个开源架构(WaLa、Make-a-Shape)上,于真实数据集(GSO、SimJEB)中,以DDPM和DDIM采样方式均在89.9%-100%的形状中复现熔毁。通过前向传播追踪发现,熔毁受点分布均匀性影响,经点云编码器传递,并由扩散主干中一次早期去噪交叉注意力写操作决定。扩散轨迹集合在该写操作附近表现出对称性破缺,符合逆过程分岔特征。通过匹配幅度对照实验,证实模型决策依赖于方向性低秩子空间中的扰动漂移。基于此,提出测试时控制方法PowerRemap,重塑局部写操作的奇异谱以抑制漂移,在WaLa上实现98.3%的救援率,Make-a-Shape上达84.6%。结果将电路级注意力机制与轨迹级失败解释相联系,展示机制分析如何指导条件扩散变换器行为。
原文摘要 · Abstract (English)
Sparse point clouds are a common input modality for 3D surface reconstruction, including in safety-critical settings such as surgical navigation and autonomous perception. Recent point-cloud-conditioned 3D diffusion transformers achieve state-of-the-art results in this regime by leveraging learned priors. We show that these models can fail catastrophically under realistic input variation, and present a mechanistic case study of why. We identify a failure mode we call Meltdown: tiny on-surface perturbations to a sparse input point cloud can fracture the reconstructed output into hundreds of disconnected pieces. Adversarial search recovers Meltdown in 89.9-100% of shapes across the two open-weight state-of-the-art architectures we study (WaLa, Make-a-Shape) on real-world datasets (GSO, SimJEB) and under both DDPM and DDIM sampling. We trace Meltdown along the forward pass: it is governed by how uniformly the points are distributed on the surface, faithfully transduced through the point-cloud encoder, and committed by a single early-denoising cross-attention write in the diffusion backbone. Diffusion-trajectory ensembles exhibit symmetry-breaking near this commit step, consistent with a bifurcation of the reverse process. Through a suite of matched-magnitude controls, we show that the variable on which the model commits is directional, concentrated in a low-rank subspace of the write's perturbation drift. Motivated by this finding, we introduce PowerRemap, a test-time control that reshapes the singular spectrum of the localized write to suppress this drift, with rescue rates of 98.3% on WaLa and 84.6% on Make-a-Shape. Together, these results link a circuit-level cross-attention mechanism to a trajectory-level account of the failure, demonstrating how mechanistic analysis can explain and guide behavior in conditional diffusion transformers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。