arXiv:2512.03939cs.CVcs.RO2025-12被引 6

不重训练,用注意力图自动识别动态区域并抑制误差传播。

MUT3R: Motion-aware Updating Transformer for Dynamic 3D Reconstruction

  • 通过分析注意力图发现动态区域被自然弱化,提取隐含运动线索。
  • 在推理阶段早期用门控机制抑制动态内容,提升时序一致性与姿态鲁棒性。
  • 无需微调模型,适用于实时动态3D重建场景。

近期基于状态的循环神经网络在静态3D重建中取得显著进展,但在动态场景下易受运动伪影影响,非刚性区域会干扰空间记忆与图像特征间的注意力传播。通过分析状态与图像标记更新机制,我们发现跨层自注意力图聚合呈现一致模式:动态区域被自然弱化,暴露出预训练变压器已编码但未显式利用的隐含运动线索。受此启发,我们提出MUT3R,一种无需训练的框架,在推理时利用注意力导出的运动线索,抑制变换器前几层中的动态内容。我们的注意力级门控模块在伪影沿特征层级传播前即抑制动态区域的影响。值得注意的是,我们不进行重训练或微调;让预训练变压器自主诊断其自身运动线索并自我修正。该早期调控在流式场景中稳定几何推理,显著提升多个动态基准上的时序一致性和相机位姿鲁棒性,为实现运动感知的流式重建提供简单且免训练的新路径。

原文摘要 · Abstract (English)

Recent stateful recurrent neural networks have achieved remarkable progress on static 3D reconstruction but remain vulnerable to motion-induced artifacts, where non-rigid regions corrupt attention propagation between the spatial memory and image feature. By analyzing the internal behaviors of the state and image token updating mechanism, we find that aggregating self-attention maps across layers reveals a consistent pattern: dynamic regions are naturally down-weighted, exposing an implicit motion cue that the pretrained transformer already encodes but never explicitly uses. Motivated by this observation, we introduce MUT3R, a training-free framework that applies the attention-derived motion cue to suppress dynamic content in the early layers of the transformer during inference. Our attention-level gating module suppresses the influence of dynamic regions before their artifacts propagate through the feature hierarchy. Notably, we do not retrain or fine-tune the model; we let the pretrained transformer diagnose its own motion cues and correct itself. This early regulation stabilizes geometric reasoning in streaming scenarios and leads to improvements in temporal consistency and camera pose robustness across multiple dynamic benchmarks, offering a simple and training-free pathway toward motion-aware streaming reconstruction.

3D重建运动感知注意力机制免训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。