arXiv:2603.22893cs.CV2026-03被引 2

SLARM统一动态场景重建与语义理解,支持自然语言查询。

SLARM: Streaming and Language-Aligned Reconstruction Model for Dynamic Scenes

  • 通过高阶运动建模和无流监督训练,实现复杂运动捕捉。
  • 相比现有方法,运动精度提升21%,重建PSNR提高1.6 dB。
  • 适合需要实时流式推理与语义交互的动态场景应用。

我们提出SLARM,一种前馈模型,统一了动态场景重建、语义理解与实时流式推理。SLARM通过高阶运动建模捕捉复杂非均匀运动,仅基于可微渲染训练,无需光流监督。同时,SLARM从LSeg中提取语义特征,获得与语言对齐的表示,支持自然语言语义查询;语义与几何的紧密耦合进一步提升了动态重建的准确性和鲁棒性。此外,SLARM采用基于窗口的因果注意力处理图像序列,实现稳定低延迟的流式推理,且不累积内存开销。在统一框架下,SLARM在动态估计、渲染质量与场景解析上均达当前最优,相比现有方法运动精度提升21%,重建PSNR提升1.6 dB,分割mIoU提升20%。

原文摘要 · Abstract (English)

We propose SLARM, a feed-forward model that unifies dynamic scene reconstruction, semantic understanding, and real-time streaming inference. SLARM captures complex, non-uniform motion through higher-order motion modeling, trained solely on differentiable renderings without any flow supervision. Besides, SLARM distills semantic features from LSeg to obtain language-aligned representations. This design enables semantic querying via natural language, and the tight coupling between semantics and geometry further enhances the accuracy and robustness of dynamic reconstruction. Moreover, SLARM processes image sequences using window-based causal attention, achieving stable, low-latency streaming inference without accumulating memory cost. Within this unified framework, SLARM achieves state-of-the-art results in dynamic estimation, rendering quality, and scene parsing, improving motion accuracy by 21%, reconstruction PSNR by 1.6 dB, and segmentation mIoU by 20% over existing methods.

动态重建语言对齐实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。