针对无人机视频动态重建中的局部动态不均问题,提出自适应锚点特征聚合方法。
AdaAnchor4D: Anchor-Conditioned Spatiotemporal Feature Aggregation for Monocular UAV 4D Reconstruction

- 基于锚点条件的特征聚合,动态适配不同区域的时序状态
- 在三个数据集上实现更清晰的动态细节与实时渲染速度
- 适合需要高精度动态场景重建的研究者和工程应用
单目无人机视频为复杂城市场景的动态重建提供了宝贵观测。然而,这类场景存在显著的时空异质性:不同区域具有不同的时序活动模式,部分动态区域的运动状态还会随时间演变。尽管基于分解共享时空特征场的动态高斯方法在以物体为中心或相对紧凑的场景中实现了高效准确的重建,但其普遍采用的固定平面特征组合机制难以适应无人机场景中局部动态的异质性,常导致鬼影伪影和模糊的动态细节。为此,本文提出AdaAnchor4D,一种用于单目无人机动态场景重建的自适应锚点变形框架。核心是锚点条件特征聚合(ACFA),通过锚点特定的聚合嵌入和时序信息自适应地聚合共享时空特征,使不同局部单元获得与其局部状态和时序特性相匹配的动态表示。解耦局部几何形变(DLGD)将锚点状态形变与局部高斯几何形变分离,密度自适应坐标扭曲(DACW)根据轴向锚点分布重参数化特征查询坐标,缓解了非均匀几何采样与均匀网格参数化之间的不匹配问题。在UAV-Arc4D、VisDrone和UAVDT上的实验表明,AdaAnchor4D在保持实时渲染性能的同时,优于代表性动态高斯方法的渲染质量。代码将公开发布。
原文摘要 · Abstract (English)
Monocular UAV videos provide valuable observations for dynamic reconstruction of complex urban scenes. However, such scenes exhibit pronounced spatiotemporal heterogeneity: different regions follow distinct temporal activity patterns, while the motion states of some dynamic regions may further evolve over time. Although dynamic Gaussian methods based on decomposed shared spatiotemporal feature fields have achieved efficient and accurate reconstruction in object-centric or relatively compact scenes, their commonly adopted fixed plane-wise feature combination mechanisms are less suited to the heterogeneous local dynamics of UAV scenes, often leading to ghosting artifacts and blurred dynamic details. To address this challenge, we propose AdaAnchor4D, an adaptive anchor deformation framework for monocular UAV dynamic scene reconstruction. At its core, Anchor-Conditioned Feature Aggregation (ACFA) adaptively aggregates shared spatiotemporal features using anchor-specific aggregation embeddings and temporal information, allowing different local units to obtain dynamic representations tailored to their local and temporal states. Decoupled Local Geometry Deformation (DLGD) separates anchor-state deformation from local Gaussian geometry deformation, while Density-Adaptive Coordinate Warping (DACW) reparameterizes feature-query coordinates according to the axis-wise anchor distributions, alleviating the mismatch between non-uniform geometric sampling and uniform grid parameterization. Experiments on UAV-Arc4D, VisDrone, and UAVDT show that AdaAnchor4D achieves higher rendering quality than representative dynamic Gaussian methods while maintaining real-time rendering performance. The code will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。