arXiv:2508.01259cs.CV2025-08AAAI被引 5

提出STDNet网络,解决视频深度超分辨率中的长尾分布问题。

SpatioTemporal Difference Network for Video Depth Super-Resolution

  • 设计空间与时间差异分支,动态对齐特征以缓解非平滑区域误差
  • 在多个数据集上超越现有方法,显著提升复杂区域重建精度
  • 适合关注视频深度重建质量提升的研究者与工程师

深度超分辨率已取得显著进展,引入多帧信息进一步提升了重建质量。然而统计分析表明,视频深度超分辨率仍受明显长尾分布影响,主要体现在空间非平滑区域和时间变化区域。为此,我们提出一种新型时空差异网络(STDNet),包含两个核心分支:空间差异分支和时间差异分支。在空间差异分支中,引入空间差异机制,动态对齐RGB特征与学习到的空间差异表示,实现帧内RGB-D特征聚合以校准深度。在时间差异分支中,设计时间差异策略,优先将相邻帧的时空变化信息传播至当前深度帧,利用时间差异表示实现对时间长尾区域的精确运动补偿。在多个数据集上的大量实验结果证明了所提方法的有效性,性能优于现有方法。

原文摘要 · Abstract (English)

Depth super-resolution has achieved impressive performance, and the incorporation of multi-frame information further enhances reconstruction quality. Nevertheless, statistical analyses reveal that video depth super-resolution remains affected by pronounced long-tailed distributions, with the long-tailed effects primarily manifesting in spatial non-smooth regions and temporal variation zones. To address these challenges, we propose a novel SpatioTemporal Difference Network (STDNet) comprising two core branches: a spatial difference branch and a temporal difference branch. In the spatial difference branch, we introduce a spatial difference mechanism to mitigate the long-tailed issues in spatial non-smooth regions. This mechanism dynamically aligns RGB features with learned spatial difference representations, enabling intra-frame RGB-D aggregation for depth calibration. In the temporal difference branch, we further design a temporal difference strategy that preferentially propagates temporal variation information from adjacent RGB and depth frames to the current depth frame, leveraging temporal difference representations to achieve precise motion compensation in temporal long-tailed areas. Extensive experimental results across multiple datasets demonstrate the effectiveness of our STDNet, outperforming existing approaches.

深度估计视频重建超分辨率差异网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。