arXiv:2603.14377cs.CV2026-03

提出无需对齐的协同注意力框架,解决动态场景下HDR视频重建的鬼影问题。

LoCAtion: Long-time Collaborative Attention Framework for High Dynamic Range Video Reconstruction

论文配图:LoCAtion: Long-time Collaborative Attention Framework for High Dynamic Range Video Reconstruction
图 1 · 摘自论文原文
  • 用协同注意力替代刚性对齐,从特征路由角度重构HDR视频。
  • 在复杂动态场景中实现更优视觉质量与时间稳定性,无鬼影和闪烁。
  • 适合需要高动态范围视频稳定重建的实时应用开发者。

现有高动态范围(HDR)视频重建方法受限于脆弱的对齐-融合范式。尽管显式空间对齐在受控环境下能恢复精细细节,但在非受限动态场景中成为严重瓶颈。强制跨不可预测运动和变化曝光进行刚性对齐,导致注册误差转化为严重鬼影伪影与时间闪烁。本文重新思考这一传统前提,提出LoCAtion:一种长时协同注意力框架,将HDR视频生成从脆弱的空间扭曲任务转变为鲁棒的无对齐协同特征路由问题。基于新范式,架构显式解耦高度纠缠的重建任务:不强行扭曲邻近帧,而是以连续中曝光主干为锚点,利用协同注意力动态提取并注入未对齐曝光中的可靠辐射信息。此外,引入可学习全局序列求解器,通过双向上下文与长程时序建模,在全序列传播校正信号与结构特征,天然保障整段视频一致性并消除抖动。大量实验表明,LoCAtion在视觉质量与时间稳定性上达到领先水平,兼具高精度与计算效率。

原文摘要 · Abstract (English)

Prevailing High Dynamic Range (HDR) video reconstruction methods are fundamentally trapped in a fragile alignment-and-fusion paradigm. While explicit spatial alignment can successfully recover fine details in controlled environments, it becomes a severe bottleneck in unconstrained dynamic scenes. By forcing rigid alignment across unpredictable motions and varying exposures, these methods inevitably translate registration errors into severe ghosting artifacts and temporal flickering. In this paper, we rethink this conventional prerequisite. Recognizing that explicit alignment is inherently vulnerable to real-world complexities, we propose LoCAtion, a Long-time Collaborative Attention framework that reformulates HDR video generation from a fragile spatial warping task into a robust, alignment-free collaborative feature routing problem. Guided by this new formulation, our architecture explicitly decouples the highly entangled reconstruction task. Rather than struggling to rigidly warp neighboring frames, we anchor the scene on a continuous medium-exposure backbone and utilize collaborative attention to dynamically harvest and inject reliable irradiance cues from unaligned exposures. Furthermore, we introduce a learned global sequence solver. By leveraging bidirectional context and long-range temporal modeling, it propagates corrective signals and structural features across the entire sequence, inherently enforcing whole-video coherence and eliminating jitter. Extensive experiments demonstrate that LoCAtion achieves state-of-the-art visual quality and temporal stability, offering a highly competitive balance between accuracy and computational efficiency.

HDR视频协同注意力无对齐重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。