arXiv:2511.07479cs.CVcs.AI2025-11

用视觉Transformer还原折叠视频,突破传统相机动态范围限制

Modulo Video Recovery via Selective Spatiotemporal Vision Transformer

  • 提出选择性时空视觉变压器,聚焦关键区域提升重建效率
  • 在8比特折叠视频上实现高质量重建,性能超越现有方法
  • 适合做高动态范围成像、视频恢复的科研与工程人员

传统图像传感器动态范围有限,导致高动态范围(HDR)场景下出现饱和。模态相机通过将入射辐射能折叠到有限范围内来解决此问题,但需专用解包裹算法恢复原始信号。与扩展动态范围的HDR恢复不同,模态恢复旨在从折叠采样中重建真实数值。尽管模态成像已提出十余年,其进展缓慢,尤其在现代深度学习应用方面。本文表明标准HDR方法不适用于模态恢复。而视觉变换器能捕捉全局依赖和时空关系,对解包裹至关重要。然而,现有变换器架构需创新技术适配模态恢复。为此,我们提出首个用于模态视频重建的深度学习框架——选择性时空视觉变换器(SSViT)。SSViT采用标记选择策略,提升效率并聚焦关键区域。实验验证,该方法可在8位折叠视频上生成高质量重构结果,在模态视频恢复任务中达到当前最优性能。

原文摘要 · Abstract (English)

Conventional image sensors have limited dynamic range, causing saturation in high-dynamic-range (HDR) scenes. Modulo cameras address this by folding incident irradiance into a bounded range, yet require specialized unwrapping algorithms to reconstruct the underlying signal. Unlike HDR recovery, which extends dynamic range from conventional sampling, modulo recovery restores actual values from folded samples. Despite being introduced over a decade ago, progress in modulo image recovery has been slow, especially in the use of modern deep learning techniques. In this work, we demonstrate that standard HDR methods are unsuitable for modulo recovery. Transformers, however, can capture global dependencies and spatial-temporal relationships crucial for resolving folded video frames. Still, adapting existing Transformer architectures for modulo recovery demands novel techniques. To this end, we present Selective Spatiotemporal Vision Transformer (SSViT), the first deep learning framework for modulo video reconstruction. SSViT employs a token selection strategy to improve efficiency and concentrate on the most critical regions. Experiments confirm that SSViT produces high-quality reconstructions from 8-bit folded videos and achieves state-of-the-art performance in modulo video recovery.

视频恢复视觉Transformer模态成像动态范围

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。