用事件相机提升低帧率RGB-D视频的插帧质量
UniRED: Unified RGB-D Video Frame Interpolation with Event Guidance

- 融合RGB、深度与事件数据,统一建模多模态信息
- 在公开数据集上实现更清晰的图像细节和更准确的深度结构
- 适合需要高精度动态场景重建的研究者
高帧率RGB-D视频对运动分析、动态场景理解与三维重建等任务至关重要。但受硬件限制,实际RGB-D相机通常帧率较低,难以捕捉快速变化。现有插帧方法在纯RGB上表现良好,但在RGB-D场景中常导致边界模糊、伪影明显及几何不一致。仅靠两帧边界图像估计运动,在复杂动态场景中本质欠约束。事件相机提供超高速异步测量,可补足密集运动线索。本文提出一种统一的多模态插帧框架,联合利用RGB外观、深度几何与事件时间线索。首先提取并融合三类信号,再通过运动基优化估计双向光流(用于RGB)与Z轴深度细化(用于深度),最后通过双向变形与软融合生成目标帧。此外,构建首个三模态数据集以缓解训练数据稀缺问题。在公开基准与自建数据集上的实验表明,本方法在RGB插帧中保持更高保真度,深度插帧中具有更强几何精度。
原文摘要 · Abstract (English)
High frame-rate RGB-D videos are crucial for a variety of downstream tasks, including motion analysis, dynamic scene understanding, and 3D reconstruction. However, due to hardware and sensing constraints, practical RGB-D cameras are typically limited to low frame rates, making it difficult to capture rapid scene dynamics. Existing video interpolation methods have achieved strong performance on RGB data, but they are not readily applicable to RGB-D scenarios, where they often yield blurry boundaries, visible artifacts, and degraded geometric consistency. Furthermore, motion estimation from only two boundary frames is inherently under-constrained in complex dynamic scenes. Event cameras, by contrast, provide asynchronous measurements with ultra-high temporal resolution, offering dense motion cues. In this paper, we propose a unified multimodal framework for RGB-D video interpolation that jointly exploits RGB appearance, depth geometry, and event-based temporal cues. Specifically, it first extracts and fuses RGB, depth and event cues, then estimates bidirectional flow with motion basis refinement for RGB and Z-axial refinement for depth, and finally synthesizes the target RGB-D frame via bidirectional warping and soft blending. In addition, we construct a new RGB-D-Event dataset to alleviate the scarcity of tri-modal training data. Extensive experiments on a public benchmark and the proposed dataset demonstrate that our method achieves superior photometric fidelity for RGB interpolation and stronger geometric accuracy for depth interpolation than existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。