用廉价摄像头实现每秒4800帧的高速视频捕捉
RnGCam: High-speed video from rolling & global shutter measurements
- 融合滚动快门与全局快门传感器数据,提升时空分辨率
- 在仿真中优于现有压缩视频方法,硬件成本低10倍
- 适合需要高帧率、低成本高速成像的应用场景
压缩视频捕获将短时高速视频编码为单个测量值,通过低速传感器采集后计算重建。以往方法依赖昂贵硬件,且仅适用于稀疏场景。本文提出RnGCam系统,利用消费级低速滚动快门(RS)和全局快门(GS)传感器融合测量,实现千赫兹级帧率视频。RS传感器搭配伪随机光学器件(扩散器),实现空间多路复用;GS传感器使用传统镜头。前者提供高时间分辨率、低空间细节,后者相反。采用隐式神经表示(INR)重建方法,分别建模静态与动态场景成分,并显式正则化动态变化。仿真结果表明,该方法显著优于以往滚动快门压缩视频技术及先进帧插值算法。硬件验证在双相机系统上完成,以低于以往系统10倍的成本,实现4800帧/秒、230帧的密集场景高速视频生成。
原文摘要 · Abstract (English)
Compressive video capture encodes a short high-speed video into a single measurement using a low-speed sensor, then computationally reconstructs the original video. Prior implementations rely on expensive hardware and are restricted to imaging sparse scenes with empty backgrounds. We propose RnGCam, a system that fuses measurements from low-speed consumer-grade rolling-shutter (RS) and global-shutter (GS) sensors into video at kHz frame rates. The RS sensor is combined with a pseudorandom optic, called a diffuser, which spatially multiplexes scene information. The GS sensor is coupled with a conventional lens. The RS-diffuser provides low spatial detail and high temporal detail, complementing the GS-lens system's high spatial detail and low temporal detail. We propose a reconstruction method using implicit neural representations (INR) to fuse the measurements into a high-speed video. Our INR method separately models the static and dynamic scene components, while explicitly regularizing dynamics. In simulation, we show that our approach significantly outperforms previous RS compressive video methods, as well as state-of-the-art frame interpolators. We validate our approach in a dual-camera hardware setup, which generates 230 frames of video at 4,800 frames per second for dense scenes, using hardware that costs $10\times$ less than previous compressive video systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。