arXiv:2508.09068eess.IVcs.CV2025-08中稿 · manuscript, 13 pag…

构建多相机数据集,公平比较帧插值与视图合成效果

A new dataset and comparison for multi-camera frame synthesis

  • 用自研密集线性相机阵列采集多视角图像,实现跨方法公平对比
  • 真实图像上深度学习方法仅略胜传统方法,3D高斯点云反而低3.5 dB PSNR
  • 合成场景中3D高斯点云优于帧插值近5 dB,适合虚拟内容生成研究者

现有帧插值与视图合成方法虽目标相似(基于时/空邻近帧生成新帧),但数据集侧重不同:帧插值多关注单相机时序变化,视图合成偏向立体深度估计。这导致两者难以直接比较。本文构建了基于自研密集线性相机阵列的新型多相机数据集,用于公平评估经典与深度学习帧插值器与视图合成方法(3D Gaussian Splatting)在视图插值任务上的表现。结果显示,在真实图像上,深度学习方法未显著超越传统方法,3D高斯点云反而落后达3.5 dB PSNR;而在合成场景中,3D高斯点云以95%置信度领先帧插值算法近5 dB PSNR。

原文摘要 · Abstract (English)

Many methods exist for frame synthesis in image sequences but can be broadly categorised into frame interpolation and view synthesis techniques. Fundamentally, both frame interpolation and view synthesis tackle the same task, interpolating a frame given surrounding frames in time or space. However, most frame interpolation datasets focus on temporal aspects with single cameras moving through time and space, while view synthesis datasets are typically biased toward stereoscopic depth estimation use cases. This makes direct comparison between view synthesis and frame interpolation methods challenging. In this paper, we develop a novel multi-camera dataset using a custom-built dense linear camera array to enable fair comparison between these approaches. We evaluate classical and deep learning frame interpolators against a view synthesis method (3D Gaussian Splatting) for the task of view in-betweening. Our results reveal that deep learning methods do not significantly outperform classical methods on real image data, with 3D Gaussian Splatting actually underperforming frame interpolators by as much as 3.5 dB PSNR. However, in synthetic scenes, the situation reverses -- 3D Gaussian Splatting outperforms frame interpolation algorithms by almost 5 dB PSNR at a 95% confidence level.

帧插值视图合成多相机3DGS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。