arXiv:2510.17037eess.IV2025-10

无需训练的低复杂度视图合成失真估计,适合大基线场景。

A Low-Complexity View Synthesis Distortion Estimation Method for 3D Video with Large Baseline Considerations

  • 通过纹理-深度联合分类区分静止与非静止区域,减少误判。
  • 设计基线距离指标,有效补偿大基线导致的失真。
  • 基于区域融合策略,精准预测多视角合成失真,适合3D视频应用。

基于深度图像的渲染是3D视频系统中的关键视图合成算法,可从纹理图和深度图生成虚拟视图。在交互式自由视角视频和3D视频流等实时应用中,高效的视图合成失真估计模型对资源优化至关重要。然而,现有方法通常计算量大、需参数训练,或在大基线配置下表现不佳。本文提出一种新颖的、低复杂度且无需训练的方法,可在不实际渲染的情况下准确估计合成视图的失真。主要贡献包括:(1)一种纹理-深度联合分类方法,能精确分离纹理图中的局部静止与非静止区域,避免仅依赖纹理方法的误判;(2)设计新型基线距离指标,用于补偿大基线配置引起的失真;(3)基于区域的融合估计策略,几何识别重叠区、单视图区和互遮挡区,从两个参考视图预测不同合成条件下的失真。在标准MPEG 3D视频序列上的实验验证了该方法在大基线场景下的高精度与高效性。该方法使3D内容采集时可采用更灵活的摄像机布局,并准确预测复杂几何配置下的合成质量。

原文摘要 · Abstract (English)

Depth-image-based rendering is a key view synthesis algorithm in 3D video systems, which enables the synthesis of virtual views from texture images and depth maps. An efficient view synthesis distortion estimation model is critical for optimizing resource allocation in real-time applications such as interactive free-viewpoint video and 3D video streaming services. However, existing estimation methods are often computationally intensive, require parameter training, or performance poorly in challenging large baseline configurations. This paper presents a novel, low-complexity, and training-free method to accurately estimate the distortion of synthesized views without performing the actual rendering process. Key contributions include: (1) A joint texture-depth classification method that accurately separates texture image into locally stationary and non-stationary regions, which mitigates misclassifications by using texture-only methods. (2) A novel baseline distance indicator is designed for the compensation scheme for distortions caused by large baseline configurations. (3) A region-based blending estimation strategy that geometrically identifies overlapping, single-view, and mutual disocclusion regions, predicting distortion in synthesized views from two reference views with differing synthesis conditions. Experiments on standard MPEG 3D video sequences validate the proposed method's high accuracy and efficiency, especially in large baseline configurations. This method enables more flexible camera arrangements in 3D content acquisition by accurately predicting synthesis quality under challenging geometric configurations.

3D视频视图合成失真估计大基线

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。