arXiv:2603.05507cs.CVcs.GR2026-03

用Transformer模型实时修复多摄像头3D流的缺失纹理,保持画面连贯性。

Transformer-Based Inpainting for Real-Time 3D Streaming in Sparse Multi-Camera Setups

  • 基于时空嵌入的Transformer网络,跨帧保持纹理一致性。
  • 在实时约束下,图像与视频指标均优于现有方法。
  • 模块化设计适配不同摄像头配置,支持实时运行。

从多摄像头获取高质量3D流对增强现实与虚拟现实应用至关重要。由于实时性限制,视图数量有限,导致渲染图像中存在信息缺失和表面不完整问题。现有方法通常依赖简单启发式填充,易产生不一致或视觉伪影。本文提出一种面向应用场景的新型图像后处理修复方法,独立于底层表示,用于填补新视角渲染后的纹理空洞。该方法采用多视图感知的Transformer架构,结合时空嵌入,确保帧间一致性并保留细节。其分辨率无关设计可适配不同摄像头布局,自适应补丁选择策略平衡推理速度与质量,实现实时性能。我们在相同实时约束下对比了最先进的修复技术,结果表明,本模型在图像与视频评估指标上均达到最佳质量-速度权衡,显著优于现有方法。

原文摘要 · Abstract (English)

High-quality 3D streaming from multiple cameras is crucial for immersive experiences in many AR/VR applications. The limited number of views - often due to real-time constraints - leads to missing information and incomplete surfaces in the rendered images. Existing approaches typically rely on simple heuristics for the hole filling, which can result in inconsistencies or visual artifacts. We propose to complete the missing textures using a novel, application-targeted inpainting method independent of the underlying representation as an image-based post-processing step after the novel view rendering. The method is designed as a standalone module compatible with any calibrated multi-camera system. For this we introduce a multi-view aware, transformer-based network architecture using spatio-temporal embeddings to ensure consistency across frames while preserving fine details. Additionally, our resolution-independent design allows adaptation to different camera setups, while an adaptive patch selection strategy balances inference speed and quality, allowing real-time performance. We evaluate our approach against state-of-the-art inpainting techniques under the same real-time constraints and demonstrate that our model achieves the best trade-off between quality and speed, outperforming competitors in both image and video-based metrics.

3D流图像修复Transformer实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。