arXiv:2410.07987cs.CV2024-10被引 1

统一视觉场景理解与3D虚拟合成,提升系统灵活性与协同性。

A transition towards virtual representations of visual scenes

  • 构建统一架构,融合视觉与语义数据处理
  • 支持3D虚拟合成的自适应场景描述
  • 适用于虚拟/增强现实等多场景应用

视觉场景理解是计算机视觉中的基础任务,旨在从视觉数据中提取有意义的信息。传统方法依赖于针对特定应用场景设计的独立专用算法,导致在需要处理视觉与语义数据的复杂系统设计中显得繁琐。尤其在虚拟现实与增强现实应用日益增多的背景下,这一问题更为突出。当系统需通过自动视觉场景理解生成精确且语义连贯的场景描述,并用于驱动3D虚拟合成时,缺乏灵活性与统一框架的问题愈发明显。为缓解此问题及其固有挑战,本文提出一种新架构,面向3D虚拟合成,实现可适配、统一且一致的解决方案。同时,展示了该架构在多个应用领域的实用性,并构建了概念验证系统以证明其实际可用性。

原文摘要 · Abstract (English)

Visual scene understanding is a fundamental task in computer vision that aims to extract meaningful information from visual data. It traditionally involves disjoint and specialized algorithms for different tasks that are tailored for specific application scenarios. This can be cumbersome when designing complex systems that include processing of visual and semantic data extracted from visual scenes, which is even more noticeable nowadays with the influx of applications for virtual or augmented reality. When designing a system that employs automatic visual scene understanding to enable a precise and semantically coherent description of the underlying scene, which can be used to fuel a visualization component with 3D virtual synthesis, the lack of flexibility and unified frameworks become more prominent. To alleviate this issue and its inherent problems, we propose an architecture that addresses the challenges of visual scene understanding and description towards a 3D virtual synthesis that enables an adaptable, unified and coherent solution. Furthermore, we expose how our proposition can be of use into multiple application areas. Additionally, we also present a proof of concept system that employs our architecture to further prove its usability in practice.

场景理解3D生成虚拟合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。