arXiv:2409.18236cs.CVcs.LG2024-09被引 3

通过细胞可见性建模,提升点云视频视场预测精度与效率

Spatial Visibility and Temporal Dynamics: Revolutionizing Field of View Prediction in Adaptive Point Cloud Video Streaming

  • 从细胞可见性出发重构视场预测,实现3D数据级精准传输决策
  • 相比顶尖模型,长期可见性预测误差降低50%,支持百万点实时处理
  • 适合需要低延迟、高精度点云流媒体的系统开发者与研究者

视场(FoV)自适应流媒体能显著降低沉浸式点云视频(PCV)的带宽需求,仅传输用户视场内的可见点。传统方法多基于六自由度(6DoF)轨迹进行视场预测,再转换为点可见性。此类方法未显式考虑视频内容对用户注意力的影响,且视场到点可见性的转换常存在误差且耗时。本文从细胞可见性视角重新构建点云视频视场预测问题,实现基于预测可见性分布的3D数据级精准传输决策。我们提出一种新颖的空间可见性与对象感知图模型,利用历史3D可见性数据,融合空间感知、邻近细胞相关性与遮挡信息,预测未来细胞可见性。该模型显著提升长期细胞可见性预测性能,在超过一百万点的点云视频上保持超过30fps的实时性能,相比当前最优模型,预测均方误差损失降低最高达50%。

原文摘要 · Abstract (English)

Field-of-View (FoV) adaptive streaming significantly reduces bandwidth requirement of immersive point cloud video (PCV) by only transmitting visible points in a viewer's FoV. The traditional approaches often focus on trajectory-based 6 degree-of-freedom (6DoF) FoV predictions. The predicted FoV is then used to calculate point visibility. Such approaches do not explicitly consider video content's impact on viewer attention, and the conversion from FoV to point visibility is often error-prone and time-consuming. We reformulate the PCV FoV prediction problem from the cell visibility perspective, allowing for precise decision-making regarding the transmission of 3D data at the cell level based on the predicted visibility distribution. We develop a novel spatial visibility and object-aware graph model that leverages the historical 3D visibility data and incorporates spatial perception, neighboring cell correlation, and occlusion information to predict the cell visibility in the future. Our model significantly improves the long-term cell visibility prediction, reducing the prediction MSE loss by up to 50% compared to the state-of-the-art models while maintaining real-time performance (more than 30fps) for point cloud videos with over 1 million points.

点云视频视场预测可见性建模实时传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。