无需训练的360°视频流系统,靠物体吸引眼球自动预测视角。
Training-Free Adaptive 360-degree Video Streaming via Semantic Potential Fields
- 用语义物体生成引力场模拟操作员视线,实现零样本视角预测。
- 在3600次模拟中QoE达2.71,接近最优方案,延迟仅1.01毫秒。
- 适合对可靠性要求高的远程操控场景,无需用户数据预训练。
面向远程操控的自适应360°视频流面临两个耦合挑战:在不确定注视模式下的视口预测,以及在波动无线信道下的码率自适应。尽管深度强化学习方法能实现高用户体验质量(QoE),但其可解释性差且依赖离线训练,限制了在安全关键系统中的部署。我们提出OrbitStream,一种无训练框架,将视口预测建模为引力视口预测(GVP)问题,其中语义物体生成势场吸引操作员视线,并采用饱和式比例-微分(PD)控制器进行缓冲区调节。在富含物体的远程操控轨迹上,OrbitStream实现了94.7%的零样本视口预测准确率,接近轨迹外推基线(约98.5%)。在3,600次蒙特卡洛模拟中,其QoE为2.71,排名第二(优于FastMPC的1.84),仅次于BOLA-E的2.80,决策延迟仅1.01毫秒,重缓冲极少。
原文摘要 · Abstract (English)
Adaptive 360° video streaming for teleoperation faces two coupled challenges: viewport prediction under uncertain gaze patterns and bitrate adaptation over fluctuating wireless channels. While Deep Reinforcement Learning (DRL) methods achieve high Quality of Experience (QoE), their lack of interpretability and dependence on offline training limit deployment in safety-critical systems. We propose OrbitStream, a training-free framework that formulates viewport prediction as a Gravitational Viewport Prediction (GVP) problem, where semantic objects generate potential fields that attract operator gaze, and employs a Saturation-Based Proportional-Derivative (PD) Controller for buffer regulation. On object-rich teleoperation traces, OrbitStream achieves 94.7% zero-shot viewport prediction accuracy without user-specific profiling, approaching trajectory-extrapolation baselines (~98.5%). Across 3,600 Monte Carlo simulations, it ranks second among 12 algorithms (QoE 2.71 vs. BOLA-E's 2.80), outperforming FastMPC (1.84), with 1.01 ms decision latency and minimal rebuffering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。