为远程驾驶场景定制视频质量评估模型,提升安全可靠性。
Beyond VMAF: Towards Application-Specific Metrics for Teleoperation Video

- 基于远程驾驶数据重新训练VMAF模型,适配特定应用场景。
- 客观评分与人眼评价相关性提升,RMSE降15%,MAD降27%。
- 适合自动驾驶远程操控、视频质量评测等高安全要求领域。
自动驾驶虽取得显著进展,但仍需人工介入。远程操作提供了一种可扩展的解决方案,使操作员无需亲临现场即可支持车辆运行。在此背景下,视频传输成为操作员获取情境感知的主要来源,视频质量直接影响安全与任务表现。一项在线研究中,参与者对Zenseact数据集中的压缩视频序列进行主观质量评分,并用于重新训练视频多方法评估融合(VMAF)模型,得到一个专为远程操作优化的变体。该重训模型相较于原始4K VMAF在人类评分预测上表现更优:RMSE从10.36降至8.83,MAD从8.71降至6.38,分别改善15%和27%。结果表明,引入领域特定数据可增强通用质量度量在高安全性应用中的预测能力。同时,也发现部分异常情况:某些视频在关键驾驶区域出现明显退化,但客观得分仍偏高。
原文摘要 · Abstract (English)
Automated driving has made remarkable progress, yet situations still arise where human intervention is necessary. Teleoperation provides a scalable solution to address such cases, enabling remote operators to support vehicles without being physically present. In this context, video transmission forms the operator's primary source of situational awareness, making video quality a decisive factor for both safety and task performance. In an online study, participants rated compressed video sequences from the Zenseact Dataset and provided subjective quality ratings. These ratings were then used to retrain the Video Multi-Method Assessment Fusion (VMAF) model, yielding an adapted variant tailored to teleoperation. The retrained model demonstrated improved alignment with human ratings compared to the original 4K VMAF. In particular, RMSE decreased from 10.36 to 8.83, and MAD from 8.71 to 6.38, corresponding to improvements of 15% and 27%, respectively. These results highlight that incorporating domain-specific data can enhance the predictive power of established quality metrics in safety-critical applications. At the same time, Outlier cases emerged in which videos received high objective scores despite noticeable degradations in regions critical for the driving task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。