arXiv:2605.01187eess.IVcs.AR2026-05

NVENC在黑曜石架构上提升编码质量,但延迟飙升400%难用于实时通信。

Evolution of NVENC Efficiency: A Longitudinal Analysis of HQ and UHQ Tuning Efficiency, Latency and Energy Trade-offs

  • 对比帕斯卡到黑曜石世代,分析硬件编码器效率与延迟的演变。
  • 黑曜石架构在标准模式下提升5.94%编码效率,UHQ模式最高达22.79%。
  • UHQ模式引入7个B帧,导致延迟超400%,适合点播而非实时交互。

随着上行密集型应用的迅速发展,视频编码需在高率失真(RD)效率与极低延迟间取得平衡。本文对从帕斯卡到新兴黑曜石架构的NVIDIA硬件编码(NVENC)进行了纵向性能分析。特别评估了新型“超高质量”(UHQ)调优模式与标准低延迟配置的可行性。结果表明,尽管黑曜石架构打破了历史效率瓶颈,在标准模式下实现5.94%的BD-Rate提升,UHQ模式更高达22.79%,但这些增益带来严重系统级代价。揭示出UHQ作为混合流水线运行,将复杂度迁移至CUDA核心,并强制采用激进的时间结构(最多7个B帧),使端到端延迟增加超过400%,显卡板级功耗上升达40%。因此,虽然UHQ成功弥合了与软件编码器的质量差距,其致命的串行化延迟使其不适用于交互式实时通信,更适合用于点播(VoD)转码场景。

原文摘要 · Abstract (English)

The rapid expansion of uplink-intensive applications necessitates video coding solutions that balance high Rate-Distortion (RD) efficiency with ultra-low latency. This paper presents a longitudinal performance analysis of NVIDIA hardware encoding (NVENC), spanning from Pascal to the emerging Blackwell generation. We specifically evaluate the operational viability of the new "Ultra High Quality" (UHQ) tuning mode against standard low-latency configurations. Our results demonstrate that while the Blackwell architecture breaks historical efficiency plateaus, achieving a 5.94% BD-Rate gain in standard modes and up to 22.79% in UHQ modes, these gains incur severe system-level penalties. We reveal that UHQ operates as a hybrid pipeline, offloading complexity to CUDA cores and enforcing aggressive temporal structures (up to 7 B-frames) that increase end-to-end latency by over 400% and GPU board power consumption by up to 40%. Consequently, while UHQ successfully bridges the quality gap with software encoders, its prohibitive serialization delay renders it unsuitable for interactive real-time communications, positioning it instead as a specialized solution for Video-on-Demand (VoD) transcoding.

视频编码NVENC延迟优化黑曜石架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。