arXiv:2603.00938cs.CVcs.AI2026-03被引 1

首个针对HDR用户生成视频的多模态质量评估模型,解决高动态范围视频感知难题。

Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videos

  • 设计HDR感知视觉编码器与基于强化学习的优化框架,提升对高动态范围特征的敏感度。
  • 在4.4万段视频、超150万条评分数据上验证,性能超越现有方法。
  • 适合视频质量评估、内容优化及多模态大模型研究者参考。

高动态范围(HDR)用户生成视频正迅速在社交平台普及,但多数感知视频质量评估(VQA)系统仍针对标准动态范围(SDR)。HDR具有更高位深、广色域和更广亮度范围,会暴露近黑压缩、高光截断、条带效应和曝光闪烁等失真,加剧UGC内容缺陷并挑战SDR模型。为推动进展,我们构建了Beyond8Bits——一个包含4.4万段视频、来自6500个来源、超过150万条众包评分的大规模主观数据集,涵盖多样场景、采集条件和压缩设置。我们进一步提出首个用于HDR-UGC VQA的多模态大语言模型HDR-Q:(i) 设计新颖的HDR感知视觉编码器以生成HDR敏感嵌入;(ii) 提出HDR-aware Policy Optimization(HAPO),通过基于HDR-SDR对比KL约束和高斯加权回归奖励,引导模型依赖HDR输入进行细粒度平均意见得分(MOS)校准。在Beyond8Bits及公开的HDR-VQA基准上,HDR-Q均达到最先进性能。

原文摘要 · Abstract (English)

High Dynamic Range (HDR) user-generated (UGC) videos are rapidly proliferating across social platforms, yet most perceptual video quality assessment (VQA) systems remain tailored to Standard Dynamic Range (SDR). HDR has a higher bit depth, wide color gamut, and elevated luminance range, exposing distortions such as near-black crushing, highlight clipping, banding, and exposure flicker that amplify UGC artifacts and challenge SDR models. To catalyze progress, we curate Beyond8Bits, a large-scale subjective dataset of 44K videos from 6.5K sources with over 1.5M crowd ratings, spanning diverse scenes, capture conditions, and compression settings. We further introduce HDR-Q, the first Multimodal Large Language Model (MLLM) for HDR-UGC VQA. We propose (i) a novel HDR-aware vision encoder to produce HDR-sensitive embeddings, and (ii) HDR-Aware Policy Optimization (HAPO), an RL finetuning framework that anchors reasoning to HDR cues. HAPO augments GRPO via an HDR-SDR contrastive KL that encourages token reliance on HDR inputs and a Gaussian weighted regression reward for fine-grained MOS calibration. Across Beyond8Bits and public HDR-VQA benchmarks, HDR-Q delivers state-of-the-art performance.

HDR视频质量评估多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。