arXiv:2605.01219cs.MMcs.CV2026-05中稿 · ICIP 2026, 6 pages…

让音视频质量评估学会判断哪路信号更可信,自动弱化失真严重的信号。

Multimodal Confidence Modeling in Audio-Visual Quality Assessment

论文配图:Multimodal Confidence Modeling in Audio-Visual Quality Assessment
图 1 · 摘自论文原文
  • 引入模态置信度估计,动态评估音视频各自可靠性
  • 通过置信度引导的混合器提升与人类评分的相关性
  • 适用于真实场景中音视频失真不对称的情况

音视频质量评估(AVQA)在流媒体、远程会议和沉浸式媒体中至关重要。现实中,音视频失真常呈非对称性:一端严重受损而另一端完好。现有方法将音视频视为同等可靠,导致融合时错误放大不可靠信号。本文提出MCM-AVQA框架,显式估计各模态置信度,并将其注入专用音视频混合器以实现跨模态注意力。该混合器采用帧级、置信度引导的通道注意力,抑制低置信度信号,保留高置信度信息,有效维持时间维度上的失真模式。视觉置信度模块将帧级伪影概率转换为平滑的片段级置信分数;音频置信度模块基于语音质量线索推断,无需干净参考。多个基准测试结果表明,MCM-AVQA及其置信度引导混合器显著提升与人类平均意见分的相关性,在真实非对称失真下表现更可解释。

原文摘要 · Abstract (English)

Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortions are often asymmetric, where one modality may be severely degraded while the other remains clean. Still, most contemporary AVQA metrics treat audio and video as equally reliable, causing confidence-unaware fusion to emphasize unreliable signals. This paper proposes MCM-AVQA, a multimodal confidence-aware AVQA framework that explicitly estimates modality-specific confidence and injects it into a dedicated audio-visual mixer for cross-modal attention. The Audio-Visual Mixer utilizes frame-level, confidence-guided channel attention to gate fusion, modulating feature interaction between modalities so that high-confidence streams dominate while unreliable inputs are suppressed, preserving temporal degradation patterns. A multi-head visual confidence estimator turns frame-level artifact probabilities into temporally smoothed, clip-level visual confidence scores, while an audio confidence module derives confidence from speech-quality cues without requiring a clean reference. Experiments on multiple AVQA benchmarks show that MCM-AVQA, and specifically its confidence-guided Audio-Visual Mixer, improve correlation with human mean opinion scores and yield more interpretable behavior under real-world asymmetric audio-visual distortions.

音视频评估置信度建模多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。