arXiv:2411.05794cs.MMcs.SD2024-11被引 6

提出新评估指标,更准确衡量多媒体质量模型性能。

Beyond Correlation: Evaluating Multimedia Quality Models with the Constrained Concordance Index

  • 引入约束一致性指数CCI,考虑评分者差异与置信区间重叠。
  • 在小样本和评分组变异场景下,评估结果比传统方法更可靠。
  • 适合关注主观评价可信度的研究者与模型开发者。

本研究聚焦多媒体质量模型的评估问题,指出主观均值意见分(MOS)评分因评分者不一致和偏差等因素存在固有不确定性。传统相关性指标如皮尔逊相关系数(PCC)、斯皮尔曼等级相关系数(SRCC)和肯德尔τ系数(KTAU)未能充分考虑这些不确定性,导致模型性能评估失准。为此,本文提出约束一致性指数(CCI),通过考虑MOS差异的统计显著性,并排除置信区间重叠的比较对,克服现有指标局限。在语音与图像质量评估等多个领域开展的全面实验表明,CCI在样本量小、评分群体差异大及范围受限等场景下,能提供更稳健、更准确的模型评估结果。研究发现,纳入评分者主观性并聚焦统计显著性配对,可显著提升多媒体质量预测模型的评估框架可靠性。该工作不仅揭示了主观评分不确定性被忽视的问题,也推动了评估方法学的进步。

原文摘要 · Abstract (English)

This study investigates the evaluation of multimedia quality models, focusing on the inherent uncertainties in subjective Mean Opinion Score (MOS) ratings due to factors like rater inconsistency and bias. Traditional statistical measures such as Pearson's Correlation Coefficient (PCC), Spearman's Rank Correlation Coefficient (SRCC), and Kendall's Tau (KTAU) often fail to account for these uncertainties, leading to inaccuracies in model performance assessment. We introduce the Constrained Concordance Index (CCI), a novel metric designed to overcome the limitations of existing metrics by considering the statistical significance of MOS differences and excluding comparisons where MOS confidence intervals overlap. Through comprehensive experiments across various domains including speech and image quality assessment, we demonstrate that CCI provides a more robust and accurate evaluation of instrumental quality models, especially in scenarios of low sample sizes, rater group variability, and restriction of range. Our findings suggest that incorporating rater subjectivity and focusing on statistically significant pairs can significantly enhance the evaluation framework for multimedia quality prediction models. This work not only sheds light on the overlooked aspects of subjective rating uncertainties but also proposes a methodological advancement for more reliable and accurate quality model evaluation.

质量评估主观评价统计方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。