提出BinoMOS模型,为多媒体质量评估设定主观与客观评分的合理预期上限。
Bounds on Agreement between Subjective and Objective Measurements
- 基于二项分布建模主观评分,推导出客观评分相关性与误差的理论边界。
- 模型能复现MOS值离散性及投票数对评分的影响,与18个实测数据高度吻合。
- 无需原始方差数据也能估算边界,适合评估缺乏详细统计信息的测试。
客观多媒体质量评估通常通过与主观评分比较来评判,常用皮尔逊相关系数(PCC)或均方误差(MSE)。但主观测试存在噪声,追求PCC=1.0或MSE=0.0既不现实也不可重复。本文在基本假设下推导出主观测试中可期待的PCC与MSE理论边界,其依赖于主观评分方差。当测试提供方差信息时,边界计算简便,称为“全数据驱动”。对于无方差信息的情况,提出两种替代方案:使用其他测试的方差数据,或采用主观评分的二项分布模型(BinoVotes),进而构建出名为BinoMOS的均值意见分模型。该模型自然捕获了MOS值的离散性及其随每文件投票数的变化特性,为边界计算提供所需方差信息。在18个主观测试数据上的对比显示,模型推导的边界与实测边界高度一致。该方法使研究者能在未提供方差信息的测试中,合理预估可能达到的PCC和MSE水平。
原文摘要 · Abstract (English)
Objective estimators of multimedia quality are often judged by comparing estimates with subjective "truth data," most often via Pearson correlation coefficient (PCC) or mean-squared error (MSE). But subjective test results contain noise, so striving for a PCC of 1.0 or an MSE of 0.0 is neither realistic nor repeatable. Numerous efforts have been made to acknowledge and appropriately accommodate subjective test noise in objective-subjective comparisons, typically resulting in new analysis frameworks and figures-of-merit. We take a different approach. By making only basic assumptions, we derive bounds on PCC and MSE that can be expected for a subjective test. Consistent with intuition, these bounds are functions of subjective vote variance. When a subjective test includes vote variance information, the calculation of the bounds is easy, and in this case we say the resulting bounds are "fully data-driven." We provide two options for calculating bounds in cases where vote variance information is not available. One option is to use vote variance information from other subjective tests that do provide such information, and the second option is to use a model for subjective votes. Thus we introduce a binomial-based model for subjective votes (BinoVotes) that naturally leads to a mean opinion score (MOS) model, named BinoMOS, with multiple unique desirable properties. BinoMOS reproduces the discrete nature of MOS values and its dependence on the number of votes per file. This modeling provides vote variance information required by the PCC and MSE bounds and we compare this modeling with data from 18 subjective tests. The modeling yields PCC and MSE bounds that agree very well with those found from the data directly. These results allow one to set expectations for the PCC and MSE that might be achieved for any subjective test, even those where vote variance information is not available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。