提出对抗性测试方法评估图像质量评分中的异常值检测效果
Robustness and accuracy of mean opinion scores with hard and soft outlier detection
- 用进化优化生成对抗样本,模拟最坏情况下的异常评分
- 验证多种方法在极端干扰下的表现差异,发现部分方法易被攻破
- 设计两种低复杂度高鲁棒的新算法,适合实际应用
在图像和视频质量的主观评价中,观察者对选定刺激进行评分或比较。在计算这些刺激的平均意见分数(MOS)前,建议识别并处理可能提供不可靠评分的异常值。目前已有多种方法用于此目的,部分已标准化,通常基于统计学,并通过引入人工异常评分(如随机点击者)进行测试。然而,缺乏可靠且全面的方法来对比分析异常值检测算法的性能。为填补这一空白,本文提出并应用一种经验性最坏情况分析作为通用解决方案。该方法通过进化优化对异常值检测算法进行黑盒对抗攻击,使评分尺度值相对于真实值产生最大扭曲。我们将该分析应用于绝对类别评分中的若干硬性和软性异常值检测方法,展示了它们在压力测试中的不同表现。此外,我们提出了两种低复杂度、最坏情况表现优异的新异常值检测方法。对抗攻击与数据分析的软件已公开。
原文摘要 · Abstract (English)
In subjective assessment of image and video quality, observers rate or compare selected stimuli. Before calculating the mean opinion scores (MOS) for these stimuli from the ratings, it is recommended to identify and deal with outliers that may have given unreliable ratings. Several methods are available for this purpose, some of which have been standardized. These methods are typically based on statistics and sometimes tested by introducing synthetic ratings from artificial outliers, such as random clickers. However, a reliable and comprehensive approach is lacking for comparative performance analysis of outlier detection methods. To fill this gap, this work proposes and applies an empirical worst-case analysis as a general solution. Our method involves evolutionary optimization of an adversarial black-box attack on outlier detection algorithms, where the adversary maximizes the distortion of scale values with respect to ground truth. We apply our analysis to several hard and soft outlier detection methods for absolute category ratings and show their differing performance in this stress test. In addition, we propose two new outlier detection methods with low complexity and excellent worst-case performance. Software for adversarial attacks and data analysis is available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。