传统评分方法已不足,新模型需懂上下文、能推理、跨模态。
Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality
- 引入上下文感知、可解释推理和多模态对齐能力
- 主张用带背景信息与理由的新数据集替代单一分数评价
- 适合关注模型可信度与人类对齐的研究者
本文指出,尽管均值意见分(MOS)曾是多媒体质量评估的基础,但其将丰富、情境敏感的人类判断简化为单一标量,掩盖了语义失败、用户意图及判断依据。现代质量评估模型必须整合三项相互关联的能力:(1) 上下文感知,以适应任务目标与观看条件;(2) 推理能力,生成基于证据的可解释判断理由;(3) 多模态融合,利用视觉-语言模型对齐感知与语义线索。文章批判当前以MOS为中心的基准局限性,提出改革路线:构建含上下文元数据与专家理由的更丰富数据集,开发评估语义一致性、推理真实性和上下文敏感性的新指标。通过将质量评估重构为情境化、可解释、多模态建模任务,推动建立更鲁棒、符合人类认知且可信的评估体系。
原文摘要 · Abstract (English)
This position paper argues that Mean Opinion Score (MOS), while historically foundational, is no longer sufficient as the sole supervisory signal for multimedia quality assessment models. MOS reduces rich, context-sensitive human judgments to a single scalar, obscuring semantic failures, user intent, and the rationale behind quality decisions. We contend that modern quality assessment models must integrate three interdependent capabilities: (1) context-awareness, to adapt evaluations to task-specific goals and viewing conditions; (2) reasoning, to produce interpretable, evidence-grounded justifications for quality judgments; and (3) multimodality, to align perceptual and semantic cues using vision-language models. We critique the limitations of current MOS-centric benchmarks and propose a roadmap for reform: richer datasets with contextual metadata and expert rationales, and new evaluation metrics that assess semantic alignment, reasoning fidelity, and contextual sensitivity. By reframing quality assessment as a contextual, explainable, and multimodal modeling task, we aim to catalyze a shift toward more robust, human-aligned, and trustworthy evaluation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。