揭示人类标注差异对多模态模型评估的影响,发现大模型在主观任务中未必更优。
Rethinking Ground Truth: A Case Study on Human Label Variation in MLLM Benchmarking
- 设计新评估协议,区分人类标注一致与不一致场景
- 大模型在高一致标注上表现好,但在分歧大时反而不如小模型
- 提醒研究者:仅用共识标签会高估模型能力,适合内容审核场景
人类标注差异(HLV)在多模态大模型(MLLM)基准测试中长期被忽视。本文提出一种新评估协议,明确考虑人类标注的一致性与分歧性两种条件。针对Gemma 3和Qwen 2.5 VL两个先进MLLM家族,在社交媒体内容分类数据集的非聚合人工标注上进行测试。结果表明:大模型在高一致性子集上表现更优,但在人类意见分歧较高时,常弱于中等规模模型,说明参数量无法决定模型对模糊性和主观性的敏感度。这表明仅依赖共识标签的基准会夸大模型在主观领域的能力,纳入标注差异可提供更真实、稳健的评估,尤其适用于内容审核流程。
原文摘要 · Abstract (English)
Human Label Variation (HLV), i.e. systematic differences among annotators' judgments, remains underexplored in benchmarks despite rapid progress in large language model (LLM) development. We address this gap by introducing an evaluation protocol for multimodal large language model (MLLM) benchmarking that explicitly accounts for two conditions: (1) human label agreement and (2) disagreement. We apply this protocol to two state-of-the-art MLLM families (Gemma 3, Qwen 2.5 VL) using non-aggregated human annotations from a social media content classification dataset. Across tasks, we find that larger models tend to perform best on high-agreement subsets, yet often underperform medium-sized models when human disagreement is high, indicating that parameter count alone does not determine sensitivity to ambiguity and subjectivity. These results show that benchmarks based solely on consensus labels can overstate model capabilities in such domains and that incorporating human label variation yields more realistic and robust assessments of MLLMs in content moderation pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。