arXiv:2506.05142cs.CL2025-06

对比大模型与人类对文本错误严重性的判断,发现多数模型在颜色错误上误判严重性。

Do Large Language Models Judge Error Severity Like Humans?

  • 构建图文双模态实验框架,测试四类语义错误的严重性评估
  • 人类认为颜色和类型错误更严重,而多数大模型高估颜色错误
  • 仅通义千问(Doubao)接近人类判断,但未完全匹配;深思-V3表现最佳

大型语言模型(LLMs)被广泛用于自然语言生成的自动评估,但其是否能准确模拟人类对错误严重性的判断仍不明确。本研究系统比较了人类与大模型在含受控语义错误的图像描述中的判断差异。扩展van Miltenburg等人(2020)的实验框架,涵盖单模态(仅文本)与多模态(文本+图像)场景,评估年龄、性别、服装类型和颜色四类错误。结果表明,人类对不同错误类型的严重性判断存在差异,视觉上下文显著增强了颜色和类型错误的感知严重性。值得注意的是,多数大模型对性别错误打分偏低,却对颜色错误打分过高,与人类判断相反——人类认为两者均严重,但原因不同。这暗示模型可能内化了社会规范影响性别判断,却缺乏对颜色的感知基础。仅一个模型(Doubao)接近人类排序,但仍无法清晰区分错误类型。令人意外的是,单模态模型DeepSeek-V3在双模态条件下均达到最高人类对齐度,优于现有主流多模态模型。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used as automated evaluators in natural language generation, yet it remains unclear whether they can accurately replicate human judgments of error severity. In this study, we systematically compare human and LLM assessments of image descriptions containing controlled semantic errors. We extend the experimental framework of van Miltenburg et al. (2020) to both unimodal (text-only) and multimodal (text + image) settings, evaluating four error types: age, gender, clothing type, and clothing colour. Our findings reveal that humans assign varying levels of severity to different error types, with visual context significantly amplifying perceived severity for colour and type errors. Notably, most LLMs assign low scores to gender errors but disproportionately high scores to colour errors, unlike humans, who judge both as highly severe but for different reasons. This suggests that these models may have internalised social norms influencing gender judgments but lack the perceptual grounding to emulate human sensitivity to colour, which is shaped by distinct neural mechanisms. Only one of the evaluated LLMs, Doubao, replicates the human-like ranking of error severity, but it fails to distinguish between error types as clearly as humans. Surprisingly, DeepSeek-V3, a unimodal LLM, achieves the highest alignment with human judgments across both unimodal and multimodal conditions, outperforming even state-of-the-art multimodal models.

大模型评估错误判断多模态人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。