arXiv:2608.15828cs.CLcs.AI2026-08中稿 · INLG 2026

提出六维框架,解析隐喻解释质量的多维度结构。

A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations

论文配图:A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations
图 1 · 摘自论文原文
  • 基于认知理论拆解隐喻解释质量为六个维度
  • 11200次标注显示人类判断存在系统性分歧
  • 自动评估可部分还原人类判断结构,适合优化

当前隐喻解释评价主要依赖整体质量评分,难以揭示评价结构及人类判断的一致性与差异。本文提出一个基于认知动机的多维框架,将解释质量分解为六个理论驱动的维度。在一项密集标注研究(11,200次评分)中发现:(i) 解释质量确为多维结构;(ii) 标注者分歧具有系统性而非随机性;(iii) 六个维度可聚为一个共享簇和两个独立判断轴。探索性可行性研究进一步表明,标准自动评估流程可部分复现该结构,能较好预测最具区分性的维度,其误差与人类分歧呈相关性。结果表明,多维评价比整体评分提供更丰富的诊断信息,且开放生成任务的自动评估应以能否保留人类判断结构作为衡量标准。

原文摘要 · Abstract (English)

Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is structured or where human judgments agree and diverge. We introduce a cognitively motivated framework that decomposes metaphor explanation quality into six theoretically grounded dimensions. In a dense annotation study (11,200 ratings), we find that: {\bfseries(i)} explanation quality is genuinely multidimensional; {\bfseries(ii)} annotator disagreement is systematic rather than random; and {\bfseries(iii)} the six dimensions collapse into a shared cluster and two independent axes of judgment. An exploratory feasibility study further shows that a standard automatic evaluation pipeline can recover parts of this structure, predicting the most discriminative dimensions well while its errors correlate human (dis)agreement. Together, these results suggest that multidimensional evaluation offers richer diagnostic insight than holistic ratings, and that automatic evaluators for open-ended generation tasks should be judged on how well they preserve the structure of human judgment.

隐喻理解多维评价自动评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。