arXiv:2509.02855cs.CLcs.CY2025-09Conference of the …

评估大模型生成的开放性观点与专家意见的相似性,提出可复用的评测框架

IDEAlign: Comparing Ideas of Large Language Models to Domain Expert

  • 通过'找不同'任务模拟专家判断,构建人类标注基准
  • 多数相似度方法表现不佳,LLM作为裁判仅提升11~18%
  • 适合教育领域开放评语评估,助力大模型负责任部署

大语言模型(LLMs)被越来越多用于生成开放式、解释性注释,但目前缺乏经过验证且可扩展的观点级相似性度量。本文(一)将大模型注释的内容评估确立为核心、未被充分研究的任务;(二)提出IDEAlign方法,通过‘找不同’任务捕捉专家相似性判断;(三)将多种相似性方法(文本嵌入、主题模型、LLM作为评判者)与人工评分进行对比。在两个真实教育数据集(如数学推理解读与反馈生成)上的应用表明,大多数度量方法无法捕捉专家认为有意义的细微相似维度。虽然以大模型为裁判的方法表现最佳(相比其他方法提升11~18%),但仍未能达到专家水平,因此适合作为筛选工具而非人工审核替代品。本研究揭示了大规模评估开放性大模型注释的难度,并将IDEAlign定位为该任务的可复用基准协议,以指导大模型负责任地部署。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to produce open-ended, interpretive annotations, yet there is no validated, scalable measure of idea-level similarity to expert annotations. We (i) introduce the content evaluation of LLM annotations as a core, understudied task, (ii) propose IDEAlign for capturing expert similarity judgments via pick-the-odd-one-out tasks, and (iii) benchmark various similarity methods (text embeddings, topic models, and LLM-as-a-judge) against these human ratings. Applying this approach to two real-world educational datasets (e.g., interpreting math reasoning and feedback generation), we find that most metrics fail to capture the nuanced dimensions of similarity meaningful to experts. LLM-as-a-judge performs best (11~18% improvement over other methods) but still falls short of expert alignment, making it useful as a triage tool rather than a substitute for human review. Our work demonstrates the difficulty of evaluating open-ended LLM annotations at scale, and positions IDEAlign as a reusable protocol for benchmarking on this task to help guide responsible deployment of LLMs.

大模型评估专家对齐教育应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。