arXiv:2602.21059cs.HCcs.CL2026-02中稿 · ACM CHI conference…被引 3

为评估学术问答中大模型错误,构建了专家级评价框架。

An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems

  • 基于科学家实际评估习惯,提炼20类错误模式。
  • 通过10位专家验证,发现框架能识别遗漏问题。
  • 适合需要严谨评估的科研人员与评测工具开发者。

大语言模型正在改变学术搜索与摘要等任务,但其可靠性仍存疑。现有评估方法多依赖自动化指标,虽高效可扩展,却缺乏上下文敏感性,无法反映科学家真实评估方式。本文与领域专家合作,通过分析68组问答对,采用主题分析法识别出7类共20种错误模式,并在10位额外科学家中进行情境化验证,证实该框架不仅能捕捉专家自然识别的错误,还能帮助发现此前忽略的问题。科学家常使用技术精度检验、价值判断及自我评估等系统性策略。研究探讨了支持专家评估的前景,提出个性化、基于框架的工具,可适配不同专家的评估习惯与专业水平。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are transforming scholarly tasks like search and summarization, but their reliability remains uncertain. Current evaluation metrics for testing LLM reliability are primarily automated approaches that prioritize efficiency and scalability, but lack contextual nuance and fail to reflect how scientific domain experts assess LLM outputs in practice. We developed and validated a schema for evaluating LLM errors in scholarly question-answering systems that reflects the assessment strategies of practicing scientists. In collaboration with domain experts, we identified 20 error patterns across seven categories through thematic analysis of 68 question-answer pairs. We validated this schema through contextual inquiries with 10 additional scientists, which showed not only which errors experts naturally identify but also how structured evaluation schemas can help them detect previously overlooked issues. Domain experts use systematic assessment strategies, including technical precision testing, value-based evaluation, and meta-evaluation of their own practices. We discuss implications for supporting expert evaluation of LLM outputs, including opportunities for personalized, schema-driven tools that adapt to individual evaluation patterns and expertise levels.

大模型评估学术问答专家系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。