arXiv:2509.26205cs.AI2025-09被引 1

为大模型生成结果设计了人本评估框架,提升AI与人类协作质量。

Human-Centered Evaluation of RAG outputs: a framework and questionnaire for human-AI collaboration

  • 基于实用维度框架设计12维问卷,聚焦用户意图与信息可验证性。
  • 人类与大模型在数值评分上一致性差,但模型解释有助于辅助判断。
  • 优化后的问卷更关注文本结构和真实场景需求,适合人机协同研究。

检索增强生成(RAG)系统在面向用户的场景中日益普及,但其输出的人本化评估仍缺乏系统性方法。本文基于Gienapp的实用维度框架,构建了一套以人类为中心的评估问卷,涵盖12个评估维度。通过多轮对查询-生成结果对的评分及语义讨论,持续迭代优化问卷,并结合人类评估员与人-大模型协作组的反馈。结果显示,尽管大语言模型(LLMs)在识别度量描述和尺度标签方面表现稳定,但在检测文本格式变化时存在不足;而人类评估者难以严格聚焦于度量描述与标签。虽然模型提供的评分与解释具有参考价值,但其数值评分与人类评分之间缺乏一致性。最终问卷扩展了初始框架,强化了对用户意图、文本组织结构和信息可验证性的关注。

原文摘要 · Abstract (English)

Retrieval-augmented generation (RAG) systems are increasingly deployed in user-facing applications, yet systematic, human-centered evaluation of their outputs remains underexplored. Building on Gienapp's utility-dimension framework, we designed a human-centred questionnaire that assesses RAG outputs across 12 dimensions. We iteratively refined the questionnaire through several rounds of ratings on a set of query-output pairs and semantic discussions. Ultimately, we incorporated feedback from both a human rater and a human-LLM pair. Results indicate that while large language models (LLMs) reliably focus on metric descriptions and scale labels, they exhibit weaknesses in detecting textual format variations. Humans struggled to focus strictly on metric descriptions and labels. LLM ratings and explanations were viewed as a helpful support, but numeric LLM and human ratings lacked agreement. The final questionnaire extends the initial framework by focusing on user intent, text structuring, and information verifiability.

人机协作评估框架RAGLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。