从用户感受出发,评估大模型在科研协作中的真实表现。
Research-Oriented Human-Centric Evaluation for Foundation Models
- 构建三维度人本评估框架,覆盖解题能力、信息质量与交互体验。
- 604次跨学科实验显示,用户偏好与模型表现存在显著差异。
- 大模型无法准确模仿人类主观判断,真人评估不可替代。
当前大模型评估多聚焦客观基准,如知识覆盖和推理准确性,常忽略人机协作中的主观体验。为此,我们提出面向研究的人本评估框架,从问题解决能力、信息质量与交互体验三个核心维度捕捉用户感知,提供细粒度的多模态科研场景评估方法。我们在多个学科开展604次人工评估会话,使用限时开放任务收集丰富主观反馈,揭示模型能力与用户偏好。此外,通过大模型作为评判者实验发现,即使先进模型也难以准确复现人类主观判断,凸显第一人称评估的不可替代性。项目链接:https://github.com/yijinguo/Human-Centric-Evaluation。
原文摘要 · Abstract (English)
Most current evaluations of foundation models focus on objective benchmarks, such as knowledge coverage and reasoning accuracy, often overlooking users' subjective experiences in human-AI collaboration. To address this gap, we propose a research-oriented Human-Centric Evaluation framework. It captures user perceptions across three core dimensions: problem-solving ability, information quality, and interaction experience, providing a structured, fine-grained approach to understanding how users evaluate and respond to model behavior in multi-modal research contexts. We conduct 604 human evaluation sessions across various disciplines, involving recent advanced foundation models. Through open-ended, time-limited collaborative tasks, we gather rich subjective assessments that highlight model capabilities and user preferences. Additionally, we perform an LLM-as-a-judge experiment and find that even sophisticated models struggle to accurately replicate human subjective judgment, emphasizing the irreplaceable value of first-person human assessment. Our project link is https://github.com/yijinguo/Human-Centric-Evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。