将对话记录与用户目标、结果反馈绑定,让LLM评估更贴近真实使用。
MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models

- 把完整对话、用户目标和任务成果一起记录,作为核心评估信号。
- 26人用两周测试,23.1%的任务失败率,多轮对话失败率是单轮的2.5倍。
- 适合关注真实用户体验、评估效果落地的研究者和开发者。
评测套件在控制任务中评估模型能力;大规模对话语料库捕捉自然使用场景但缺乏用户反馈;界面内反馈机制记录满意度但无任务目标。三者共同留下关键空白:现有基础设施无法常规关联交互轨迹与用户定义的目标成果。我们提出MonitrLLM,一个开源的社区中心化大语言模型评估框架,将完整对话转录、用户报告的任务意图和结果评估三者统一为首要评估信号,而非可选元数据。为验证该方法价值,我们开展了为期两周的可行性试点,26名大学生使用ChatGPT,收集了206份包含完整对话记录的评估报告。结果显示,尽管平均满意度高达4.19/5,但用户任务失败率仍达23.1%;多轮对话失败率是单轮的2.5倍,表明长时间交互更多反映困难而非参与度。我们进一步讨论结合直接用户反馈与观察数据对鲁棒评估的意义,以及实现这一目标的基础设施可能性。
原文摘要 · Abstract (English)
Benchmark suites assess model capability on controlled tasks; large-scale conversation corpora capture naturalistic use without user feedback; and in-interface feedback mechanisms record satisfaction without task purpose. Together, they leave a critical gap in LLM evaluation: no existing infrastructure routinely links interaction trajectories to user-defined outcomes. We introduce MonitrLLM, open-source infrastructure for community-centered LLM evaluations that links full conversation transcripts to user-reported task intent and outcome assessments, treating all three as primary evaluative signals rather than optional metadata. To demonstrate the value of this approach, we conducted a two-week feasibility pilot with 26 college students using ChatGPT, collecting 206 evaluation reports with full conversation transcripts. The findings from our pilot demonstrate the value of connecting conversation trajectories with user-reported outcomes. For instance, despite reporting high average satisfaction (4.19/5) with their LLM interactions, participants also experience a substantial 23.1% failure rate on their goal tasks. We also find that multi-turn conversations are reported as failing at 2.5 times the rate of single-turn exchanges, a pattern that reframes extended interaction as a signal of difficulty rather than engagement. We conclude by discussing the value of incorporating direct user feedback with observational data for robust LLM evaluations, and the possibilities for infrastructure that enables this goal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。