arXiv:2509.11206cs.HCcs.AI2025-09被引 2

让大模型评估从模糊打分变成可追溯的细粒度分析

Evalet: Evaluating Large Language Models through Functional Fragmentation

  • 将生成内容拆解为功能片段,分析其在评估中的作用
  • 用户研究显示能发现48%更多评估偏差
  • 适合需要深入理解模型输出问题的研究者与工程师

实践者越来越多地使用大语言模型(LLMs)通过“大模型作为评判者”方法来评估生成式AI的输出。然而,这些方法仅给出整体评分,掩盖了具体影响评估的因素。我们提出功能碎片化方法,将每个输出分解为关键片段,并解析每个片段相对于评估标准所承担的修辞功能——揭示出值得关注的元素及其如何满足或阻碍用户目标。我们实现了Evalet系统,通过可视化多个输出的片段级功能,支持对评估结果的检查、评分和比较。一项用户研究(N=10)发现,尽管从业者难以验证整体评分,但本方法帮助他们识别出48%更多的评估不一致问题。这增强了他们对大模型评估结果的信任,并使其能更有效地发现模型输出中的可操作性问题。我们的工作推动大模型评估从量化分数转向对模型行为的定性、细粒度分析。

原文摘要 · Abstract (English)

Practitioners increasingly rely on Large Language Models (LLMs) to evaluate generative AI outputs through "LLM-as-a-Judge" approaches. However, these methods produce holistic scores that obscure which specific elements influenced the assessments. We propose functional fragmentation, a method that dissects each output into key fragments and interprets the rhetoric functions that each fragment serves relative to evaluation criteria -- surfacing the elements of interest and revealing how they fulfill or hinder user goals. We instantiate this approach in Evalet, an interactive system that visualizes fragment-level functions across many outputs to support inspection, rating, and comparison of evaluations. A user study (N=10) found that, while practitioners struggled to validate holistic scores, our approach helped them identify 48% more evaluation misalignments. This helped them calibrate trust in LLM evaluations and rely on them to find more actionable issues in model outputs. Our work shifts LLM evaluation from quantitative scores toward qualitative, fine-grained analysis of model behavior.

大模型评估细粒度分析交互式系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。