把用户凭感觉测试大模型的方法变成可复现的评估流程
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
- 将用户自定义测试内容和评价标准作为核心机制
- 实验证明个性化测试能改变模型优劣排序
- 适合关注真实使用体验而非分数的开发者
评估大模型难度高,因为基准分数常无法反映其实际用途。用户更依赖‘感觉测试’:基于自身工作流的非正式体验,如在编码任务中对比模型表现。尽管普遍,这种测试往往随意且难以系统分析或复现。本文通过分析两项实证数据——用户评估习惯调查及博客与社交媒体中的真实模型对比报告——将感觉测试形式化为两步过程:用户个性化测试内容与评价标准。随后提出一个概念验证评估流水线,生成个性化提示并用用户感知的主观标准比较模型输出。在编码基准测试中,结合个性化提示与用户感知评价可改变模型偏好结果,揭示了感觉测试的实际作用。研究建议,形式化的感觉测试可成为连接基准分数与真实体验的有效方法。
原文摘要 · Abstract (English)
Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness. Instead, users often rely on ``vibe-testing'': informal experience-based evaluation, such as comparing models on coding tasks related to their own workflow. While prevalent, vibe-testing is often too ad hoc and unstructured to analyze or reproduce at scale. In this work, we study how vibe-testing works in practice and then formalize it to support systematic analysis. We first analyze two empirical resources: (1) a survey of user evaluation practices, and (2) a collection of in-the-wild model comparison reports from blogs and social media. Based on these resources, we formalize vibe-testing as a two-part process: users personalize both what they test and how they judge responses. We then introduce a proof-of-concept evaluation pipeline that follows this formulation by generating personalized prompts and comparing model outputs using user-aware subjective criteria. In experiments on coding benchmarks, we find that combining personalized prompts and user-aware evaluation can change which model is preferred, reflecting the role of vibe-testing in practice. These findings suggest that formalized vibe-testing can serve as a useful approach for bridging benchmark scores and real-world experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。