实证发现离线评估无法反映大模型在真实对话中的个性化行为差异。
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
- 通过对比离线测试与800名用户的真实对话数据,揭示模型响应受上下文影响
- 相同问题在不同用户会话中产生显著不同的回答结果
- 提醒研究者关注个性化对模型评估的干扰,适合评估方法设计者参考
标准的离线语言模型评估——即模型进行一系列独立、无状态的推理——无法捕捉大模型在实际使用中的表现,因为个性化会从根本上改变模型行为。例如,同一个基准问题在同一个模型上,可能在无状态系统中、某个用户的聊天会话中或另一个用户的聊天会话中产生截然不同的回答。本文通过实证研究展示了这一现象:对比了800名ChatGPT和Gemini真实用户在实际对话中提出基准问题和其他问题的表现,与离线评估结果进行分析,证实个性化显著影响输出一致性。
原文摘要 · Abstract (English)
Standard offline evaluations for language models -- a series of independent, state-less inferences made by models -- fail to capture how language models actually behave in practice, where personalization fundamentally alters model behavior. For instance, identical benchmark questions to the same language model can produce markedly different responses when prompted to a state-less system, in one user's chat session, or in a different user's chat session. In this work, we provide empirical evidence showcasing this phenomenon by comparing offline evaluations to field evaluations conducted by having 800 real users of ChatGPT and Gemini pose benchmark and other provided questions to their chat interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。