测试大模型能否按复杂度生成多版本回答,发现效果普遍不稳定。
Explain Like I'm 5 or Whatever I Choose: Evaluating the Interactive Potential of Language Model Responses

- 设计新评估框架,让模型对同一问题生成不同语言复杂度的回答。
- 4个模型在98个科学问题上表现不佳,最佳仅46%正确调整复杂度。
- 适合关注人机交互与可解释性评估的研究者参考。
大型语言模型在科学信息检索任务中的评估正转向以用户为中心,如进行实时或多轮交互评估。然而,现有评估仍假设使用单一静态聊天界面,而随着模型集成到新界面,评估需引入界面特异性标准。我们基于16名参与者的形成性研究,提出一种新评估框架:测试模型对同一查询生成沿语言复杂度轴可解释的多个回答,灵感来自以人为本设计中的直接操作界面。我们评估了GPT-5.1、GPT-5 mini、Claude Sonnet 4.5 + Thinking和DeepSeek-V3.1,针对98个科学问题生成5个不同复杂度的回答。尽管各模型能变化复杂度,但多数变化不一致,表现最佳的Claude Sonnet 4.5仅在46%的情况下正确方向地调整复杂度。该结果在更大样本量和不同复杂度层级下依然成立。
原文摘要 · Abstract (English)
Evaluations of large language models (LLMs) in scientific information seeking tasks have become increasingly use-centric, such as conducting live or multi-turn evaluations with real users. These evaluations still assume a single, static chat interface, but as models are integrated into new interfaces, evaluations must shift to incorporate interface-specific criteria. We propose a new evaluation framework based on a formative study with $16$ participants that tests models' ability to generate multiple responses to one query that differ along an interpretable axis of language (language complexity), inspired by direct manipulation interfaces from human-centered design literature. We evaluate GPT-5.1, GPT-5 mini, Claude Sonnet 4.5 + Thinking, and DeepSeek-V3.1 by generating 5 responses at different levels of language complexity for $98$ scientific queries. While models vary complexity across responses, most changes remain inconsistent, with the best performing model (Claude Sonnet 4.5) only shifting reliable complexity measures in the correct direction $46\%$ of the time. Our findings hold with increased sample size and alternative complexity levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。