发现用户提问方式影响模型回答,提出可量化的表达风格测量方法
How You Ask Shapes What You Get: A Theory-Seeded Measurement of Articulation in Advice-Seeking LLM Conversations

- 从1.6万条求助对话中提取可解释的表达风格因子
- 发现六分之一提问者用长但信息少的方式提问,模型回应更短且不追问
- 提出按表达风格分层评估的基准设计新思路,适合模型评估研究者
用户以不同方式提出相同的求助请求:有的详细说明约束条件,有的仅模糊表达需求。以往研究将这种差异视为需平均掉的噪声;本文将其视为输入分布中稳定可测的结构。我们探究表达方式(如何提问)是否形成与主题(问什么)分离的潜在维度,并与语言模型响应相关联。从公开聊天数据集(WildChat、LMSYS、ShareChat)收集的16,447条求助提示中提取可解释特征,复现了少数几组跨训练/测试划分及跨语料库稳定的表达风格因子。由于该结构基本独立于主题,其所定义群体跨越主题,常规主题或任务评估难以察觉。这些因子揭示了几种常见表达风格,其中一种突出:约六分之一最大语料中的提示为长篇但信息贫乏,对应模型生成更短、更模糊的回答,且不主动要求澄清——尽管正是这种不明确才应触发追问。该对比在每个主题组和长度分位数内均成立,且非仅因不明确本身:另一种同样不明确的风格却会引发澄清问题。两名独立人工标注者复现了这一差异。我们认为评测应按表达风格分层,并提供所提取结构作为测量工具。
原文摘要 · Abstract (English)
Users articulate the same advice-seeking request in different ways: some specify detailed constraints, others gesture at a vague need. Prior work treats this variation as noise to be averaged away; we instead treat it as a stable, measurable structure in the input distribution. We ask whether articulation (how people ask) forms latent dimensions separable from topic (what they ask about), and whether it is associated with how language models respond. We extract interpretable features from 16,447 advice-seeking prompts pooled from public chat corpora (WildChat, LMSYS, and ShareChat) and recover a small set of latent articulation factors that replicate across train/test splits and across corpora. Because this structure is largely separable from topic, the populations it defines cut across topics and stay invisible to topic- or task-based evaluation. The factors define a handful of recurring articulation styles, one of which stands out: a long-form but information-poor style, roughly one in six prompts in the largest corpus, where models return shorter, vaguer answers and do not ask for clarification even though under-specification is exactly the condition that warrants it. The contrast holds within every topic group and length quintile, and is not under-specification alone -- a second, equally under-specified style does draw clarifying questions. Two independent human annotators reproduce this contrast. We argue that benchmarks should stratify on articulation, and we offer the extracted structure as a measurement instrument for doing so.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。