arXiv:2607.22392cs.IR2026-07被引 1

对话中最后的提示词无法代表用户全部意图,常遗漏历史信息。

The Prompt Is Not the Query: How Request State Evolves Across Multi-Turn AI Conversations

  • 用可观察的请求状态替代隐含意图,分析多轮对话中用户需求演变。
  • 最终提示仅包含36%左右的对话词汇,超半数对话中遗漏关键信息维度。
  • 最终提示常新增未出现过的请求维度,说明它仍是动态更新而非总结。

AI搜索评估通常将提示视为稳定查询,可独立计数、分类与重放。但对话中每一轮都可能添加约束、修正假设、要求证据或引用早期内容,使该分析单位失效。本文引入可观察的‘对话条件请求状态’替代隐含意图,并测量其在用户各轮中的分布。研究复用先前对话研究的冻结规则与受控样本:670条英文商业多轮对话(发现-验证设计)和来自1,389名参与者的7,463条公开PRISM对话。在商业语料中,最终提示包含会话唯一用户侧词汇的中位数为35.6%;在PRISM中为36.4%。在68.4%和74.3%的对话中,最终提示所含词汇不超过一半。更重要的是,透明规则在历史中检测到至少一个请求状态维度,但在最终提示中未检测到的情况分别占50.3%和44.8%。在携带维度的对话中,最终提示完整再现所有维度的比例仅为26.1%和26.2%。同时,最终提示在17.9%和19.3%的对话中引入了此前未见的维度,表明终点并非总结或参考,而是又一次状态更新。长度匹配的零模型对照显示,词汇覆盖率低主要源于回合长度,因此结果反映信息可用性而非语义漂移。类别结果支持以会话级为单位进行AI搜索评估,但不估计历史对模型回答的因果影响。

原文摘要 · Abstract (English)

AI-search evaluation commonly treats a prompt as a stable query that can be counted, classified, and replayed in isolation. A conversation makes that unit of analysis questionable: each user turn can add a constraint, revise an assumption, request evidence, or refer to alternatives established earlier. We replace latent "intent" with an observable construct, conversation-conditioned request state, and measure how that state is distributed across user turns. The analysis reuses frozen rules and the governed cohort of a preceding conversation study: 670 English commercial multi-turn conversations in a discovery-replication design and 7,463 public PRISM conversations from 1,389 participants. In the commercial corpus, the final prompt contains a median 35.6% of the session's unique user-side content vocabulary; in PRISM, the median is 36.4%. The final prompt contains at most half of that vocabulary in 68.4% and 74.3% of conversations, respectively. More importantly, transparent rules detect at least one request-state dimension in history but not in the final prompt in 50.3% of commercial conversations and 44.8% of PRISM conversations. Among dimension-bearing conversations, the final prompt reproduces the full observed dimension set in only 26.1% and 26.2%. At the same time, the final prompt adds a previously unseen dimension in 17.9% and 19.3%, showing that the endpoint is neither a summary nor merely a reference: it is often another state update. Length-matched nulls show that low lexical coverage is largely a consequence of turn length, so vocabulary results are interpreted as information availability, not semantic drift. The categorical results support session-level measurement for AI search. They do not estimate the causal effect of history on model answers.

对话系统请求状态多轮交互提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。