arXiv:2608.02556cs.IR2026-08

对话上下文对AI回答影响显著,单独看最后一条消息会误导判断。

Beyond the Final Prompt: Measuring the Effect of Within-Conversation Context on AI Answers

  • 用完整对话、仅末句、压缩前缀三组对比,验证上下文作用
  • 完整对话与孤立末句在44.7%案例中产生实质性差异,满意度高0.49分
  • 压缩前缀可减少差异,但仍有近1/3回答仍不同,适合关注对话理解的研究者

评估中常将孤立的最终用户消息视为查询,但在对话中有效请求可能分散在多轮交互中。本研究从受控商业语料库和公开PRISM数据集抽取180个英文多轮对话,固定最终消息与模型,生成三种回答:完整角色标注对话、仅末句、末句加不超过160词的前缀重建。由独立评分模型在随机标签下评估,主要指标为是否产生用户行为改变的实质性差异。经逆概率加权后,完整对话与孤立末句答案在44.7%案例中存在实质性差异(95%置信区间33.8%至56.1%),完整对话满意度高0.49分(0.32至0.67)。加入压缩前缀后,实质性差异率降至30.8%(20.2%至42.1%),满意度差距缩小至0.01分(-0.12至0.13)。顺序重排测试显示48例一致性达91.7%,肯德尔系数为0.83。研究聚焦单次对话内的上下文,不涉及跨对话记忆。

原文摘要 · Abstract (English)

An isolated final user message is often treated as the query in evaluations of AI systems. In a conversation, however, the actionable request may be distributed across preceding turns. We directly test whether that omitted within-conversation context changes answers. For each of 180 English multi-turn conversations sampled from a governed commercial corpus and the public PRISM dataset, we hold the final user message and requested answer model constant while generating three answers: one from the full role-labelled conversation, one from the final message alone, and one from the final message plus a prefix-only reconstruction capped at 160 words. A separately requested judge model evaluates answers under randomized labels. The prespecified primary endpoint is a material difference that could change what the user does, rather than a difference in style or detail. After inverse-probability weighting to the eligible cohorts, the full-conversation and isolated-final answers differ materially in 44.7% of cases (95% bootstrap CI 33.8% to 56.1%). Full-conversation answers score 0.49 points higher on a 0 to 4 request-satisfaction scale (0.32 to 0.67). Adding the compressed prefix reduces the material-difference rate to 30.8% (20.2% to 42.1%), a 13.9-point reduction (4.9% to 24.1%), and reduces the mean satisfaction gap to 0.01 points (-0.12 to 0.13). Yet compression is not equivalent to the complete dialogue context: almost one third of answers remain materially different. An order-swapped repeat on 48 cases yields 91.7% agreement and kappa = 0.83 for the primary decision. The study concerns preceding turns in the same conversation and does not test persistent memory across separate conversations.

对话理解评估方法上下文影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。