模型更信任工具生成的结果,而非纯文本,但文本也能诱导错误采纳。
Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5

- 测试语言模型在虚假信息中对工具结果与纯文本的信任差异。
- 工具结果使错误代码采纳率达14/24,纯文本达60/60,但差距随实验条件变化。
- 结果提示工具输出虽有影响,但并非不可替代,关键在呈现方式。
语言模型越来越多地从自身生成的内容中读取信息,导致之前写入的声明可能被误当作检索结果。我们测试了在合成查询任务中,携带未经验证声明的消息包是否会影响模型的回答。使用Claude Opus 5,在无目标声明时,错误代码采纳率为0/24;当助手先前声明目标时为0/22;当工具结果记录命名目标时为14/24;当该结果带有十字段元数据包装且未标记验证状态时为15/24。工具结果在11/12个支持性试验中采纳正确代码,但在14/24个非支持性试验中仍采纳错误代码,排除了固定输出令牌偏差的影响。一项注册过的重复实验确认了工具结果与助手声明之间的差距(7/24 对 0/24),单侧Fisher精确检验p = 0.0047。然而,跨四天进行的两次实验中,工具结果采纳率从14/24降至7/24。第二项注册研究加入实时文本对照:两者均提前公布并置于同一用户回合末尾,仅交换绑定关系。结果显示,纯文本条件下错误代码采纳率为60/60,工具结果为57/60,原“结果优先”优势不成立(p = 1)。研究结论表明,工具结果确有影响,但其效果并非不可替代,本实验未发现结果包比宣布的内联文本具有更大行为权重。研究基于单一模型、一种合成任务模板及一个API接口。
原文摘要 · Abstract (English)
Language-model systems increasingly read from stores they also write to, so a claim that was merely written earlier can return looking retrieved. We tested whether the message package carrying an unsupported assignment changes which answer a model gives in a synthetic lookup task. Claude Opus 5 selected a color code for a named item or abstained. In an exploratory four-arm study, false-code adoption was 0/24 with no target claim, 0/22 scorable trials when a prior assistant assertion named the target, 14/24 when a tool-result record named it, and 15/24 when that result used a ten-field metadata wrapper that marked it unchecked. The tool-result arm selected the record's code in 11/12 supported trials and 14/24 unsupported trials, ruling out a fixed output-token bias while leaving substantial planted-token heterogeneity. A document-preregistered replication reproduced the tool-result versus assistant-assertion gap, 7/24 against 0/24, one-sided Fisher exact p = 0.0047. The tool-result rate nevertheless fell from 14/24 to 7/24 across runs made four days apart. A second preregistered study gave the earlier comparison a live text control: both records were announced in advance and placed in the same final user turn, then target binding was swapped between the linked tool result and later inline JSON. Inline text was sufficient for false-code adoption in 60/60 trials; the tool-result condition produced 57/60, so the registered result-first superiority criterion failed, p = 1. The result does not show that tool results have no effect. It shows that native tool-result placement was not necessary and that this experiment did not find greater behavioral weight for the result package than for announced inline text. The findings concern a single model on one synthetic task template, accessed through one API.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。