arXiv:2607.20827cs.AI2026-07被引 4

检测大模型在工具选择中对证据来源权威性的敏感度。

Auditing Provenance Sensitivity in LLM Agent Action Selection

论文配图:Auditing Provenance Sensitivity in LLM Agent Action Selection
图 1 · 摘自论文原文
  • 设计针对性授权审计,区分每项工具与论点的上下文来源权限。
  • 5.4%的竞争性决策受非授权证据影响,2.4%在削弱可信证据时仍保留错误行为。
  • 揭示模型易受文本来源权威性暗示影响,适合安全与可信推理研究者参考。

LLM代理从混合用户请求、工具输出、检索记录、记忆和不可信文本的上下文中选择工具和论据。即使证据相关但未经授权,其决策仍可能正确,因此无需仅基于允许的证据做出判断。本文提出一种目标特定的授权审计方法,对每个工具和论点目标分别标注上下文因素。主要测试在任务、命题、立场和策略固定时,仅改变命题来源权威性的行为变化;同时通过削弱有效证据并分析上下文子集交互作为辅助定位诊断。在450个受控的下一步行动任务及多个开源权重的LLM家族中,可信与不可信变体在竞争案例中产生不同动作的比例为5.4%,支持案例中为1.7%。在受控降级条件下,未经授权的竞争行为在全正确、混合错误、干净正确模式中占2.4%,置信区间为2.1%至3.0%。这些是受控压力测试下的发生率,非实际部署中的普遍性。模型会响应文本来源权威性线索,但无法阻止不可信证据对其行为的影响。

原文摘要 · Abstract (English)

LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records, memory, and untrusted text. Evidence can be relevant without being authorized to determine a decision, so a correct action need not be grounded only in permitted evidence. We introduce a target-specific authorization audit that labels context factors separately for each tool and argument target. Its primary test holds the task, proposition, position, and policy fixed while changing only the proposition's source authority. We then test behavior when valid evidence is weakened and use context-subset interactions as a secondary localization diagnostic. Across 450 controlled next-action tasks and multiple open-weight LLM families, trusted and untrusted variants produce different actions in 5.4 percent of competing cases versus 1.7 percent of supporting cases. Under controlled degradation, unauthorized competition is retained in a full-correct, mixed-error, clean-correct pattern in 2.4 percent of comparisons, with a 95 percent confidence interval from 2.1 to 3.0 percent. These are controlled stress-set rates, not deployment prevalence. The models respond to textual source-authority cues, but this does not prevent untrusted evidence from influencing their actions.

大模型安全证据可信度代理系统审计机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。