arXiv:2603.23433cs.AI2026-03被引 1

AI决策环境正被悄悄优化,让机器更好判断却不影响人类。

Mecha-nudges for Machines

  • 用比特数统一衡量环境信息对AI的影响力
  • 电商平台信息对AI的可用性提升0.143比特,人类感知无变化
  • 现象已普遍存在,但人类难以察觉,适合关注AI安全与治理的研究者

AI代理正成为互联网上的主动决策者。当它们与人类共享同一决策环境时,环境本身可被调整以影响其行为。我们称之为“机械助推”(mecha-nudging):在不损害人类决策体验的前提下,通过改变选项呈现方式系统性影响AI代理。为量化这一现象,我们融合经济学中的贝叶斯劝说与计算机科学中的$\mathcal{V}$-可用信息框架,获得统一的比特单位,用于跨干预、上下文与模型评估环境变化。我们在六百万个Etsy商品数据上应用该框架,发现自ChatGPT发布后,商品信息中可供预测代理筛选决策的机器可用信息增加了0.143比特(最大可能增加为0.355比特)。该变化在不同提示、令牌选择、标注模型及微调架构下均稳健存在;在受控文本对照组中未出现;且远超通用大模型重写的影响。相反,人类实验显示其可用信息几乎无变化。结果首次提供了大规模实证证据,表明系统性机械助推已在现实中发生,但尚未被察觉。

原文摘要 · Abstract (English)

AI agents are becoming active decision-makers on the Internet. As they make decisions in the same environments as humans, the environments themselves can change to influence them. We call this $\textit{mecha-nudging}$: changes to how choices are presented that systematically influence AI agents without materially degrading the decision environment for humans. To measure this phenomenon, we combine two frameworks -- Bayesian persuasion from economics and $\mathcal{V}$-usable information from computer science -- to get a common unit (bits) for quantifying how environments change across a wide range of interventions, contexts, and models. We apply this framework to over six million Etsy listings and find that, after ChatGPT's release, listings contain significantly more machine-usable information for predicting agent curation decisions, increasing by 0.143 bits out of a maximum possible increase of 0.355. This shift is robust across prompts, token choices, labeling models, and fine-tuning architectures; absent in a regulated-text placebo; and far larger than the effect of generic LLM rewriting. In contrast, a human study finds little to no change in human-usable information. Our results provide the first large-scale evidence that systematic mecha-nudging is already occurring in the wild, but going unnoticed.

AI治理机器学习信息设计大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。