arXiv:2608.24662cs.AIcs.CL2026-08

模型输出的偏见可能来自部署时的隐形干预,而非模型本身。

The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models

  • 揭示部署时通过概率重分配实现的隐性内容引导机制
  • 证明仅凭观察无法确定行为偏见来自模型还是推理层
  • 提醒监管需区分模型审计与系统整体审计,尤其对广告和合规

生成式语言模型的评估常将政治立场、品牌倾向等行为特征归因于模型权重、对齐训练或提示工程。但这种理解忽略了模型输出背后的多层部署系统。现代推理栈支持在不修改模型参数的前提下,实时干预生成内容。本文研究推理阶段的内容框架偏见:通过非参数化运行时调整,使生成文本向特定机构、意识形态或商业立场倾斜。我们形式化了‘推理归因问题’,并证明在黑箱观察下,行为等效的系统可能由截然不同的模型参数与推理策略组合构成,因此无法唯一确定偏见来源。进一步指出‘概率放置’是一种隐蔽的商业影响模式——通过系统性地重新分配概率质量,使广告内容看似自然出现,区别于显式的生成广告拍卖机制。最后讨论其对行为审计、推理溯源、可信计算、加密认证、欧盟《人工智能法案》与《数字服务法》及广告披露原则的影响。强调治理必须区分‘模型审计’与‘最终发言系统的审计’。

原文摘要 · Abstract (English)

Evaluations of generative language models frequently interpret observable behavioral traits, such as political stance, brand inclination, and normative framing, as manifestations of model weights, post-training alignment, or prompting. This interpretation risks conflating a foundation model with the multi-layered production system through which its outputs are ultimately served. Modern inference stacks support runtime interventions capable of modifying generation while model parameters remain frozen. We examine inference-time framing bias: systematic runtime steering of generated text toward institutional, ideological, or commercial frames without requiring changes to the underlying model parameters. We formalize the Inference Attribution Problem and establish an observational non-identifiability result showing that, under black-box observation alone, behaviorally equivalent deployed systems may arise from structurally distinct combinations of model parameters and inference policies. Consequently, observed behavioral bias does not uniquely identify the architectural layer responsible for it. We further characterize Probability Placement as a deployment pattern in which undisclosed commercial influence is embedded within an ostensibly organic assistant response through systematic probability-mass reallocation, distinguishing it from explicit token-auction mechanisms for generative advertising. Finally, we discuss implications for behavioral auditing, inference provenance, confidential computing, cryptographic attestation, the EU AI Act, the Digital Services Act, and advertising-disclosure principles. We argue that governance of generative systems must increasingly distinguish between auditing a model and auditing the deployed system that ultimately speaks.

大模型治理推理偏见可解释性合规审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。