arXiv:2606.18424stat.OTcs.AI2026-06

用变分框架建模大模型生成与监管的博弈,揭示安全与自由的权衡机制。

A Variational Framework for LLM Generator-Regulator Games

  • 将生成与监管建模为鞍点问题,通过熵正则吉布斯分布描述输出分布
  • 平衡效用、熵、合规性与检测概率,实现多目标优化
  • 适用于内容过滤、钓鱼防御等场景,适合安全与可控生成研究者

本文构建了一个用于受控语言生成的变分框架。从自回归标记采样出发,推导出完整消息的诱导分布,并将其与熵正则化的吉布斯定律相关联。监管被建模为最优判别器,其对偶值为f-散度,生成器与监管者的互动被形式化为一个鞍点问题。该框架适用于内容审核、审查过滤、AI欺骗检测、合规审计、网络钓鱼防御和操控控制,其中监管关注的是可能消息的分布而非单一输出。均衡状态阐明了效用、熵、监管一致性与有限长度可检测性之间的权衡。两个有限词汇量案例研究——审查过滤与钓鱼防御——展示了如何通过效用、熵、散度、接收端评分和检测概率来评估理论有效性。

原文摘要 · Abstract (English)

This paper develops a variational framework for regulated language generation. Starting from autoregressive token sampling, we derive the induced distribution over complete messages and relate it to an entropy-regularized Gibbs law. Regulation is modeled as an optimal discriminator whose convex-dual value is an f-divergence, and the generator-regulator interaction is formulated as a saddle-point problem. The framework applies to moderation, censorship, AI deception detection, compliance auditing, phishing defense, and manipulation control, where regulation concerns a distribution over possible messages rather than a single output. The equilibrium clarifies the tradeoff among utility, entropy, regulatory alignment, and finite-length detectability. Two finite-vocabulary case studies, censorship filtering and phishing defense, illustrate how the theory can be evaluated through utility, entropy, divergence, receiver-side scores, and detection probability.

语言生成生成对抗安全可控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。