arXiv:2604.19431cs.LOcs.AI2026-04被引 1

用计数世界逻辑分析生成式AI输出中的偏见演化过程

Counting Worlds Branching Time Semantics for post-hoc Bias Mitigation in generative AI

论文配图:Counting Worlds Branching Time Semantics for post-hoc Bias Mitigation in generative AI
图 1 · 摘自论文原文
  • 引入计数世界语义,形式化描述生成过程中的每一步可能输出
  • 可验证保护属性的概率分布是否公平,预测后续输出的偏差风险
  • 适合研究生成模型公平性与后验偏见缓解的学者和工程师

生成式AI系统会放大训练数据中的偏见。尽管已有多种推理阶段的缓解策略,但大多为经验性方法,缺乏形式化保证。本文提出CTL-F,一种用于分析生成输出序列中偏见的分支时间逻辑。该逻辑采用计数世界语义,每个世界代表生成过程某一时刻的可能输出,并引入模态算子,可用于验证当前输出序列是否符合目标保护属性的概率分布,预测新增输出时是否仍处于可接受偏差范围内,以及判断需删除多少输出才能恢复公平性。通过一个有偏图像生成的简化案例,展示了CTL-F公式如何在输出序列的不同阶段表达具体的公平性要求。

原文摘要 · Abstract (English)

Generative AI systems are known to amplify biases present in their training data. While several inference-time mitigation strategies have been proposed, they remain largely empirical and lack formal guarantees. In this paper we introduce CTLF, a branching-time logic designed to reason about bias in series of generative AI outputs. CTLF adopts a counting worlds semantics where each world represents a possible output at a given step in the generation process and introduces modal operators that allow us to verify whether the current output series respects an intended probability distribution over a protected attribute, to predict the likelihood of remaining within acceptable bounds as new outputs are generated, and to determine how many outputs are needed to remove in order to restore fairness. We illustrate the framework on a toy example of biased image generation, showing how CTLF formulas can express concrete fairness properties at different points in the output series.

生成模型偏见缓解形式验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。