提出自回归生成模型信用归属的两大障碍,揭示现有方法为何难实现公平归因。
Barriers to Counterfactual Credit Attribution for Autoregressive Models
- 分析自回归模型在反事实信用归属下的组合失效问题
- 证明信用补救方法在输出长度增长时查询复杂度指数级上升
- 为生成模型公平归因提供理论边界,适合关注模型伦理的研究者
生成式AI正在改变对先前工作给予信用的传统做法。理想的生成模型应对其输出有显著依赖的工作给予信用。反事实信用归属(Counterfactual Credit Attribution, CCA)是一种形式化该目标的技术条件,是差分隐私的一种宽松版本,由Livni等人于2024年在PAC学习框架中提出。本文首次研究了生成模型中的CCA问题,聚焦自回归模型在部署时对数据集(如RAG数据库)进行信用归因的情形。我们发现了两种自然实现方式的障碍:首先,对底层下一个词预测器施加CCA并不能保证整个模型满足CCA——CCA不具备自回归组合性(不同于差分隐私);其次,我们提出一种称为‘补救’(retrofitting)的方法,即对不具信用归属能力的模型添加信用。在弱最优性要求下,我们证明了基于黑盒访问原模型的补救方法存在下界:其查询复杂度随输出长度呈指数级增长。
原文摘要 · Abstract (English)
Generative AI disrupts the practice of giving credit to work that came before. Ideally, a generative model would give credit to any work on which its output depends in a significant way. \emph{Counterfactual credit attribution} (CCA) is a technical condition formalizing this goal--a relaxation of differential privacy--recently introduced by Livni, Moran, Nissim, and Pabbaraju [2024] who studied it in the PAC learning setting. We initiate the study of CCA generative models. Specifically, we consider autoregressive models giving credit to a deployment-time dataset (e.g., a RAG database). We uncover barriers to two natural approaches to CCA autoregressive models. First, we show that imposing CCA on the underlying next-token predictor does not guarantee that the model is CCA: CCA does not compose autoregressively (unlike DP). Second, we consider a different approach to building CCA models which we call \emph{retrofitting}. Retrofitting takes a model that does not attribute credit, and adds credit onto it. We prove a lower bound for CCA retrofitting under a weak optimality requirement. Given black-box access to the starting model, retrofitting requires query complexity exponential in the length of the model's outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。