arXiv:2509.02133cs.CL2025-09被引 2

用宪法约束生成过程,让大模型输出更公平包容。

AMBEDKAR-A Multi-level Bias Elimination through a Decoding Approach with Knowledge Augmentation for Robust Constitutional Alignment of Language Models

  • 用小模型生成、大模型按宪法校验,反向实现公平推理
  • 在不训练模型的前提下,降低偏见最高达26.41%
  • 专为印度社会的种姓与宗教偏见设计,适合本土化部署

大型语言模型可能无意中反映训练数据中的社会偏见,导致有害或歧视性输出。我们在一系列模型上对印度语境下的实证评估显示,种姓与宗教相关的偏见尤为突出。然而,现有大多数缓解策略以西方为中心,难以应对这些本地化特征。我们提出 AMBEDKAR 框架,灵感源自印度宪法之父阿姆倍伽尔的平等理念,引导大模型输出符合第14至17条宪法精神的公平、中立与包容内容。该方法引入一个基于印度人工智能宪法的推理时解码层,仅在推理阶段应用,不更新模型参数。通过推测性解码算法,在生成过程中主动减少种姓主义与宗派偏见。此缓解层直接嵌入解码过程,无需修改模型内部结构,显著降低重训练带来的计算与基础设施成本。我们将推测性解码重新定义为一种公平机制:小语言模型(SLM)作为潜在偏见生成器,而宪法引导的大语言模型(LLM)充当验证者。并非加速生成,而是强制生成路径的抗偏见性。这一角色反转催生了‘以推测实现公平’的新范式。实验表明,相比基线,该方法可实现最高26.41%的偏见绝对降低。代码、数据集与结果已公开于 https://anonymous.4open.science/r/AMBEDKAR-983B/

原文摘要 · Abstract (English)

Large Language Models (LLMs) can inadvertently reflect societal biases present in their training data, leading to harmful or prejudiced outputs. In the Indian context, our empirical evaluations across a suite of models reveal that biases around caste and religion are particularly salient. Yet, most existing mitigation strategies are Western-centric and fail to address these local nuances. We propose AMBEDKAR, a framework inspired by the egalitarian vision of Dr B. R. Ambedkar, architect of the Indian Constitution, to guide LLM outputs toward fairness, neutrality, and inclusion in line with Articles 14 to 17. Our approach introduces a Constitution-Aware Decoding Layer, guided by the AI Constitution of India and applied only at inference time, without any parameter updates to the base model. We incorporate a speculative decoding algorithm that proactively reduces casteist and communal bias during generation. This mitigation layer operates directly within the decoding process, avoiding changes to model internals and lowering the computational and infrastructural costs associated with retraining. We reinterpret speculative decoding not merely as an efficiency tool but as a mechanism for fairness. In this framework, a Small Language Model (SLM) acts as a potentially biased generator, while a constitutionally guided Large Language Model (LLM) serves as the verifier. Rather than accelerating generation, the LLM enforces bias-robust trajectories in the SLM outputs. This inversion of roles gives rise to a fairness-by-speculation paradigm. Our approach yields an absolute reduction of bias up to 26.41 percent compared to baseline. Our source code, datasets, and results are available at https://anonymous.4open.science/r/AMBEDKAR-983B/

偏见消除宪法对齐推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。