arXiv:2510.18914cs.CLcs.AI2025-10ACL被引 3

提出动态可逆剪枝方法,实时调节模型偏见行为。

Fairness Evaluation and Inference Level Mitigation in LLMs

  • 基于上下文感知激活,动态掩码特定神经元以抑制偏见。
  • 在多轮跨语言对话中保持连贯性,减少有害内容传播。
  • 适用于真实场景的实时公平性调控,适合对话系统开发者。

大型语言模型常在其内部表示中表现出不良行为,导致公平性下降、不一致漂移、有害内容放大及对话中不当模式的传播。尽管训练阶段或数据驱动的方法试图缓解这些问题,但存在计算成本高、部署后不可逆且难以适应新对话情境等局限。剪枝方法提供了灵活透明的去偏途径,但多数现有方法为静态剪枝,一旦神经元被移除便无法随上下文变化恢复。为此,本文提出一种动态、可逆的剪枝框架,通过检测上下文感知的神经元激活,在生成过程中自适应地施加掩码,调节其影响。该推理时解决方案实现了细粒度、内存感知的偏见缓解,在保留知识的前提下,提升了多语言单轮与多轮对话中的连贯性,支持现实对话人工智能中的动态公平性控制。

原文摘要 · Abstract (English)

Large language models often display undesirable behaviors embedded in their internal representations, undermining fairness, inconsistency drift, amplification of harmful content, and the propagation of unwanted patterns during extended dialogue and conversations. Although training-time or data-centric methods attempt to reduce these effects, they are computationally expensive, irreversible once deployed, and slow to adapt to new conversational contexts. Pruning-based methods provide a flexible and transparent way to reduce bias by adjusting the neurons responsible for certain behaviors. However, most existing approaches are static; once a neuron is removed, the model loses the ability to adapt when the conversation or context changes. To address this, we propose a dynamic, reversible, pruning-based framework that detects context-aware neuron activations and applies adaptive masking to modulate their influence during generation. Our inference-time solution provides fine-grained, memory-aware mitigation with knowledge-preserved, more coherent behavior across multilingual single- and multi-turn dialogues, enabling dynamic fairness control in real-world conversational AI.

大模型公平性剪枝对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。