研究大模型长期记忆如何累积并传播隐性偏见,提出动态标记机制有效抑制偏见扩散。
How Implicit Bias Accumulates and Propagates in LLM Long-term Memory
- 构建决策型隐性偏见基准,量化九个社会领域中的偏见演化
- 实验证明偏见随时间加剧并在不同领域间传播,现有方法效果有限
- 提出动态记忆标记法,在写入阶段干预,显著降低偏见积累与跨域传播
长期记忆机制使大语言模型在长时间交互中保持连贯性与个性化,但也带来公平性新风险。本文研究隐性偏见(即微妙的统计偏差)在具备长期记忆的模型中如何累积与传播。为此,我们引入决策型隐性偏见(DIB)基准,包含3,776个跨九个社会领域的决策场景,用于量化长期决策过程中的偏见。通过真实长周期模拟框架,评估六种先进大模型与三种代表性记忆架构在DIB上的表现,发现模型隐性偏见并非静态,而是随时间增强,并在无关领域间传播。进一步分析表明,静态系统提示基线仅能提供短暂且有限的去偏效果。为此,我们提出动态记忆标记(DMT),一种在记忆写入时强制公平约束的代理干预方法。大量实验显示,DMT显著减少偏见累积,有效遏制跨领域偏见传播。
原文摘要 · Abstract (English)
Long-term memory mechanisms enable Large Language Models (LLMs) to maintain continuity and personalization across extended interaction lifecycles, but they also introduce new and underexplored risks related to fairness. In this work, we study how implicit bias, defined as subtle statistical prejudice, accumulates and propagates within LLMs equipped with long-term memory. To support systematic analysis, we introduce the Decision-based Implicit Bias (DIB) Benchmark, a large-scale dataset comprising 3,776 decision-making scenarios across nine social domains, designed to quantify implicit bias in long-term decision processes. Using a realistic long-horizon simulation framework, we evaluate six state-of-the-art LLMs integrated with three representative memory architectures on DIB and demonstrate that LLMs' implicit bias does not remain static but intensifies over time and propagates across unrelated domains. We further analyze mitigation strategies and show that a static system-level prompting baseline provides limited and short-lived debiasing effects. To address this limitation, we propose Dynamic Memory Tagging (DMT), an agentic intervention that enforces fairness constraints at memory write time. Extensive experimental results show that DMT substantially reduces bias accumulation and effectively curtails cross-domain bias propagation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。