arXiv:2606.01755cs.AIcs.CL2026-06

让个性化大模型在不同群体间保持事实一致性,避免偏见。

TriAlign: Towards Universal Truth Consistency in Personalized LLM Alignment

论文配图:TriAlign: Towards Universal Truth Consistency in Personalized LLM Alignment
图 1 · 摘自论文原文
  • 用多智能体强化学习建模不同社会群体,联合优化事实准确、一致与个性化。
  • 在多个基准上显著降低群体间事实差异,同时提升任务表现和用户适配度。
  • 首个面向个性化模型的事实一致性对齐框架,适合关注公平性的研究者。

个性化大语言模型能适应用户的偏好与社会属性,但可能导致跨社会群体的普遍事实不一致,部分群体在客观任务中持续获得不准确的回答。现有对齐方法或忽略个性化,或仅关注主观偏好对齐,严重忽视普遍事实的公平性与一致性。为此,我们提出真理不变对齐(TIA),旨在确保个性化模型在不同社会群体间保持事实一致的同时保留个性化能力。我们提出TriAlign,首个离线多智能体强化学习框架用于TIA,将每个社会群体视为一个智能体进行交互。TriAlign通过公平性感知目标函数和显式不一致惩罚项,联合优化普遍事实准确性、跨群体一致性与个性化程度。在多个基准上的实验表明,相比强基线,TriAlign在三者之间实现更优平衡,显著减少群体间事实差异,同时提升客观任务性能与个性化质量。

原文摘要 · Abstract (English)

Personalized large language models adapt responses to users' preferences and social attributes, but can introduce substantial universal truth inconsistencies across social groups, where some groups systematically receive less accurate responses on objective tasks. Existing alignment methods either ignore personalization or mainly focus on subjective preference alignment, largely overlooking fairness and consistency in universal truths. To address this gap, we study Truth-Invariant Alignment (TIA), an alignment problem for personalized LLMs that aims to ensure universal truths remain consistent across social groups while preserving personalization. We propose TriAlign, the first offline multi-agent reinforcement learning (MARL) framework for TIA, where each social group is modeled as an agent interacting. TriAlign jointly optimizes universal truth accuracy, cross-group truth consistency, and personalization through a fairness-aware objective and an explicit inconsistency penalty. Experiments across diverse benchmarks demonstrate that TriAlign achieves a stronger balance among these three objectives than strong baselines, reducing universal truth disparities across social groups while improving both objective task performance and personalization quality.

个性化公平性对齐多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。