不微调模型,用角色分歧实现跨文化对齐。
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement

- 用世界价值观调查构建各国角色代理,以分歧为信号
- 20国测试中减少10%-24%的文化偏差,无需修改权重
- 适合需快速适配多元文化场景的AI服务部署
大语言模型在涉及道德判断的决策中日益重要,但其隐含偏好并非文化中立。现有方法或需各国细调数据与预算,或依赖可访问模型内部的白盒条件,而商业API不具备这些条件。本文聚焦黑盒、仅使用公开数据的现实场景,发现国内社会人口学分歧才是主要引导信号,而非共识。提出DISCA(基于分歧引导的文化对齐推理方法),将每个国家视为由基于世界价值观调查的角色代理组成的评议团,将其分歧转化为有界、抗损失的逻辑值修正。在20个国家和7个开源模型(2B–70B)上,DISCA在≥3.8B的6个模型上使MultiTP的文化偏差降低10%-24%,开放问答场景降低2%-7%,且不更改任何模型权重。结果表明,推理时校准是替代微调以满足全球道德偏好长尾需求的可扩展方案。
原文摘要 · Abstract (English)
Large language models increasingly mediate decisions that turn on moral judgement, yet a growing body of evidence shows that their implicit preferences are not culturally neutral. Existing cultural alignment methods either require per-country preference data and fine-tuning budgets or assume white-box access to model internals that commercial APIs do not expose. In this work, we focus on this realistic black-box, public-data-only regime and observe that within-country sociodemographic disagreement, not consensus, is the primary steering signal. We introduce DISCA (Disagreement-Informed Steering for Cultural Alignment), an inference-time method that instantiates each country as a panel of World-Values-Survey-grounded persona agents and converts their disagreement into a bounded, loss-averse logit correction. Across 20 countries and 7 open-weight backbones (2B--70B), DISCA reduces cultural misalignment on MultiTP by 10--24% on the six backbones >=3.8B, and 2--7% on open-ended scenarios, without changing any weights. Our results suggest that inference-time calibration is a scalable alternative to fine-tuning for serving the long tail of global moral preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。