让大模型更准确预测不同文化下人们对选择题的回答分布。
Evidence-based Distributional Alignment for Large Language Models
- 基于世界价值观调查数据检索证据,构建结构化回答分布预测框架。
- 在跨文化场景下,预测分布与真实分布的差异降低44%(相对改进)。
- 适合做跨文化社会调研、政策分析的学者和机构使用。
分布对齐使大语言模型能够预测目标群体在多个选项间的回答分布,而非将分歧压缩为单一共识答案。然而,现有基于LLM的分布预测在文化或领域迁移下常不稳定:基于标记得分的估计易受选项表述微小变化影响;基于采样的方法成本高且对提示和解码设置敏感;直接生成的分布常出现校准偏差。本文提出Evi-DA,一种基于证据的对齐方法,提升模型在领域和文化迁移下的分布估计保真度与鲁棒性。给定目标国家和多选题,Evi-DA检索相关世界价值观调查条目及其回答分布,为每个选项预测粗粒度的Welzel值签名,并以结构化形式推断该国条件下的回答分布。采用两阶段训练流程,强化学习优化基于调查的奖励,鼓励准确的中间值预测、忠实的最终分布、规范的结构化输出及减少文化偏见。在域内与域外基准测试中,使用多种开源模型骨干,Evi-DA相较强基线平均相对改进达44%,显著降低预测分布与真实分布间的Jensen-Shannon散度。
原文摘要 · Abstract (English)
Distributional alignment enables large language models (LLMs) to predict how a target population distributes its responses across answer options, rather than collapsing disagreement into a single consensus answer. However, existing LLM-based distribution prediction is often unstable and degrades under cultural and domain shift. Token score-based estimates can change with minor option wording or formatting, response sampling-based estimates are expensive and sensitive to prompts and decoding settings, and directly generated distributions are frequently miscalibrated. We propose Evi-DA, an evidence-based alignment technique that improves the fidelity and robustness of LLM-based distribution estimation under domain and cultural shift. Given a target country and a multiple-choice question, Evi-DA retrieves related World Values Survey items and their answer distributions, predicts a coarse Welzel value signature for each option, and infers the country-conditioned answer distribution in a structured format. We train the LLMs using a two-stage pipeline, where reinforcement learning optimizes survey-derived rewards that encourage accurate intermediate value predictions, faithful final distributions, well-formed structured outputs, and reduced cultural bias. Across in-domain and out-of-domain benchmarks and multiple open-source backbones, Evi-DA reduces Jensen-Shannon divergence between predicted and gold distributions relative to strong baselines, with average relative improvements of up to 44%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。