arXiv:2511.13359cs.AI2025-11AAAI

用推理能力让大模型更懂不同文化,提升跨文化适应性。

Reasoning Shapes Alignment: Investigating Cultural Alignment in Large Reasoning Models with Cultural Norms

  • 通过调查数据自动挖掘文化规范,驱动模型推理对齐
  • 推理强的模型在文化对齐上提升更明显,效果显著
  • 适合研究跨文化AI对齐与价值观多样性的学者

大型推理模型的强大推理能力使其能通过深思熟虑理解并应用安全政策,从而提升安全性。除了安全,模型还需反映多元文化背景下的价值观差异。本文提出文化规范驱动的文化对齐框架(CNCA),利用模型的推理能力对齐文化规范。具体提出三种方法,从有限调查数据中自动挖掘文化规范,并探索其有效应用方式。考察两种对齐范式:上下文内对齐(将文化规范显式融入用户提示)与基于微调的方法(通过增强的思维链训练数据内化规范)。全面实验表明,具备更强推理能力的模型在文化规范挖掘与利用中受益更大。研究强调,通过文化感知对齐策略,推理模型可更好体现人类价值观多样性。

原文摘要 · Abstract (English)

The advanced reasoning capabilities of Large Reasoning Models enable them to thoroughly understand and apply safety policies through deliberate thought processes, thereby improving the models' safety. Beyond safety, these models must also be able to reflect the diverse range of human values across various cultures. This paper presents the Cultural Norm-based Cultural Alignment (CNCA) framework, which enables models to leverage their powerful reasoning ability to align with cultural norms. Specifically, we propose three methods to automatically mine cultural norms from limited survey data and explore ways to effectively utilize these norms for improving cultural alignment. Two alignment paradigms are examined: an in-context alignment method, where cultural norms are explicitly integrated into the user context, and a fine-tuning-based method, which internalizes norms through enhanced Chain-of-Thought training data. Comprehensive experiments demonstrate the effectiveness of these methods, highlighting that models with stronger reasoning capabilities benefit more from cultural norm mining and utilization. Our findings emphasize the potential for reasoning models to better reflect diverse human values through culturally informed alignment strategies.

文化对齐推理模型价值观多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。