用微调模型比较五大宗教的伦理思维差异,发现训练来源决定道德立场。
Six Llamas: Comparative Religious Ethics Through LoRA-Adapted Language Models
- 用LoRA微调六版LLaMA模型,分别学习基督教、伊斯兰教等五宗教经典文本。
- 不同宗教模型在17个伦理问题上表现出可区分的推理模式,且高共识问题一致性达100%。
- 适合研究宗教伦理、文化认知差异或想用AI做跨文明比较的学者。
我们提出六头骆驼(Six Llamas),一项对比研究,检验经不同宗教文本微调的大语言模型是否编码出系统性差异的伦理推理模式。构建了六个基于Meta-Llama-3.1-8B的变体:一个未修改的对照组,以及五个仅在基督教、伊斯兰教、犹太教、印度教或佛教圣典与神学文本上训练的LoRA适配模型。所有六种模型均使用相同的17个标准化伦理提示集进行测试,涵盖道德困境、博弈论情景、公共政策问题及道德心理自我评估。为评估鲁棒性和可重复性,采用十种温度设置的多温度采样设计。计算响应一致性指标、模型间两两同意率、四个提示领域内的温度敏感系数及运行间稳定性分析。结果显示,LoRA适配模型的伦理推理模式(a)与基础模型系统性不同,(b)与训练传统对应的道德逻辑一致,(c)在道德哲学空间中呈现可解释的结构维度,(d)核心伦理立场在高共识难题中对温度变化保持稳定。电车难题在所有模型和温度下达到100%一致性;(e)在有争议的道德领域,高温度下传统特异性差异加剧;(f)基础模型整体响应一致性最高(均值88.3%),表明LoRA适配引入了传统特异性信号并增加了采样敏感性。本研究为使用差异化训练语言模型作为文化与伦理分析工具提供了概念验证,并明确了可证伪性标准及后续扩展方向。
原文摘要 · Abstract (English)
We present Six Llamas, a comparative study examining whether large language models fine-tuned on distinct religious corpora encode systematically different patterns of ethical reasoning. Six variants of Meta-Llama-3.1-8B are constructed: one unmodified control and five LoRA-adapted models trained exclusively on the sacred and theological texts of Christianity, Islam, Judaism, Hinduism, or Buddhism. All six models are probed with an identical battery of 17 standardized ethical prompts spanning moral dilemmas, game-theoretic scenarios, public policy questions, and moral-psychological self-assessments. To assess robustness and reproducibility, we implement a multi-temperature sampling design spanning ten temperature settings. We compute response consistency metrics, pairwise inter-model agreement rates, temperature sensitivity coefficients across four prompt domains, and run-to-run stability analyses. Findings show that LoRA-adapted models produce ethical reasoning patterns that are (a) systematically differentiated from the base model, (b) consistent with the moral logics of their training traditions, (c) structured along interpretable dimensions in moral-philosophical space, (d) core ethical positions remain stable across temperature variations for high-consensus dilemmas. The Trolley Problem achieves 100% consistency across all models and temperatures, while (e) tradition-specific divergence intensifies at higher temperatures in morally contested domains, and (f) the base model exhibits the highest overall response consistency (mean 88.3%), suggesting LoRA adaptation introduces both tradition-specific signal and increased sampling sensitivity. The study offers a proof-of-concept for the condensate comparative method using differentially trained language models as instruments for cultural and ethical analysis and identifies specific criteria for falsification and planned extensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。