metaphor 会引发大模型跨领域推理偏差,影响决策可靠性
Metaphors are a Source of Cross-Domain Misalignment of Large Reasoning Models
- 用隐喻干预训练过程,发现模型产生跨领域推理偏差
- 隐喻触发深层特征激活,导致推理路径偏离正确方向
- 可检测隐喻相关特征,适合关注模型安全性的研究者
早期研究表明隐喻影响人类决策,这引发疑问:当大型语言模型(LLMs)的训练数据包含大量隐喻时,是否也会干扰其推理路径?本文在新兴偏差问题背景下研究该现象,发现训练数据中的隐喻确实导致了大模型推理输出的跨域偏差。通过在持续预训练和微调中引入基于隐喻的干预,模型表现出不同程度的跨域偏差。在重新对齐设置下也观察到类似效应。进一步分析表明,隐喻与大模型中潜在特征的激活相关。通过监测这些特征,我们设计了一个高精度检测器,能有效预测偏差内容。
原文摘要 · Abstract (English)
Earlier research has shown that metaphors influence human decision-making, raising the question of whether metaphors also influence large language models (LLMs)' reasoning pathways, given that their training data contain a large number of metaphors. In this work, we investigate the problem in the scope of the emergent misalignment problem, where LLMs can generalize patterns learned from misaligned content in one domain to another domain. We find strong evidence that metaphors in training data contribute to cross-domain misalignment in LLMs' reasoning outputs. With metaphor-based interventions during continued pre-training and fine-tuning for inducing misalignment, models exhibit significantly different degrees of emergent cross-domain misalignment. We also observe similar effects in re-alignment settings. As we further investigate this phenomenon, we find that metaphors are linked to the activation of latent features in large reasoning models. By monitoring these latent features, we design a detector that predicts misaligned content with high accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。