用人类道德推理的脑数据微调大模型,提升其类人决策能力。
Inducing Human-like Biases in Moral Reasoning Language Models
- 用行为与脑活动数据联合微调BERT等模型
- 模型在伦理任务上准确率提升,但脑活动相似度未显著改善
- 验证了脑数据对模型类人推理的有限促进作用
本文研究在道德推理任务中,通过人类行为数据和/或脑数据(fMRI)微调大型语言模型(LLMs)后的对齐程度(BrainScore)。我们对BERT、RoBERTa、DeBERTa等模型在ETHICS基准的行为数据、Koster-Hale等人(2013)的道德推理fMRI数据,或两者结合的数据上进行微调。评估模型在ETHICS基准上的准确率及模型激活与fMRI数据之间的BrainScore。结果显示,大模型在两项指标上表现更优,但经脑数据微调后,BrainScore并未显著提升。
原文摘要 · Abstract (English)
In this work, we study the alignment (BrainScore) of large language models (LLMs) fine-tuned for moral reasoning on behavioral data and/or brain data of humans performing the same task. We also explore if fine-tuning several LLMs on the fMRI data of humans performing moral reasoning can improve the BrainScore. We fine-tune several LLMs (BERT, RoBERTa, DeBERTa) on moral reasoning behavioral data from the ETHICS benchmark [Hendrycks et al., 2020], on the moral reasoning fMRI data from Koster-Hale et al. [2013], or on both. We study both the accuracy on the ETHICS benchmark and the BrainScores between model activations and fMRI data. While larger models generally performed better on both metrics, BrainScores did not significantly improve after fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。