微调语言模型提升对道德模糊场景的理解与人类判断的一致性
Fine-Tuning Language Models for Ethical Ambiguity: A Comparative Study of Alignment with Human Responses
- 通过对比人类标注,用两个精选数据集评估模型在道德困境中的判断能力
- 微调后模型在交叉熵和狄利克雷分数上显著提升,尤其在复杂情境中表现增强
- 微调后的Mistral-7B性能接近GPT-4o,但整体仍落后于BERT/RoBERTa
语言模型常因无法处理语义模糊而误解人类意图,尤其在道德模糊场景中表现更差。本文基于Scruples项目的两个数据集——DILEMMAS(用于比较不同道德情境)和ANECDOTES(用于分析单个叙事)——评估了三类模型(Llama-3.1-8b、Zephyr-7b-beta、Mistral-7b)与人类判断的对齐度。通过提取模型对各选项的概率并对比人工标注,发现微调后模型在交叉熵与狄利克雷评分上均有显著提升,其中狄利克雷分数改善尤为明显。值得注意的是,微调后的Mistral-7B-Instruct-v0.3性能已接近GPT-4o。然而,在交叉熵得分上,所有实验模型仍不及BERT和RoBERTa。该方法通过优化文本到文本格式下的分布理解,有效提升了模型在复杂决策中的表现与人类判断的一致性,凸显了进一步研究伦理推理机制的必要性。
原文摘要 · Abstract (English)
Language models often misinterpret human intentions due to their handling of ambiguity, a limitation well-recognized in NLP research. While morally clear scenarios are more discernible to LLMs, greater difficulty is encountered in morally ambiguous contexts. In this investigation, we explored LLM calibration to show that human and LLM judgments are poorly aligned in such scenarios. We used two curated datasets from the Scruples project for evaluation: DILEMMAS, which involves pairs of distinct moral scenarios to assess the model's ability to compare and contrast ethical situations, and ANECDOTES, which presents individual narratives to evaluate the model's skill in drawing out details, interpreting, and analyzing distinct moral scenarios. Model answer probabilities were extracted for all possible choices and compared with human annotations to benchmark the alignment of three models: Llama-3.1-8b, Zephyr-7b-beta, and Mistral-7b. Significant improvements were observed after fine-tuning, with notable enhancements in both cross-entropy and Dirichlet scores, particularly in the latter. Notably, after fine-tuning, the performance of Mistral-7B-Instruct-v0.3 was on par with GPT-4o. However, the experimental models that were examined were all still outperformed by the BERT and RoBERTa models in terms of cross-entropy scores. Our fine-tuning approach, which improves the model's understanding of text distributions in a text-to-text format, effectively enhances performance and alignment in complex decision-making contexts, underscoring the need for further research to refine ethical reasoning techniques and capture human judgment nuances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。