arXiv:2607.21279cs.CL2026-07中稿 · the 4th Internatio…

构建首个专用于指令微调的统一道德价值数据集。

A Unified Moral-Value Dataset for Instruction Tuning

论文配图:A Unified Moral-Value Dataset for Instruction Tuning
图 1 · 摘自论文原文
  • 将多个现有道德数据集合并并转换为指令-响应格式。
  • 混合使用该数据集训练可保持通用任务性能,且提升道德任务表现。
  • 适合研究大模型对齐与伦理行为训练的研究者使用。

大型语言模型(LLMs)发展迅速,已成为日常生活中有价值的工具。然而,如何使LLMs与特定人类价值观对齐仍是开放问题。近期研究表明,指令微调在零样本任务中具有潜力,可能是解决价值对齐的有效方法。然而,尽管已有众多指令微调数据集,它们大多未专门针对道德情境和行为设计。本文构建了一个可直接用于指令微调的统一道德价值数据集,该数据集通过整合现有道德价值数据集并转换为指令-响应格式而成。实验表明,在通用任务数据集与本数据集混合训练时,通用任务性能得以保留,同时初步观察到混合比例对道德任务性能的影响。本工作为指令微调提供了一个道德价值数据集,为后续对齐研究提供了重要资源。数据集已开源:https://huggingface.co/datasets/teohzzh/value-for-instruction-tuning。

原文摘要 · Abstract (English)

Large language models (LLMs) have developed rapidly and become valuable tools in everyday life. However, how to align LLMs to a particular set of human values is still an open problem. Recent studies show that instruction tuning has strong potential for zero-shot tasks and may serve as an effective approach to addressing value alignment. Nevertheless, although many datasets for instruction tuning already exist, they are not specifically designed around moral scenarios and behaviors. We construct a unified moral-value dataset that can be directly used for instruction tuning. This dataset is built upon existing moral-value datasets by merging them into a unified corpus and converting them into an instruction-response format. We show that training on a mixed dataset combining general task datasets with our dataset preserves general-task performance, and we report preliminary observations on how the mixing ratio affects value-oriented task performance. Our work provides a moral-value dataset for instruction tuning and offers a useful resource for further alignment research. The dataset is available at https://huggingface.co/datasets/teohzzh/value-for-instruction-tuning.

指令微调道德对齐数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。