用合成文档训练让AI具备动物共情,效果优于传统方法
Alignment midtraining for animals
- 用3000份合成文档进行中段训练,提升AI动物共情能力
- 在26题评估中达77%正确率,远超指令微调的40%
- 适合关注伦理对齐与可持续价值训练的研究者
我们研究了通过中段训练使用合成文档实现价值对齐的鲁棒性,以动物共情这一重要且与现有对齐工作正交的价值为目标。为评估共情推理能力,我们开发并公开发布动物道德判断规范(ANIMA),一个包含26个问题、覆盖13个伦理维度的评估数据集及可检查的评测工具。在ANIMA上,使用3000份文档训练可达到77%的准确率,显著高于指令微调的40%,且泛化至人类共情,未在标准安全基准或能力测试中出现退化。然而,后续无关的指令微调会削弱该干预效果,优势在5000样本后消失。探索性结果表明,基于文档的价值干预可能需要显式保存策略,才能在典型训练流程中保持有效。
原文摘要 · Abstract (English)
We investigate the robustness of value alignment via midtraining with synthetic documents, using animal compassion as a value that is both important in its own right and orthogonal to existing alignment efforts. To evaluate compassionate reasoning, we develop and publicly release Animal Norms In Moral Assessment (ANIMA), a 26-question evaluation spanning 13 ethical dimensions, publicly available as a dataset and Inspect evaluation. On ANIMA, training with 3000 documents achieves 77% compared to 40% for instruction-tuning approaches, with generalization to human compassion and no degradation in standard safety benchmarks or capabilities. However, subsequent unrelated instruction-tuning degrades the intervention, with the advantage disappearing after 5000 samples. Our exploratory results suggest document-based value interventions may require explicit preservation strategies to remain effective through typical training pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。