arXiv:2604.13076cs.CLcs.AI2026-04被引 3

用合成文档训练让AI具备动物共情,效果优于传统方法

Alignment midtraining for animals

  • 用3000份合成文档进行中段训练,提升AI动物共情能力
  • 在26题评估中达77%正确率,远超指令微调的40%
  • 适合关注伦理对齐与可持续价值训练的研究者

我们研究了通过中段训练使用合成文档实现价值对齐的鲁棒性,以动物共情这一重要且与现有对齐工作正交的价值为目标。为评估共情推理能力,我们开发并公开发布动物道德判断规范(ANIMA),一个包含26个问题、覆盖13个伦理维度的评估数据集及可检查的评测工具。在ANIMA上,使用3000份文档训练可达到77%的准确率,显著高于指令微调的40%,且泛化至人类共情,未在标准安全基准或能力测试中出现退化。然而,后续无关的指令微调会削弱该干预效果,优势在5000样本后消失。探索性结果表明,基于文档的价值干预可能需要显式保存策略,才能在典型训练流程中保持有效。

原文摘要 · Abstract (English)

We investigate the robustness of value alignment via midtraining with synthetic documents, using animal compassion as a value that is both important in its own right and orthogonal to existing alignment efforts. To evaluate compassionate reasoning, we develop and publicly release Animal Norms In Moral Assessment (ANIMA), a 26-question evaluation spanning 13 ethical dimensions, publicly available as a dataset and Inspect evaluation. On ANIMA, training with 3000 documents achieves 77% compared to 40% for instruction-tuning approaches, with generalization to human compassion and no degradation in standard safety benchmarks or capabilities. However, subsequent unrelated instruction-tuning degrades the intervention, with the advantage disappearing after 5000 samples. Our exploratory results suggest document-based value interventions may require explicit preservation strategies to remain effective through typical training pipelines.

价值对齐动物共情中段训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。