arXiv:2506.20949cs.AIcs.CL2025-06ACL被引 3

用长期模拟评估大模型建议的潜在社会危害,提升安全对齐效果。

Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation

  • 通过宏观时间尺度模拟,预测模型建议的社会传播后果。
  • 新数据集上比基线提升20%以上,安全基准平均胜率超70%。
  • 适合关注长时风险、政策与医疗领域安全的AI研究者。

随着基于语言模型的智能体在公共政策、医疗等高风险决策中影响日益增大,确保其有益性需理解其建议的长远影响。本文提出一个概念验证框架,通过宏观时间尺度模拟模型建议在社会系统中的传播路径,实现更稳健的安全对齐。为此,我们构建了包含100个间接伤害场景的数据集,用于测试模型对看似无害提示引发非明显负面后果的预见能力。实验表明,该方法在新数据集上性能提升超过20%,且在现有安全基准(AdvBench、SafeRLHF、WildGuardMix)上平均胜率超过70%,验证了面向长期安全意识对齐的可行性。

原文摘要 · Abstract (English)

Given the growing influence of language model-based agents on high-stakes societal decisions, from public policy to healthcare, ensuring their beneficial impact requires understanding the far-reaching implications of their suggestions. We propose a proof-of-concept framework that projects how model-generated advice could propagate through societal systems on a macroscopic scale over time, enabling more robust alignment. To assess the long-term safety awareness of language models, we also introduce a dataset of 100 indirect harm scenarios, testing models' ability to foresee adverse, non-obvious outcomes from seemingly harmless user prompts. Our approach achieves not only over 20% improvement on the new dataset but also an average win rate exceeding 70% against strong baselines on existing safety benchmarks (AdvBench, SafeRLHF, WildGuardMix), suggesting a promising direction for safer agents.

大模型安全长时风险对齐评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。