用1900道有毒数学题测试大模型,发现可安全对齐且不丢解题能力。
SafeMath: Safe Solutions for Unsafe Math Word Problems
- 构建1900道含敏感内容的数学题数据集,保留合理推理结构
- 验证主流大模型在有毒题目下易生成有害输出,安全与正确性可兼得
- 提出SafeMath对齐方法,提升安全性同时保持甚至提升解题准确率
近期研究揭示大语言模型可能被对抗性或看似无害的输入操纵,导致有害、偏见或违规输出。本文探讨一个未受重视的问题:有害或有毒的数学应用题。我们发现,以自然语言叙述形式呈现的数学问题,可能成为传播偏见、不道德或心理有害内容的隐蔽媒介,尤其在涉及儿童的教育场景中风险更高。为系统研究此现象,我们引入ToxicGSM数据集,包含1900个算术问题,其背景嵌入有害或敏感内容,但数学推理任务仍定义明确。基于该数据集,我们审计现有大模型行为,并分析安全防护与数学正确性之间的权衡。进一步提出SafeMath——一种安全对齐技术,可在降低有害输出的同时维持,甚至在某些情况下提升数学推理性能。结果强调将语言危害与数学推理解耦的重要性,并表明有效安全对齐无需以准确性为代价。
原文摘要 · Abstract (English)
Recent research points toward LLMs being manipulated through adversarial and seemingly benign inputs, resulting in harmful, biased, or policy-violating outputs. In this paper, we study an underexplored issue concerning harmful and toxic mathematical word problems. We show that math questions, particularly those framed as natural language narratives, can serve as a subtle medium for propagating biased, unethical, or psychologically harmful content, with heightened risks in educational settings involving children. To support a systematic study of this phenomenon, we introduce ToxicGSM, a dataset of 1.9k arithmetic problems in which harmful or sensitive context is embedded while preserving mathematically well-defined reasoning tasks. Using this dataset, we audit the behaviour of existing LLMs and analyse the trade-offs between safety enforcement and mathematical correctness. We further propose SafeMath -- a safety alignment technique that reduces harmful outputs while maintaining, and in some cases improving, mathematical reasoning performance. Our results highlight the importance of disentangling linguistic harm from math reasoning and demonstrate that effective safety alignment need not come at the cost of accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。