给大模型打疫苗,用纠错对提升抗谣言能力
Just as Humans Need Vaccines, So Do Models: Model Immunization to Combat Falsehoods
- 用标注的假话-纠正对做有监督微调,注入少量'疫苗剂量'
- 真相问答准确率提升12点,假信息拒绝率提高30点
- 适合关注模型安全与可信度的研究者和开发者
大型语言模型传播错误信息,不仅因记忆虚假事实,更因学习了使谎言更具说服力的语言模式,如模棱两可、虚假预设和虚构引用。我们提出模型免疫:在真实数据中注入小剂量(5%至10%的词元)的精心筛选的(错误陈述,纠正)配对进行有监督微调。与事后过滤或基于偏好的对齐不同,该方法对已标注的错误信息施加直接的负向监督。在四个开源权重模型族中,该方法将TruthfulQA准确率提升12分,假信息拒绝率提升30分,同时保持整体模型能力。我们进一步提出关键设计要求,包括剂量、标注、隔离和多样性,并倡导建立标准化疫苗语料库与评估基准以衡量泛化能力。这些发现表明免疫是负责任大模型开发中一项实用且可扩展的组件。
原文摘要 · Abstract (English)
Large language models (LLMs) reproduce misinformation not by memorizing false facts alone, but by learning the linguistic patterns that make falsehoods persuasive, such as hedging, false presuppositions, and fabricated citations. We propose model immunization, a training paradigm based on supervised fine-tuning over curated (false claim, correction) pairs, injected as small vaccine doses (5 to 10% of tokens) alongside truthful data. Unlike post-hoc filtering or preference-based alignment, immunization introduces direct negative supervision on labeled falsehoods. Across four open weight model families, this approach improves TruthfulQA accuracy by 12 points and increases misinformation rejection rates by 30 points, while preserving overall model capability. We further outline key design requirements, including dosage, labeling, quarantine, and diversity and advocate for standardized vaccine corpora and benchmarks to evaluate generalization. These findings position immunization as a practical and scalable component of responsible LLM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。