arXiv:2409.06241cs.LGcs.AI2024-09NAACL被引 5

让大模型从多角度思考,提升推理准确性和安全性

DiPT: Enhancing LLM reasoning through diversified perspective-taking

  • 引入多元视角机制,增强模型对问题的理解深度
  • 在改写问题上推理准确率显著提升,且输出更安全
  • 适用于现有推理方法的升级,适合希望提升模型鲁棒性的研究者

现有提升语言模型推理能力的方法通常只探索单一解题路径,易出错。受社会科学中换位思考的启发,本文提出DiPT,一种通过显式引入多样化视角来补充现有推理方法的新范式。该方法使模型在推理阶段能更深入理解问题背景,识别最优解题路径。同时,提供一种通用的数据驱动AI方案,用于增强已有数据质量以优化微调效果。实验表明,DiPT可灵活融入单路径推理方法,显著提升其在改写问题上的推理性能与稳定性;并有效维持模型在对抗性‘越狱’提示下的安全输出;此外,使用包含多元视角的数据微调,相比原始数据可明显增强模型推理能力。

原文摘要 · Abstract (English)

Existing work on improving language model reasoning typically explores a single solution path, which can be prone to errors. Inspired by perspective-taking in social studies, this paper introduces DiPT, a novel approach that complements current reasoning methods by explicitly incorporating diversified viewpoints. This approach allows the model to gain a deeper understanding of the problem's context and identify the most effective solution path during the inference stage. Additionally, it provides a general data-centric AI recipe for augmenting existing data to improve their quality for fine-tuning. Our empirical results demonstrate that DiPT can be flexibly integrated into existing methods that focus on a single reasoning approach, enhancing their reasoning performance and stability when presented with paraphrased problems. Furthermore, we illustrate improved context understanding by maintaining the model's safe outputs against "jailbreaking" prompts intentionally designed to bypass safeguards built into deployed models. Lastly, we show that fine-tuning with data enriched with diverse perspectives can boost the reasoning capabilities of the model compared to fine-tuning with raw data alone.

大模型推理多视角安全增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。