arXiv:2501.13080cs.CLcs.CR2025-01被引 24

通过思维链微调提升大模型判别恶意输入的效率与可信度。

Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment

  • 用少量数据微调思维链,让大模型生成带推理过程的判题结果。
  • 在有限数据下显著提升对多种恶意输入的检测准确率与鲁棒性。
  • 适合需要高安全性的对话系统部署,尤其关注可解释性防御场景。

大型语言模型(LLMs)在各类应用中展现出强大能力,尤其在对话式AI产品中价值显著。确保其安全与可靠性至关重要,以防范恶意用户交互带来的风险与声誉损失。本文系统研究了不同LLM作为输入防护屏障时,对其思维链(CoT)输出进行微调与对齐的有效性。我们利用少量训练数据,探索多种微调方法,使模型充当代理防御机制,识别恶意输入并提供推理依据,从而防止对话代理被滥用。通过严格评估不同策略在多样对抗性与恶意查询上的泛化能力,实验表明:即使在数据资源受限的情况下,针对多样化有害输入设计的对齐流程仍具潜力。这些技术显著提升了对话AI系统的安全性,并为部署更安全、可信的AI交互提供了可行框架。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated powerful capabilities that render them valuable in different applications, including conversational AI products. It is paramount to ensure the security and reliability of these products by mitigating their vulnerabilities towards malicious user interactions, which can lead to the exposure of great risks and reputational repercussions. In this work, we present a comprehensive study on the efficacy of fine-tuning and aligning Chain-of-Thought (CoT) responses of different LLMs that serve as input moderation guardrails. We systematically explore various tuning methods by leveraging a small set of training data to adapt these models as proxy defense mechanisms to detect malicious inputs and provide a reasoning for their verdicts, thereby preventing the exploitation of conversational agents. We rigorously evaluate the efficacy and robustness of different tuning strategies to generalize across diverse adversarial and malicious query types. Our experimental results outline the potential of alignment processes tailored to a varied range of harmful input queries, even with constrained data resources. These techniques significantly enhance the safety of conversational AI systems and provide a feasible framework for deploying more secure and trustworthy AI-driven interactions.

大模型安全思维链输入防护对抗防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。