用少量数据让大模型更好识别低资源语言的恶意请求
Unlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training Data
- 基于推理与跨语言对齐,提升多语言安全检测可解释性
- 仅用1000样本在六种语言上表现超越更大模型
- 适合关注低资源语言安全与模型可解释性的研究者
大语言模型虽能力增强,但恶意请求风险上升,亟需有效防护。现有方法多依赖分类器,缺乏可解释性且在低资源语言上表现差。本文提出ConsistentGuard,一种基于推理的多语言防护机制,通过推理增强可解释性,通过语言对齐促进知识迁移。仅用1,000个训练样本,在三个数据集上的六种语言中均优于使用更多数据训练的大模型,展现优异性能、可解释性与泛化能力。同时构建多语言评测扩展基准并开源代码,推动后续研究。
原文摘要 · Abstract (English)
Recent advances in LLMs have enhanced AI capabilities, but also increased the risk posed by malicious requests, highlighting the need for effective LLM safeguards to detect such queries. Existing approaches largely rely on classifier-based methods that lack interpretability and perform poorly on low-resource languages. To address these limitations, we propose ConsistentGuard, a novel reasoning-based multilingual safeguard, which enhances explainability via reasoning and boosts knowledge transfer between languages through alignment. With only 1,000 training samples, our method demonstrates superior performance on three datasets across six languages, outperforming larger models trained with significantly more data, and exhibits strong interpretability and generalization ability. We also contribute a multilingual benchmark extension and release our codes to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。