arXiv:2605.10862cs.CL2026-05中稿 · ICDE 2026被引 3

用规则解释检索增强大模型输出,提升安全性验证效率

RUBEN: Rule-Based Explanations for Retrieval-Augmented LLM Systems

论文配图:RUBEN: Rule-Based Explanations for Retrieval-Augmented LLM Systems
图 1 · 摘自论文原文
  • 通过剪枝策略找到最简规则集,覆盖所有推理路径
  • 规则集可识别安全训练漏洞,检测对抗提示攻击效果
  • 适合关注大模型可解释性与安全评估的研究者

本文提出RUBEN,一种交互式工具,用于在数据驱动应用中发现最小规则集,以解释检索增强型大语言模型(LLM)的输出。我们采用新颖的剪枝策略,高效识别出能够涵盖所有其他规则的最小规则集。此外,我们进一步展示了这些规则在大模型安全领域的创新应用,具体包括测试安全训练的鲁棒性以及评估对抗性提示注入的有效性。

原文摘要 · Abstract (English)

This paper demonstrates RUBEN, an interactive tool for discovering minimal rules to explain the outputs of retrieval-augmented large language models (LLMs) in data-driven applications. We leverage novel pruning strategies to efficiently identify a minimal set of rules that subsume all others. We further demonstrate novel applications of these rules for LLM safety, specifically to test the resiliency of safety training and effectiveness of adversarial prompt injections.

可解释性大模型安全规则挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。