构建对话合规检测基准,验证大模型判官的漏洞识别能力
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems
- 用可控缺陷注入生成带精确标签的对话数据
- 现役大模型在复杂规则检测中表现不佳,准确率不足60%
- 小模型微调后超越大模型,适合企业级应用
随着大语言模型在企业场景中作为任务导向代理的广泛应用,确保其严格遵守复杂的领域特定操作规范至关重要。尽管采用大模型作为评判者是可扩展评估的有前景方案,但这类评判者在检测具体政策违规方面的可靠性仍缺乏系统研究。这主要源于缺乏系统的数据生成方法,受限于细粒度人工标注的成本高昂以及真实代理违规行为的合成困难。本文提出CompliBench,一个用于评估大模型评判者在多轮对话中检测并定位规范违规行为能力的新基准。为克服数据稀缺问题,我们开发了一种可扩展的自动化数据生成流水线,模拟用户-代理交互。通过可控缺陷注入过程,自动产生精确的违规指南和具体对话轮次的真值标签;结合对抗搜索方法,确保引入的扰动具有高度挑战性。全面评估表明,当前最先进的专有大模型在此任务上表现显著不足。此外,我们证明,在合成数据上微调的小规模判官模型优于领先的大模型,并能在未见业务领域良好泛化,凸显该流水线作为训练鲁棒生成奖励模型的有效基础。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are increasingly deployed as task-oriented agents in enterprise environments, ensuring their strict adherence to complex, domain-specific operational guidelines is critical. While utilizing an LLM-as-a-Judge is a promising solution for scalable evaluation, the reliability of these judges in detecting specific policy violations remains largely unexplored. This gap is primarily due to the lack of a systematic data generation method, which has been hindered by the extensive cost of fine-grained human annotation and the difficulty of synthesizing realistic agent violations. In this paper, we introduce CompliBench, a novel benchmark designed to evaluate the ability of LLM judges to detect and localize guideline violations in multi-turn dialogues. To overcome data scarcity, we develop a scalable, automated data generation pipeline that simulates user-agent interactions. Our controllable flaw injection process automatically yields precise ground-truth labels for the violated guideline and the exact conversation turn, while an adversarial search method ensures these introduced perturbations are highly challenging. Our comprehensive evaluation reveals that current state-of-the-art proprietary LLMs struggle significantly with this task. In addition, we demonstrate that a small-scale judge model fine-tuned on our synthesized data outperforms leading LLMs and generalizes well to unseen business domains, highlighting our pipeline as an effective foundation for training robust generative reward models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。