arXiv:2412.16974cs.AIcs.CL2024-12被引 13

构建拒绝行为分类体系,实现对大模型拒绝行为的自动分析与审计

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs

  • 提出16类拒绝行为分类框架,覆盖不可为与不应为两类
  • 构建超8600条人工标注数据集,含合成数据共16万例用于训练
  • 支持黑箱模型拒绝行为审计,助力安全可控的大模型优化

拒绝行为——大语言模型(LLMs)拒绝或未能完全执行用户指令的现象——对AI安全与能力至关重要,尤其在减少幻觉方面。这类行为主要在指令微调(IFT)和基于人类反馈的强化学习(RLHF)阶段习得。然而,现有拒绝行为的分类体系与评估数据集不足,多仅关注“不应为”类而忽略“不可为”类,且缺乏对黑箱模型输出中拒绝内容的审计工具。本文提出一个完整框架:(a) 16类拒绝行为分类体系;(b) 来自公开IFT/RLHF数据集的8600余条人工标注实例;(c) 每类8000例的合成数据集;(d) 经训练的拒绝分类器。该工作使黑箱模型拒绝行为的精确审计成为可能,并支持对大规模IFT/RLHF数据集中的拒绝模式进行自动分析,有助于策略性调整模型拒绝行为,推动更安全可靠的大型语言模型发展。

原文摘要 · Abstract (English)

Refusals - instances where large language models (LLMs) decline or fail to fully execute user instructions - are crucial for both AI safety and AI capabilities and the reduction of hallucinations in particular. These behaviors are learned during post-training, especially in instruction fine-tuning (IFT) and reinforcement learning from human feedback (RLHF). However, existing taxonomies and evaluation datasets for refusals are inadequate, often focusing solely on should-not-related (instead of cannot-related) categories, and lacking tools for auditing refusal content in black-box LLM outputs. We present a comprehensive framework for classifying LLM refusals: (a) a taxonomy of 16 refusal categories, (b) a human-annotated dataset of over 8,600 instances from publicly available IFT and RLHF datasets, (c) a synthetic dataset with 8,000 examples for each refusal category, and (d) classifiers trained for refusal classification. Our work enables precise auditing of refusal behaviors in black-box LLMs and automatic analyses of refusal patterns in large IFT and RLHF datasets. This facilitates the strategic adjustment of LLM refusals, contributing to the development of more safe and reliable LLMs.

拒绝行为大模型安全分类框架黑箱审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。