让大模型输出更可信:通过词元概率提升多标签安全分类的可解释性
Token-Level Marginalization for Multi-Label LLM Classifiers
- 基于词元逻辑斯蒂值估算每类置信度,解决生成模型无直接概率的问题
- 在合成数据集上验证,该方法显著提升分类结果的可解释性和可靠性
- 适用于需精细内容审核的场景,如平台安全策略制定与错误分析
本文针对生成式语言模型在多标签内容安全分类中缺乏直观类别置信度的问题,提出并评估三种新的词元级概率估计方法。尽管像LLaMA Guard这样的模型能有效识别不安全内容及其类别,但其生成架构天然不提供类级别概率,导致难以评估模型置信度和解释性能,影响动态阈值设定与细粒度错误分析。研究通过在人工生成、严格标注的数据集上进行大量实验,证明利用词元逻辑斯蒂值可显著提升生成式分类器的可解释性与可靠性,支持更精细的内容安全治理。
原文摘要 · Abstract (English)
This paper addresses the critical challenge of deriving interpretable confidence scores from generative language models (LLMs) when applied to multi-label content safety classification. While models like LLaMA Guard are effective for identifying unsafe content and its categories, their generative architecture inherently lacks direct class-level probabilities, which hinders model confidence assessment and performance interpretation. This limitation complicates the setting of dynamic thresholds for content moderation and impedes fine-grained error analysis. This research proposes and evaluates three novel token-level probability estimation approaches to bridge this gap. The aim is to enhance model interpretability and accuracy, and evaluate the generalizability of this framework across different instruction-tuned models. Through extensive experimentation on a synthetically generated, rigorously annotated dataset, it is demonstrated that leveraging token logits significantly improves the interpretability and reliability of generative classifiers, enabling more nuanced content safety moderation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。