arXiv:2502.20285cs.LGstat.ML2025-02ICML被引 11

为大模型输出风险设计可证明对齐的校准框架,保障极端不良输出可控。

Conformal Tail Risk Control for Large Language Model Alignment

  • 基于分位数加权损失的鲁棒风险控制方法
  • 在高置信度下有效约束大模型的极端劣质输出
  • 适用于对安全性要求高的场景,如内容审核

大型语言模型(LLM)在社会中的广泛应用,对其性能可靠性提出了更高要求。尤其在对风险敏感的应用中,需重点关注意外发生的劣质输出(如毒性回答、冒犯性语言等)这类尾部事件。由于人工标注成本高昂,通用评分模型被用于自动化量化这些尾部事件,但由此可能引入人机评分机制间的偏差。本文提出一种轻量级黑箱模型校准框架,可保证人机评分对齐,并具备理论保障。该框架通过连接共形风险控制与传统统计学中的L-统计量,实现对任意由损失分布分位数加权构成的风险度量的严格控制,在高置信度下有效抑制偏差。我们通过全面实验验证了该框架在缓解人机错位问题上的有效性。

原文摘要 · Abstract (English)

Recent developments in large language models (LLMs) have led to their widespread usage for various tasks. The prevalence of LLMs in society implores the assurance on the reliability of their performance. In particular, risk-sensitive applications demand meticulous attention to unexpectedly poor outcomes, i.e., tail events, for instance, toxic answers, humiliating language, and offensive outputs. Due to the costly nature of acquiring human annotations, general-purpose scoring models have been created to automate the process of quantifying these tail events. This phenomenon introduces potential human-machine misalignment between the respective scoring mechanisms. In this work, we present a lightweight calibration framework for blackbox models that ensures the alignment of humans and machines with provable guarantees. Our framework provides a rigorous approach to controlling any distortion risk measure that is characterized by a weighted average of quantiles of the loss incurred by the LLM with high confidence. The theoretical foundation of our method relies on the connection between conformal risk control and a traditional family of statistics, i.e., L-statistics. To demonstrate the utility of our framework, we conduct comprehensive experiments that address the issue of human-machine misalignment.

大模型对齐风险控制共形推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。