教AI自判可信度,该自动处理还是交给人类。
LLM Performance Predictors: Learning When to Escalate in Hybrid Human-AI Moderation Systems
- 用模型输出的概率、熵值等构建可信度预测器。
- 实测可降低人工审核量,同时提升判断准确率。
- 适合需要安全可控的AI内容审核场景。
随着大语言模型(LLM)越来越多地融入人机协同的内容审核系统,核心挑战在于判断何时可信赖其输出,何时需转交人工审核。本文提出一种新型监督式不确定性量化框架,通过学习基于大语言模型性能预测因子(LPPs)的元模型,包括对数概率、熵值以及新提出的不确定性归因指标。实验表明,该方法可在真实人机工作流中实现成本敏感的择优分类:对高风险案例进行升级处理,其余则自动化完成。在多模态、多语言审核任务上,对包括Gemini、GPT、Llama、Qwen在内的多种先进LLM进行测试,结果均显著优于现有不确定性估计方法,在准确率与成本权衡上表现更优。此外,LPPs还提升了可解释性,揭示了模型失效的具体情境(如内容模糊或政策定义不清)。本研究建立了一个可扩展、负责任的不确定性感知型人机协同审核框架。
原文摘要 · Abstract (English)
As LLMs are increasingly integrated into human-in-the-loop content moderation systems, a central challenge is deciding when their outputs can be trusted versus when escalation for human review is preferable. We propose a novel framework for supervised LLM uncertainty quantification, learning a dedicated meta-model based on LLM Performance Predictors (LPPs) derived from LLM outputs: log-probabilities, entropy, and novel uncertainty attribution indicators. We demonstrate that our method enables cost-aware selective classification in real-world human-AI workflows: escalating high-risk cases while automating the rest. Experiments across state-of-the-art LLMs, including both off-the-shelf (Gemini, GPT) and open-source (Llama, Qwen), on multimodal and multilingual moderation tasks, show significant improvements over existing uncertainty estimators in accuracy-cost trade-offs. Beyond uncertainty estimation, the LPPs enhance explainability by providing new insights into failure conditions (e.g., ambiguous content vs. under-specified policy). This work establishes a principled framework for uncertainty-aware, scalable, and responsible human-AI moderation workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。