arXiv:2505.14933cs.LG2025-05

让机器学习模型学会识别未知输入,提升开放世界下的可靠性。

Foundations of Unknown-aware Machine Learning

  • 提出未知感知学习框架,无需标注的外部数据即可训练模型识别新类别。
  • 通过合成未知样本提升模型对分布外数据的检测能力,显著降低误判率。
  • 适用于大模型安全防护,适合关注AI可靠性和安全的研究者与工程师。

在开放世界部署中确保机器学习模型的可靠性和安全性是人工智能安全的核心挑战。本论文从标准神经网络到大型语言模型(LLMs)等现代基础模型,构建了算法与理论基础,应对分布不确定性与未知类别的可靠性问题。传统学习范式(如经验风险最小化)假设训练与推理分布一致,常导致对分布外(OOD)输入产生过度自信的预测。本文提出新型联合优化框架,在保证分布内准确率的同时提升对未知数据的可靠性。核心贡献包括未知感知学习框架,使模型能在无标注OOD数据的情况下识别并处理新输入;提出新的异常样本合成方法(VOS、NPOS、DREAM-OOD),用于训练时生成有信息量的未知样本。在此基础上,提出SAL框架,利用未标注的真实世界数据增强在实际部署条件下的OOD检测能力。结果表明,大量未标注数据可被有效利用以识别和适应意外输入,并提供形式化可靠性保障。论文还将可靠学习扩展至基础模型,开发了针对LLMs的幻觉检测工具HaloScope、防御多模态模型恶意提示的MLLMGuard,以及用于净化人类反馈数据的去噪方法,以改善对齐效果。这些工作推动未知感知学习成为新范式,有望以最少人工投入提升AI系统的可靠性。

原文摘要 · Abstract (English)

Ensuring the reliability and safety of machine learning models in open-world deployment is a central challenge in AI safety. This thesis develops both algorithmic and theoretical foundations to address key reliability issues arising from distributional uncertainty and unknown classes, from standard neural networks to modern foundation models like large language models (LLMs). Traditional learning paradigms, such as empirical risk minimization (ERM), assume no distribution shift between training and inference, often leading to overconfident predictions on out-of-distribution (OOD) inputs. This thesis introduces novel frameworks that jointly optimize for in-distribution accuracy and reliability to unseen data. A core contribution is the development of an unknown-aware learning framework that enables models to recognize and handle novel inputs without labeled OOD data. We propose new outlier synthesis methods, VOS, NPOS, and DREAM-OOD, to generate informative unknowns during training. Building on this, we present SAL, a theoretical and algorithmic framework that leverages unlabeled in-the-wild data to enhance OOD detection under realistic deployment conditions. These methods demonstrate that abundant unlabeled data can be harnessed to recognize and adapt to unforeseen inputs, providing formal reliability guarantees. The thesis also extends reliable learning to foundation models. We develop HaloScope for hallucination detection in LLMs, MLLMGuard for defending against malicious prompts in multimodal models, and data cleaning methods to denoise human feedback used for better alignment. These tools target failure modes that threaten the safety of large-scale models in deployment. Overall, these contributions promote unknown-aware learning as a new paradigm, and we hope it can advance the reliability of AI systems with minimal human efforts.

未知感知模型可靠性大模型安全OOD检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。