解决机器学习在分布偏移下的可靠性问题,提升模型鲁棒性与可信度。
Trustworthy Machine Learning under Distribution Shifts
- 区分三类分布偏移:扰动、域、模态,系统研究其影响机制。
- 从鲁棒性、可解释性、自适应性三方面构建可信学习框架。
- 适用于需要高可靠性的实际场景,如医疗、自动驾驶。
机器学习(ML)作为人工智能(AI)的核心,推动了从视觉识别到跨模态对齐的多项突破。然而,分布偏移仍是制约其可靠性与通用性的根本难题,尤其在模型泛化时引发信任危机。本文聚焦于「分布偏移下的可信机器学习」,系统分析三类常见分布偏移:(1)扰动偏移,(2)域偏移,(3)模态偏移;并从(1)鲁棒性、(2)可解释性、(3)自适应性三个维度深入探讨可信性。基于此,提出有效解决方案与基础性洞见,旨在提升模型在效率、适应性与安全性等关键问题上的表现。
原文摘要 · Abstract (English)
Machine Learning (ML) has been a foundational topic in artificial intelligence (AI), providing both theoretical groundwork and practical tools for its exciting advancements. From ResNet for visual recognition to Transformer for vision-language alignment, the AI models have achieved superior capability to humans. Furthermore, the scaling law has enabled AI to initially develop general intelligence, as demonstrated by Large Language Models (LLMs). To this stage, AI has had an enormous influence on society and yet still keeps shaping the future for humanity. However, distribution shift remains a persistent ``Achilles' heel'', fundamentally limiting the reliability and general usefulness of ML systems. Moreover, generalization under distribution shift would also cause trust issues for AIs. Motivated by these challenges, my research focuses on \textit{Trustworthy Machine Learning under Distribution Shifts}, with the goal of expanding AI's robustness, versatility, as well as its responsibility and reliability. We carefully study the three common distribution shifts into: (1) Perturbation Shift, (2) Domain Shift, and (3) Modality Shift. For all scenarios, we also rigorously investigate trustworthiness via three aspects: (1) Robustness, (2) Explainability, and (3) Adaptability. Based on these dimensions, we propose effective solutions and fundamental insights, meanwhile aiming to enhance the critical ML problems, such as efficiency, adaptability, and safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。