剖析Transformer在高风险场景下的信任危机与潜在缺陷
In Transformer We Trust? A Perspective on Transformer Architecture Failure Modes
- 从可解释性、鲁棒性等角度系统评估Transformer可靠性
- 发现其在医疗、气候、核科学等关键领域存在共性结构漏洞
- 为安全关键应用提供可信部署的挑战与研究方向
Transformer架构已革新多个领域,涵盖自然语言处理、计算机视觉及科学计算。然而,其在医疗、自动驾驶、气候建模、材料发现、药物研发、核科学和机器人等高风险领域的广泛应用,要求对其可信度有更深入的理解。本文系统评估了Transformer在可解释性、可解释性、对抗攻击鲁棒性、公平性与隐私保护方面的表现。通过跨领域分析,覆盖自然语言处理、计算机视觉、机器人学、医学、地球科学、材料科学、流体动力学、核科学与自动定理证明等领域,揭示了其在关键应用场景中的结构性脆弱性、领域特异性风险及未解的研究挑战,为推动其可靠部署提供了综合洞见。
原文摘要 · Abstract (English)
Transformer architectures have revolutionized machine learning across a wide range of domains, from natural language processing to scientific computing. However, their growing deployment in high-stakes applications, such as computer vision, natural language processing, healthcare, autonomous systems, and critical areas of scientific computing including climate modeling, materials discovery, drug discovery, nuclear science, and robotics, necessitates a deeper and more rigorous understanding of their trustworthiness. In this work, we critically examine the foundational question: \textitHow trustworthy are transformer models?} We evaluate their reliability through a comprehensive review of interpretability, explainability, robustness against adversarial attacks, fairness, and privacy. We systematically examine the trustworthiness of transformer-based models in safety-critical applications spanning natural language processing, computer vision, and science and engineering domains, including robotics, medicine, earth sciences, materials science, fluid dynamics, nuclear science, and automated theorem proving; highlighting high-impact areas where these architectures are central and analyzing the risks associated with their deployment. By synthesizing insights across these diverse areas, we identify recurring structural vulnerabilities, domain-specific risks, and open research challenges that limit the reliable deployment of transformers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。