arXiv:2512.10058cs.AIcs.CY2025-12中稿 · presentation at IA…被引 3

揭示AI安全与伦理研究的分裂现状,呼吁融合技术与伦理以构建更可靠的AI系统。

Mind the Gap! Pathways Towards Unifying AI Safety and Ethics Research

  • 通过6442篇论文的网络分析,发现安全与伦理研究高度割裂。
  • 80%合作集中在单一领域内,仅5%论文承担了超85%的跨领域连接。
  • 建议建立共享基准与混合方法,推动技术安全与伦理融合。

尽管人工智能研究多聚焦于能力扩展,但快速发展使构建无害、对齐系统的反制工作愈发紧迫。然而,对齐研究已分化为两大平行路径:以高阶智能、欺骗行为和生存风险为核心的安全部门,以及关注现实危害、社会偏见和生产流程缺陷的伦理部门。两者虽均呼吁增加对齐投入,却对“对齐”定义存在分歧。我们基于2020–2025年十二个主流机器学习与自然语言处理会议的6,442篇论文,进行文献计量与合作者网络分析,发现超过80%的合作发生在安全或伦理内部,跨领域连接高度集中:约5%的论文贡献了超过85%的桥梁链接。移除少数关键连接者会显著加剧隔离,表明跨学科交流依赖少数个体而非广泛协作。结果表明,安全与伦理的分裂不仅是概念性的,更是制度性的,影响研究议程、政策制定与学术场所。我们主张通过共享基准、跨机构平台与混合方法,整合技术安全与规范伦理,以构建既稳健又公正的AI系统。

原文摘要 · Abstract (English)

While much research in artificial intelligence (AI) has focused on scaling capabilities, the accelerating pace of development makes countervailing work on producing harmless, "aligned" systems increasingly urgent. Yet research on alignment has diverged along two largely parallel tracks: safety--centered on scaled intelligence, deceptive or scheming behaviors, and existential risk--and ethics--focused on present harms, the reproduction of social bias, and flaws in production pipelines. Although both communities warn of insufficient investment in alignment, they disagree on what alignment means or ought to mean. As a result, their efforts have evolved in relative isolation, shaped by distinct methodologies, institutional homes, and disciplinary genealogies. We present a large-scale, quantitative study showing the structural split between AI safety and AI ethics. Using a bibliometric and co-authorship network analysis of 6,442 papers from twelve major ML and NLP conferences (2020-2025), we find that over 80% of collaborations occur within either the safety or ethics communities, and cross-field connectivity is highly concentrated: roughly 5% of papers account for more than 85% of bridging links. Removing a small number of these brokers sharply increases segregation, indicating that cross-disciplinary exchange depends on a handful of actors rather than broad, distributed collaboration. These results show that the safety-ethics divide is not only conceptual but institutional, with implications for research agendas, policy, and venues. We argue that integrating technical safety work with normative ethics--via shared benchmarks, cross-institutional venues, and mixed-method methodologies--is essential for building AI systems that are both robust and just.

AI安全伦理对齐跨学科研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。