arXiv:2511.18633cs.AIcs.LG2025-11

用结构主义哲学分析神经网络表征,揭示其隐含的哲学立场。

Bridging Philosophy and Machine Learning: A Structuralist Framework for Classifying Neural Network Representations

  • 基于结构主义哲学构建分类框架,分析模型表征的本体论倾向。
  • 发现二十年研究中普遍倾向结构唯心主义,表征依赖模型架构与训练。
  • 为机器学习与科学哲学交叉研究提供清晰概念工具,适合跨学科研究者。

机器学习模型日益成为表征系统,但其内部结构背后的哲学假设却鲜受审视。本文提出一种结构主义决策框架,用于分类机器学习研究中关于神经网络表征的隐含本体论承诺。通过改进的PRISMA协议,对过去二十年表征学习与可解释性领域的文献进行系统综述,并以三个源自结构主义科学哲学的层级标准(实体消解、结构来源、存在方式)分析五篇有影响力论文。结果表明,研究普遍倾向于结构唯心主义:所学表征被视为由架构、数据先验和训练动态塑造的模型依赖构造。消解性与非消解性结构主义立场被选择性采用,而结构实在论则显著缺失。该框架澄清了可解释性、涌现与认识信任等争论中的概念矛盾,为未来哲学与机器学习的跨学科工作提供了严谨基础。

原文摘要 · Abstract (English)

Machine learning models increasingly function as representational systems, yet the philosoph- ical assumptions underlying their internal structures remain largely unexamined. This paper develops a structuralist decision framework for classifying the implicit ontological commitments made in machine learning research on neural network representations. Using a modified PRISMA protocol, a systematic review of the last two decades of literature on representation learning and interpretability is conducted. Five influential papers are analysed through three hierarchical criteria derived from structuralist philosophy of science: entity elimination, source of structure, and mode of existence. The results reveal a pronounced tendency toward structural idealism, where learned representations are treated as model-dependent constructions shaped by architec- ture, data priors, and training dynamics. Eliminative and non-eliminative structuralist stances appear selectively, while structural realism is notably absent. The proposed framework clarifies conceptual tensions in debates on interpretability, emergence, and epistemic trust in machine learning, and offers a rigorous foundation for future interdisciplinary work between philosophy of science and machine learning.

结构主义表征学习哲学与AI可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。