深度神经网络仍占主导,量子与脉冲网络难撼其地位。
The Enduring Dominance of Deep Neural Networks: A Critical Analysis of the Fundamental Limitations of Quantum Machine Learning and Spiking Neural Networks
- 分析量子与脉冲网络的底层限制,指出其难以替代DNN
- 实证显示优化后的DNN在性能与能效上均优于二者
- 适合关注实际落地、模型效率与工程可行性的研究者
近期量子机器学习(QML)和脉冲神经网络(SNNs)虽引发热潮,声称可实现指数级加速与类脑能效,但本文认为它们短期内难以取代深度神经网络(DNN)。QML受限于酉操作约束、测量导致态坍缩、贫瘠高原问题及高测量开销,且当前噪声中等规模量子硬件条件有限,正则化不足导致过拟合风险,与机器学习泛化目标不匹配。SNN因离散脉冲机制,表达能力受限,难以处理长程依赖与语言语义编码;其模拟大脑的目标反而引入认知偏差、工作记忆有限与学习缓慢等固有低效。所谓能效优势也被夸大:经量化的优化DNN在真实条件下能耗低于SNN。此外,SNN训练需时间展开,计算开销巨大。相较之下,DNN凭借高效反向传播、鲁棒正则化及基于大语言模型(LRMs)的推理时扩展,通过强化学习与蒙特卡洛树搜索(MCTS)实现自进化,缓解数据稀缺。例如xAI的Grok-4 Heavy与gpt-oss-120b(仅1200亿参数,单80GB GPU部署)已达到或超越主流工业模型性能。专用集成电路(ASIC)进一步放大其效率优势。因此,尽管QML与SNN或可担任特定混合角色,但DNN仍是当前人工智能发展的主导实践范式。
原文摘要 · Abstract (English)
Recent advancements in QML and SNNs have generated considerable excitement, promising exponential speedups and brain-like energy efficiency to revolutionize AI. However, this paper argues that they are unlikely to displace DNNs in the near term. QML struggles with adapting backpropagation due to unitary constraints, measurement-induced state collapse, barren plateaus, and high measurement overheads, exacerbated by the limitations of current noisy intermediate-scale quantum hardware, overfitting risks due to underdeveloped regularization techniques, and a fundamental misalignment with machine learning's generalization. SNNs face restricted representational bandwidth, struggling with long-range dependencies and semantic encoding in language tasks due to their discrete, spike-based processing. Furthermore, the goal of faithfully emulating the brain might impose inherent inefficiencies like cognitive biases, limited working memory, and slow learning speeds. Even their touted energy-efficient advantages are overstated; optimized DNNs with quantization can outperform SNNs in energy costs under realistic conditions. Finally, SNN training incurs high computational overhead from temporal unfolding. In contrast, DNNs leverage efficient backpropagation, robust regularization, and innovations in LRMs that shift scaling to inference-time compute, enabling self-improvement via RL and search algorithms like MCTS while mitigating data scarcity. This superiority is evidenced by recent models such as xAI's Grok-4 Heavy, which advances SOTA performance, and gpt-oss-120b, which surpasses or approaches the performance of leading industry models despite its modest 120-billion-parameter size deployable on a single 80GB GPU. Furthermore, specialized ASICs amplify these efficiency gains. Ultimately, QML and SNNs may serve niche hybrid roles, but DNNs remain the dominant, practical paradigm for AI advancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。