简单模型更抗干扰,深层模型更贴近大脑,双赛道冠军方案揭示视觉智能新规律。
Navigating Simply, Aligning Deeply: Winning Solutions for Mouse vs. AI 2025
- 轻量两层CNN加门控单元,95.4%得分实现强鲁棒性
- 16层深网络仅1780万参数,神经预测性能达顶尖水平
- 训练20万步时效果最佳,非越练越好,适合生物视觉研究者
视觉鲁棒性与神经对齐仍是构建类生物视觉系统的关键挑战。本文呈现了团队 HCMUS_TheFangs 在 NeurIPS 2025 Mouse vs. AI:鲁棒视觉觅食竞赛中两个赛道的优胜方案。在赛道1(视觉鲁棒性)中,我们证明架构简化结合特定组件可实现优异泛化能力,采用轻量级两层CNN,辅以门控线性单元(GLU)和观测归一化,最终取得95.4%的得分。在赛道2(神经对齐)中,我们设计了一种类似ResNet的深层结构,含16个卷积层,使用基于GLU的门控机制,仅1780万参数即达到顶峰神经预测性能。通过对60K至114万训练步数间十组模型检查点的系统分析,发现训练时长与性能呈非单调关系,约20万步时表现最优。通过全面消融实验与失败案例分析,揭示了为何简单架构在视觉鲁棒性上更优,而更深模型因容量更大,在神经对齐上表现更佳。结果挑战了关于模型复杂度的传统认知,为开发鲁棒且生物启发的视觉智能体提供了实用指导。
原文摘要 · Abstract (English)
Visual robustness and neural alignment remain critical challenges in developing artificial agents that can match biological vision systems. We present the winning approaches from Team HCMUS_TheFangs for both tracks of the NeurIPS 2025 Mouse vs. AI: Robust Visual Foraging Competition. For Track 1 (Visual Robustness), we demonstrate that architectural simplicity combined with targeted components yields superior generalization, achieving 95.4% final score with a lightweight two-layer CNN enhanced by Gated Linear Units and observation normalization. For Track 2 (Neural Alignment), we develop a deep ResNet-like architecture with 16 convolutional layers and GLU-based gating that achieves top-1 neural prediction performance with 17.8 million parameters. Our systematic analysis of ten model checkpoints trained between 60K to 1.14M steps reveals that training duration exhibits a non-monotonic relationship with performance, with optimal results achieved around 200K steps. Through comprehensive ablation studies and failure case analysis, we provide insights into why simpler architectures excel at visual robustness while deeper models with increased capacity achieve better neural alignment. Our results challenge conventional assumptions about model complexity in visuomotor learning and offer practical guidance for developing robust, biologically-inspired visual agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。