arXiv:2511.14440cs.CV2025-11被引 4

模拟婴儿视觉发育过程,让机器学得更鲁棒的视觉能力。

Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines

  • 用灰度-彩色、模糊-清晰、连续时间三阶段模拟婴儿视觉训练模型。
  • 在10个数据集上均提升识别鲁棒性,且行为类似婴儿视崖反应。
  • 结合早期视觉训练的模型在域内域外任务中表现更优,适合视觉智能研究。

新生儿视觉具有低分辨率、色彩退化和时间连续性的特点,随发育逐步清晰。为探究这种阶段性‘视觉饮食’的生态优势,我们在物体中心视频上训练自监督学习(SSL)模型,施加灰度到彩色(C)、模糊到清晰(A)、保留时间连续性(T)的约束,统称CATDiet。评估涵盖10个数据集,包括干净与受损图像识别、纹理-形状冲突测试、轮廓识别、深度顺序分类及视崖范式。所有CATDiet变体在仅使用物体中心视频训练的情况下,均表现出更强的物体识别鲁棒性。令人惊讶的是,模型还展现出生物对齐的发展模式,如神经可塑性变化类似猕猴V1区突触密度变化,行为类似婴儿视崖反应。基于此,CombDiet在标准训练前先用CATDiet初始化,并保持时间连续性。在物体中心或头戴式婴儿视角视频上训练后,其在域内与域外物体识别和深度感知上均优于标准SSL。结果表明,婴儿视觉发展过程为理解机器鲁棒视觉智能的涌现提供了有力的逆向工程框架。所有代码、数据和模型均已开源于Github。

原文摘要 · Abstract (English)

Newborns perceive the world with low-acuity, color-degraded, and temporally continuous vision, which gradually sharpens as infants develop. To explore the ecological advantages of such staged "visual diets", we train self-supervised learning (SSL) models on object-centric videos under constraints that simulate infant vision: grayscale-to-color (C), blur-to-sharp (A), and preserved temporal continuity (T)-collectively termed CATDiet. For evaluation, we establish a comprehensive benchmark across ten datasets, covering clean and corrupted image recognition, texture-shape cue conflict tests, silhouette recognition, depth-order classification, and the visual cliff paradigm. All CATDiet variants demonstrate enhanced robustness in object recognition, despite being trained solely on object-centric videos. Remarkably, models also exhibit biologically aligned developmental patterns, including neural plasticity changes mirroring synaptic density in macaque V1 and behaviors resembling infants' visual cliff responses. Building on these insights, CombDiet initializes SSL with CATDiet before standard training while preserving temporal continuity. Trained on object-centric or head-mounted infant videos, CombDiet outperforms standard SSL on both in-domain and out-of-domain object recognition and depth perception. Together, these results suggest that the developmental progression of early infant visual experience offers a powerful reverse-engineering framework for understanding the emergence of robust visual intelligence in machines. All code, data, and models are available at Github.

视觉智能自监督学习发育模拟鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。