让AI模仿人类视觉发育过程,显著提升对形状的识别能力。
Adopting a human developmental visual diet yields robust, shape-based AI vision
- 用人类视觉发展规律设计训练课程,引导AI学习
- 模型对形状依赖度达当前最高,抽象形状识别超越现有水平
- 适合追求鲁棒性与人机一致性的AI视觉研究者
尽管人工智能系统规模不断扩张,但人工视觉与人类视觉之间仍存在显著差异。与人类不同,当前的AI严重依赖纹理特征而非形状信息,对图像畸变不鲁棒,极易受到对抗攻击,且在复杂背景下难以识别简单抽象形状。为此,本文借鉴人类从婴儿期到成年期的视觉发展过程,将多年研究成果整合为一种新型的“人类视觉发育饮食”(DVD)用于AI视觉训练。通过这一类比人类发展的渐进式训练路径,考虑视觉敏锐度、对比敏感度和色彩感知的发展,所生成的模型在各项稳健视觉测试中均更接近人类表现,展现出迄今为止最强的形状依赖性,抽象形状识别能力突破现有水平,并对图像失真和对抗攻击具有更强抵抗力。结果表明,通过优化学习路径而非单纯扩大数据量,可实现更高效、更安全、更接近人类的视觉系统。
原文摘要 · Abstract (English)
Despite years of research and the dramatic scaling of artificial intelligence (AI) systems, a striking misalignment between artificial and human vision persists. Contrary to humans, AI relies heavily on texture-features rather than shape information, lacks robustness to image distortions, remains highly vulnerable to adversarial attacks, and struggles to recognise simple abstract shapes within complex backgrounds. To close this gap, here we take inspiration from how human vision develops from early infancy into adulthood. We quantified visual maturation by synthesising decades of research into a novel developmental visual diet (DVD) for AI vision. Guiding AI systems through this human-inspired curriculum, which considers the development of visual acuity, contrast sensitivity, and colour, produces models that better align with human behaviour on every hallmark of robust vision tested, yielding the strongest reported reliance on shape information to date, abstract shape recognition beyond the state of the art, and higher resilience to image corruptions and adversarial attacks. Our results thus demonstrate that robust AI vision can be achieved by guiding how a model learns, not merely how much it learns, offering a resource-efficient route toward safer and more human-like artificial visual systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。