arXiv:2505.21589cs.CVcs.AI2025-05

用视觉错觉数据集暴露AI解释性局限,揭示关键感知线索影响模型判断。

Do you see what I see? An Ambiguous Optical Illusion Dataset exposing limitations of Explainable AI

  • 构建含动物对的视觉错觉数据集,聚焦凝视方向与眼神线索的影响。
  • 发现眼神方向等细微特征显著降低模型准确率,最高下降达18%。
  • 适合研究视觉偏差、人机感知对齐及可解释AI的学者使用。

从不确定性量化到真实世界物体检测,机器学习算法在自动驾驶、医疗诊断等安全关键领域的重要性日益凸显。模糊数据在多个机器学习领域扮演关键角色,而视觉错觉为理解人类与机器感知局限提供了独特视角。尽管如此,相关数据集仍十分稀缺。本文提出一个新型视觉错觉数据集,包含交织的动物配对,旨在引发感知歧义。我们识别出通用视觉概念,特别是凝视方向与眼神线索,这些微妙特征显著影响模型性能。通过引入感知歧义,研究结果强调了视觉学习中概念的重要性,并为探究偏见与人机视觉对齐提供基础。为提升通用性,我们系统化生成涵盖不同概念的错觉图像,相关内容见偏见缓解部分。数据集已上传至Kaggle(https://kaggle.com/datasets/693bf7c6dd2cb45c8a863f9177350c8f9849a9508e9d50526e2ffcc5559a8333),源代码见GitHub(https://github.com/KDD-OpenSource/Ambivision.git)。

原文摘要 · Abstract (English)

From uncertainty quantification to real-world object detection, we recognize the importance of machine learning algorithms, particularly in safety-critical domains such as autonomous driving or medical diagnostics. In machine learning, ambiguous data plays an important role in various machine learning domains. Optical illusions present a compelling area of study in this context, as they offer insight into the limitations of both human and machine perception. Despite this relevance, optical illusion datasets remain scarce. In this work, we introduce a novel dataset of optical illusions featuring intermingled animal pairs designed to evoke perceptual ambiguity. We identify generalizable visual concepts, particularly gaze direction and eye cues, as subtle yet impactful features that significantly influence model accuracy. By confronting models with perceptual ambiguity, our findings underscore the importance of concepts in visual learning and provide a foundation for studying bias and alignment between human and machine vision. To make this dataset useful for general purposes, we generate optical illusions systematically with different concepts discussed in our bias mitigation section. The dataset is accessible in Kaggle via https://kaggle.com/datasets/693bf7c6dd2cb45c8a863f9177350c8f9849a9508e9d50526e2ffcc5559a8333. Our source code can be found at https://github.com/KDD-OpenSource/Ambivision.git.

视觉错觉可解释AI感知偏差机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。