arXiv:2412.09765cs.CVcs.HC2024-12ICLR被引 2

用神经网络选图并增强图像,让人类学分类快准稳

L-WISE: Boosting Human Visual Category Learning Through Model-Based Image Selection and Enhancement

  • 用模型预测图片难易度,挑难图重点练
  • 对难图加干扰提升识别率,准确率提高33%-72%
  • 适合医学图像等专业领域新手快速上手

当前领先的视觉通路神经网络模型在视觉分类任务中表现出与人类高度一致的行为。我们发现,这些模型生成的图像扰动能显著提升人类对真实类别报告的准确性。同时,模型可直接用于预测个体图像的人类正确反应比例,提供一种简单的人类对齐的图像难度估计器。基于此,我们提出一种学习增强方法:(i) 根据模型估计的识别难度选择图像;(ii) 对新手学习者使用有助于识别的图像扰动。结果表明,结合这两种策略后,受试者在未修改的保留测试图像上的分类准确率相对对照组提升33%-72%,且训练时间缩短20%-23%(尽管训练次数相同)。该方法在细粒度分类任务及临床相关的组织病理学和皮肤镜图像任务中均有效。据我们所知,这是首个通过增强类别特异性图像特征来提升人类视觉学习性能的神经网络应用。

原文摘要 · Abstract (English)

The currently leading artificial neural network models of the visual ventral stream - which are derived from a combination of performance optimization and robustification methods - have demonstrated a remarkable degree of behavioral alignment with humans on visual categorization tasks. We show that image perturbations generated by these models can enhance the ability of humans to accurately report the ground truth class. Furthermore, we find that the same models can also be used out-of-the-box to predict the proportion of correct human responses to individual images, providing a simple, human-aligned estimator of the relative difficulty of each image. Motivated by these observations, we propose to augment visual learning in humans in a way that improves human categorization accuracy at test time. Our learning augmentation approach consists of (i) selecting images based on their model-estimated recognition difficulty, and (ii) applying image perturbations that aid recognition for novice learners. We find that combining these model-based strategies leads to categorization accuracy gains of 33-72% relative to control subjects without these interventions, on unmodified, randomly selected held-out test images. Beyond the accuracy gain, the training time for the augmented learning group was also shortened by 20-23%, despite both groups completing the same number of training trials. We demonstrate the efficacy of our approach in a fine-grained categorization task with natural images, as well as two tasks in clinically relevant image domains - histology and dermoscopy - where visual learning is notoriously challenging. To the best of our knowledge, our work is the first application of artificial neural networks to increase visual learning performance in humans by enhancing category-specific image features.

视觉学习神经网络医学图像图像增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。