arXiv:2412.20761cs.CV2024-12被引 1

发现同一类图片中有些更易记住,影响AI识别与学习效果。

Unforgettable Lessons from Forgettable Images: Intra-Class Memorability Matters in Computer Vision

  • 提出细粒度记忆度评分ICMscore,量化同类别图像的记忆难易程度。
  • 高记忆度图像反而降低AI识别与持续学习性能,低记忆度图像表现更好。
  • 可控制图像记忆度的扩散模型,适用于内容编辑与视觉增强场景。

我们提出‘类内记忆度’概念,即同一类别中某些图像比其他图像更易被记住。为探究其成因,设计并开展人类行为实验:参与者需识别序列中重复出现的图像。据此提出新颖的类内记忆度得分(ICMscore),将重复间隔时间纳入计算。构建包含5000余张图像、10个物体类别的类内记忆度数据集(ICMD),基于2000名参与者响应生成评分。后续实验证明,利用该数据集训练的AI模型在记忆度预测、图像识别、持续学习及可控图像编辑任务中均具价值。令人意外的是,高记忆度图像会损害图像识别与持续学习表现,而低记忆度图像提升性能。进一步对先进图像扩散模型进行微调,使其能通过遮蔽语义对象,成功调节图像记忆度。本研究揭示了最易/最难记图像背后的精细视觉特征,为计算机视觉应用开辟新路径。所有代码、数据与模型将公开发布。

原文摘要 · Abstract (English)

We introduce intra-class memorability, where certain images within the same class are more memorable than others despite shared category characteristics. To investigate what features make one object instance more memorable than others, we design and conduct human behavior experiments, where participants are shown a series of images, and they must identify when the current image matches the image presented a few steps back in the sequence. To quantify memorability, we propose the Intra-Class Memorability score (ICMscore), a novel metric that incorporates the temporal intervals between repeated image presentations into its calculation. Furthermore, we curate the Intra-Class Memorability Dataset (ICMD), comprising over 5,000 images across ten object classes with their ICMscores derived from 2,000 participants' responses. Subsequently, we demonstrate the usefulness of ICMD by training AI models on this dataset for various downstream tasks: memorability prediction, image recognition, continual learning, and memorability-controlled image editing. Surprisingly, high-ICMscore images impair AI performance in image recognition and continual learning tasks, while low-ICMscore images improve outcomes in these tasks. Additionally, we fine-tune a state-of-the-art image diffusion model on ICMD image pairs with and without masked semantic objects. The diffusion model can successfully manipulate image elements to enhance or reduce memorability. Our contributions open new pathways in understanding intra-class memorability by scrutinizing fine-grained visual features behind the most and least memorable images and laying the groundwork for real-world applications in computer vision. We will release all code, data, and models publicly.

记忆度图像识别扩散模型持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。