arXiv:2604.13883cs.CVcs.LG2026-04被引 1

让机器像人一样根据上下文理解图像,准确率提升15%。

Context Sensitivity Improves Human-Machine Visual Alignment

论文配图:Context Sensitivity Improves Human-Machine Visual Alignment
图 1 · 摘自论文原文
  • 用锚图作为上下文,动态计算图像相似性
  • 在奇偶判断任务中准确率最高提升15%
  • 适用于原始和人类对齐的视觉基础模型

现代机器学习模型通常将输入表示为高维嵌入空间中的固定点,尽管这种方法在众多下游任务中表现强劲,但其处理信息的方式与人类存在根本差异。人类会持续适应环境,以高度依赖上下文的方式表征物体及其关系。为弥合这一差距,我们提出一种从神经网络嵌入中进行上下文敏感相似性计算的方法,并应用于以锚图作为同时上下文的三元组奇一出任务。引入上下文建模使我们在奇一出任务上的准确率相比无上下文感知模型最高提升15%,且该提升在原始和“人类对齐”的视觉基础模型上均保持一致。

原文摘要 · Abstract (English)

Modern machine learning models typically represent inputs as fixed points in a high-dimensional embedding space. While this approach has been proven powerful for a wide range of downstream tasks, it fundamentally differs from the way humans process information. Because humans are constantly adapting to their environment, they represent objects and their relationships in a highly context-sensitive manner. To address this gap, we propose a method for context-sensitive similarity computation from neural network embeddings, applied to modeling a triplet odd-one-out task with an anchor image serving as simultaneous context. Modeling context enables us to achieve up to a 15% improvement in odd-one-out accuracy over a context-insensitive model. We find that this improvement is consistent across both original and "human-aligned" vision foundation models.

视觉对齐上下文感知模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。