让AI理解任务描述,预测人看图时的注意力焦点。
TDSal: Task-Based Top-Down Saliency Prediction Model

- 用自然语言描述任务,引导模型生成注意力图。
- 在不同任务下,注意力分布更贴近真实人类行为。
- 适合需要理解用户意图的视觉系统设计者。
视觉显著性旨在预测图像中吸引人类视觉注意的区域。大多数显著性模型假设自由观看条件,但人类注意力常受明确任务目标影响。本文提出一种基于任务的自上而下显著性预测模型,通过自然语言任务描述来调节视觉注意。该模型生成随任务变化的显著性图,反映不同观察意图下的注意力转移。定量与定性分析表明,引入显式任务语义能更准确地建模有目的的视觉注意。
原文摘要 · Abstract (English)
Visual saliency aims to predict the regions of an image most likely to attract human visual attention. While most saliency models assume free-viewing conditions, human attention is often shaped by explicit task goals. In this work, we address task-driven saliency prediction by proposing a model that conditions visual attention on natural-language task descriptions. The model produces task-dependent saliency maps that reflect how attention shifts under different viewing intents. Through quantitative and qualitative analysis, we show that incorporating explicit task semantics enables more faithful modeling of goal-directed visual attention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。