统一跨模态注意力建模,让AI像人一样看懂图像、视频和声音。
Attend to Anything: Foundation Model for Unified Human Attention Modeling

- 用语言提示与双曲空间嵌入构建从通用到具体的注意力层级。
- 在16个基准上平均性能提升6%,视频推理速度加快4倍。
- 适合做多模态感知、视频理解或需要通用注意力模型的研究者。
现有注意力建模方法在模态、场景和任务间高度碎片化,即使模型规模与数据量不断增长,仍主要局限于特定场景与任务,难以在真实应用中有效泛化。为此,我们提出 Attend to Anything 模型(AAM),一个统一图像、视频与音视频任务的多模态基础模型。AAM 将注意力重新定义为具有通用到具体层次的认知蕴含关系,通过语言提示结合双曲空间中的分层嵌入实现。为统一静态图像与动态视频注意力,引入流体动力学视角,将视频帧注意力建模为由 Fokker--Planck 方程支配的扩散时间演化过程。在16个基准上的实验表明,AAM 在各类场景下平均性能优于当前最优方法6%,同时实现约4倍的视频推理加速。结果证明 AAM 为未来注意力与显著性相关任务提供了原则性基础。数据集与代码将在 https://github.com/wz-zhao/Attend-to-Anything 发布。
原文摘要 · Abstract (English)
Existing human attention (saliency) modeling methods persist as highly fragmented across modalities, scenes, and task formulations. Consequently, even with increasing model capacity and data scale, current models predominantly remain scene-dependent and task-specific, failing to practically generalize in real-world applications. To address the fundamental limitations, we present the Attend to Anything Model (AAM), a multi-modal foundation model that unifies attention modeling across various image, video, and audio-visual tasks and scenes. AAM reformulates attention as a cognitive entailment relationship organized in a general-to-specific hierarchy, implemented through language prompts with hierarchical embeddings in hyperbolic space. Furthermore, to unify static image and dynamic video attention, we adopt a fluid-dynamics perspective, formulating video-frame attention as a diffusive temporal evolution governed by the Fokker--Planck equation. Extensive experiments on 16 benchmarks demonstrate that AAM consistently outperforms state-of-the-art methods by an average of 6\% across various scenarios, while achieving approximately a 4$\times$ speedup in video inference. Overall, these results demonstrate that AAM provides a principled foundation for future research on attention and saliency-related tasks. The dataset and code will be available at https://github.com/wz-zhao/Attend-to-Anything.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。