用强化学习建模个人化眼动,精准预测广告视频中多属性注视点。
PRE-MAP: Personalized Reinforced Eye-tracking Multimodal LLM for High-Resolution Multi-Attribute Point Prediction
- 基于多属性用户画像与强化学习优化眼动模型,捕捉个体差异。
- 在4500人参与的SPA-ADV数据集上,点位预测精度显著优于现有方法。
- 适合关注个性化视觉注意力、广告内容优化的研究者与工程师。
视觉选择性注意由个体偏好驱动,连接主观认知机制与客观视觉元素,引导动态视觉场景的语义解读与层级处理。然而,现有模型与数据集大多忽略主观认知多样性对注视行为的影响。传统显著性预测模型通常采用分割方法,依赖低分辨率图像生成热图并上采样至原生分辨率,难以捕捉个性化注意力模式。此外,多模态大模型(MLLM)易产生幻觉,在涉及多点预测的任务中严格遵循预期格式成本高昂,精确点定位困难。为此,我们构建了大规模广告视频眼动数据集SPA-ADV,涵盖486个视频及超过4500名年龄与性别各异的参与者。同时提出PRE-MAP模型,通过强化学习优化的眼动机制,结合多属性用户画像,实现个性化注视点预测。为确保MLLM输出格式正确且空间准确,引入一致性组相对策略优化(C-GRPO),借鉴眼动点与多属性画像的变异性。在SPA-ADV及其他基准上的实验验证了该方法的有效性。代码与数据集已开源。
原文摘要 · Abstract (English)
Visual selective attention, driven by individual preferences, regulates human prioritization of visual stimuli by bridging subjective cognitive mechanisms with objective visual elements, thereby steering the semantic interpretation and hierarchical processing of dynamic visual scenes. However, existing models and datasets predominantly neglect the influence of subjective cognitive diversity on fixation behavior. Conventional saliency prediction models, typically employing segmentation approaches, rely on low-resolution imagery to generate saliency heatmaps, subsequently upscaled to native resolutions, which limiting their capacity to capture personalized attention patterns. Furthermore, MLLMs are constrained by factors such as hallucinations, making it very costly to strictly adhere to the expected format in tasks involving multiple point predictions, and achieving precise point positioning is challenging. To address these limitations, we present Subjective Personalized Attention for Advertisement Videos, namely SPA-ADV, a large-scale multimodal dataset capturing gaze behaviors from over 4,500 participants varying in age and gender with 486 videos. Furthermore, we propose PRE-MAP, a novel eye-tracking saliency model that characterizes Personalized visual disparities through Reinforcement learning-optimized Eye-tracking, built upon MLLMs and guided by Multi-Attribute user profiles to predict Points. To ensure MLLMs produce prediction points that are both format-correct and spatially accurate, we introduce Consistency Group Relative Policy Optimization (C-GRPO), inspired by the variability in eye movement points and Multi-Attribute profiles. Extensive experiments on SPA-ADV and other benchmarks demonstrate the effectiveness of our approach. The code and dataset are available at \href{https://github.com/mininglamp-MLLM/PRE-MAP}{this URL}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。