解决视觉语言模型在跨市场广告偏好预测中的关键上下文低估问题
GeoReward: Mitigating Contextual Variable Overestimation in Vision-Language Models for Cross-Market Preference Prediction

- 设计三重机制增强对稀疏但关键的地域上下文感知
- 在多国广告数据集上显著提升跨市场偏好预测准确率
- 适合需要精准地域化内容生成的广告与推荐系统应用
视觉语言模型在多模态任务中表现优异,但存在一种隐蔽而严重的问题:过度依赖主导的视觉-文本线索,忽视稀疏却决策关键的上下文变量。我们称此为上下文变量过估计(CVE),在跨地理市场广告偏好预测中尤为明显。例如,当比较针对不同国家定制的产品图像时,模型常输出一致结果,忽略真实存在的地区差异。这是由于高密度信号(如产品属性、密集图像块)掩盖了少数关键的市场特定标记。为此,我们构建了一个包含多国真实广告创意及其点击率数据的新多模态数据集,并提出GeoReward奖励模型,通过三种专设机制:(1)市场感知检索增强,(2)上下文引导的视觉调制,(3)选择性敏感度损失来缓解此问题。实验表明,该框架有效减轻了CVE,优于现有基线。本工作不仅揭示了视觉语言模型对主导感知特征的系统性偏差,也为稀疏上下文决定决策的应用提供了针对性解决方案。
原文摘要 · Abstract (English)
Vision-language models excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate dominant visual-textual cues while underestimating sparse but decision-critical contextual variables. This issue, which we term Contextual Variable Overestimation (CVE), becomes particularly evident in real-world applications such as predicting advertisement image preferences across diverse geographic markets. For instance, when a VLM is asked to choose between two product images tailored for different countries, it often defaults to a consistent output, ignoring ground-truth regional variations. This collapse occurs because pervasive high-volume signals, such as product attributes and dense image patches, overwhelm the few but critical tokens that encode market-specific context. To address CVE, we first collect a new multimodal dataset of real advertising creatives and their click-through performance across multiple countries. We then introduce GeoReward, a reward model designed to predict ad image preferences across diverse geographic markets. GeoReward integrates three purpose-built mechanisms: (1) Market-Aware Retrieval Augmentation, (2) Context-Guided Visual Modulation, (3) Selective Sensitivity Loss. Furthermore, we demonstrate how GeoReward can guide the fine-tuning of RL for a VLM to generate background designs for text-to-image models, producing market-aware advertising creatives. Experiments validate that our framework mitigates CVE and outperforms existing baselines. This work not only diagnoses a systematic bias in VLMs toward dominant perceptual features but also delivers a targeted solution for applications where sparse contextual variables govern decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。