构建多模态广告框架数据集,助力识别油气行业绿色洗牌行为
A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection
- 基于社交媒体视频构建跨国家多模态标注数据集
- GPT-4.1检测环境信息达79%准确率,最佳模型仅46%
- 适合研究能源公关传播与视觉语言模型评估的学者
企业投入巨资开展公共关系活动以塑造积极品牌形象,但言行常不一致。例如,油气公司通过宣传气候友好举措被指存在“绿色洗牌”行为。为系统理解此类宣传的框架及其演变,我们构建了一个专家标注的多模态视频广告基准数据集,涵盖来自Facebook和YouTube的50多家公司及倡导组织在20个国家的广告内容。数据集提供13类框架标注,专为视觉-语言模型(VLMs)评估设计,区别于以往纯文本框架数据集。基线实验显示:GPT-4.1在识别环境信息上达79% F1分数,而最优模型在绿色创新框架识别上仅46% F1。研究还揭示了当前模型面临的挑战,如隐含框架识别、长视频处理及文化背景理解。该数据集推动能源领域战略传播的多模态分析研究。
原文摘要 · Abstract (English)
Companies spend large amounts of money on public relations campaigns to project a positive brand image. However, sometimes there is a mismatch between what they say and what they do. Oil & gas companies, for example, are accused of "greenwashing" with imagery of climate-friendly initiatives. Understanding the framing, and changes in framing, at scale can help better understand the goals and nature of public relations campaigns. To address this, we introduce a benchmark dataset of expert-annotated video ads obtained from Facebook and YouTube. The dataset provides annotations for 13 framing types for more than 50 companies or advocacy groups across 20 countries. Our dataset is especially designed for the evaluation of vision-language models (VLMs), distinguishing it from past text-only framing datasets. Baseline experiments show some promising results, while leaving room for improvement for future work: GPT-4.1 can detect environmental messages with 79% F1 score, while our best model only achieves 46% F1 score on identifying framing around green innovation. We also identify challenges that VLMs must address, such as implicit framing, handling videos of various lengths, or implicit cultural backgrounds. Our dataset contributes to research in multimodal analysis of strategic communication in the energy sector.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。