arXiv:2501.13950cs.CV2025-01

构建百万级烟产品数据集与模型,助力烟草广告智能监管

DEFEND: A Large-scale 1M Dataset and Foundation Model for Tobacco Addiction Prevention

  • 用多模态增强与视觉一致性机制提升烟品识别能力
  • 分类准确率达83.1%,零样本识别达45.6%
  • 适合公共健康监管与政策制定者使用

面对烟草广告在社交媒体上前所未有的创新速度,传统监测手段却停滞不前。本文提出Tobacco-1M数据集,包含一百万张烟产品图像,涵盖75个产品类别,并构建了DEFEND基础模型以实现烟品理解。该模型融合特征增强模块、局部-全局视觉一致性机制与图像-文本对齐策略,显著提升多模态表征能力。实验表明,其在产品分类任务中达到83.1%准确率,在视觉问答任务中达73.8%;在未见产品类别上仍保持45.6%的零样本识别准确率。该成果为监管机构和公共卫生研究者提供了高效工具,有望革新烟草管控与公共健康监测方式。

原文摘要 · Abstract (English)

While tobacco advertising innovates at unprecedented speed, traditional surveillance methods remain frozen in time, especially in the context of social media. The lack of large-scale, comprehensive datasets and sophisticated monitoring systems has created a widening gap between industry advancement and public health oversight. This paper addresses this critical challenge by introducing Tobacco-1M, a comprehensive dataset of one million tobacco product images with hierarchical labels spanning 75 product categories, and DEFEND, a novel foundation model for tobacco product understanding. Our approach integrates a Feature Enhancement Module for rich multimodal representation learning, a Local-Global Visual Coherence mechanism for detailed feature discrimination, and an Enhanced Image-Text Alignment strategy for precise product characterization. Experimental results demonstrate DEFEND's superior performance, achieving 83.1% accuracy in product classification and 73.8% in visual question-answering tasks, outperforming existing methods by significant margins. Moreover, the model exhibits robust zero-shot learning capabilities with 45.6% accuracy on novel product categories. This work provides regulatory bodies and public health researchers with powerful tools for monitoring emerging tobacco products and marketing strategies, potentially revolutionizing approaches to tobacco control and public health surveillance.

烟瘾防控图像识别基础模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。