arXiv:2410.19896cs.CV2024-10被引 3

用流注意力融合视觉与文本,精准分析短视频中的烟草内容。

FLAASH: Flow-Attention Adaptive Semantic Hierarchical Fusion for Multi-Modal Tobacco Content Analysis

  • 基于流网络思想设计分层融合机制,动态捕捉跨模态交互
  • 在MTCAD数据集上分类准确率与F1值显著优于现有方法
  • 适合关注数字媒体烟草监管的公共卫生与AI研究者

社交媒体上烟草相关内容的泛滥给公共健康监测带来严峻挑战。本文提出一种名为FLAASH的新型多模态深度学习框架,用于全面分析烟草相关视频内容。该框架通过受流网络理论启发的分层融合机制,解决短视频中视觉与文本信息整合的复杂性。核心创新包括:流注意力机制以捕捉跨模态细微互动、自适应加权方案平衡不同层级贡献、门控机制选择性强化关键特征。该方法能有效处理从产品展示到使用场景等多样内容。我们在大型社交媒体烟草视频数据集MTCAD上评估了FLAASH,结果表明其在分类准确率、F1分数和时间一致性上均显著优于现有方法。此外,在标准视频问答数据集上也展现出强泛化能力,超越当前主流模型。本工作推动了公共健康与人工智能的交叉应用,为数字媒体中烟草推广的智能分析提供有效工具。

原文摘要 · Abstract (English)

The proliferation of tobacco-related content on social media platforms poses significant challenges for public health monitoring and intervention. This paper introduces a novel multi-modal deep learning framework named Flow-Attention Adaptive Semantic Hierarchical Fusion (FLAASH) designed to analyze tobacco-related video content comprehensively. FLAASH addresses the complexities of integrating visual and textual information in short-form videos by leveraging a hierarchical fusion mechanism inspired by flow network theory. Our approach incorporates three key innovations, including a flow-attention mechanism that captures nuanced interactions between visual and textual modalities, an adaptive weighting scheme that balances the contribution of different hierarchical levels, and a gating mechanism that selectively emphasizes relevant features. This multi-faceted approach enables FLAASH to effectively process and analyze diverse tobacco-related content, from product showcases to usage scenarios. We evaluate FLAASH on the Multimodal Tobacco Content Analysis Dataset (MTCAD), a large-scale collection of tobacco-related videos from popular social media platforms. Our results demonstrate significant improvements over existing methods, outperforming state-of-the-art approaches in classification accuracy, F1 score, and temporal consistency. The proposed method also shows strong generalization capabilities when tested on standard video question-answering datasets, surpassing current models. This work contributes to the intersection of public health and artificial intelligence, offering an effective tool for analyzing tobacco promotion in digital media.

多模态视频分析公共健康注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。