ALF模型融合多模态数据,精准理解广告主行为与意图。
ALF: Advertiser Large Foundation Model for Multi-Modal Advertiser Understanding
- 通过对比学习和多任务优化统一建模文本、图像、视频等数据。
- 在欺诈检测等任务中召回率提升超40个百分点,精度达99.8%。
- 适合需要高精度广告审核与相似广告主识别的场景。
我们提出ALF(Advertiser Large Foundation model),一种用于跨文本、图像、视频和结构化数据模态理解广告主行为与意图的多模态Transformer架构。通过对比学习与多任务优化,ALF构建了统一的广告主表征,捕捉内容与行为模式。模型在欺诈检测、违规识别及广告主相似匹配等关键任务上达到领先性能。实际部署中,ALF显著提升真实业务指标:例如,在一项关键政策上召回率提升超过40个百分点,另一项任务精度达99.8%。其有效性源于新颖的多模态变换、样本间注意力机制、谱归一化投影及校准概率输出。
原文摘要 · Abstract (English)
We present ALF (Advertiser Large Foundation model), a multi-modal transformer architecture for understanding advertiser behavior and intent across text, image, video, and structured data modalities. Through contrastive learning and multi-task optimization, ALF creates unified advertiser representations that capture both content and behavioral patterns. Our model achieves state-of-the-art performance on critical tasks including fraud detection, policy violation identification, and advertiser similarity matching. In production deployment, ALF demonstrates significant real-world impact by delivering simultaneous gains in both precision and recall, for instance boosting recall by over 40 percentage points on one critical policy and increasing precision to 99.8% on another. The architecture's effectiveness stems from its novel combination of multi-modal transformations, inter-sample attention mechanism, spectrally normalized projections, and calibrated probabilistic outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。