arXiv:2412.16215cs.CVcs.AI2024-12被引 5

用文本描述+跨模态嵌入实现零样本广告图像审核,无需标注数据。

Zero-Shot Image Moderation in Google Ads with LLM-Assisted Textual Descriptions and Cross-modal Co-embeddings

  • 用大模型生成政策相关的文本描述,构建审核基准。
  • 通过图文相似度匹配,零样本识别违规广告图像。
  • 适合需要快速适配新政策的平台内容审核场景。

我们提出一种可扩展且灵活的广告图像内容审核方法,应对谷歌广告中海量、多样化内容及不断变化的审核策略挑战。该方法利用人工标注的文本描述和跨模态文本-图像共嵌入,实现对违规广告图像的零样本分类,无需大量监督训练数据或人工标注。通过大语言模型(LLMs)和用户经验,系统生成并优化一套全面的文本描述,以表征政策要求。推理时,输入图像与文本描述间的共嵌入相似度作为违规检测的可靠信号,实现高效且可适应的广告内容审核。评估结果表明,该框架显著提升了违规内容的检出能力。

原文摘要 · Abstract (English)

We present a scalable and agile approach for ads image content moderation at Google, addressing the challenges of moderating massive volumes of ads with diverse content and evolving policies. The proposed method utilizes human-curated textual descriptions and cross-modal text-image co-embeddings to enable zero-shot classification of policy violating ads images, bypassing the need for extensive supervised training data and human labeling. By leveraging large language models (LLMs) and user expertise, the system generates and refines a comprehensive set of textual descriptions representing policy guidelines. During inference, co-embedding similarity between incoming images and the textual descriptions serves as a reliable signal for policy violation detection, enabling efficient and adaptable ads content moderation. Evaluation results demonstrate the efficacy of this framework in significantly boosting the detection of policy violating content.

图像审核零样本跨模态LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。