arXiv:2509.15241cs.CVcs.CL2025-09被引 1

用母子模型框架一键检测广告多模态合规性,效率提升31倍。

M-PACE: Mother Child Framework for Multimodal Compliance

  • 母子MLLM架构:强模型评估弱模型输出,统一处理图文信息。
  • 实测每图成本降至0.0005美元,比原方案低31倍,精度相当。
  • 专为广告合规设计,适合需自动化审核的平台和品牌方。

确保多模态内容符合品牌、法律或平台合规标准日益复杂。传统框架依赖分散的多阶段流水线,分别处理图像分类、文本提取、语音转录、手工检查与规则合并,导致运维负担重、难扩展且难以动态适应新规。随着多模态大语言模型(MLLM)的发展,统一处理视觉与文本输入的通用框架成为可能。为此,我们提出多模态无参数合规引擎(M-PACE),支持单次遍历评估超过15项合规属性。为支持结构化评测,我们构建了一个人工标注基准,包含模拟真实挑战场景的增强样本,如视觉遮挡和辱骂词注入。M-PACE采用母子MLLM结构,证明强模型评估弱模型输出可显著减少人工审核依赖,实现质量控制自动化。分析显示,推理成本降低超31倍;最优配置(母模型为Gemini 2.0 Flash,子模型由其选择)每图成本仅0.0005美元,相较同精度的Gemini 2.5 Pro(0.0159美元)大幅下降,验证了该框架在广告数据实际部署中的实时性与成本效益。

原文摘要 · Abstract (English)

Ensuring that multi-modal content adheres to brand, legal, or platform-specific compliance standards is an increasingly complex challenge across domains. Traditional compliance frameworks typically rely on disjointed, multi-stage pipelines that integrate separate modules for image classification, text extraction, audio transcription, hand-crafted checks, and rule-based merges. This architectural fragmentation increases operational overhead, hampers scalability, and hinders the ability to adapt to dynamic guidelines efficiently. With the emergence of Multimodal Large Language Models (MLLMs), there is growing potential to unify these workflows under a single, general-purpose framework capable of jointly processing visual and textual content. In light of this, we propose Multimodal Parameter Agnostic Compliance Engine (M-PACE), a framework designed for assessing attributes across vision-language inputs in a single pass. As a representative use case, we apply M-PACE to advertisement compliance, demonstrating its ability to evaluate over 15 compliance-related attributes. To support structured evaluation, we introduce a human-annotated benchmark enriched with augmented samples that simulate challenging real-world conditions, including visual obstructions and profanity injection. M-PACE employs a mother-child MLLM setup, demonstrating that a stronger parent MLLM evaluating the outputs of smaller child models can significantly reduce dependence on human reviewers, thereby automating quality control. Our analysis reveals that inference costs reduce by over 31 times, with the most efficient models (Gemini 2.0 Flash as child MLLM selected by mother MLLM) operating at 0.0005 per image, compared to 0.0159 for Gemini 2.5 Pro with comparable accuracy, highlighting the trade-off between cost and output quality achieved in real time by M-PACE in real life deployment over advertising data.

多模态合规检测MLLM自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。