arXiv:2605.20584cs.CV2026-05

用多模态模型自动识别应用内容评级描述,提升审核准确性。

QwenSafe: Multimodal Content Rating Description Identification via Preference-Aligned VLMs

论文配图:QwenSafe: Multimodal Content Rating Description Identification via Preference-Aligned VLMs
图 1 · 摘自论文原文
  • 融合文本与截图信息,通过偏好对齐的视觉语言模型判断敏感内容
  • 在12类苹果评级中召回率最高提升111.8%,显著优于现有模型
  • 适合平台内容审核、AI安全评估等场景的开发者和研究者

移动应用市场要求开发者披露标准化的内容评级描述(CRDs),以告知用户潜在敏感或受限内容。由于应用内容具有多模态特性,涵盖文本描述和视觉界面,确保披露的准确性和一致性仍具挑战。本文提出QwenSafe,一种视觉语言模型(VLM),通过联合推理应用元数据与截图,自动识别苹果定义的CRDs。为支持可扩展训练,我们构建了metadata2CRD数据构造管道,结合应用描述、截图与正式描述定义生成对齐的问答对。采用监督微调后进行直接偏好优化(DPO),使模型预测与跨视觉与文本模态的描述特异性证据和解释对齐。我们在12个苹果定义的内容评级描述上评估QwenSafe,对比Qwen3-VL、LLaVA-1.6和Gemini-2.5-Flash等先进视觉语言模型。QwenSafe在二分类任务中持续领先,正类召回率分别提升111.8%、36.1%和2.1%。结果表明,描述感知的多模态对齐显著提升自动化内容分类效果,凸显视觉语言模型在支撑可扩展、一致的内容评级方面的潜力。

原文摘要 · Abstract (English)

Mobile app marketplaces require developers to disclose standardized content rating descriptors (CRDs) to inform users about potentially sensitive or restricted content. Ensuring the accuracy and consistency of these disclosures remains challenging due to the multimodal nature of app content, which spans textual descriptions and visual interfaces. In this paper, we present QwenSafe, a Vision-Language Model (VLM) designed to automatically identify the presence of Apple-defined CRDs by jointly reasoning over app metadata and screenshots. To enable scalable training for this task, we introduce metadata2CRD, a data-construction pipeline that synthesizes descriptor-aligned question-answer pairs by combining app descriptions, screenshots, and formal descriptor definitions. We adapt Qwen3-VL-8B using supervised fine-tuning followed by Direct Preference Optimization (DPO) to align model predictions with descriptor-specific evidence and explanations across visual and textual modalities. We evaluate QwenSafe on 12 Apple-defined content rating descriptors and compare it against state-of-the-art vision-language models, including Qwen3-VL, LLaVA-1.6, and Gemini-2.5-Flash. QwenSafe consistently outperforms all baselines in binary CRD classification, achieving improvements in positive-class recall of 111.8%, 36.1%, and 2.1%, respectively. Our results demonstrate that descriptor-aware multimodal alignment substantially improves automated content classification and highlights the potential of vision-language models to support scalable and consistent content rating in mobile app marketplaces.

多模态内容审核视觉语言模型应用安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。