首个中文多模态避检内容评测基准,专为电商场景设计。
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection
- 构建真实电商场景下的多模态避检内容评测集
- 26个模型均出现误判,顶尖模型仍易受骗
- 分步推理+规则分类能显著提升检测准确率
电商平台日益依赖大语言模型(LLMs)和视觉语言模型(VLMs)识别非法或误导性商品内容。然而,这些模型仍易受避检内容攻击——即通过词拆分、委婉表达或图像裁剪等手段刻意隐藏违规信息,同时保留违规意图。有效检测需同时具备理解复杂规则与推断被伪装的多模态输入真实意图的能力。尽管已有研究分别关注规则推理与避检检测,但尚无统一评测框架整合二者。为此,我们提出EVADE-Bench,首个由专家标注的中文多模态避检内容评测基准,专为真实电商场景设计。对26个开源与闭源的LLM/VLM进行评估发现,即使先进模型也频繁误判。我们进一步证明,更清晰的规则分类可显著提升预测一致性并减少误报,凸显评测设计的关键作用。探索多智能体分解策略——将视觉描述与逻辑推理分给不同智能体处理——取得显著性能提升。
原文摘要 · Abstract (English)
E-commerce platforms increasingly rely on Large Language Models (LLMs) and Vision Language Models (VLMs) to detect illicit or misleading product content. However, these models remain vulnerable to evasive content, which refers to inputs that have been deliberately modified through techniques such as word splitting, euphemistic language, or image cropping to conceal policy violations while still conveying prohibited claims. Crucially, detecting such content requires a model to simultaneously master two capabilities: accurately comprehending complex rules, and correctly inferring the true intent behind deliberately obfuscated multimodal inputs. While prior work has separately explored LLM reasoning over complex rules and LLM-based detection of evasive content, no existing benchmark combines both within a unified evaluation framework. This gap is particularly consequential in e-commerce, where accurate moderation demands that both capabilities operate in concert. To address this gap, we introduce EVADE-Bench, the first expert-curated Chinese multimodal benchmark specifically designed to evaluate LLMs and VLMs on evasive content detection in real-world e-commerce scenarios. Our comprehensive evaluation of 26 open- and closed-source LLMs and VLMs reveals that even state-of-the-art models frequently misclassify evasive samples. We further demonstrate that clearer rule categorization significantly improves model prediction consistency and reduces false predictions, highlighting the critical role of benchmark design in enabling reliable evaluation. To explore paths for performance improvement, we investigate the feasibility of multi-agent decomposition for multimodal reasoning, wherein visual description and logical inference are decoupled into separate agents, and find that this strategy yields notable accuracy gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。