arXiv:2412.13794cs.CLcs.AI2024-12被引 2

用图文联合分析识别皮条客,助力打击人口贩卖

MATCHED: Multimodal Authorship-Attribution To Combat Human Trafficking in Escort-Advertisement Data

  • 构建图文并茂的皮条客数据集,支持跨模态身份识别
  • 多模态模型比单模态提升识别准确率,尤其在分布外数据上表现更稳
  • 适合执法机构用于追踪贩运网络,也适用于反人口贩卖研究

人口贩卖仍是严峻问题,犯罪分子常利用在线皮条广告匿名宣传受害者。现有检测方法多依赖文本分析,忽视了广告中常配有的图文信息。为此,我们构建了MATCHED数据集,包含来自美国四个地理区域七座城市的Backpage平台采集的27,619条唯一文本描述和55,115张唯一图像。本研究全面评估了纯文本、纯视觉及多模态基线模型在供应商识别与验证任务中的表现,采用多任务联合训练目标,在分布内与分布外(OOD)数据上均实现更优分类与检索性能。融合多模态特征进一步提升效果,捕捉文本与图像间的互补模式。尽管文本仍为主导模态,视觉信息提供了风格线索,增强模型表现。然而,如CLIP、BLIP2等对齐策略因图文语义重合度低而失效,端到端多模态训练更具鲁棒性。结果表明,多模态作者归属(MAA)具有打击人口贩卖的潜力,为执法机构提供可靠工具以关联广告、瓦解贩运网络。

原文摘要 · Abstract (English)

Human trafficking (HT) remains a critical issue, with traffickers increasingly leveraging online escort advertisements (ads) to advertise victims anonymously. Existing detection methods, including Authorship Attribution (AA), often center on text-based analyses and neglect the multimodal nature of online escort ads, which typically pair text with images. To address this gap, we introduce MATCHED, a multimodal dataset of 27,619 unique text descriptions and 55,115 unique images collected from the Backpage escort platform across seven U.S. cities in four geographical regions. Our study extensively benchmarks text-only, vision-only, and multimodal baselines for vendor identification and verification tasks, employing multitask (joint) training objectives that achieve superior classification and retrieval performance on in-distribution and out-of-distribution (OOD) datasets. Integrating multimodal features further enhances this performance, capturing complementary patterns across text and images. While text remains the dominant modality, visual data adds stylistic cues that enrich model performance. Moreover, text-image alignment strategies like CLIP and BLIP2 struggle due to low semantic overlap and vague connections between the modalities of escort ads, with end-to-end multimodal training proving more robust. Our findings emphasize the potential of multimodal AA (MAA) to combat HT, providing LEAs with robust tools to link ads and disrupt trafficking networks.

多模态分析身份识别反人口贩卖图文对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。