用大模型和梯度提升法识别图文与短视频中的性别歧视,效果受特征设计影响大。
Multimodal Sexism Identification and Characterization using Large Language Models and Gradient Boosting

- 针对图片和视频设计多模态特征,融合文本、视觉与大模型语义线索。
- 图片中针对性语义特征提升识别效果,视频则对特征维度敏感且噪声干扰大。
- 适合关注多模态偏见检测、内容安全或社会计算的研究者与开发者。
我们提交了AILS-NTUA系统参与CLEF 2026 EXIST 2026实验室任务,针对表情包(任务2)和短时视频(任务3)的多模态性别歧视识别与表征。系统采用基于梯度提升回归模型的特征工程化后融合流水线及分层后处理。对于表情包,结合视觉、文本、人口统计、生物特征及大模型生成的语义指标,捕捉刻板印象、物化、讽刺与厌女等高层线索;对于视频,则研究特征选择、帧级视觉表示、OCR文本特征、声学描述符及传感器元数据的影响。开发结果显示,聚焦的大模型语义线索可提升表情包识别性能,而视频表现高度依赖特征维度与跨模态噪声。尽管精简特征选择在开发集上表现更优,但官方测试结果表明该结论无法完全外推至未见数据,未经过滤的表示反而泛化能力更强。总体而言,研究强调了针对静态表情包进行精准语义特征工程的重要性,以及在噪声密集的短视视频场景下对更鲁棒时序建模的需求。
原文摘要 · Abstract (English)
We present the AILS-NTUA submission to the EXIST 2026 Lab at CLEF, addressing multimodal sexism identification and characterization in memes (Task 2) and short-form videos (Task 3). Our system follows a feature-engineered late-fusion pipeline built around gradient-boosted regression models and hierarchical post-processing. For memes, we combine visual, textual, demographic, biometric, and LLM-derived semantic indicators designed to capture high-level cues such as stereotyping, objectification, irony, and misogyny. For videos, we investigate the effect of feature selection, frame-based visual representations, OCR-based textual features, acoustic descriptors, and sensor-derived metadata. Development results show that focused LLM-derived semantic cues improve meme sexism identification, while video performance is highly sensitive to feature dimensionality and cross-modal noise. For videos, development results favor compact feature selection, but official test results show that this conclusion does not fully transfer to unseen data, where the unfiltered representation generalizes better. Overall, our findings highlight the usefulness of targeted semantic feature engineering for static memes and the need for more robust temporal modeling in noisy short-form video settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。