arXiv:2505.22084cs.CV2025-05被引 2

构建首个电影多模态物化视角数据集,助力识别影视中的性别刻板印象。

MObyGaze: a film dataset of multimodal objectification densely annotated by experts

  • 基于心理学与影视研究,定义物化5个子构念和11个跨模态概念
  • 标注20部电影43小时视频,生成6072个细粒度片段标注
  • 提供可复现的模型评估框架,适用于性别偏见检测研究

刻画和量化音视频叙事内容中的性别表征差异,是理解屏幕上的刻板印象如何延续的关键。本文关注高阶的物化概念,向机器学习领域提出新任务:识别并量化电影中由视觉、语音、音频等多模态构成的复杂时序物化模式。基于影视研究与心理学,我们构建了一个结构化术语库,包含5个子构念,覆盖11个跨3种模态的概念。我们推出多模态物化凝视(MObyGaze)数据集,涵盖20部电影,由专家对自由划定片段进行密集标注,共生成6072个标注段落,总时长43小时,实现细粒度定位与分类。我们提出新的视频理解任务,验证了多模态物化检测的可行性,并分析数据与模型偏差,提出改进方向。通过两个应用示例,展示丰富概念标注如何提升模型可靠性与可解释性。代码与数据集已公开,符合Croissant格式:https://github.com/husky-helen/MObyGaze。

原文摘要 · Abstract (English)

Characterizing and quantifying gender representation disparities in audiovisual storytelling contents is necessary to grasp how stereotypes may perpetuate on screen. In this article, we consider the high-level construct of objectification and introduce a new AI task to the ML community: characterize and quantify complex multimodal (visual, speech, audio) temporal patterns producing objectification in films. Building on film studies and psychology, we define the construct of objectification in a structured thesaurus involving 5 sub-constructs manifesting through 11 concepts spanning 3 modalities. We introduce the Multimodal Objectifying Gaze (MObyGaze) dataset, made of 20 movies annotated densely by experts for objectification levels and concepts over freely delimited segments: it amounts to 6072 segments over 43 hours of video with fine-grained localization and categorization. We formulate new video interpretation tasks, show the feasibility of multimodal objectification detection, and analyze data and model bias to propose improvements. We exemplify two applications of MObyGaze, showing how rich concept annotation can improve model reliability and explainability. We make our code and our dataset available to the community and described in the Croissant format: https://github.com/husky-helen/MObyGaze.

多模态分析性别偏见影视数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。