研究多模态对话结构,发现女性角色更常被当作听者而非发起者。
Multimodal Conversation Structure Understanding
- 构建新数据集TV-MMPC,标注对话角色与话题线程
- 匿名化后模型性能显著下降,证明角色识别难
- 女性角色虽发言比例相当,但更常被设定为听众或旁观者
尽管多模态大语言模型在对话任务中表现优异,但其对对话结构(如角色分工与话题延续)的理解仍不充分。本文提出一套新任务,并发布新标注数据集TV-MMPC以支持多模态对话结构理解研究。评估显示,所有多模态LLM均优于启发式基线,但即使最佳模型在人物身份匿名化后性能仍大幅下降。此外,对350,842条电视问答(TVQA)语句的社会语言学分析发现:女性角色的发言频率与其说话时间成正比,但她们作为对话接收方或旁观者的概率是男性的1.2倍;当存在旁观者时,对话风格从个人化转向社交化。
原文摘要 · Abstract (English)
While multimodal large language models (LLMs) excel at dialogue, whether they can adequately parse the structure of conversation -- conversational roles and threading -- remains underexplored. In this work, we introduce a suite of tasks and release TV-MMPC, a new annotated dataset, for multimodal conversation structure understanding. Our evaluation reveals that while all multimodal LLMs outperform our heuristic baseline, even the best-performing model we consider experiences a substantial drop in performance when character identities of the conversation are anonymized. Beyond evaluation, we carry out a sociolinguistic analysis of 350,842 utterances in TVQA. We find that while female characters initiate conversations at rates in proportion to their speaking time, they are 1.2 times more likely than men to be cast as an addressee or side-participant, and the presence of side-participants shifts the conversational register from personal to social.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。