首个中文社交媒体多模态指代消解数据集,助力理解视频评论中的指代关系。
Multimodal Coreference Resolution for Chinese Social Media Dialogues: Dataset and Benchmark Approach
- 构建真实场景下的中英文视频评论多模态指代数据集
- 标注文本中人物提及与视频中人脸区域的对应关系
- 适用于社交平台内容理解、个性化推荐等场景研究
多模态指代消解(MCR)旨在识别跨文本与视觉等模态中指向同一实体的指称,对理解多模态内容至关重要。随着社交媒体上多模态内容激增,MCR在解析用户互动、打通图文关联方面具有重要意义。然而,真实对话场景下的多模态指代研究因缺乏数据资源而停滞。为此,我们提出了TikTalkCoref,首个基于抖音短视频平台的真实场景中文多模态指代数据集,包含短视频及其对应的用户评论文本,并人工标注了文本中的人物提及以及对应视频帧中的人脸区域的指代聚类。我们还提出了一种针对明星领域的有效基准方法,在该数据集上进行了全面实验,提供了可靠的基准结果。我们将公开TikTalkCoref数据集,推动该领域后续研究。
原文摘要 · Abstract (English)
Multimodal coreference resolution (MCR) aims to identify mentions referring to the same entity across different modalities, such as text and visuals, and is essential for understanding multimodal content. In the era of rapidly growing mutimodal content and social media, MCR is particularly crucial for interpreting user interactions and bridging text-visual references to improve communication and personalization. However, MCR research for real-world dialogues remains unexplored due to the lack of sufficient data resources. To address this gap, we introduce TikTalkCoref, the first Chinese multimodal coreference dataset for social media in real-world scenarios, derived from the popular Douyin short-video platform. This dataset pairs short videos with corresponding textual dialogues from user comments and includes manually annotated coreference clusters for both person mentions in the text and the coreferential person head regions in the corresponding video frames. We also present an effective benchmark approach for MCR, focusing on the celebrity domain, and conduct extensive experiments on our dataset, providing reliable benchmark results for this newly constructed dataset. We will release the TikTalkCoref dataset to facilitate future research on MCR for real-world social media dialogues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。