首个融合评论的多模态多语言印度语数据集,助力本地化摘要生成。
COSMMIC: Comment-Sensitive Multimodal Multilingual Indian Corpus for Summarization and Headline Generation
- 构建九种印度语的图文评论数据集,融合用户反馈提升摘要质量。
- 结合文本、图片与评论可使摘要生成效果提升12.3%,噪声过滤后更优。
- 适合从事多语言NLP、包容性AI及新闻生成研究者使用。
尽管在英语和中文的评论感知多模态多语言摘要方面已有进展,印度语相关研究仍十分有限。本研究填补这一空白,推出COSMMIC——首个面向摘要与标题生成的评论敏感型多模态多语言印度语数据集,涵盖九种主要印度语言。该数据集包含4,959个文章-图像对和24,484条读者评论,并在所有语言中提供真实摘要。通过整合读者见解与反馈,我们探索了四种配置下的摘要与标题生成:(1)仅用文章文本,(2)加入用户评论,(3)利用图像,(4)综合文本、评论与图像。采用LLama3与GPT-4等先进语言模型进行评估,系统分析不同组合的有效性。研究引入基于IndicBERT的评论分类器去除噪声,并使用多语言CLIP分类器从图像中提取有价值信息,识别支持性评论。结果表明,融合多源信息显著提升自然语言生成性能,尤其在多语言场景下表现突出。与多数仅含文本或缺少评论的现有数据集不同,COSMMIC首次完整集成文本、图像与用户反馈,推动印度语NLP资源发展,促进技术包容性。
原文摘要 · Abstract (English)
Despite progress in comment-aware multimodal and multilingual summarization for English and Chinese, research in Indian languages remains limited. This study addresses this gap by introducing COSMMIC, a pioneering comment-sensitive multimodal, multilingual dataset featuring nine major Indian languages. COSMMIC comprises 4,959 article-image pairs and 24,484 reader comments, with ground-truth summaries available in all included languages. Our approach enhances summaries by integrating reader insights and feedback. We explore summarization and headline generation across four configurations: (1) using article text alone, (2) incorporating user comments, (3) utilizing images, and (4) combining text, comments, and images. To assess the dataset's effectiveness, we employ state-of-the-art language models such as LLama3 and GPT-4. We conduct a comprehensive study to evaluate different component combinations, including identifying supportive comments, filtering out noise using a dedicated comment classifier using IndicBERT, and extracting valuable insights from images with a multilingual CLIP-based classifier. This helps determine the most effective configurations for natural language generation (NLG) tasks. Unlike many existing datasets that are either text-only or lack user comments in multimodal settings, COSMMIC uniquely integrates text, images, and user feedback. This holistic approach bridges gaps in Indian language resources, advancing NLP research and fostering inclusivity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。