构建1000篇学术论文多模态数据集,提升关键词提取效果
Building a Multimodal Dataset of Academic Paper for Keyword Extraction
- 构建含文本、图像、音频的多模态学术论文数据集
- 融合多模态文本可显著提升关键词提取准确率
- 适合研究多模态信息融合与学术文本分析的学者
以往关键词提取主要依赖文本数据,忽视图像与音频模态中的视觉和听觉信息,导致信息丰富度不足且遗漏潜在关联,限制了模型表征学习能力与预测精度。当前针对关键词提取任务的多模态数据集尤为稀缺,阻碍了相关研究进展。为此,本研究构建了一个包含1000个样本的学术论文多模态数据集,每个样本包含论文正文、图像、音频及关键词。基于无监督与有监督的关键词提取方法,分别使用论文文本、图像中提取的文本以及音频转录文本进行实验,旨在探究不同模态信息及多模态融合对关键词提取性能的影响。实验结果表明,不同模态的文本在模型中表现出各异特征;将论文文本、图像文本与音频文本拼接后,能有效提升学术论文关键词提取性能。
原文摘要 · Abstract (English)
Up to this point, keyword extraction task typically relies solely on textual data. Neglecting visual details and audio features from image and audio modalities leads to deficiencies in information richness and overlooks potential correlations, thereby constraining the model's ability to learn representations of the data and the accuracy of model predictions. Furthermore, the currently available multimodal datasets for keyword extraction task are particularly scarce, further hindering the progress of research on multimodal keyword extraction task. Therefore, this study constructs a multimodal dataset of academic paper consisting of 1000 samples, with each sample containing paper text, images, audios and keywords. Based on unsupervised and supervised methods of keyword extraction, experiments are conducted using textual data from papers, as well as text extracted from images and audio. The aim is to investigate the differences in performance in keyword extraction task with respect to different modal information and the fusion of multimodal information. The experimental results indicate that text from different modalities exhibits distinct characteristics in the model. The concatenation of paper text, image text and audio text can effectively enhance the keyword extraction performance of academic papers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。