构建近30年医学信息学会议论文数据集,支持研究趋势分析
Decoding MIE: A Novel Dataset Approach Using Topic Extraction and Affiliation Parsing
- 基于文本排序算法解析机构信息,提取4606篇论文元数据
- 发现数字对象标识符使用模式与作者数据不一致现象
- 适合做医学信息学长期趋势、合作网络等研究的学者使用
医学信息学文献的快速扩张给研究趋势的整合与分析带来挑战。本研究基于医学信息学欧洲会议(MIE)论文集,利用Triple-A软件处理了1996年以来发表在《健康技术与信息研究》期刊系列中的4,606篇论文的元数据和摘要。通过文本排序算法实现机构信息解析,构建了包含引文趋势、作者归属及提取主题的结构化数据集,以JSON格式提供。分析显示,数字对象标识符(DOI)使用存在模式变化,作者数据存在不一致性,并曾出现短暂的语言多样性。该数据集为医学信息学领域的长期研究趋势追踪、合作网络分析和深入的文献计量研究提供了支持。
原文摘要 · Abstract (English)
The rapid expansion of medical informatics literature presents significant challenges in synthesizing and analyzing research trends. This study introduces a novel dataset derived from the Medical Informatics Europe (MIE) Conference proceedings, addressing the need for sophisticated analytical tools in the field. Utilizing the Triple-A software, we extracted and processed metadata and abstract from 4,606 articles published in the "Studies in Health Technology and Informatics" journal series, focusing on MIE conferences from 1996 onwards. Our methodology incorporated advanced techniques such as affiliation parsing using the TextRank algorithm. The resulting dataset, available in JSON format, offers a comprehensive view of bibliometric details, extracted topics, and standardized affiliation information. Analysis of this data revealed interesting patterns in Digital Object Identifier usage, citation trends, and authorship attribution across the years. Notably, we observed inconsistencies in author data and a brief period of linguistic diversity in publications. This dataset represents a significant contribution to the medical informatics community, enabling longitudinal studies of research trends, collaboration network analyses, and in-depth bibliometric investigations. By providing this enriched, structured resource spanning nearly three decades of conference proceedings, we aim to facilitate novel insights and advancements in the rapidly evolving field of medical informatics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。