找论文方法别只看摘要,中后段正文更关键。
Which Sections of a Research Paper Best Reveal Its Research Methods? Evidence from Library and Information Science
- 按文章位置切分全文,重点分析中后段
- 中后段和结尾部分方法信息最丰富
- 结合引用信息能显著提升识别效果
研究方法是学术论文知识贡献的核心载体。自动进行多标签方法分类可支持方法检索、综述生成与科研情报分析等知识服务。现有研究多依赖标题和摘要,但摘要中的方法信息有限,而使用全文又面临篇幅过长与信息冗余的问题。为此,本文提出一种基于物理位置的段落组合策略,利用来自图书馆与信息科学领域三本代表性期刊(JASIST、LISR、JDoc)的1,954篇全文标注语料,评估不同段落及其组合在多种模型上的分类性能。实验结果表明,方法信息在全文中分布不均,中后段及末尾段落具有更强的区分能力。此外,结合参考文献元数据与跨段落组合策略能有效提升分类效果。
原文摘要 · Abstract (English)
Research methods are essential carriers of knowledge contribution in academic papers. Automatic multi-label classification of research methods can support knowledge services such as method retrieval, review generation, and research intelligence analysis. While existing studies primarily rely on titles and abstracts, abstracts often provide only limited methodological information, whereas utilizing full-text content faces challenges related to excessive length and information redundancy. Therefore, this paper proposes a segment combination strategy by partitioning the full-text content according to its physical postion. Using an annotated corpus of 1,954 full-text articles from three representative journals in Library and Information Science (JASIST, LISR, and JDoc), we evaluate the classification performance of various segments and their combinations across multiple models. Experimental results indicate that methodological information is distributed unevenly within the full-text content, with the middle-to-late and final segments exhibiting greater discriminative power. Furthermore, integrating bibliographic metadata with cross-segment combination strategies effectively enhances classification performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。