用大模型自动解析论文引用信息,助力全球科研公平
Citation Parsing and Analysis with Language Models
- 用开源大模型自动标注引文的各个组成部分
- 最小模型Qwen3-0.6B在32次尝试内完成高精度解析
- 适合关注科研公平、引文网络研究的学者
为解决知识生产与传播中的全球不平等问题,亟需工具帮助期刊理解知识流动。当前缺乏此类工具,导致对全球南方地区知识共享网络了解不足,进而使南方学者被排除在索引服务之外,固化殖民式学术格局。为支持全球范围的引文网络追踪,我们研究了开源语言模型在将手稿引文标注为可索引格式方面的表现。我们构建了一个包含预印本和已发表论文的明文与标注引文配对数据集,并评估多个开源语言模型的标注能力。结果显示,即使未经微调,当前语言模型在识别引文各组成部分方面已达到很高准确率,优于现有最先进方法。此外,最小模型Qwen3-0.6B在2^5=32次尝试内即可高精度解析所有字段,表明后训练有望生成小型、鲁棒的引文解析模型。此类工具可显著提升引文网络的准确性,从而改善研究索引与发现,并推动元科学进一步发展。
原文摘要 · Abstract (English)
A key type of resource needed to address global inequalities in knowledge production and dissemination is a tool that can support journals in understanding how knowledge circulates. The absence of such a tool has resulted in comparatively less information about networks of knowledge sharing in the Global South. In turn, this gap authorizes the exclusion of researchers and scholars from the South in indexing services, reinforcing colonial arrangements that de-center and minoritize those scholars. In order to support citation network tracking on a global scale, we investigate the capacity of open-weight language models to mark up manuscript citations in an indexable format. We assembled a dataset of matched plaintext and annotated citations from preprints and published research papers. Then, we evaluated a number of open-weight language models on the annotation task. We find that, even out of the box, today's language models achieve high levels of accuracy on identifying the constituent components of each citation, outperforming state-of-the-art methods. Moreover, the smallest model we evaluated, Qwen3-0.6B, can parse all fields with high accuracy in $2^5$ passes, suggesting that post-training is likely to be effective in producing small, robust citation parsing models. Such a tool could greatly improve the fidelity of citation networks and thus meaningfully improve research indexing and discovery, as well as further metascientific research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。