arXiv:2601.14429cs.DLcs.AI2026-01

用大模型自动评估交通领域开放科学现状,发现代码数据共享率极低。

Measuring the State of Open Science in Transportation Using Large Language Models

  • 用大语言模型自动提取论文中代码与数据共享信息
  • 仅5%定量研究共享代码,4%共享数据,3%两者皆有
  • 适合关注科研透明度与政策制定的研究者参考

开放科学倡议已增强多个领域的科研诚信并加速研究进展,但其在交通研究中的实践状态仍缺乏深入考察。开放科学的关键特征——数据与代码可获取性——因领域复杂性难以提取。以往研究或受限于人工分析的低效率而规模小,或依赖大规模文献计量方法而损失上下文细节。本文提出一种自动、可扩展的特征提取流程,利用大语言模型(LLMs)识别交通研究中代码与数据的可用性,并通过人工标注数据集和评分者一致性分析验证其性能。该流程应用于2019至2024年间发表于《Transportation Research Part》系列期刊的10,724篇论文。结果显示,仅5%的定量论文共享代码仓库,4%共享数据仓库,约3%同时共享二者,且趋势在不同期刊、主题与地区间存在差异。分析未发现提供数据与代码的论文在引用量或审稿时长上具有显著差异,表明开放科学努力与传统学术评价指标之间存在脱节。因此,推动此类实践可能需期刊与资助机构实施结构性干预,以弥补作者激励不足的问题。本研究开发的管道可轻松扩展至其他期刊,是实现交通研究中开放科学实践自动化测量与监测的关键一步。

原文摘要 · Abstract (English)

Open science initiatives have strengthened scientific integrity and accelerated research progress across many fields, but the state of their practice within transportation research remains under-investigated. Key features of open science, defined here as data and code availability, are difficult to extract due to the inherent complexity of the field. Previous work has either been limited to small-scale studies due to the labor-intensive nature of manual analysis or has relied on large-scale bibliometric approaches that sacrifice contextual richness. This paper introduces an automatic and scalable feature-extraction pipeline to measure code and data availability in transportation research. We employ Large Language Models (LLMs) for this task and validate their performance against a manually curated dataset and through an inter-rater agreement analysis. We applied this pipeline to examine 10,724 research articles published in the Transportation Research Part series of journals between 2019 and 2024. Our analysis found that only 5% of quantitative papers shared a code repository, 4% of quantitative papers shared a data repository, and about 3% of papers shared both, with trends differing across journals, topics, and geographic regions. We found no significant difference in citation counts or review duration between papers that provided data and code and those that did not, suggesting a misalignment between open science efforts and traditional academic metrics. Consequently, encouraging these practices will likely require structural interventions from journals and funding agencies to supplement the lack of direct author incentives. The pipeline developed in this study can be readily scaled to other journals, representing a critical step toward the automated measurement and monitoring of open science practices in transportation research.

开放科学交通研究大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。