arXiv:2502.07263cs.SIcs.CL2025-02被引 3

分析160万篇论文代码,发现科研团队存在隐性分工。

Hidden Division of Labor in Scientific Teams Revealed Through 1.6 Million LaTeX Files

  • 通过分析LaTeX文件中的作者宏,追踪写作贡献
  • 揭示部分作者专攻概念性章节,部分专注技术细节
  • 为科研署名制度提供数据支持,适合政策制定者

科学奖励体系依赖个体贡献认定,但合著论文掩盖了具体分工。传统代理指标(如作者顺序、职级)加剧偏见,而贡献声明多为自报且覆盖有限。本研究基于1991至2023年间来自200万科学家的160万篇LaTeX文件,构建首个大规模写作贡献数据集。经自报声明(准确率0.87)、作者顺序模式、学科规范及Overleaf记录验证(斯皮尔曼相关系数0.6,p<0.05),数据可靠。结合显式章节信息,发现科研团队中存在隐性分工:部分作者主攻引言与讨论等概念性章节,另一些则集中于方法与实验等技术性内容。这是首次大规模揭示科学团队中的隐性劳动分配,挑战现有署名惯例,为机构信用分配政策提供依据。

原文摘要 · Abstract (English)

Recognition of individual contributions is fundamental to the scientific reward system, yet coauthored papers obscure who did what. Traditional proxies-author order and career stage-reinforce biases, while contribution statements remain self-reported and limited to select journals. We construct the first large-scale dataset on writing contributions by analyzing author-specific macros in LaTeX files from 1.6 million papers (1991-2023) by 2 million scientists. Validation against self-reported statements (precision = 0.87), author order patterns, field-specific norms, and Overleaf records (Spearman's rho = 0.6, p < 0.05) confirms the reliability of the created data. Using explicit section information, we reveal a hidden division of labor within scientific teams: some authors primarily contribute to conceptual sections (e.g., Introduction and Discussion), while others focus on technical sections (e.g., Methods and Experiments). These findings provide the first large-scale evidence of implicit labor division in scientific teams, challenging conventional authorship practices and informing institutional policies on credit allocation.

科研协作署名机制数据分析量化研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。