arXiv:2606.24897cs.DLcs.CL2026-06被引 2

四款生物医学文献接口对特殊字符保留率极低,影响大模型训练与科研分析。

Invisible to humans, visible to machines: a preregistered audit of Unicode fidelity across four biomedical bibliographic APIs

  • 对比四个文献接口与原始XML,检测特殊字符保留情况
  • 标点符号仅0.6%被保留,特殊空格完全丢失
  • 结果揭示接口差异显著,影响文本分析与大模型训练

生物医学文本挖掘、科学计量学及大语言模型训练语料构建均依赖于文献接口返回的摘要内容与原文一致。本预注册审计(OSF osf.io/269b5)以PubMed Central JATS XML为基准,测试了四个常用公共API(PubMed E-utilities、Crossref、OpenAlex、Semantic Scholar)在2024年开放获取数据集(约70万条)中随机抽取的4,000篇英文研究论文的字符保真度。重点检查四类预设字符:排版标点、数学/科学符号、希腊字母、特殊空白符。结果显示,只有数学符号和希腊字母保留率超过95%;而PubMed的AbstractText字段仅0.6%(95% CI 0.3–1.0%)保留排版标点,OpenAlex对特殊空白符保留率为0%(0.0–0.4%),两者均满足预设损失阈值。盲测归因于字符替换与索引序列化错误。此外,Crossref有24.6%的论文未返回摘要(覆盖率75.4%,95% CI 74.1–76.7%),且集中在爱思唯尔与ACS出版社(均为0%)。字符级保真度高度依赖接口,且未公开,同一来源文本经不同接口呈现不同表层特征,直接影响分词敏感型计量分析、语料构建及基于字符的LLM写作检测。

原文摘要 · Abstract (English)

Biomedical text mining, scientometrics, and the construction of training corpora for biomedical large language models (LLMs) all assume that the abstract text returned by a bibliographic API faithfully reproduces the published abstract. This pre-registered audit (OSF osf.io/269b5) tests that assumption for four widely used public APIs (PubMed E-utilities, Crossref, OpenAlex, Semantic Scholar) against PubMed Central (PMC) JATS XML as a common ground truth. From a complete enumeration of the PMC Open Access subset for 2024 (about 700,000 records), a simple random sample of 4,000 English-language research articles was drawn; for each, we recorded whether Unicode characters from four pre-specified classes present in the JATS abstract (typographic punctuation, mathematical/scientific symbols, Greek letters, special whitespace) were preserved by each API. Two systematic, deterministic losses met the pre-registered criterion (upper 95% CI bound below 5%): the PubMed AbstractText field preserved typographic punctuation in only 0.6% of eligible abstracts (95% CI 0.3-1.0%), and OpenAlex preserved special whitespace in 0% (0.0-0.4%). A blinded mechanism audit attributed the first loss to character substitution and the second to inverted-index serialization. Mathematical symbols and Greek letters were preserved faithfully (over 95%) by all four APIs. Separately, Crossref returned no abstract for 24.6% of papers (coverage 75.4%, 95% CI 74.1-76.7%), concentrated in specific publishers (Elsevier and ACS: 0%). Character-level fidelity is therefore API-dependent and undocumented: the same publisher-deposited JATS text carries different surface signatures depending on the serving API, with direct consequences for tokenization-sensitive bibliometrics, corpus construction, and character-level indicators of LLM-assisted writing.

文献接口字符保真LLM训练科学计量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。