用AI自动清理科研摘要中的无关信息,提升文本分析准确性。
Cleaning English Abstracts of Scientific Publications
- 基于语言模型自动识别并移除摘要中的版权、注释等冗余内容
- 清理后摘要的相似度排序更准确,嵌入向量信息密度更高
- 开源易集成,适合需要清洗学术文本的研究者
科学摘要常被用作研究内容与主题焦点的代理指标。然而,大量已发表摘要包含额外信息,如出版商版权声明、章节标题、作者注释、注册信息以及文献计量或书目元数据,这些可能扭曲下游分析,尤其是涉及文档相似性或文本嵌入的任务。我们提出一个开源且易于集成的语言模型,可自动识别并清除英文科学摘要中的此类杂乱内容。实验表明,该模型兼具保守性与精确性,能有效改变清理后摘要的相似度排名,并提升标准长度嵌入向量的信息含量。
原文摘要 · Abstract (English)
Scientific abstracts are often used as proxies for the content and thematic focus of research publications. However, a significant share of published abstracts contains extraneous information-such as publisher copyright statements, section headings, author notes, registrations, and bibliometric or bibliographic metadata-that can distort downstream analyses, particularly those involving document similarity or textual embeddings. We introduce an open-source, easy-to-integrate language model designed to clean English-language scientific abstracts by automatically identifying and removing such clutter. We demonstrate that our model is both conservative and precise, alters similarity rankings of cleaned abstracts and improves information content of standard-length embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。