arXiv:2606.24346cs.IRcs.CL2026-06

用网页文本构建石油工程领域检索数据集,提升搜索效果。

PETRA: Transforming Web Text for Petroleum-Engineering Domain Adaptation

论文配图:PETRA: Transforming Web Text for Petroleum-Engineering Domain Adaptation
图 1 · 摘自论文原文
  • 从公开网络文本中筛选并生成高质量石油工程语料
  • 使检索指标nDCG从0.703提升至0.763,基准测试提升44%
  • 适合做领域适配与密集检索的科研人员参考

石油工程领域的检索面临强泛化模型的标注缺口:相关证据存在于公共网络文本中,但领域相关性标签稀缺。为此,我们提出PETRA——一个大规模石油工程文本检索适配数据集与流程,将嘈杂的公共网络数据转化为经过筛选的领域语料和合成监督信号,用于密集检索与重排序。PETRA包含136万条精选文本块,约20亿词元量级,约85.9万条嵌入训练样本(来自约22.4万个锚点),以及约40万条由教师模型评分的重排序候选样本。其构建融合高召回能源领域筛选、准确率达98.4%的领域分类器、基于文本块的查询生成、LLM撰写的难例负样本及检索挖掘的候选列表。通过分数融合,第一阶段在域内nDCG从0.703提升至0.763。重排序适配使公开地球科学基准提升44%相对性能,六项推理密集型任务提升23%。实验表明,合成标签训练精度高并不预示检索性能提升;仅当检索挖掘数据被重构为推断时候选分布采样的教师评分候选列表后,才能有效提升效果。

原文摘要 · Abstract (English)

Petroleum-engineering search exposes a supervision gap for strong general retrievers: relevant evidence exists in public web text, but domain relevance labels are scarce. To address this gap, we propose PETRA, a large-scale Petroleum Engineering Text for Retrieval Adaptation dataset and pipeline that converts noisy public web data into a curated domain corpus and synthetic supervision for dense retrieval and reranking. PETRA contains 1.36M curated chunks, approximately 2B token equivalents, $\approx$859k, embedding training rows from $\approx$224k anchors, and roughly 400k teacher-scored reranker candidate rows. Its construction combines high-recall energy-domain curation, an energy-domain classifier with 98.4% test accuracy, chunk-grounded query generation, LLM-written hard negatives, and retrieval-mined candidate lists. PETRA improves first-stage in-domain Normalized Discounted Cumulative Gain (nDCG) from 0.703 to 0.763 through score fusion. Reranker adaptation improves the public Earth Science benchmark by 44% relative and a six-task reasoning-intensive panel by 23%. Failed training recipes show that high train-holdout accuracy on synthetic labels does not predict retrieval gains; retrieval-mined data helps only after being repackaged as teacher-scored candidate lists sampled from the inference-time candidate distribution.

信息检索领域适配数据构建石油工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。