arXiv:2605.04257cs.LG2026-05

构建4383组冷喷涂实验数据集,解决文献数据难用问题。

HUGO-CS: A Hybrid-Labeled, Uncertainty-Aware, General-Purpose, Observational Dataset for Cold Spray

论文配图:HUGO-CS: A Hybrid-Labeled, Uncertainty-Aware, General-Purpose, Observational Dataset for Cold Spray
图 1 · 摘自论文原文
  • 融合大模型与人工校验,高效提取文献中的实验数据。
  • 数据集含144个特征,规模超前一数据集30倍。
  • 适合材料制造、数据驱动建模研究者使用。

冷喷涂因固态成形能力日益广泛用于部件修复与制造,但工艺优化因参数高度耦合且缺乏大规模可机读数据而困难。现有文献虽有大量实验,但结果报告不统一(常以表格和图表呈现),单位不一致,限制了规模化利用。为此,本文提出HUGO-CS,一个从1,124篇文献中提取的4,383组冷喷涂实验数据集,包含144个特征,规模超过此前最大数据集(137样本)30倍。通过完全人工提取,每篇文档平均耗时91分钟,设计并应用混合标注、不确定性感知、通用性观测提取框架HUGO,结合自动化大模型标注与针对性人工修正,实现高效准确的数据抽取。HUGO引入分层风险缓解机制(HRM),将高风险预测结果转交人工复核,低风险记录自动标记。后期处理对分类描述归一化、将粉末化学成分映射为结构化连续组成,并统一来源单位。4,383组实验中,1,765组经人工标注,构成高质量基准子集,可用于基准测试、误差分析与更高保真度数据点。所有代码及完整数据集已按CC-BY协议开源于https://github.com/sprice134/HUGO。

原文摘要 · Abstract (English)

Cold spraying is an increasingly common approach for repairing and manufacturing components due to its solid-state manufacturing capabilities. However, process optimization remains difficult due to many interdependent parameters and the lack of large-scale, machine-readable data to support modeling. While the scientific literature contains many relevant experiments, results are inconsistently reported (often in tables and figures) and use non-uniform units, limiting utilization at scale. To address these limitations, this work presents HUGO-CS, a literature-derived dataset of 4,383 cold-spray experiments with 144 features from 1,124 sources, exceeding the previous largest dataset (137 samples) by 30x. With completely manual extraction requiring an average of 91 minutes per document, this work designs and leverages a Hybrid-labeled, Uncertainty-aware, General-purpose, Observational extraction framework, called HUGO, to support this extraction. HUGO combines automated LLM-based labeling with targeted manual label refinement to handle this experimental result extraction process from scientific literature. To balance labeling efficiency with extraction accuracy, HUGO introduces a Hierarchical Risk Mitigation (HRM) to route LLM outputs with a high risk of potential errors for manual review, while retaining low-risk records as auto-labeled. Lastly, HUGO post-processing consolidates categorical descriptors, maps reported feedstock chemistries into structured continuous compositions, and normalizes units across sources. Of the 4,383 reported experiments, 1,765 are hand-labeled, providing a high-quality labeled subset for benchmarking, error analysis, and higher-fidelity data points. All code to replicate this work, along with the complete HUGO-CS dataset, are released under a CC-BY license at https://github.com/sprice134/HUGO.

数据集冷喷涂文献挖掘材料制造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。