首个面向非洲工业机械的低资源数据集,含真实来源的设备运行与故障记录。
Nigeria Machinery: A Low-Resource Industrial Dataset with a Domain-Grounded Reasoning Layer
- 基于真实公开数据构建89条机器记录,覆盖28个指标
- 通过领域对齐推理层使94条提示全部匹配实际数据
- 适合做非洲工业场景下的语言模型训练与可解释性研究
非洲经济中的工业机械公开、可模型使用的数据极为稀缺,难以支持量化分析或数值任务的语言模型训练。本文发布两部分内容:一是尼日利亚机械使用与故障数据集,包含2006至2025年间制造业和油气行业共89条机器级记录,涵盖28个指标,每条均有公开来源并附代码本;二是从稀疏数值中构建链式思维(CoT)推理示例的方法,生成94组带提示、完成和推理轨迹的数据行。每行明确标注原始指标、子行业、年份及来源。该数据适配由Adaption Labs完成。我们指出,此前语言模型构建数据时常出现虽数字匹配但领域无关的问题。经修正后,领域对齐提示比例由78中的1提升至94中的94,所有检索答案均准确对应源数据值(84/84)。数据、推理层及逐行溯源文件已按CC-BY-4.0开源。需注意:仅有89条记录,17个指标仅一次观测,属参考与种子数据集,多数推理为单步检索而非多步计算。
原文摘要 · Abstract (English)
There is relatively little, public, and model-ready data on industrial machinery for African economies. This makes it hard to do quantitative analysis or to train language models on numeric tasks grounded in that setting. We release two things to help with part of this problem. The first is the Nigeria Machinery Usage and Failures Dataset: 89 machine-level records across 28 indicators, covering Nigeria's manufacturing and oil and gas sectors from 2006 to 2025. Every record names a public source and is decoded by a codebook. The second is a method for building chain-of-thought (CoT) reasoning examples from these sparse numeric values. The result is 94 prompt, completion, and reasoning-trace rows. In every row, the prompt names the real indicator, subsector, year, and source of the record it comes from. The data adaptation work was carried out by Adaption Labs. Along the way we describe a problem that is common when language models are used to build datasets. The prompts can match the real numbers while saying nothing about the real domain. We show that fixing this raises the share of domain-grounded prompts from 1 out of 78 in an earlier release to 94 out of 94, and that every retrieval answer now matches its source value (84 out of 84). We release the data, the reasoning layer, and a per-row provenance file under CC-BY-4.0. We are clear about the limits. With 89 records and 17 indicators that have only one observation, this is a reference and seed dataset, not a large training set. Most reasoning rows are retrieval rather than multi-step computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。