arXiv:2609.07334cs.AI2026-09

用检索增强的上下文学习,自动从产品手册生成工业资产数字档案。

AAS-RAIL: Improving Information Extraction for Asset Administration Shells through Retrieval-Augmented In-Context Learning

论文配图:AAS-RAIL: Improving Information Extraction for Asset Administration Shells through Retrieval-Augmented In-Context Learning
图 1 · 摘自论文原文
  • 从相似资产档案中检索提取助手,动态适配企业命名习惯。
  • 相比传统提示方法,信息抽取准确率提升30.4%至52.4%。
  • 无需微调模型,适合工业界快速构建定制化数字档案。

资产行政壳(AAS)是工业4.0和数字产品护照的核心,提供工业资产的标准数字表示。尽管制造商已保存大量技术文档,但从异构文档结构中提取信息生成AAS实例仍费时费力,因涉及企业特有术语和格式。本文提出AAS-RAIL,一种基于检索增强的上下文学习(RAIL)的信息抽取方法,利用大语言模型(LLMs)自动从PDF产品手册生成AAS。该方法不依赖固定少样本示例,而是检索相似AAS实例中的模型生成提取辅助,实现针对具体实例的上下文学习(ICL),使模型能适应企业特定命名与格式,且无需微调。核心贡献在于动态选择每份手册对应的公司特有AAS示例,以语义检索结合结构化抽取替代静态提示。在多个工业产品手册数据集上,采用开源与闭源权重的LLMs进行评估,结果表明RAIL在各类模型上均显著优于传统少样本提示,相对提升达30.4%-52.4%,证明其对定制化AAS生成的有效性。

原文摘要 · Abstract (English)

The Asset Administration Shell (AAS) is a cornerstone of Industry 4.0 and the Digital Product Passport, providing standardized digital representations of industrial assets. While manufacturers already maintain extensive technical product documentation, generating AAS instances from existing product datasheets remains a labor-intensive task because technical information is extracted from heterogeneous document structures and often involves company-specific terminology and conventions. In this work, we present AAS-RAIL, a retrieval-augmented information extraction (IE) approach that automatically generates Asset Administration Shells from PDF product datasheets using large language models (LLMs). Instead of relying on a fixed set of few-shot examples, the proposed retrieval-augmented in-context learning (RAIL) approach retrieves LLM-generated extraction helpers from similar Asset Administration Shells to provide instance-specific in-context learning (ICL). This enables the model to adapt its extraction behavior to company-specific naming conventions and formatting styles without fine-tuning. Our core contribution is the dynamic selection of company-specific AAS examples for each datasheet, replacing static prompting with an extraction pipeline that adapts to instances and combines semantic retrieval and structured information extraction. The proposed approach is evaluated on a collection of industrial product datasheets using a selection of open- and closed-weight LLMs. Experimental results show that RAIL consistently improves extraction quality over conventional few-shot prompting, yielding relative improvements of 30.4-52.4%. These results demonstrate that our approach provides an effective improvement for company-specific AAS generation.

信息抽取工业4.0大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。