用主动学习+贝叶斯模型,减少工程预测中的检测次数。
Active learning for regression in engineering populations: A risk-informed approach
- 结合主动学习与分层贝叶斯模型,优先获取关键数据。
- 减少30%以上检测次数,同时保持预测精度。
- 适合需要长期维护的工业设备预测场景。
回归是数据驱动型工程应用中的基础预测任务,涉及连续变量间的映射学习。在诸多工程场景(如结构健康监测)中,特征-标签数据有限,制约了传统监督学习的效果。本文提出一种结合主动学习与分层贝叶斯建模的方法,以应对数据稀缺问题。主动学习通过资源高效的方式选择性获取特征-标签对,本研究采用风险导向策略,利用与工程决策(如检测与维护)相关的上下文信息。分层贝叶斯模型使多个相关回归任务能在群体层面联合学习,捕捉局部与全局效应。该方法实现跨系统的知识共享,一个系统获取的信息可提升整体预测性能。在加工工具群体的实验案例中,目标为预测工件表面粗糙度。基于回归结果定义检测维护流程,并构建主动学习算法。与无信息引导的标签获取方式及独立建模相比,所提方法在预期成本上表现更优——在保持预测性能的同时,显著减少所需检测次数。
原文摘要 · Abstract (English)
Regression is a fundamental prediction task common in data-centric engineering applications that involves learning mappings between continuous variables. In many engineering applications (e.g.\ structural health monitoring), feature-label pairs used to learn such mappings are of limited availability which hinders the effectiveness of traditional supervised machine learning approaches. The current paper proposes a methodology for overcoming the issue of data scarcity by combining active learning with hierarchical Bayesian modelling. Active learning is an approach for preferentially acquiring feature-label pairs in a resource-efficient manner. In particular, the current work adopts a risk-informed approach that leverages contextual information associated with regression-based engineering decision-making tasks (e.g.\ inspection and maintenance). Hierarchical Bayesian modelling allow multiple related regression tasks to be learned over a population, capturing local and global effects. The information sharing facilitated by this modelling approach means that information acquired for one engineering system can improve predictive performance across the population. The proposed methodology is demonstrated using an experimental case study. Specifically, multiple regressions are performed over a population of machining tools, where the quantity of interest is the surface roughness of the workpieces. An inspection and maintenance decision process is defined using these regression tasks which is in turn used to construct the active-learning algorithm. The novel methodology proposed is benchmarked against an uninformed approach to label acquisition and independent modelling of the regression tasks. It is shown that the proposed approach has superior performance in terms of expected cost -- maintaining predictive performance while reducing the number of inspections required.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。