帮助非专家用目标建模法明确机器学习数据需求
Data Requirement Goal Modeling for Machine Learning Systems
- 构建数据需求目标模型(DRGM),整合文献中的常见挑战
- 通过定制机制调整任务、指标与目标权重,适配具体问题
- 适合缺乏数据经验的开发者和需求工程师使用
机器学习已广泛应用于各类系统中,其训练依赖于数据与算法。由于数据在机器学习开发中的关键作用,评估数据属性质量并确保其满足特定要求至关重要。本文提出一种面向非专家的数据需求引导方法,基于白皮书调研构建数据需求目标模型(DRGM),涵盖跨项目通用任务;再结合灰度文献洞察,设计定制化机制以调整任务、关键绩效指标(KPI)及目标重要性。生成的模型可支持用户使用GRL评估策略对比不同数据集。通过两个真实项目案例验证,该方法识别出的数据需求与实际项目一致,证明了框架的实用性与有效性。所提定制机制与DRGM有助于非专家针对特定机器学习问题确定合适数据需求,并评估数据集优劣。未来建议开发基于聊天机器人界面的工具实现自动建模。
原文摘要 · Abstract (English)
Machine Learning (ML) has been integrated into various software and systems. Two main components are essential for training an ML model: the training data and the ML algorithm. Given the critical role of data in ML system development, it has become increasingly important to assess the quality of data attributes and ensure that the data meets specific requirements before its utilization. This work proposes an approach to guide non-experts in identifying data requirements for ML systems using goal modeling. In this approach, we first develop the Data Requirement Goal Model (DRGM) by surveying the white literature to identify and categorize the issues and challenges faced by data scientists and requirement engineers working on ML-related projects. An initial DRGM was built to accommodate common tasks that would generalize across projects. Then, based on insights from both white and gray literature, a customization mechanism is built to help adjust the tasks, KPIs, and goals' importance of different elements within the DRGM. The generated model can aid its users in evaluating different datasets using GRL evaluation strategies. We then validate the approach through two illustrative examples based on real-world projects. The results from the illustrative examples demonstrate that the data requirements identified by the proposed approach align with the requirements of real-world projects, demonstrating the practicality and effectiveness of the proposed framework. The proposed dataset selection customization mechanism and the proposed DRGM are helpful in guiding non-experts in identifying the data requirements for machine learning systems tailored to a specific ML problem. This approach also aids in evaluating different dataset alternatives to choose the optimum dataset for the problem. For future work, we recommend implementing tool support to generate the DRGM based on a chatbot interface.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。