用可视化分析指导大模型生成数据,提升文本分类模型效果。
iGAiVA: Integrated Generative AI and Visual Analytics in a Machine Learning Workflow for Text Classification
- 结合视觉分析与大模型生成,精准定位数据缺陷并补救。
- 针对数据分布不均问题,合成数据后模型准确率显著提升。
- 集成工具iGAiVA支持全流程文本分类开发,适合数据工程师和算法研究员。
在构建文本分类的机器学习模型时,常面临数据分布不理想的问题,尤其在新增类别应对数据或任务变化时更为突出。本文提出一种利用视觉分析(VA)引导大语言模型生成合成数据的解决方案。由于VA能帮助开发者识别数据缺陷,因此数据合成可有针对性地弥补这些不足。我们探讨了不同类型的数据缺陷,描述了支持其识别的多种视觉分析技术,并验证了定向数据合成对提升模型准确率的有效性。此外,我们开发了软件工具iGAiVA,将四类机器学习任务映射至四个视觉分析视图,实现生成式AI与视觉分析在文本分类模型开发与优化流程中的深度融合。
原文摘要 · Abstract (English)
In developing machine learning (ML) models for text classification, one common challenge is that the collected data is often not ideally distributed, especially when new classes are introduced in response to changes of data and tasks. In this paper, we present a solution for using visual analytics (VA) to guide the generation of synthetic data using large language models. As VA enables model developers to identify data-related deficiency, data synthesis can be targeted to address such deficiency. We discuss different types of data deficiency, describe different VA techniques for supporting their identification, and demonstrate the effectiveness of targeted data synthesis in improving model accuracy. In addition, we present a software tool, iGAiVA, which maps four groups of ML tasks into four VA views, integrating generative AI and VA into an ML workflow for developing and improving text classification models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。