arXiv:2506.21623cs.CLcs.LG2025-06被引 2

用专家训练+合成数据提升投诉文本分类准确率

Performance of diverse evaluation metrics in NLP-based assessment and text generation of consumer complaints

  • 结合专家训练的分类器与生成对抗网络合成数据
  • 显著提升分类性能,降低数据采集成本
  • 适合需高精度文本评估的消费者权益研究者

机器学习已显著推动文本分类发展,实现对复杂非结构化文本的自动化理解与分类。然而,在消费者投诉这类自然语言中捕捉细微语义差异和上下文变化仍具挑战。本研究通过引入基于人类经验训练的算法,有效识别判断消费者救济资格所需的微妙语义差异。同时提出利用生成对抗网络进行合成数据生成,并经专家标注优化。结合专家训练的分类器与高质量合成数据,旨在显著提升机器学习分类器性能,降低数据获取成本,改善文本分类任务中的整体评估指标与鲁棒性。

原文摘要 · Abstract (English)

Machine learning (ML) has significantly advanced text classification by enabling automated understanding and categorization of complex, unstructured textual data. However, accurately capturing nuanced linguistic patterns and contextual variations inherent in natural language, particularly within consumer complaints, remains a challenge. This study addresses these issues by incorporating human-experience-trained algorithms that effectively recognize subtle semantic differences crucial for assessing consumer relief eligibility. Furthermore, we propose integrating synthetic data generation methods that utilize expert evaluations of generative adversarial networks and are refined through expert annotations. By combining expert-trained classifiers with high-quality synthetic data, our research seeks to significantly enhance machine learning classifier performance, reduce dataset acquisition costs, and improve overall evaluation metrics and robustness in text classification tasks.

文本分类合成数据消费者投诉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。