arXiv:2602.11168cs.CLcs.AI2026-02中稿 · 2025 IEEE Internat…

用多模型融合与生成数据提升联合国可持续发展目标文本分类效果

Enhancing SDG-Text Classification with Combinatorial Fusion Analysis and Generative AI

  • 通过组合多个模型并利用认知多样性提升分类性能
  • 达到96.73%准确率,优于单一模型
  • 适合政策分析、知识发现等需跨领域理解的场景

自然语言处理中的文本分类与主题发现广泛应用于信息检索、知识发现、政策制定和决策支持。然而,在类别不明确、难以区分或相互关联的情况下仍具挑战性。社会分析依赖大量文本数据,可从中受益。本文通过整合多个模型的智能,增强基于联合国可持续发展目标(SDGs)的文本分类。采用组合融合分析(CFA)方法,结合排名得分特征(RSC)函数与认知多样性(CD),提升分类器表现。利用生成式AI合成训练数据,并应用CFA进行分类。实验显示CFA达到96.73%性能,优于最优单个模型。与人类领域专家结果对比表明,模型智能与专家意见可互补并相互增强。

原文摘要 · Abstract (English)

(Natural Language Processing) NLP techniques such as text classification and topic discovery are very useful in many application areas including information retrieval, knowledge discovery, policy formulation, and decision-making. However, it remains a challenging problem in cases where the categories are unavailable, difficult to differentiate, or are interrelated. Social analysis with human context is an area that can benefit from text classification, as it relies substantially on text data. The focus of this paper is to enhance the classification of text according to the UN's Sustainable Development Goals (SDGs) by collecting and combining intelligence from multiple models. Combinatorial Fusion Analysis (CFA), a system fusion paradigm using a rank-score characteristic (RSC) function and cognitive diversity (CD), has been used to enhance classifier methods by combining a set of relatively good and mutually diverse classification models. We use a generative AI model to generate synthetic data for model training and then apply CFA to this classification task. The CFA technique achieves 96.73% performance, outperforming the best individual model. We compare the outcomes with those obtained from human domain experts. It is demonstrated that combining intelligence from multiple ML/AI models using CFA and getting input from human experts can, not only complement, but also enhance each other.

文本分类生成模型可持续发展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。