用大模型挖掘事件模式,解决新类别发现中的标签不均衡问题。
Generalized Category Discovery in Event-Centric Contexts: Latent Pattern Mining with LLMs
- 通过大模型提取事件模式,提升聚类与分类的一致性。
- 在不均衡数据上实现最高12.58%的H-score提升。
- 适合处理复杂叙事文本的新类别发现任务。
通用类别发现(GCD)旨在使用仅标注已知类别的部分标签数据,对已知和未知类别进行分类。尽管现有文本GCD方法在基准测试中表现优异,但在真实场景下的验证仍不足。本文提出面向事件中心场景的GCD(EC-GCD),其特征为长篇复杂叙述和高度不平衡的类别分布,带来两大挑战:(1)因主观标准导致聚类与分类分组不一致;(2)少数类别的对齐不公平。为此,我们提出PaMA框架,利用大语言模型(LLMs)提取并优化事件模式,以改善聚类与类别间的对齐。此外,设计了排序-过滤-挖掘流水线,确保各类别原型在不平衡情况下仍具平衡代表性。在两个EC-GCD基准数据集(包括新构建的诈骗报告数据集)上的评估表明,PaMA相较先前方法最高提升12.58%的H-score,同时在基础GCD数据集上保持良好泛化能力。
原文摘要 · Abstract (English)
Generalized Category Discovery (GCD) aims to classify both known and novel categories using partially labeled data that contains only known classes. Despite achieving strong performance on existing benchmarks, current textual GCD methods lack sufficient validation in realistic settings. We introduce Event-Centric GCD (EC-GCD), characterized by long, complex narratives and highly imbalanced class distributions, posing two main challenges: (1) divergent clustering versus classification groupings caused by subjective criteria, and (2) Unfair alignment for minority classes. To tackle these, we propose PaMA, a framework leveraging LLMs to extract and refine event patterns for improved cluster-class alignment. Additionally, a ranking-filtering-mining pipeline ensures balanced representation of prototypes across imbalanced categories. Evaluations on two EC-GCD benchmarks, including a newly constructed Scam Report dataset, demonstrate that PaMA outperforms prior methods with up to 12.58% H-score gains, while maintaining strong generalization on base GCD datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。