针对少数类表现差的不平衡数据,提出动态优化的集成方法
CAMO: A Class-Aware Minority-Optimized Ensemble for Robust Language Model Evaluation on Imbalanced Data

- 基于投票分布与置信度校准,动态增强少数类预测
- 在两个极端不平衡数据集上,宏观F1得分领先所有基线
- 适用于大模型与小模型,适配不同模型特性
现实世界分类任务常受类别不平衡影响,传统集成方法偏向多数类,导致少数类性能下降和整体F1分数降低。本文提出一种名为CAMO(Class-Aware Minority-Optimized)的独特集成方法,通过分层流程结合投票分布、置信度校准与模型间不确定性分析,动态提升被低估类别的表现,同时保留并强化其预测结果。我们在两个高度不平衡、领域特定的基准数据集——DIAR-AI/Emotion与ternary BEA 2025上验证了CAMO的有效性。对比七种成熟集成算法,在八种语言模型(三款LLM与五款SLM)的零样本与微调设置下进行测试。使用优化后的模型时,CAMO始终获得最高的严格宏观F1分数,确立新基准。其优势与模型适配性协同作用,表明最佳集成选择依赖于模型特性。这证明CAMO是一种可靠且领域无关的不平衡分类框架。
原文摘要 · Abstract (English)
Real-world categorization is severely hampered by class imbalance because traditional ensembles favor majority classes, which lowers minority performance and overall F1-score. We provide a unique ensemble technique for imbalanced problems called CAMO (Class-Aware Minority-Optimized).Through a hierarchical procedure that incorporates vote distributions, confidence calibration, and inter model uncertainty, CAMO dynamically boosts underrepresented classes while preserving and amplifying minority forecasts. We verify CAMO on two highly unbalanced, domain-specific benchmarks: the DIAR-AI/Emotion dataset and the ternary BEA 2025 dataset. We benchmark against seven proven ensemble algorithms using eight different language models (three LLMs and five SLMs) under zero-shot and fine-tuned settings .With refined models, CAMO consistently earns the greatest strict macro F1-score, setting a new benchmark. Its benefit works in concert with model adaptation, showing that the best ensemble choice depends on model properties .This proves that CAMO is a reliable, domain-neutral framework for unbalanced categorization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。