arXiv:2501.17479cs.LGcs.AI2025-01Conference of the …被引 4

通过指纹聚类与自适应加权,提升大模型在复杂任务中的表现

DFPE: A Diverse Fingerprint Ensemble for Enhancing LLM Performance

  • 基于响应模式聚类不同大模型,挖掘互补性
  • 按学科过滤低效模型,整体准确率提升3%,学科级提升5%
  • 适合需要高鲁棒性的多领域语言理解场景

大型语言模型在多种自然语言处理任务中表现出色,但在多样化或复杂领域往往难以全面优异。我们提出一种新型集成方法——多样指纹集成(DFPE),利用多个大模型的互补优势实现更稳健的性能。该方法包括:(1) 根据响应“指纹”模式对模型进行聚类;(2) 在每个子任务层面采用分位数过滤机制剔除表现不佳的模型;(3) 根据各模型在对应子任务上的验证准确率动态分配权重。在 Massive Multitask Language Understanding (MMLU) 基准测试中,DFPE 相较于最优单个模型,整体准确率提升3%,学科级准确率提升5%。该方法增强了大模型的鲁棒性和泛化能力,凸显了模型选择、多样性保持与性能驱动加权在复杂多维语言理解任务中的有效性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown remarkable capabilities across various natural language processing tasks but often struggle to excel uniformly in diverse or complex domains. We propose a novel ensemble method - Diverse Fingerprint Ensemble (DFPE), which leverages the complementary strengths of multiple LLMs to achieve more robust performance. Our approach involves: (1) clustering models based on response "fingerprints" patterns, (2) applying a quantile-based filtering mechanism to remove underperforming models at a per-subject level, and (3) assigning adaptive weights to remaining models based on their subject-wise validation accuracy. In experiments on the Massive Multitask Language Understanding (MMLU) benchmark, DFPE outperforms the best single model by 3% overall accuracy and 5% in discipline-level accuracy. This method increases the robustness and generalization of LLMs and underscores how model selection, diversity preservation, and performance-driven weighting can effectively address challenging, multi-faceted language understanding tasks.

大模型集成性能优化多任务理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。