arXiv:2503.06085cs.CL2025-03AAAI被引 1

从贝叶斯视角改进预训练模型,提升非独立同分布文本理解能力

Multi-Attribute Multi-Grained Adaptation of Pre-Trained Language Models for Text Understanding from Bayesian Perspective

  • 提出多属性多粒度框架M2A,融合非独立同分布特征
  • 在隐式非独立同分布数据上表现更优,大模型效果更显著
  • 轻量级设计降低不确定性,适合大规模预训练模型优化

当前神经网络常通过多领域学习或属性注入机制,利用非独立同分布(non-IID)信息提升文本理解性能,但其影响程度及对预训练语言模型(PLMs)的作用仍不明确。本文从贝叶斯视角重新审视non-IID信息是否能提升PLMs性能,揭示并整合non-IID与IID特征。为此,提出多属性多粒度的PLM适配框架M2A,结合多属性与多粒度视角,在轻量化条件下缓解不确定性。在多个主流文本理解数据集上的实验表明,M2A在数据隐含non-IID且PLM规模较大时表现更优。

原文摘要 · Abstract (English)

Current neural networks often employ multi-domain-learning or attribute-injecting mechanisms to incorporate non-independent and identically distributed (non-IID) information for text understanding tasks by capturing individual characteristics and the relationships among samples. However, the extent of the impact of non-IID information and how these methods affect pre-trained language models (PLMs) remains unclear. This study revisits the assumption that non-IID information enhances PLMs to achieve performance improvements from a Bayesian perspective, which unearths and integrates non-IID and IID features. Furthermore, we proposed a multi-attribute multi-grained framework for PLM adaptations (M2A), which combines multi-attribute and multi-grained views to mitigate uncertainty in a lightweight manner. We evaluate M2A through prevalent text-understanding datasets and demonstrate its superior performance, mainly when data are implicitly non-IID, and PLMs scale larger.

预训练模型文本理解贝叶斯方法非独立同分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。