arXiv:2501.15876cs.CLcs.AI2025-01被引 1

用伪标签和模型融合提升句子嵌入性能,效果显著。

Optimizing Sentence Embedding with Pseudo-Labeling and Model Ensembles: A Hierarchical Framework for Enhanced NLP Tasks

  • 构建分层框架,融合多模型与外部数据增强
  • 在多个数据集上准确率与F1值大幅提升
  • 适合需要高质量句子表示的NLP任务研究者

句子嵌入在自然语言处理中至关重要,但提升性能同时保持可靠性仍具挑战。本文提出一种结合伪标签生成与模型集成的框架,利用SimpleWiki、Wikipedia和BookCorpus等外部数据确保训练数据一致性。框架包含编码层、精炼层和集成预测层,采用ALBERT-xxlarge、RoBERTa-large和DeBERTa-large模型,通过交叉注意力融合外部上下文,并使用同义词替换与反向翻译等数据增强技术提升数据多样性。实验表明,相比基础模型,该方法在多个任务上显著提升准确率与F1分数,验证了交叉注意力与数据增强的有效性。本工作为句子嵌入优化提供了有效路径,也为未来NLP研究奠定基础。

原文摘要 · Abstract (English)

Sentence embedding tasks are important in natural language processing (NLP), but improving their performance while keeping them reliable is still hard. This paper presents a framework that combines pseudo-label generation and model ensemble techniques to improve sentence embeddings. We use external data from SimpleWiki, Wikipedia, and BookCorpus to make sure the training data is consistent. The framework includes a hierarchical model with an encoding layer, refinement layer, and ensemble prediction layer, using ALBERT-xxlarge, RoBERTa-large, and DeBERTa-large models. Cross-attention layers combine external context, and data augmentation techniques like synonym replacement and back-translation increase data variety. Experimental results show large improvements in accuracy and F1-score compared to basic models, and studies confirm that cross-attention and data augmentation make a difference. This work presents an effective way to improve sentence embedding tasks and lays the groundwork for future NLP research.

句子嵌入模型融合数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。