arXiv:2507.12782cs.CL2025-07

用大模型数据增强小模型的否定理解能力

Learning Robust Negation Text Representations

  • 从大模型中蒸馏多种否定和模糊表达数据
  • 微调后在否定理解任务上显著提升
  • 适用于文本编码器和大模型,提升鲁棒性

尽管自回归大语言模型广泛应用,小型文本编码器仍在需要丰富上下文表示的文本理解任务中扮演重要角色。否定是一种关键语义功能,但现有方法仍未能有效捕捉,影响依赖文本嵌入的下游应用。本文提出一种策略,通过使用多样化的否定与模糊表达模式,从大语言模型中蒸馏数据,以增强文本编码器的否定鲁棒性。采用标准对比学习策略微调基于BERT的强大模型,在保持通用基准上竞争力的同时,显著提升否定理解能力。此外,该方法也可适配大语言模型,在否定相关基准上实现性能改进。

原文摘要 · Abstract (English)

Despite rapid adoption of autoregressive large language models, smaller text encoders still play an important role in text understanding tasks that require rich contextualized representations. Negation is an important semantic function that is still not properly captured by such methods, affecting many downstream applications relying on text embeddings. We propose a strategy to improve negation robustness of text encoders, by distilling data from large language models using diverse patterns of negation and hedging. We adopt a standard contrastive learning strategy to finetune a strong BERT-based model, and observe large improvement in negation understanding capabilities while maintaining competitive performance on general benchmarks. In addition, we also show that our method can be adapted to LLMs, leading to improved performance on negation benchmarks.

文本编码否定理解模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。