arXiv:2506.00418cs.CLcs.AI2025-06ACL被引 1

提出双去偏框架,提升文本生成中噪声示例的检测能力。

Dual Debiasing for Noisy In-Context Learning for Text Generation

  • 用合成邻域修正困惑度估计,消除标注与模型知识带来的偏差。
  • 在高噪声比例下仍能准确识别样本清洁度,性能接近全干净数据集。
  • 适合需要可靠示例筛选的复杂文本生成任务,尤其抗噪能力强。

上下文学习(ICL)依赖高质量的示例,而现有方法通过局部困惑度排序识别噪声标注,假设噪声样本的困惑度更高。但当噪声比例高、多数示例存在缺陷时,该假设失效。本文重新审视文本生成场景下的困惑度范式,揭示困惑度中的两大偏差来源:标注本身与大语言模型(LLM)固有的领域知识。为此,提出双去偏框架,利用合成邻域显式校正困惑度估计,构建鲁棒的样本清洁度评分。该指标可无惧整体语料噪声水平,准确反映样本绝对清洁度。大量实验证明,本方法在噪声检测上表现更优,最终的ICL性能接近完全清洁示例库的表现,且在极高噪声比下依然稳健。

原文摘要 · Abstract (English)

In context learning (ICL) relies heavily on high quality demonstrations drawn from large annotated corpora. Existing approaches detect noisy annotations by ranking local perplexities, presuming that noisy samples yield higher perplexities than their clean counterparts. However, this assumption breaks down when the noise ratio is high and many demonstrations are flawed. We reexamine the perplexity based paradigm for text generation under noisy annotations, highlighting two sources of bias in perplexity: the annotation itself and the domain specific knowledge inherent in large language models (LLMs). To overcome these biases, we introduce a dual debiasing framework that uses synthesized neighbors to explicitly correct perplexity estimates, yielding a robust Sample Cleanliness Score. This metric uncovers absolute sample cleanliness regardless of the overall corpus noise level. Extensive experiments demonstrate our method's superior noise detection capabilities and show that its final ICL performance is comparable to that of a fully clean demonstration corpus. Moreover, our approach remains robust even when noise ratios are extremely high.

上下文学习去偏文本生成噪声检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。