arXiv:2503.19650cs.CLcs.AI2025-03ACL被引 1

用合成数据训练模型,细粒度检测大模型幻觉。

HausaNLP at SemEval-2025 Task 3: Towards a Fine-Grained Model-Aware Hallucination Detection

  • 基于自然语言推理与合成数据微调ModernBERT模型
  • 置信度与幻觉存在性相关性达0.422,但重叠率仅0.032
  • 适合关注多语言幻觉检测与模型可信性研究者

本文报告了我们在多语言幻觉与相关过生成错误共享任务(MU-SHROOM)中的研究成果,该任务聚焦于识别大语言模型(LLMs)输出中的幻觉及过生成错误。任务要求在14种语言中检测出构成幻觉的具体文本片段。为应对该任务,我们旨在提供对英语中幻觉发生及其严重程度的细致、模型感知理解。通过使用自然语言推理并基于400个样本的合成数据微调ModernBERT模型,取得了0.032的交并比(IoU)和0.422的相关性得分。结果表明,模型置信度与实际幻觉存在之间具有中等正相关性。然而,较低的IoU值表明预测的幻觉范围与真实标注之间的重叠程度有限。这一表现符合预期,因幻觉常以微妙形式出现,依赖上下文,准确界定其边界极具挑战。

原文摘要 · Abstract (English)

This paper presents our findings of the Multilingual Shared Task on Hallucinations and Related Observable Overgeneration Mistakes, MU-SHROOM, which focuses on identifying hallucinations and related overgeneration errors in large language models (LLMs). The shared task involves detecting specific text spans that constitute hallucinations in the outputs generated by LLMs in 14 languages. To address this task, we aim to provide a nuanced, model-aware understanding of hallucination occurrences and severity in English. We used natural language inference and fine-tuned a ModernBERT model using a synthetic dataset of 400 samples, achieving an Intersection over Union (IoU) score of 0.032 and a correlation score of 0.422. These results indicate a moderately positive correlation between the model's confidence scores and the actual presence of hallucinations. The IoU score indicates that our model has a relatively low overlap between the predicted hallucination span and the truth annotation. The performance is unsurprising, given the intricate nature of hallucination detection. Hallucinations often manifest subtly, relying on context, making pinpointing their exact boundaries formidable.

幻觉检测大模型多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。