arXiv:2508.05179cs.CL2025-08ACL被引 1

用提示工程与微调模型检测问答中的虚假文本片段

ATLANTIS at SemEval-2025 Task 3: Detecting Hallucinated Text Spans in Question Answering

  • 结合提示工程与合成数据微调,实现文本级幻觉检测
  • 在西班牙语任务中取得第一,在英语、德语中表现优异
  • 适合关注大模型生成安全与可信度的研究者

本文介绍了ATLANTIS团队在SemEval-2025任务3中的贡献,聚焦于检测问答系统中的幻觉文本片段。大型语言模型(LLMs)虽显著提升了自然语言生成能力,但仍易产生错误或误导性内容。为此,我们探索了有无外部上下文的多种方法,包括使用少样本提示的LLM、基于标记级别的分类,以及在合成数据上微调的LLM。值得注意的是,我们的方法在西班牙语任务中排名第一,在英语和德语任务中获得具有竞争力的成绩。该工作凸显了整合相关上下文对缓解幻觉的重要性,并展示了微调模型与提示工程的潜力。

原文摘要 · Abstract (English)

This paper presents the contributions of the ATLANTIS team to SemEval-2025 Task 3, focusing on detecting hallucinated text spans in question answering systems. Large Language Models (LLMs) have significantly advanced Natural Language Generation (NLG) but remain susceptible to hallucinations, generating incorrect or misleading content. To address this, we explored methods both with and without external context, utilizing few-shot prompting with a LLM, token-level classification or LLM fine-tuned on synthetic data. Notably, our approaches achieved top rankings in Spanish and competitive placements in English and German. This work highlights the importance of integrating relevant context to mitigate hallucinations and demonstrate the potential of fine-tuned models and prompt engineering.

幻觉检测大模型提示工程多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。