首个针对波斯语习语幻觉的评测基准,揭示大模型在隐喻表达上的严重缺陷。
FFE-Hallu:Hallucinations in Fixed Figurative Expressions:Benchmark of Idioms and Proverbs in the Persian Language
- 构建波斯语习语幻觉评测集,涵盖生成、识别与翻译三类任务
- 六款主流多语言模型均难以区分真实与虚构习语,翻译中幻觉率高
- 揭示大模型缺乏文化语境理解,适合研究语言幻觉与跨文化生成的学者
隐喻语言,特别是固定隐喻表达(FFEs),如习语和谚语,持续给大语言模型(LLMs)带来挑战。与字面表达不同,FFEs具有文化根基,非组合性且固定,极易产生隐喻幻觉。我们定义隐喻幻觉为生成或认可听起来像习语但并不存在于目标语言中的表达。本文提出FFEHallu,首个针对波斯语的全面隐喻幻觉评测基准,包含600个精心设计的样本,覆盖三类任务:(i) 从语义生成习语,(ii) 在四类受控构造下检测伪造习语,(iii) 英文到波斯语的习语互译。评估六款前沿多语言模型发现,其在隐喻能力与文化契合度上存在系统性缺陷。尽管GPT4.1在拒绝伪造习语和召回真实习语方面表现相对较好,多数模型仍无法可靠区分真实与高质量伪造表达,且在跨语言翻译中频繁产生幻觉。结果揭示当前大模型处理隐喻语言的显著差距,凸显构建针对性评测基准以评估和缓解隐喻幻觉的必要性。
原文摘要 · Abstract (English)
Figurative language, particularly fixed figurative expressions (FFEs) such as idioms and proverbs, poses persistent challenges for large language models (LLMs). Unlike literal phrases, FFEs are culturally grounded, largely non-compositional, and conventionally fixed, making them especially vulnerable to figurative hallucination. We define figurative hallucination as the generation or endorsement of expressions that sound idiomatic and plausible but do not exist as authentic figurative expressions in the target language. We introduce FFEHallu, the first comprehensive benchmark for evaluating figurative hallucination in LLMs, with a focus on Persian, a linguistically rich yet underrepresented language. FFEHallu consists of 600 carefully curated instances spanning three complementary tasks: (i) FFE generation from meaning, (ii) detection of fabricated FFEs across four controlled construction categories, and (iii) FFE to FFE translation from English to Persian. Evaluating six state of the art multilingual LLMs, we find systematic weaknesses in figurative competence and cultural grounding. While models such as GPT4.1 demonstrate relatively strong performance in rejecting fabricated FFEs and retrieving authentic ones, most models struggle to reliably distinguish real expressions from high quality fabrications and frequently hallucinate during cross lingual translation. These findings reveal substantial gaps in current LLMs handling of figurative language and underscore the need for targeted benchmarks to assess and mitigate figurative hallucination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。