arXiv:2605.29523cs.LG2026-05被引 1

首个面向韩语金融多轮RAG幻觉检测的基准,填补领域空白。

K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance

论文配图:K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance
图 1 · 摘自论文原文
  • 基于真实韩语金融文档构建多轮对话并注入幻觉。
  • 强模型仍难识别细粒度金融错误,拒答能力最弱。
  • 适用于金融AI安全评估与韩语大模型优化研究者。

大型语言模型(LLMs)通过检索增强生成(RAG)推动金融自动化,但幻觉仍是高风险场景部署的关键障碍。现有基准集中于单轮、英语任务,未能覆盖韩语金融领域的多轮交互特性与语言法规细节。本文提出K-FinHallu,首个针对韩语金融多轮RAG幻觉检测的基准。我们基于真实韩语金融文档构建多轮对话,并依据上下文可回答性提出的分层分类法注入幻觉,明确支持合理拒答。对前沿及开源大模型进行幻觉检测评测发现,即使最强模型在细粒度金融诊断和拒答行为上仍表现不佳。在训练集上微调8B模型可达到与前沿模型相当的性能,但合理拒答仍是所有模型中最薄弱环节。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have advanced financial automation through Retrieval-Augmented Generation (RAG), yet hallucinations remain a critical barrier to deployment in high-stakes environments. Existing benchmarks focus on single-turn, English-centric tasks, leaving the multi-turn dynamics and linguistic-regulatory nuances of the Korean financial domain unaddressed. We introduce K-FinHallu, the first benchmark for hallucination detection in multi-turn Korean financial RAG. We construct multi-turn dialogues from authentic Korean financial documents and inject hallucinations under a proposed hierarchical taxonomy based on context answerability that explicitly accounts for justified abstention. Benchmarking frontier and open-source LLMs as hallucination detectors, we find that even the strongest models struggle with fine-grained financial diagnostics and refusal behavior. While fine-tuning an 8B model on our training split yields performance competitive with frontier LLMs, justified abstention remains the weakest axis across all evaluated models.

幻觉检测金融AI多轮对话韩语NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。