针对韩语金融领域构建评估基准,检验大模型在知识、推理与安全上的表现。
KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding
- 基于半自动化流程生成超1000道韩语金融题,覆盖知识、法律推理与毒性内容。
- 不同模型在准确率与输出安全性上差异显著,存在明显权衡关系。
- 专为韩国金融监管与语言环境设计,适合评估高风险场景下的AI可靠性。
我们提出KFinEval-Pilot,一个专为评估大语言模型在韩语金融领域表现而设计的基准套件。针对现有以英语为中心的评测体系的局限性,该基准包含超过1,000道经筛选的问题,涵盖金融知识、法律推理与金融毒性三个关键维度。通过结合GPT-4生成提示与专家验证的半自动化流程,确保内容领域相关性与事实准确性。我们评估了多种代表性大模型,发现不同模型在任务准确率与输出安全性之间存在显著性能差异,凸显大模型在高风险金融应用中推理与安全性的持续挑战。该基准基于真实金融应用场景,契合韩国监管与语言背景,可作为开发更安全可靠金融AI系统的早期诊断工具。
原文摘要 · Abstract (English)
We introduce KFinEval-Pilot, a benchmark suite specifically designed to evaluate large language models (LLMs) in the Korean financial domain. Addressing the limitations of existing English-centric benchmarks, KFinEval-Pilot comprises over 1,000 curated questions across three critical areas: financial knowledge, legal reasoning, and financial toxicity. The benchmark is constructed through a semi-automated pipeline that combines GPT-4-generated prompts with expert validation to ensure domain relevance and factual accuracy. We evaluate a range of representative LLMs and observe notable performance differences across models, with trade-offs between task accuracy and output safety across different model families. These results highlight persistent challenges in applying LLMs to high-stakes financial applications, particularly in reasoning and safety. Grounded in real-world financial use cases and aligned with the Korean regulatory and linguistic context, KFinEval-Pilot serves as an early diagnostic tool for developing safer and more reliable financial AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。