arXiv:2605.10241cs.CLcs.LG2026-05被引 1

构建韩语金融对话语料库,提升银行客服NLU模型性能

Building Korean linguistic resource for NLU data generation of banking app CS dialog system

  • 基于韩语银行应用评论提取三类语义模式并建模为局部语法图
  • 使用生成数据训练的DIET+KorBERT模型在意图识别上达95%准确率
  • 适合需要韩语金融对话系统训练数据的研究者与开发者

自然语言理解(NLU)是任务导向型对话系统的核心,但需要大量标注数据以覆盖多样化的用户表达。本文报告了名为FIAD(Financial Annotated Dataset)的韩语金融领域语言资源的构建,并用于生成银行客服场景的韩语标注训练数据。通过对银行应用评论语料的实证分析,我们识别出韩语请求句中的三种语言模式:主题(实体、特征)、事件和话语标记符,并将其表示为局部语法图(LGG),以生成涵盖多种意图和实体的标注数据。为评估该资源的实用性,我们在FIAD生成的数据上训练了多个模型:仅用DIET(意图:0.91 / 主题[实体+特征]:0.83)、DIET+HANBERT(I:0.94/T:0.85)、DIET+KoBERT(I:0.94/T:0.86)和DIET+KorBERT(I:0.95/T:0.84),用于提取各类语义项,结果表明其在实际任务中具有显著效果。

原文摘要 · Abstract (English)

Natural language understanding (NLU) is integral to task-oriented dialog systems, but demands a considerable amount of annotated training data to increase the coverage of diverse utterances. In this study, we report the construction of a linguistic resource named FIAD (Financial Annotated Dataset) and its use to generate a Korean annotated training data for NLU in the banking customer service (CS) domain. By an empirical examination of a corpus of banking app reviews, we identified three linguistic patterns occurring in Korean request utterances: TOPIC (ENTITY, FEATURE), EVENT, and DISCOURSE MARKER. We represented them in LGGs (Local Grammar Graphs) to generate annotated data covering diverse intents and entities. To assess the practicality of the resource, we evaluate the performances of DIET-only (Intent: 0.91 /Topic [entity+feature]: 0.83), DIET+ HANBERT (I:0.94/T:0.85), DIET+ KoBERT (I:0.94/T:0.86), and DIET+ KorBERT (I:0.95/T:0.84) models trained on FIAD-generated data to extract various types of semantic items.

韩语NLU金融对话数据生成语义解析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。