arXiv:2410.01100cs.CL2024-10NAACL

打造韩语动词知识探索工具,助力自然语言处理研究。

Unlocking Korean Verbs: A User-Friendly Exploration into the Verb Lexicon

  • 构建用户友好的网页界面,整合动词的句法框架信息
  • 将动词搭配模式与例句精准对齐,提升语义理解
  • 开源解析库,简化韩语语法分析与语义标注

Sejong词典数据集提供了丰富的形态、句法和语义表示,是深入探索韩语语言信息的宝贵资源。本文设计了一款用户友好的网页界面,用于收集和整合与动词相关的语言信息,重点聚焦子分类框架(subcategorization frames)。同时,我们通过将子分类框架与对应的例句进行对齐,实现语义关联的可视化。此外,还发布了一个Python库,可简化韩语句法解析与语义角色标注任务。这些工具旨在帮助研究人员和开发者更高效地利用Sejong数据集,推动韩语自然语言处理应用的发展。

原文摘要 · Abstract (English)

The Sejong dictionary dataset offers a valuable resource, providing extensive coverage of morphology, syntax, and semantic representation. This dataset can be utilized to explore linguistic information in greater depth. The labeled linguistic structures within this dataset form the basis for uncovering relationships between words and phrases and their associations with target verbs. This paper introduces a user-friendly web interface designed for the collection and consolidation of verb-related information, with a particular focus on subcategorization frames. Additionally, it outlines our efforts in mapping this information by aligning subcategorization frames with corresponding illustrative sentence examples. Furthermore, we provide a Python library that would simplify syntactic parsing and semantic role labeling. These tools are intended to assist individuals interested in harnessing the Sejong dictionary dataset to develop applications for Korean language processing.

韩语处理动词框架语义标注工具开发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。