构建越南叙事文本共指消解数据集,验证大模型处理能力
Coreference Resolution for Vietnamese Narrative Texts
- 基于VnExpress新闻文本构建标注数据集
- GPT-4准确率与回复一致性显著优于GPT-3.5-Turbo
- 为低资源语言共指消解提供基准与工具评估
共指消解是自然语言处理中的关键任务,旨在识别文本中指向同一实体的不同表达。该任务对越南语尤其具挑战性,因其为低资源语言且标注数据稀缺。为此,我们利用广泛阅读的越南新闻平台VnExpress的叙事文本,构建了一个全面的标注数据集,并制定了详细的标注指南以确保一致性和准确性。此外,我们评估了大型语言模型(LLMs)如GPT-3.5-Turbo和GPT-4在该数据集上的表现。结果表明,GPT-4在准确率和回复一致性方面均显著优于GPT-3.5-Turbo,展现出更强的越南语共指消解能力,可作为更可靠的处理工具。
原文摘要 · Abstract (English)
Coreference resolution is a vital task in natural language processing (NLP) that involves identifying and linking different expressions in a text that refer to the same entity. This task is particularly challenging for Vietnamese, a low-resource language with limited annotated datasets. To address these challenges, we developed a comprehensive annotated dataset using narrative texts from VnExpress, a widely-read Vietnamese online news platform. We established detailed guidelines for annotating entities, focusing on ensuring consistency and accuracy. Additionally, we evaluated the performance of large language models (LLMs), specifically GPT-3.5-Turbo and GPT-4, on this dataset. Our results demonstrate that GPT-4 significantly outperforms GPT-3.5-Turbo in terms of both accuracy and response consistency, making it a more reliable tool for coreference resolution in Vietnamese.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。