arXiv:2510.25783cs.CLcs.AI2025-10

构建首个大规模韩语无目标立场数据集,助力低资源语言研究

LASTIST: LArge-Scale Target-Independent STance dataset

  • 基于韩语政论新闻构建56万条无目标立场标注数据
  • 支持无目标立场检测与立场演变分析等新任务
  • 开源数据集推动低资源语言立场研究

立场检测已成为人工智能领域的研究热点,但现有研究多集中于针对特定目标的立场判断,且主流基准数据集以英语为主,难以支持韩语等低资源语言的模型开发。本文提出LArge-Scale Target-Independent STance(LASTIST)数据集,从韩国两大政党发布的新闻稿中收集563,299条韩语句子并进行标注。该数据集专为无目标立场检测及历时性立场演变分析等任务设计,同时提供了数据构建过程、标注规范及先进深度学习模型训练方法。相关数据已公开发布于https://anonymous.4open.science/r/LASTIST-3721/,旨在填补韩语立场检测研究空白。

原文摘要 · Abstract (English)

Stance detection has emerged as an area of research in the field of artificial intelligence. However, most research is currently centered on the target-dependent stance detection task, which is based on a person's stance in favor of or against a specific target. Furthermore, most benchmark datasets are based on English, making it difficult to develop models in low-resource languages such as Korean, especially for an emerging field such as stance detection. This study proposes the LArge-Scale Target-Independent STance (LASTIST) dataset to fill this research gap. Collected from the press releases of both parties on Korean political parties, the LASTIST dataset uses 563,299 labeled Korean sentences. We provide a detailed description of how we collected and constructed the dataset and trained state-of-the-art deep learning and stance detection models. Our LASTIST dataset is designed for various tasks in stance detection, including target-independent stance detection and diachronic evolution stance detection. We deploy our dataset on https://anonymous.4open.science/r/LASTIST-3721/.

立场检测韩语数据无目标开源数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。