手把手教你怎么制定文本标注规范和数据集
Guidelines for the Creation of an Annotated Corpus
- 提供从制定规范到数据共享的全流程方法
- 包含定义与实例,确保标注一致性
- 适合语言学、NLP研究者构建高质量数据集
本文基于UMR TETIS成员反馈与科学文献,提出一种通用的文本标注指南与标注语料库创建方法。涵盖方法论、数据存储、共享与价值化等环节,通过定义与实例清晰展示每个步骤,为不同研究场景下语料库的构建与使用提供完整框架。
原文摘要 · Abstract (English)
This document, based on feedback from UMR TETIS members and the scientific literature, provides a generic methodology for creating annotation guidelines and annotated textual datasets (corpora). It covers methodological aspects, as well as storage, sharing, and valorization of the data. It includes definitions and examples to clearly illustrate each step of the process, thus providing a comprehensive framework to support the creation and use of corpora in various research contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。