把人工标注指南改造成大模型能懂的指令,提升自动标注效率
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
- 用大模型自我审查方式重写标注指南,使其适合模型执行
- 在NCBI疾病语料库上验证,改写后指南可有效指导大模型标注
- 适合需要大规模自动化标注的研究者和数据标注团队
本研究探讨如何将原有面向人类标注员的指南改造为适合大语言模型(LLM)执行的指令。传统指南依赖人类内化训练,而大模型需明确、结构化的指令。我们提出一种以监督为导向的指南重用方法,通过大模型自身审核过程,将原始指南转化为对大模型清晰的执行指令。以NCBI疾病语料库为案例,实验表明改写后的指南能有效引导大模型进行文本标注,同时揭示若干实际挑战。结果表明该流程具备支持规模化、低成本优化标注指南与实现自动化标注的潜力。
原文摘要 · Abstract (English)
This study investigates how existing annotation guidelines can be repurposed to instruct large language model (LLM) annotators for text annotation tasks. Traditional guidelines are written for human annotators who internalize training, while LLMs require explicit, structured instructions. We propose a moderation-oriented guideline repurposing method that transforms guidelines into clear directives for LLMs through an LLM moderation process. Using the NCBI Disease Corpus as a case study, our experiments show that repurposed guidelines can effectively guide LLM annotators, while revealing several practical challenges. The results highlight the potential of this workflow to support scalable and cost-effective refinement of annotation guidelines and automated annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。