让专家和大模型协作,自动发现文本分类中的罕见难例并优化标注体系。
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
- 专家提供初始标签框架,大模型自动标注并识别遗漏的异常样本。
- 能提炼出可复用的边缘案例描述规则,提升标注体系覆盖度。
- 适合需要高精度文本分类的领域,如医疗、法律等专业场景。
我们提出 Co-DETECT(文本分类中边缘案例的协同发现),一种融合人类专家知识与大语言模型驱动自动标注的混合主动性标注框架。该框架从领域专家提供的初始草图式标签体系和数据集出发,利用大模型对数据进行标注,并识别出初始标签体系未能充分描述的边缘案例。具体而言,Co-DETECT 能标记具有挑战性的样本,归纳出高层次、可泛化的边缘案例描述,并协助用户将边缘案例处理规则融入标签体系。这一迭代过程通过简洁且通用的标注规则,更有效地应对细微复杂的语义现象。通过大规模用户研究、定性与定量分析,验证了 Co-DETECT 的有效性。
原文摘要 · Abstract (English)
We introduce Co-DETECT (Collaborative Discovery of Edge cases in TExt ClassificaTion), a novel mixed-initiative annotation framework that integrates human expertise with automatic annotation guided by large language models (LLMs). Co-DETECT starts with an initial, sketch-level codebook and dataset provided by a domain expert, then leverages the LLM to annotate the data and identify edge cases that are not well described by the initial codebook. Specifically, Co-DETECT flags challenging examples, induces high-level, generalizable descriptions of edge cases, and assists user in incorporating edge case handling rules to improve the codebook. This iterative process enables more effective handling of nuanced phenomena through compact, generalizable annotation rules. Extensive user study, qualitative and quantitative analyses prove the effectiveness of Co-DETECT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。