用人类与大模型协作重标注文本中的说话人属性,提升多语言标注质量。
WhoSaidIt: Human-LLM Collaborative Annotation for Text-Based Multilingual Speaker-Attribute Classification
- 人类与大模型迭代互动,挖掘标注理由并聚焦分歧点重标注。
- 构建涵盖9类属性的多语言数据集WhoSaidIt,发现跨语言标注差异显著。
- 揭示大模型在说话人属性分类中的优势与局限,适合多语言社会计算研究者。
从文本中标注说话人属性本身具有模糊性,尤其在多语言场景下,人口统计与社会线索隐含且文化差异明显。我们提出一种人类-大语言模型(LLM)协作重标注框架,在资源受限条件下稳定多语言说话人属性标签。基于噪声语料,利用LLM通过与专家的迭代交互提取重复的标注依据,并采用分歧聚焦采样进行针对性重标注。借助该框架,我们构建了WhoSaidIt数据集,覆盖九类说话人属性标签。量化了原始与修订标注间的差异,基准测试了近期大模型表现,并分析显式标注理由对模型行为的影响。结果揭示了显著的跨语言标注差异,同时展现了大模型在说话人属性分类中的优势与局限。
原文摘要 · Abstract (English)
Annotating speaker attributes from text is inherently ambiguous, particularly in multilingual settings where demographic and social cues are implicit and culturally variable. We propose a human-large language model (LLM) collaborative re-annotation framework for stabilizing multilingual speaker-attribute labels under practical resource constraints. Starting from a noisy corpus, we use LLMs to surface recurring annotation rationales through iterative interaction with experts, and apply disagreement-focused sampling for targeted re-annotation. Using this framework, we construct WhoSaidIt, a multilingual dataset covering nine speaker-attribute labels. We quantify divergence between original and revised annotations, benchmark recent LLMs, and analyze the effect of explicit rationales on model behavior. Our results reveal substantial cross-lingual differences in annotation decisions and demonstrate both the strengths and limitations of LLMs in speaker-attribute classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。