arXiv:2501.08413cs.CL2025-01被引 6

用多个开源大模型组合,高效准确标注心理研究中的自由文本。

Labeling Free-text Data using Language Model Ensembles

  • 通过集成不同表现的开源大模型,模拟多人标注过程。
  • 组合模型在预测人工标注时准确率最高,且精度与召回率平衡最佳。
  • 基于语义距离的打分机制有效缓解模型间差异,适合隐私敏感场景。

自由文本在心理研究中广泛存在,能提供量化数据难以捕捉的丰富定性见解。传统由多名人工标注员对特定主题进行标注耗时费力。尽管大语言模型(LLMs)在语言处理上表现优异,但依赖闭源模型的辅助标注方法因缺乏外部使用授权,无法直接应用于自由文本。本研究提出一种本地部署的开源大模型集成框架,在隐私约束下提升对预设主题的标注效果。该框架借鉴多人标注思路,利用不同开源模型间的异质性,通过融合策略平衡模型间的一致性与差异性,采用基于主题描述与模型推理间嵌入距离的关联度评分方法进行指导。在公开的吃障碍相关Reddit数据和患者自由文本反馈数据集上进行了评估,均配有专家人工标注。结果表明:(1) 同规模模型间标注性能存在异质性,部分模型敏感度低但精确度高,另一些则反之;(2) 相较于单个模型,模型集成在预测人类标注时达到最高准确率,并实现最优的精确度-敏感度权衡;(3) 模型间的关联度评分一致性高于二元标签,说明该评分方法能有效缓解模型间标注差异。

原文摘要 · Abstract (English)

Free-text responses are commonly collected in psychological studies, providing rich qualitative insights that quantitative measures may not capture. Labeling curated topics of research interest in free-text data by multiple trained human coders is typically labor-intensive and time-consuming. Though large language models (LLMs) excel in language processing, LLM-assisted labeling techniques relying on closed-source LLMs cannot be directly applied to free-text data, without explicit consent for external use. In this study, we propose a framework of assembling locally-deployable LLMs to enhance the labeling of predetermined topics in free-text data under privacy constraints. Analogous to annotation by multiple human raters, this framework leverages the heterogeneity of diverse open-source LLMs. The ensemble approach seeks a balance between the agreement and disagreement across LLMs, guided by a relevancy scoring methodology that utilizes embedding distances between topic descriptions and LLMs' reasoning. We evaluated the ensemble approach using both publicly accessible Reddit data from eating disorder related forums, and free-text responses from eating disorder patients, both complemented by human annotations. We found that: (1) there is heterogeneity in the performance of labeling among same-sized LLMs, with some showing low sensitivity but high precision, while others exhibit high sensitivity but low precision. (2) Compared to individual LLMs, the ensemble of LLMs achieved the highest accuracy and optimal precision-sensitivity trade-off in predicting human annotations. (3) The relevancy scores across LLMs showed greater agreement than dichotomous labels, indicating that the relevancy scoring method effectively mitigates the heterogeneity in LLMs' labeling.

大模型集成文本标注隐私保护心理研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。