将行为画像标注拆解为可评估的技能单元,验证大模型与人类在不同技能上的表现差异。
Exploring and Testing Skill-Based Behavioral Profile Annotation: Human Operability and LLM Feasibility under Schema-Guided Execution
- 把标注任务拆成14个独立技能,用规则和示例定义每项技能的执行标准
- 5项技能可直接操作,4项需重注释后恢复,5项结构上无法明确界定
- GPT表现稳定但选择性可行,开源模型失败多因规则到技能的转化问题
行为画像(BP)标注难以自动化,因其需同时处理多个语言维度。本文将其视为一系列标注技能的集合,而非单一任务,从技能层面评估大模型辅助标注的可行性。基于3,134条中文隐喻性色彩词衍生语料及14维BP标注框架,构建了以技能文件驱动的标注流程:每个特征通过标注模板、决策规则和示例外部定义。两名人工标注者在300个样本子集上完成两轮仅依赖框架的标注,将技能分类为可直接操作、可经聚焦重注释恢复、或结构不明确三类。随后在相同设置下评估GPT-5.4及三个本地部署开源模型。结果表明,标注在技能层高度异质:5项可直接操作,4项可通过专注重注释恢复,5项仍结构不清。GPT-5.4在保留技能上表现可靠(准确率=0.678,κ=0.665,加权F1=0.695),但可行性具选择性而非全局。人类与GPT在技能层级难度高度一致(r=0.881),但在实例层级(r=0.016)和词汇层级(r=-0.142)无相关性,呈现‘共享分类体系,独立执行’模式。成对一致性分析表明,GPT更宜被视为独立的第三方技能声音,而非人类替代品。开源模型失败集中于从标注框架到具体技能执行的映射问题。研究建议:自动标注应以技能可行性为评价基准,而非任务级自动化。
原文摘要 · Abstract (English)
Behavioral Profile (BP) annotation is difficult to automate because it requires simultaneous coding across multiple linguistic dimensions. We treat BP annotation as a bundle of annotation skills rather than a single task and evaluate LLM-assisted BP annotation from this perspective. Using 3,134 concordance lines of 30 Chinese metaphorical color-term derivatives and a 14-feature BP schema, we implement a skill-file-driven pipeline in which each feature is externally defined through schema files, decision rules, and examples. Two human annotators completed a two-round schema-only protocol on a 300-instance validation subset, enabling BP skills to be classified as directly operable, recoverable under focused re-annotation, or structurally underspecified. GPT-5.4 and three locally deployable open-source models were then evaluated under the same setup. Results show that BP annotation is highly heterogeneous at the skill level: 5 skills are directly operable, 4 are recoverable after focused re-annotation, and 5 remain structurally underspecified. GPT-5.4 executes the retained skills with substantial reliability (accuracy = 0.678, \k{appa} = 0.665, weighted F1 = 0.695), but this feasibility is selective rather than global. Human and GPT difficulty profiles are strongly aligned at the skill level (r = 0.881), but not at the instance level (r = 0.016) or lexical-item level (r = -0.142), a pattern we describe as shared taxonomy, independent execution. Pairwise agreement further suggests that GPT is better understood as an independent third skill voice than as a direct human substitute. Open-source failures are concentrated in schema-to-skill execution problems. These findings suggest that automatic annotation should be evaluated in terms of skill feasibility rather than task-level automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。