用专家设计的提示规则提升中文隐喻识别跨数据集稳定性。
Cross-Dataset Stability of Expert-Informed Skill Prompting and Fine-Tuning for Chinese Metaphor Identification
- 用专家定义的隐喻判断流程替代调参,实现跨数据集一致表现。
- 零样本提示加规则后,在外部数据上最低误报率仅差4.08点。
- 适合追求稳定性的中文隐喻识别场景,尤其对非标注数据有效。
不同数据集间文本分布与标注政策差异会导致隐喻识别性能显著波动。本文比较固定专家指导的方法与任务特异性参数调整在中文句子级隐喻识别中的跨数据集表现。对比四种方法:BERT微调(BERT-FT)、基于QLoRA的大模型微调(LLM-FT)、直接零样本大模型提示(LLM-ZS)和冻结流程规则的零样本提示(Skill-ZS)。流程规则整合上下文意义、基本意义、对比与类比四项标准。评估涵盖CMRE Test及两个外部数据集CCIME和CMC。微调结果取三次随机种子均值,零样本结果为单一确定配置。在本域测试集上,BERT-FT达到91.76宏F1;LLM-FT外部平均分最高(83.52),而Skill-ZS接近(82.92),且在所有三数据集上最低分最高(82.64),范围最小(4.08点)。在匹配的零样本对比中,引入技能规则使各数据集隐喻预测减少,显著降低CCIME的误报,但增加CMRE Test与CMC的漏报。结果表明,专家指导的技能提示是实现更均衡跨数据集表现的补充路径,而微调仍保持本域精度优势。据我们所知,这是首个在同一跨数据集评估中对比专家流程与任务微调的中文隐喻识别研究。
原文摘要 · Abstract (English)
Metaphor-identification performance can change markedly across datasets that differ in text distribution and annotation policy. We examine whether a fixed expert-informed procedure produces a more even cross-dataset profile than task-specific parameter adaptation. Four prespecified conditions are compared for Chinese sentence-level metaphor identification: BERT fine-tuning (BERT-FT), QLoRA-based large language model fine-tuning (LLM-FT), direct zero-shot LLM prompting (LLM-ZS), and zero-shot prompting with a frozen procedural Skill (Skill-ZS). The Skill operationalizes established criteria involving contextual meaning, basic meaning, contrast, and comparison. Evaluation covers CMRE Test and two external datasets, CCIME and CMC. Fine-tuned scores are means over three seeds, whereas each zero-shot score comes from one deterministic configuration. Fine-tuning remains strongest on the native test set: BERT-FT reaches 91.76 Macro-F1. LLM-FT has the highest external mean (83.52), while Skill-ZS is close at 82.92 and has both the highest external floor (82.64) and the smallest observed range across all three datasets (4.08 points). In the matched zero-shot comparison, adding the Skill reduces metaphorical predictions on every dataset. This sharply lowers false positives on CCIME but increases false negatives on CMRE Test and CMC. The results position expert-informed Skill prompting as a complementary route to more even observed cross-dataset performance, while fine-tuning retains its advantage in native-data accuracy. To our knowledge, this is the first study to compare an expert-informed procedural Skill with task-specific fine-tuning in the same cross-dataset evaluation of Chinese sentence-level metaphor identification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。