arXiv:2509.24866cs.CLcs.AI2025-09被引 8

用大模型自动识别隐喻,三种方法效果对比。

Metaphor identification using large language models: A comparison of RAG, prompt engineering, and fine-tuning

  • 用检索增强生成、提示工程和微调三法试跑隐喻识别
  • 微调法在全文本上达中位F1 0.79,表现最优
  • 结果可辅助理论构建,适合语言学与认知研究者

隐喻是话语的普遍特征,也是研究认知、情感与意识形态的重要视角。然而,由于隐喻具有强上下文依赖性,大规模分析长期受限于人工标注。本研究探索大语言模型(LLMs)在全文隐喻识别中的自动化潜力。比较三种方法:(i) 检索增强生成(RAG),提供规则与示例让模型依规标注;(ii) 提示工程,设计任务特定指令;(iii) 微调,基于人工标注数据训练模型以优化性能。在提示工程中,测试零样本、少样本及思维链策略。结果显示,顶尖闭源大模型可实现高准确率,微调法达中位F1 0.79。人机输出对比发现,多数差异具系统性,反映隐喻理论中的经典模糊地带与概念挑战。研究建议大模型可部分自动化隐喻识别,并作为发展与完善识别协议及理论的实验平台。

原文摘要 · Abstract (English)

Metaphor is a pervasive feature of discourse and a powerful lens for examining cognition, emotion, and ideology. Large-scale analysis, however, has been constrained by the need for manual annotation due to the context-sensitive nature of metaphor. This study investigates the potential of large language models (LLMs) to automate metaphor identification in full texts. We compare three methods: (i) retrieval-augmented generation (RAG), where the model is provided with a codebook and instructed to annotate texts based on its rules and examples; (ii) prompt engineering, where we design task-specific verbal instructions; and (iii) fine-tuning, where the model is trained on hand-coded texts to optimize performance. Within prompt engineering, we test zero-shot, few-shot, and chain-of-thought strategies. Our results show that state-of-the-art closed-source LLMs can achieve high accuracy, with fine-tuning yielding a median F1 score of 0.79. A comparison of human and LLM outputs reveals that most discrepancies are systematic, reflecting well-known grey areas and conceptual challenges in metaphor theory. We propose that LLMs can be used to at least partly automate metaphor identification and can serve as a testbed for developing and refining metaphor identification protocols and the theory that underpins them.

隐喻识别大模型应用语言分析自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。