arXiv:2606.07479cs.CLcs.AI2026-06ACL

用提示词让大模型识别土耳其语习语动词结构,效果依赖演示样例设计。

Supervision versus Demonstration-Based In-Context Learning for Multiword Expression Classification

论文配图:Supervision versus Demonstration-Based In-Context Learning for Multiword Expression Classification
图 1 · 摘自论文原文
  • 用指令微调的大模型通过少量示例提示实现习语检测。
  • 三类模型在少样本下对习语的召回率最高达82.3%,但零样本时不足30%。
  • 精心设计的演示可避免模型偏见,适合语言学与低资源场景研究者。

土耳其语习语轻动词结构(LVC)因表面形式与纯字面动宾搭配相同,而功能上为部分习语化谓词,处理难度大。本文将其视为二分类任务(字面义 vs. 习语义),在人工构建的受控数据集(N=147)上评估:包含域外随机句和域内字面控制句(NLVC)作为负样本,以及习语正样本。对比了基于BERTurk的监督基线模型与三种来自不同家族的指令微调大模型,在零样本、单样本及少样本提示下的表现,并分析演示如何改变错误模式。零样本下,大模型对负样本表现良好,但习语召回率极低;单样本提示显著提升习语检测能力,但可能引入强模型特异性偏差,导致过度或不足预测;更丰富的少样本提示提升了校准度,使GPT-OSS-20B与Qwen 2.5-14B获得稳健整体性能。总体表明,土耳其语元语言分类对提示高度敏感:监督基线仍具竞争力,而经精心设计演示后,提示大模型可在习语检测上达到甚至超越基线。

原文摘要 · Abstract (English)

Turkish idiomatic light verb constructions (LVCs) are challenging for multiword expression processing because they often share the same surface form as fully literal verb-object combinations while functioning as a single, partially idiomatic predicate. We frame Turkish LVC detection as a binary classification task (literal meaning vs. idiomatic meaning) and evaluate on a manually created controlled set (N=147) with matched negatives: out-of-domain random sentences and in-domain literal controls (NLVC), alongside LVC positives. We compare a supervised Turkish encoder baseline (BERTurk with a classifier head) to three instruction-tuned LLMs from different families under zero-shot, one-shot, and few-shot prompting, and analyze how demonstrations shift error profiles. In zero-shot, LLMs perform well on negatives but show very low LVC recall. One-shot prompting sharply improves LVC detection but can induce strong, model-specific biases, leading models to overpredict or underpredict LVCs. A richer few-shot prompt improves calibration and yields robust overall performance for GPT-OSS-20B and Qwen 2.5-14B. Overall, the results highlight substantial prompt sensitivity in Turkish metalinguistic classification: the supervised baseline remains competitive, while prompted LLMs can match or exceed it on LVCs with carefully constructed demonstrations.

自然语言少样本学习习语识别大模型提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。