首个越南语隐喻理解基准,揭示大模型在文化语境下的理解短板
VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP

- 构建1636个越南谚语数据集,标注五类复杂性与七类语义主题
- 多模型评估显示顶尖模型得分不足满分一半,越语专用模型远逊于通用模型
- 发现模型常误读字面、缺词汇、忽略语用,适合研究文化理解的NLP学者
我们提出VIVID(越南语习语验证与解释深度基准),首个系统性评估越南语文化嵌入式隐喻语言理解的基准。VIVID包含1,636个习语和谚语,标注了五种复杂性特征(字面表达、语用细微差别、汉越词、生僻词汇、民间知识)和七类语义主题。我们建立生成与判别结合的评估框架,提出基于提示的LLM作为裁判方法,经人类判断验证(Cohen's kappa = 0.792)。评估八个先进模型发现显著差距:越语专用模型表现远低于多语言系统(VinaLLaMA-7B: 0.13 vs. GPT-4o: 2.46),即使最佳模型也未达满分50%。值得注意的是,少样本提示并非普遍提升性能,GPT-4o因风格过拟合出现退化。分析揭示系统性失败:字面过度解读、词汇缺失、语用扁平化,表明当前模型缺乏对隐喻文化内涵的深层理解。VIVID为推进文化丰富语境下的隐喻理解提供关键工具。
原文摘要 · Abstract (English)
We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese. VIVID comprises 1,636 idioms and proverbs annotated with five complexity traits (literal expressions, pragmatic nuances, Sino-Vietnamese terms, uncommon vocabulary, folk knowledge) and seven semantic themes. We establish an evaluation framework combining generative and discriminative tasks, proposing an LLM-as-a-Judge approach with aspect-based prompting validated against human judgment (Cohen's kappa = 0.792). Evaluating eight state-of-the-art models reveals critical gaps: Vietnamese-specialized models drastically underperform multilingual systems (VinaLLaMA-7B: 0.13 vs. GPT-4o: 2.46), and even top models achieve less than 50% of maximum scores. Notably, few-shot prompting does not universally improve performance, with GPT-4o exhibiting degradation due to stylistic overfitting. Our analysis exposes systematic failures including literal over-interpretation, lexical gaps, and pragmatic flattening, demonstrating that current models lack cultural competence for nuanced figurative interpretation. VIVID provides an essential tool for advancing figurative language understanding in culturally rich contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。