剖析大模型隐喻理解的三种机制,发现其表现可能依赖表面信号而非深层语义。
Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing
- 通过几何探测、词汇替换和句法扰动三类诊断实验分析模型隐喻处理机制。
- 模型生成解释存在语义漂移,对新隐喻的识别易受句法异常影响。
- 适合关注大模型语义理解局限性的研究人员参考。
大型语言模型在隐喻检测与解释任务中表现优异,但其行为成功背后的真实语义处理机制仍不清晰。本文通过诊断性分析,考察三个互补维度:语义属性对齐、词汇稳定性与句法敏感性。利用几何探测评估模型生成解释与参考语义属性的一致性;通过上下文变化下的词汇替换,分析隐喻与字面表达间词汇关联的稳定性;通过受控句法扰动,检验隐喻识别的敏感性。结果表明,模型生成解释可能出现相对于参考属性的语义漂移;稳定的词汇锚点在不同语境中持续存在,可能支持常规隐喻但削弱需要上下文整合的新隐喻理解;检测性能对句法异常敏感。这些发现表明,强大的行为表现可能反映异质性底层信号,提示在将隐喻基准视为稳健、集成语义理解证据时需谨慎。
原文摘要 · Abstract (English)
Large language models (LLMs) achieve strong performance on metaphor detection and interpretation tasks, yet it remains unclear what such behavioral success reveals about metaphor processing. We present a diagnostic analysis that examines the limits of behavioral evidence by probing three complementary dimensions: semantic attribute alignment, lexical invariance, and syntactic sensitivity. Using geometric probing, we assess whether model-generated interpretations align with reference semantic attributes; through context-varying substitution, we analyze the stability of lexical associations between metaphorical and literal expressions; and via controlled syntactic perturbations, we examine sensitivity in metaphor detection. Our analysis reveals that LLM-generated interpretations can exhibit semantic drift relative to reference attributes; stable lexical anchors persist across contextual conditions, potentially supporting conventional metaphors while biasing novel metaphors requiring contextual integration; and detection performance is sensitive to syntactic irregularities. These findings suggest that strong behavioral performance may reflect heterogeneous underlying signals, highlighting the need for caution when interpreting metaphor benchmarks as evidence of robust, integrated semantic understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。