探究自回归大模型对事件语义适配性的理解能力
Uncovering Autoregressive LLM Knowledge of Thematic Fit in Event Representation
- 设计多种提示策略测试模型对语义角色适配度的判断
- 在主题适配基准上达到新最好成绩,但闭源模型更依赖多步推理
- 词元输入与句子输入导致评分分布差异显著,揭示模型内部机制
主题适配估计任务衡量语义论元与给定谓词的语义角色之间的兼容性。本文通过实验探索自回归大语言模型是否具备一致且可表达的事件论元主题适配知识,采用多种提示设计,操控输入上下文、推理过程与输出形式。在主题适配基准上取得新的最佳性能,但发现封闭权重与开放权重模型对提示策略响应不同:封闭模型整体表现更优,且受益于多步推理,但在过滤与给定谓词、角色和论元不兼容的生成句方面表现较差。分析显示,词元元组输入与句子输入导致主题适配分数分布出现意外差异。
原文摘要 · Abstract (English)
The thematic fit estimation task measures semantic arguments' compatibility with a given semantic role for a given predicate. We investigate if autoregressive LLMs have consistent, expressible knowledge of event arguments' thematic fit by experimenting with various prompt designs, manipulating input context, reasoning, and output forms. We set a new state-of-the-art on thematic fit benchmarks, but show that closed and open weight LLMs respond differently to our prompting strategies: Closed models achieve better scores overall and benefit from multi-step reasoning, but they perform worse at filtering out generated sentences incompatible with the given predicate, role, and argument. Our analysis shows that lemma tuple input and sentence input result in surprisingly different thematic fit score distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。