arXiv:2605.22542cs.CL2026-05被引 1

用大模型构建词语在具体语境中的意义场景,让机器理解词的氛围和情感。

Scene Abstraction for Lexical Semantics: Structured Representations of Situated Meaning

  • 通过提示学习让大模型提取词语在不同语境中的事件、角色与环境
  • 人工标注520个实例,82.4%的人类判断准确率,比纯文本嵌入高11.8个百分点
  • 适合研究语义理解、人机交互与情感计算的学者与开发者

咖啡与茶有许多相似属性,却引发截然不同的场景、氛围与情感联想。这些情境化的词义维度真实且系统,但多数计算模型未予体现。本文提出场景抽象(Scene Abstraction)框架,用于构建词语在使用语境中所参与的解释性场景的结构化表示。每个场景包含上下文场景(事件、实体、环境)与以表达为中心的表达特征谱(参与事件、可泛化属性、唤起情绪),通过少量示例提示大语言模型实现。主要贡献包括:(1) 一种情境化词汇意义的结构化表示框架;(2) COCA-Scenes数据集,涵盖26个关键词的520个使用实例,用于区分不同场景;(3) 两项实验证据表明,场景在人类观察者间具有可识别性(82.4%准确率,较仅依赖文本嵌入提升11.8个百分点),且其场景描述比基于ATOMIC的替代方案更贴近人类对词语在上下文中的理解(三个语义维度下86.4%偏好度)。

原文摘要 · Abstract (English)

Coffee and tea share many properties, yet they evoke strikingly different situations, atmospheres, and affective associations. These situated dimensions of word meaning are real and systematic, but they remain implicit in most computational representations of lexical meaning. We propose Scene Abstraction, a framework for constructing structured representations of the interpretive scenes that words participate in across usage contexts. Each scene consists of a Contextual Scene (Events, Entities, Setting) and an expression-centered Expression Profile (Engaged events, Generalizable properties, Evoked emotions), operationalized through few-shot prompting of a large language model. Our contributions are three-fold: (1) a structured representation framework for situated lexical meaning; (2) COCA-Scenes, a dataset of 520 usage instances across 26 keywords for distinct scene identification; and (3) empirical evidence from two experiments suggesting that scenes are reliably identifiable across human observers (82.4% accuracy, +11.8 pp over text-only embeddings) and that our scene profiles more closely align with human interpretation of words in context than ATOMIC-based alternatives (86.4% preference across three semantic dimensions).

语义建模大模型应用情境理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。