研究多模态语境下如何理解陌生词义,发现关键线索与学习者背景相关。
Towards Understanding Ambiguity Resolution in Multimodal Inference of Meaning
- 通过图文配对实验,分析影响词义推断的视觉与语言特征。
- 仅部分直观特征与推断成功率强相关,需进一步挖掘预测因子。
- 评估AI对人类表现的推理能力,指明改进方向,适合教育AI研究者参考。
我们研究了一种新的外语学习场景:学习者在包含配对图像的句子中推断生词含义。通过不同图像-文本对的人类参与者实验,分析了哪些数据特征(如图像、文本)有助于参与者推断被遮蔽或不熟悉的词语意义,以及参与者的语言背景与推断成功之间的关联。结果发现,只有部分直观特征与参与者表现有强相关性,提示需要进一步探索任务成功的预测特征。同时,我们评估了人工智能系统对参与者表现的推理能力,发现了提升该能力的有前景方向。
原文摘要 · Abstract (English)
We investigate a new setting for foreign language learning, where learners infer the meaning of unfamiliar words in a multimodal context of a sentence describing a paired image. We conduct studies with human participants using different image-text pairs. We analyze the features of the data (i.e., images and texts) that make it easier for participants to infer the meaning of a masked or unfamiliar word, and what language backgrounds of the participants correlate with success. We find only some intuitive features have strong correlations with participant performance, prompting the need for further investigating of predictive features for success in these tasks. We also analyze the ability of AI systems to reason about participant performance, and discover promising future directions for improving this reasoning ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。