测试大模型对人类嗅觉描述的理解能力,发现存在明显偏差。
Sniff AI: Is My 'Spicy' Your 'Spicy'? Exploring LLM's Perceptual Alignment with Human Smell Experiences
- 设计交互任务让40人描述气味,让AI猜
- 大模型对柠檬、薄荷等气味识别率高,但对迷迭香等失败
- 揭示当前多感官交互中感知对齐的不足,适合做人机协同研究
将人工智能与人类意图对齐至关重要,但感知对齐——即AI如何理解我们所见、所听或所闻——仍鲜有研究。本文聚焦嗅觉,基于40名参与者的用户研究,探究大语言模型(LLMs)对人类气味描述的理解能力。参与者完成‘闻味并描述’的交互任务,系统根据描述尝试猜测其正在体验的气味。该实验评估了大模型在内部高维嵌入空间中对气味关系的上下文理解与表征能力。采用定量与定性方法综合评估性能。结果显示,感知对齐程度有限,模型存在偏见,倾向于正确识别柠檬和薄荷等气味,却持续无法识别迷迭香等其他气味。研究讨论了这些发现对人机对齐进展的意义,强调了在多感官体验融合中增强交互系统的机会与挑战。
原文摘要 · Abstract (English)
Aligning AI with human intent is important, yet perceptual alignment-how AI interprets what we see, hear, or smell-remains underexplored. This work focuses on olfaction, human smell experiences. We conducted a user study with 40 participants to investigate how well AI can interpret human descriptions of scents. Participants performed "sniff and describe" interactive tasks, with our designed AI system attempting to guess what scent the participants were experiencing based on their descriptions. These tasks evaluated the Large Language Model's (LLMs) contextual understanding and representation of scent relationships within its internal states - high-dimensional embedding space. Both quantitative and qualitative methods were used to evaluate the AI system's performance. Results indicated limited perceptual alignment, with biases towards certain scents, like lemon and peppermint, and continued failing to identify others, like rosemary. We discuss these findings in light of human-AI alignment advancements, highlighting the limitations and opportunities for enhancing HCI systems with multisensory experience integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。