用荒谬例子测试AI对概念的理解边界,发现模型与人类差异惊人。
Investigating Concept Alignment Using Implausible Category Members

- 用反常识类别成员(如橄榄算车?)检验AI概念边界
- 模型将词语误归为车辆、衣物等类别,蔬菜错判为水果
- 揭示潜在安全风险,适合关注AI可解释性与安全的研究者
构建具备类人日常概念理解的AI系统是实现安全可靠AI的关键。传统方法通过询问合理类别成员(如‘汽车是交通工具吗?’)探测模型知识,但易受训练数据模式影响。本文提出新策略:通过提问荒谬类别成员(如‘橄榄是交通工具吗?’)来考察模型对概念范畴边界的把握,从而探测人类习以为常的概念认知。我们基于罗施和梅尔维斯的经典心理学研究,分析AI对一组基础概念在同类别与跨类别任务中的分类表现,并与人类参与者结果对比。结果显示,模型在多个概念上与人类存在显著且出人意料的差异,例如将‘词语’归入‘车辆’或‘衣物’,将若干‘蔬菜’误判为‘水果’,并将非武器类对象错误划入‘武器’类别。这些概念错位还转化为下游任务中的异常行为,对AI安全性具有重要启示。
原文摘要 · Abstract (English)
Developing AI systems with a human-like understanding of everyday concepts is a key step towards developing safe, reliable systems whose behavior makes sense to humans. When probing concept understanding, asking questions about plausible category members (e.g., "Is a car a vehicle?") is likely to recall patterns in the model's vast training data. We pursue an alternative strategy, characterizing the boundaries of conceptual categories by asking about implausible category members (e.g., "Is an olive a vehicle?") to probe the kind of concept-level knowledge we take for granted in fellow humans. We characterize concept boundaries for a set of fundamental concepts by studying AI systems' assignments of objects to superordinate categories from a classic psychological study by Rosch and Mervis, as well as their assignments of the same objects to mismatched superordinate categories. We compare these assignments to those made by human participants on the full range of within-category and cross-category assignment tasks. Our results reveal a range of concepts for which which models differ in meaningful and surprising ways from humans, including treating "words" as belonging to categories like "vehicles" and "clothing," identifying several "vegetable" category members as "fruit," and assigning exemplars from non-weapon categories to the "weapons" category. We also demonstrate how these instances of concept misalignment translate into problematic downstream behavior with implications for AI safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。