语言模型常分不清不可能事件和不太可能事件,表现反而不如随机猜测。
Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events
- 通过分离可能性、常见性和上下文关联性,测试多个模型的判断能力
- 在特定条件下,模型给不可能句子的打分反而高于不太可能的句子
- 适合研究AI认知局限、安全评估或提示工程的读者
语言模型能否可靠区分‘可能’与‘仅不常见’的事件?我们通过剥离可能性、典型性与上下文相关性,发现尽管先前研究有正面结论,但当前模型(包括Llama 3、Gemma 2和Mistral NeMo)的表现远非稳健。事实上,在某些条件下,所有测试模型的表现均低于随机水平,反而为不可能事件(如‘汽车被刹车贴了停车罚单’)赋予比不常见事件(如‘汽车被探险家贴了停车罚单’)更高的概率。
原文摘要 · Abstract (English)
Can language models reliably predict that possible events are more likely than merely improbable ones? By teasing apart possibility, typicality, and contextual relatedness, we show that despite the results of previous work, language models' ability to do this is far from robust. In fact, under certain conditions, all models tested - including Llama 3, Gemma 2, and Mistral NeMo - perform at worse-than-chance level, assigning higher probabilities to impossible sentences such as 'the car was given a parking ticket by the brake' than to merely unlikely sentences such as 'the car was given a parking ticket by the explorer'.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。