困惑度可能误导模型选择,高自信不等于高准确
Perplexity Cannot Always Tell Right from Wrong
- 利用Transformer连续性理论证明:模型对某序列预测准确时,必存在低困惑度但错误的序列
- 分析等困惑度曲线发现:模型置信度提升需伴随准确性同步提升才能被选中
- 提醒研究者:仅用困惑度评估模型不可靠,尤其在高置信场景下
困惑度作为衡量模型输出‘意外程度’的指标,近年来广泛用于损失函数和模型质量评估。尽管已有研究指出其局限性,但多为经验观察。本文基于Transformer连续性最新成果,严谨证明:若一个紧凑的解码器仅Transformer模型能准确且自信地预测某一序列(强泛化前提),则必然存在另一个序列,其困惑度极低但模型预测错误。进一步通过解析等困惑度图,发现困惑度并非总是选出更准确的模型——模型置信度上升必须伴随准确性同步提升,新模型才可能被选中。
原文摘要 · Abstract (English)
Perplexity -- a function measuring a model's overall level of "surprise" when encountering a particular output -- has gained significant traction in recent years, both as a loss function and as a simple-to-compute metric of model quality. Prior studies have pointed out several limitations of perplexity, often from an empirical manner. Here we leverage recent results on Transformer continuity to show in a rigorous manner how perplexity may be an unsuitable metric for model selection. Specifically, we prove that, if there is any sequence that a compact decoder-only Transformer model predicts accurately and confidently -- a necessary pre-requisite for strong generalisation -- it must imply existence of another sequence with very low perplexity, but not predicted correctly by that same model. Further, by analytically studying iso-perplexity plots, we find that perplexity will not always select for the more accurate model -- rather, any increase in model confidence must be accompanied by a commensurate rise in accuracy for the new model to be selected.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。