BERT的语义关系能力可被训练诱导,非预训练自然涌现。
Relational Schemata in BERT Are Inducible, Not Emergent: A Study of Performance vs. Competence in Language Models
- 通过对比[CLS]嵌入结构与分类表现,检验关系表征
- 预训练后关系结构未显现,微调后才形成分组
- 性能好不等于理解深,需任务引导才能建模关系
尽管BERT在语义任务上表现出色,但其是否具备真正的概念理解能力仍存疑。本文通过分析概念对在分类、部分整体及功能关系中的内部表示,比较BERT在[CLS] token嵌入中的表征结构与关系分类表现。结果表明,预训练的BERT虽能实现高分类准确率,显示潜在的关系信号;但概念对仅在经过监督关系分类任务微调后,才在高维嵌入空间中按关系类型组织。这说明关系模式并非预训练中自然涌现,而是可通过任务引导诱导形成。研究揭示:行为表现强并不等同于具备结构化概念理解,模型可通过合适训练获得基于关系抽象的归纳偏置。
原文摘要 · Abstract (English)
While large language models like BERT demonstrate strong empirical performance on semantic tasks, whether this reflects true conceptual competence or surface-level statistical association remains unclear. I investigate whether BERT encodes abstract relational schemata by examining internal representations of concept pairs across taxonomic, mereological, and functional relations. I compare BERT's relational classification performance with representational structure in [CLS] token embeddings. Results reveal that pretrained BERT enables high classification accuracy, indicating latent relational signals. However, concept pairs organize by relation type in high-dimensional embedding space only after fine-tuning on supervised relation classification tasks. This indicates relational schemata are not emergent from pretraining alone but can be induced via task scaffolding. These findings demonstrate that behavioral performance does not necessarily imply structured conceptual understanding, though models can acquire inductive biases for grounded relational abstraction through appropriate training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。