分析多语言问答的语法特征,发现神经模型更擅长捕捉复杂结构。
Type and Complexity Signals in Multilingual Question Representations
- 构建跨七语种的问答数据集,标注类型与句法复杂度
- 神经探针比统计方法更好识别语言中的结构复杂性
- 适合关注多语言模型内部表征机制的研究者
本研究探究多语言Transformer模型对问题形态句法特征的表示方式。我们构建了包含七种语言的问答类型与复杂度(QTC)数据集,标注了问题类型及依赖长度、树深、词汇密度等复杂度指标。通过引入选择性控制的回归型探针评估方法,量化模型泛化能力的提升。对比冻结的Glot500-m表示、子词TF-IDF基线和微调模型的层间探针表现,结果表明:在具有显式标记的语言中,统计特征能有效分类问题;而神经探针更能捕捉细粒度的结构复杂模式。基于此结果,评估了上下文表示何时优于统计基线,以及参数更新是否削弱预训练语言信息的可用性。
原文摘要 · Abstract (English)
This work investigates how a multilingual transformer model represents morphosyntactic properties of questions. We introduce the Question Type and Complexity (QTC) dataset with sentences across seven languages, annotated with type information and complexity metrics including dependency length, tree depth, and lexical density. Our evaluation extends probing methods to regression labels with selectivity controls to quantify gains in generalizability. We compare layer-wise probes on frozen Glot500-m (Imani et al., 2023) representations against subword TF-IDF baselines, and a fine-tuned model. Results show that statistical features classify questions effectively in languages with explicit marking, while neural probes capture fine-grained structural complexity patterns better. We use these results to evaluate when contextual representations outperform statistical baselines and whether parameter updates reduce the availability of pre-trained linguistic information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。