提出可量化提示与预训练目标匹配度的指标,指导选择最佳生成模板。
The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models

- 基于深度排名差异计算掩码指数,评估提示风格适配性
- 在ATOMIC2020上验证该指标与下游生成效果正相关
- 适合低资源场景下高效提取关系知识的研究者使用
大规模预训练语言模型如T5和BERT在生成结构化知识方面表现优异,但其性能依赖于提示策略与预训练目标的一致性。本文提出掩码指数(MI),一种量化指标,用于预测知识关系更适合采用掩码式提示还是前缀式提示进行少样本生成。MI通过对比掩码与非掩码模板的DepthRank得分差异计算,提供客观的训练目标与提示对齐度评估。我们在ATOMIC2020知识库补全基准上的多种关系上评估了MI,结果表明其与下游生成性能呈正相关。这说明MI可用于选择合适的提示模板和适配策略,尤其在低资源环境下提升从预训练模型中提取关系知识的效率。
原文摘要 · Abstract (English)
Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating structured knowledge. However, their performance depends on how closely the prompting strategy matches the objectives used during pretraining. We introduce the Maskability Index (MI), a quantitative metric that estimates whether a knowledge relation is better suited to masked-style prompting or prefix-style prompting in few-shot generation. MI is computed from differences in DepthRank scores between masked and unmasked templates, providing a principled measure of objective-template alignment. We evaluate MI on a diverse set of relations from the ATOMIC2020 knowledge base completion benchmark and show that it is positively correlated with downstream generation performance. These results indicate that MI can help select appropriate prompting templates and adaptation strategies for extracting relational knowledge from pretrained language models, especially in low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。