研究大模型对高频与低频知识的回答差异,发现模型更擅长回答常见概念的定义。
TrackList: Tracing Back Query Linguistic Diversity for Head and Tail Knowledge in Open Large Language Models
- 用细粒度分析工具TrackList,追踪预训练数据对模型回答多样性的影响
- 高频概念(头)回答准确率高,低频技术术语(尾)回答质量显著下降
- 在专家文本中,模型对高频词更倾向生成同义表达,忽视冷门专业内容
大型语言模型在回答定义类问题上表现优异,但在举例、释义或改写等多样化问题上能力明显下降。本文通过TrackList这一细粒度语言学与统计分析管道,研究预训练数据对模型应对多种语言查询的影响。引入RefMed-EN英语医学数据集,包含6170个经人工标注的医学术语及其定义、别称、例证、解释或同义表述。评估结果显示,模型在定义类问题上的表现最优,而举例类问题最差;对于高频概念,模型更倾向于生成同义表达,而在专家文本中对低频技术术语的回应则显著不足。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have proven efficient in giving definition-type answers to user input queries. While for humans giving various types of answers, such as examples and paraphrases, is an easy task, LLMs struggle to provide correct answers for other than definition-type queries. In this study, we evaluated this drop in performance using TrackList, a fine-grained linguistic and statistical analysis pipeline to investigate the impact of the pre-training data on LLMs answers to diverse linguistic queries. We also introduce RefoMed-EN, an English dataset consisting of 6170 human-annotated medical terms alongside their corresponding definitions, denominations, exemplifications, explanations, or paraphrases. We studied whether the high frequency of a concept (head) or low frequency (tail) impacts the language model's performance. We evaluated the quality of the LLM's output using syntactic and semantic similarity metrics, statistical correlations and embeddings. Results showed that the LLM's task performance for definition type questions is the highest, while for the exemplification type it is the lowest. Additionally, we showed that for definition-type questions, large language models are prone to paraphrase more on popular and frequent knowledge and less on tail and technical knowledge, especially in the expert texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。