用大模型生成天体物理摘要,探究其如何编码测量得到的物理信息。
Encoding and Understanding Astrophysical Information in Large Language Model-Generated Summaries
- 通过提示工程影响大模型对物理量的编码方式
- 发现语言中特定词汇与物理统计量有强关联
- 适合关注模型解释性与科学知识编码的研究者
大语言模型在跨领域、跨模态任务中展现出良好泛化能力,并具备上下文学习能力。这促使我们思考:它们能否将通常仅来自科学测量的物理信息,以及文本描述中模糊表达的物理内容进行编码?以天体物理学为实验场景,我们通过两个核心问题展开研究:1)提示工程是否影响大模型对物理量的编码方式?2)语言中的哪些特征最有助于编码测量所代表的物理信息?研究采用稀疏自编码器从文本中提取可解释特征,分析大模型生成摘要时对物理信息的表征机制。
原文摘要 · Abstract (English)
Large Language Models have demonstrated the ability to generalize well at many levels across domains, modalities, and even shown in-context learning capabilities. This enables research questions regarding how they can be used to encode physical information that is usually only available from scientific measurements, and loosely encoded in textual descriptions. Using astrophysics as a test bed, we investigate if LLM embeddings can codify physical summary statistics that are obtained from scientific measurements through two main questions: 1) Does prompting play a role on how those quantities are codified by the LLM? and 2) What aspects of language are most important in encoding the physics represented by the measurement? We investigate this using sparse autoencoders that extract interpretable features from the text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。