评测大模型生成的气候术语定义,发现准确度一般但可助识别需标准化的词汇。
The Accuracy, Robustness, and Readability of LLM-Generated Sustainability-Related Word Definitions
- 用SBERT比对大模型与IPCC官方定义的匹配度。
- 模型平均准确率0.57-0.59,且生成文本更难读。
- 适合关注气候术语标准化的研究者参考。
统一的语言和标准定义对有效气候讨论至关重要。然而,大模型可能存在误用气候术语的问题。我们对比了GPT-4o-mini、Llama3.1 8B和Mistral 7B生成的300个定义与IPCC官方术语表的一致性,利用SBERT句向量分析其准确性、鲁棒性和可读性。结果显示,大模型平均一致性为0.57–0.59 ± 0.15,生成定义普遍比原文更难理解。模型差异主要出现在具有多重或模糊含义的术语上,表明其可能用于识别需要标准化的词汇。研究揭示了大模型在环境对话中的潜力,同时也强调必须确保其输出与既有术语体系一致,以保障表达清晰与一致。
原文摘要 · Abstract (English)
A common language with standardized definitions is crucial for effective climate discussions. However, concerns exist about LLMs misrepresenting climate terms. We compared 300 official IPCC glossary definitions with those generated by GPT-4o-mini, Llama3.1 8B, and Mistral 7B, analyzing adherence, robustness, and readability using SBERT sentence embeddings. The LLMs scored an average adherence of $0.57-0.59 \pm 0.15$, and their definitions proved harder to read than the originals. Model-generated definitions vary mainly among words with multiple or ambiguous definitions, showing the potential to highlight terms that need standardization. The results show how LLMs could support environmental discourse while emphasizing the need to align model outputs with established terminology for clarity and consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。