arXiv:2608.14631cs.AI2026-08

测试14个大模型在护肤化学领域的准确性,发现其专业能力不足且易误导用户。

Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study

论文配图:Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study
图 1 · 摘自论文原文
  • 用14个大模型测试护肤化学知识,禁用网络搜索以评估内部知识
  • 模型在定量分析和成分结构识别上表现差,通用问题回答看似合理实则浅显
  • 输出看似权威但含技术错误,更难引发用户质疑,适合研究者与开发者参考

随着消费者越来越多地依赖AI聊天机器人获取护肤建议,大型语言模型(LLMs)在化妆品化学领域的技术准确性仍缺乏充分评估。我们对14个LLMs在一系列与化妆品化学相关的主题上进行了基准测试,包括特定成分的化学性质及常见护肤场景。全程禁用网络搜索,以评估模型的内化知识而非互联网检索能力。总体表现不佳,尤其在定量推理和结构识别任务中缺陷明显。尽管模型能合理回应一般性护肤问题,但其回答普遍缺乏消费者决策所需的深度技术信息。值得注意的是,与AI对话存在风险:听起来权威但包含技术错误的输出,比明确承认不确定性的回应更不易引发用户质疑。这些发现表明,主要基于未经验证公共数据训练的通用型LLMs,目前尚不可靠作为化妆品化学信息来源。未来需在两个方面推进:使用经验证的化学与皮肤科数据进行微调,以及显著提升算法推理能力,才可能使其成为公众可用的资源。

原文摘要 · Abstract (English)

As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large Language Models (LLMs) in cosmetic chemistry remains largely under-evaluated. We benchmarked 14 LLMs on a structured set of topics related to cosmetic chemistry, including the chemical properties of specific cosmetic ingredients and common cosmetic scenarios that may be of interest to consumers. Web search was disabled throughout to assess each model's internalized knowledge rather than its internet retrieval capacity. Overall performance was poor, with the most pronounced deficits in quantitative reasoning and structural identification tasks. While models handled general skincare questions with reasonability, responses consistently lacked the technical depth required for informed consumer decision-making. Notably, conversation with AI can pose a risk: outputs that sound authoritative but contain technical errors are less likely to generate skepticism compared to responses that explicitly acknowledge uncertainty. These findings suggest that general-purpose LLMs, trained predominantly on unverified public data, are currently not reliable sources of cosmetic chemistry information. Progress on two fronts, fine-tuning verified chemical and dermatological datasets, and substantial improvements to algorithmic reasoning, will likely be needed before these tools can be considered as resources for public use.

大模型评估护肤化学可靠性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。