arXiv:2608.14622cs.AI2026-08

用专家标准评估大模型育儿建议,发现语言和风格差异显著

A Human-Centred Approach to Benchmarking LLMs for Parenting Advice

论文配图:A Human-Centred Approach to Benchmarking LLMs for Parenting Advice
图 1 · 摘自论文原文
  • 邀请育儿专家设计多维度评价标准,跨语言测试15个模型
  • 综合评分掩盖具体短板,部分模型隐含推荐特定育儿风格
  • 适合关注AI建议伦理与可解释性的开发者及政策制定者

越来越多用户使用大语言模型获取育儿建议。育儿是关键且敏感的社会领域,评估其建议需超越传统信息质量指标,纳入关系性与行为性考量。本文基于育儿专家构建的多维评价体系,采用大模型作为评判者的方法,对15个主流模型在英、中双语环境下共100个育儿场景中的表现进行评估。结果表明,综合得分可能掩盖特定维度的缺陷;不同模型隐含推荐不同育儿风格;语言类型显著影响生成内容。研究强调了评估输出可审计性的重要性,揭示了在育儿等敏感领域评估LLM建议所面临的挑战。成果为直接面向用户的育儿类应用选型与开发提供重要参考。

原文摘要 · Abstract (English)

People are increasingly using large language models (LLMs) to seek advice, including for parenting. Parenting is a critical and socially sensitive domain. Thus, evaluating advice provided by LLMs requires indicators beyond aggregated information quality benchmarks to consider relational and behavioural elements of the responses. With a multi-dimensional rubric created by parenting experts, this paper evaluates 15 LLMs across 100 parenting scenarios in 2 languages (English and Chinese), using an LLM-as-a-judge method. Results show that aggregate scores can hide rubric item-specific weaknesses, models implicitly encourage different parenting styles, and language influences responses. We highlight the importance of evaluation output auditability and challenges involved in evaluating LLM-generated advice in domains like parenting. Our findings provide important insights for selecting LLMs for direct user engagement and the development of user-facing parenting advice applications.

LLM评估育儿建议人本评估多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。