测试大模型生成健身处方的一致性,发现语义稳定但强度等关键参数波动大。
Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Using a Large Language Model
- 用6种临床场景重复生成120份处方,评估语义、结构和安全表达一致性
- 语义相似度均在0.879以上,但运动强度等量化指标变化大,抗阻训练中有10%-25%无法分类强度
- 所有处方都包含安全表述,但不同场景安全句子数量差异显著,临床病例更多
背景:大语言模型(LLMs)被探索用于生成个性化运动处方,但其在相同条件下输出的一致性尚未充分检验。目的:本研究通过重复生成设计,评估了大语言模型生成运动处方的模型内一致性。方法:使用6个临床场景,以Gemini 2.5 Flash生成运动处方(每场景20次,共n=120)。一致性从三个维度评估:(1)基于SBERT的语义一致性,(2)基于FITT原则的结构一致性,采用AI作为评判者;(3)安全性表达一致性,包括包含率与句级量化。结果:各场景语义相似度较高(平均余弦相似度:0.879–0.939),临床约束强的案例一致性更高。频率模式一致,但定量成分存在变异,尤其在运动强度方面。抗阻训练输出中10%-25%的强度表达无法分类。所有输出均包含安全性表述,但不同场景间安全句数量差异显著(H=86.18, p<0.001),临床案例产生的安全表达多于健康成人案例。结论:大语言模型生成的运动处方具有高语义一致性,但在关键定量成分上存在变异性。可靠性高度依赖提示结构,临床部署前需添加结构约束与专家验证。
原文摘要 · Abstract (English)
Background: Large language models (LLMs) have been explored as tools for generating personalized exercise prescriptions, yet the consistency of outputs under identical conditions remains insufficiently examined. Objective: This study evaluated the intra-model consistency of LLM-generated exercise prescriptions using a repeated generation design. Methods: Six clinical scenarios were used to generate exercise prescriptions using Gemini 2.5 Flash (20 outputs per scenario; total n = 120). Consistency was assessed across three dimensions: (1) semantic consistency using SBERT-based cosine similarity, (2) structural consistency based on the FITT principle using an AI-as-a-judge approach, and (3) safety expression consistency, including inclusion rates and sentence-level quantification. Results: Semantic similarity was high across scenarios (mean cosine similarity: 0.879-0.939), with greater consistency in clinically constrained cases. Frequency showed consistent patterns, whereas variability was observed in quantitative components, particularly exercise intensity. Unclassifiable intensity expressions were observed in 10-25% of resistance training outputs. Safety-related expressions were included in 100% of outputs; however, safety sentence counts varied significantly across scenarios (H=86.18, p less than 0.001), with clinical cases generating more safety expressions than healthy adult cases. Conclusions: LLM-generated exercise prescriptions demonstrated high semantic consistency but showed variability in key quantitative components. Reliability depends substantially on prompt structure, and additional structural constraints and expert validation are needed before clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。