arXiv:2504.02858cs.CLcs.LG2025-04被引 2

研究大模型生成程序员幽默的最优配置,发现温度设置和架构影响显著。

Optimizing Humor Generation in Large Language Models: Temperature Configurations and Architectural Trade-offs

  • 测试715种温度与提示组合,用五项指标评估幽默生成效果。
  • 38.7%性能差异由架构决定,低随机性(≤0.5)下73%模型表现最佳。
  • 识别出高效紧凑型与长输出专家型两类模型,适合不同场景使用。

大型语言模型在创意文本生成方面能力不断提升,但其幽默生成的系统性评估仍不充分。本研究对5类架构的13个前沿大模型进行综合分析,评估其为软件开发者生成技术相关幽默的效果。通过全因子设计测试715种温度设置与提示变体组合,采用五项加权指标(幽默质量、领域相关性、概念原创性、语气精准度、表达效率)评估输出结果。方法包含方差分析(ANOVA)、相关性研究及二次回归等严谨统计分析,识别出最优配置与架构影响。结果显示模型间性能差异显著,部分架构较基线系统提升21.8%。温度敏感性分析表明73%模型在低随机性设置(≤0.5)时表现最佳,但最优范围因架构而异。识别出两类模型集群:紧凑型高性能模型保持效率与质量平衡,冗长型专家模型需更长输出才获边际收益。统计验证显示模型架构解释了38.7%的性能变异,且幽默质量与概念原创性显著相关。研究提出实用选型与配置指南,证实温度调整与架构选择对幽默生成有效性有重要影响。成果深化了对大模型在创意技术写作中能力的理解,为开发者构建幽默生成系统提供实证配置策略。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate increasing capabilities in creative text generation, yet systematic evaluations of their humor production remain underexplored. This study presents a comprehensive analysis of 13 state-of-the-art LLMs across five architectural families, evaluating their performance in generating technically relevant humor for software developers. Through a full factorial design testing 715 unique configurations of temperature settings and prompt variations, we assess model outputs using five weighted criteria: humor quality, domain relevance, concept originality, tone precision, and delivery efficiency. Our methodology employs rigorous statistical analysis including ANOVA, correlation studies, and quadratic regression to identify optimal configurations and architectural influences. Results reveal significant performance variations across models, with certain architectures achieving 21.8% superiority over baseline systems. Temperature sensitivity analysis demonstrates that 73% of models achieve peak performance at lower stochasticity settings (<= 0.5), though optimal ranges vary substantially by architecture. We identify distinct model clusters: compact high-performers maintaining efficiency-quality balance versus verbose specialists requiring longer outputs for marginal gains. Statistical validation confirms model architecture explains 38.7% of performance variance, with significant correlations between humor quality and concept originality. The study establishes practical guidelines for model selection and configuration, demonstrating how temperature adjustments and architectural considerations impact humor generation effectiveness. These findings advance understanding of LLM capabilities in creative technical writing and provide empirically validated configuration strategies for developers implementing humor-generation systems.

幽默生成大模型温度调节架构分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。