对比四类大模型生成网页的长期表现,发现克劳德最稳定,代码量主要由模型决定。
The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking

- 在固定公开接口下,用人类评分与模型判官双标准评估4类模型8周生成结果。
- 克劳德综合得分最高,9次胜出;模型推理时间越长,质量并未提升。
- 代码行数主要受模型类型影响,无法预测社交平台曝光量,但模型判官偏宽松。
本文对2025年12月10日至2026年2月4日期间,'HTML AI Battle'项目中17次公开实验生成的68个单文件HTML结果进行为期八周的观察性对比。比较了GPT、Gemini、Grok和Claude四类推理模型,在无定制指令、无个性调优、无修复提示的统一界面协议下表现。每个输出通过浏览器渲染视频,由人工评分并结合Gemini LLM作为裁判,评估提示遵循度、功能正确性和UI质量,再按标准化社交媒体协议分发至X(Twitter)、TikTok和YouTube。同时开展两项监督预测分析:实验级模型预测24小时X平台曝光量,生成级模型预测代码冗长度。结果显示,克劳德在平均表现上最强且最稳定,在主人类加权评分中赢得9/17提示。更长的推理时间并未带来更高整体质量。Gemini作为裁判在功能正确性和总体表现上显著比人类更宽容,而自利偏见问题仍未解决。探索性曝光预测模型在跨验证后表现较弱(MAE = 46,874,R² = -0.377),而代码行数模型表现更好(仅模型家族基线即优于提示感知模型,MAE = 135.2,R² = 0.576)。总体而言,预发布的技术/音频变量不足以预测24小时X曝光,代码冗长度更多由模型家族决定而非提示内容。研究为观察性设计,受限于公共接口漂移、访问路径差异及单一人类评分员。
原文摘要 · Abstract (English)
This paper presents an eight-week observational comparison of 68 single-file HTML generations collected across 17 public experiments in the "HTML AI Battle" project between December 10, 2025 and February 4, 2026. Four reasoning model families, GPT, Gemini, Grok, and Claude, were compared under a fixed public-interface protocol with no custom instructions, no personality tuning, and no repair prompts. Each output was evaluated from a rendered browser video using human scores and a Gemini LLM-as-a-judge layer for prompt adherence, functional correctness, and UI quality, then packaged into a standardized social-media protocol spanning X (Twitter), TikTok, and YouTube. The tracker was also used for two supervised predictive analyses: an experiment-level model for 24-hour X impressions and a generation-level model for HTML verbosity. Under this protocol, Claude was the strongest and most consistent family, leading mean performance and winning 9/17 prompts under the primary human weighted score. Longer measured reasoning time was not associated with higher quality overall. Gemini as a judge was significantly more lenient than the human evaluator on functional correctness and overall performance, while stable self-favoring bias remained unresolved. The exploratory X-impressions model remained weak under post-screen cross-validation (MAE = 46,874, R^2 = -0.377), whereas the HTML-lines model performed better, with a model-family-only baseline outperforming prompt-aware alternatives (MAE = 135.2, R^2 = 0.576). Overall, selected pre-publication technical/audio variables were not sufficient to predict 24-hour X reach, while code verbosity was driven much more by model family than by prompt wording. The comparisons remain observational and are limited by public-interface drift, access-path differences, and one primary human scorer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。