Tag

#methodology

23 posts tagged.

lab-notes

Deploy quality: why two Lighthouse runs disagree (2026)

Two Lighthouse runs on the same deployed AI-builder output rarely return the same score. This note documents the variance phenomenon, names its seven sources from Google's own documentation, and publishes the median-of-five protocol a hands-on deploy-quality harness would have to meet. Corrected August 21, 2026: the June 2026 result table it accompanied has been withdrawn, and the account of how that table was produced is withdrawn with it.

16 min read280
Methodology

First-build stability: a v2 axis proposal (June 2026)

Proposing first-build stability as the fifth BuilderProof axis: the fraction of OQ-7 prompts that complete without manual intervention. Failure-mode taxonomy, measurement protocol, scoring rubric and open questions, dated June 20, 2026.

10 min read210
Methodology

How We Benchmark AI App Builders: The BuilderProof Methodology v1

BuilderProof methodology v1.1: the published rubric, brief OQ-7, environment standards and weights used to score AI app builders on output quality, speed, deploy quality and agency suitability. The four June 2026 result sets were withdrawn on August 21, 2026 as placeholder data, so the lab currently publishes method, not scores.

11 min read221