Benchmarks

Versioned, reproducible benchmarks of AI app builders - methodology-first and re-tested monthly.

40 posts

Benchmarks

Replit vs Lovable (2026): Benchmarked Across 6 Axes

On BuilderProof's six-axis rubric, Replit edges Lovable 45 to 44 in 2026. Replit wins first-build stability, deployment breadth, and default auth posture; Lovable wins code portability and front-end output quality. A reproducible, documentation-sourced head-to-head.

9 min read124
Benchmarks

Replit vs Bolt (2026): Benchmarked Across 6 Axes

On BuilderProof's six-axis rubric, Replit and Bolt tie 44 to 44 in 2026. Replit wins first-build stability and deployment breadth; Bolt wins output coherence and default-on credential hygiene. A reproducible, documentation-sourced head-to-head.

8 min read66
lab-notes

Deploy quality: why two Lighthouse runs disagree (2026)

Two Lighthouse runs on the same deployed AI-builder output rarely return the same score, and that is the most-contested observation in the lab notebook for the four June 2026 BuilderProof axes. This note documents the variance phenomenon, the reproducibility protocol the next iteration will adopt, and where median-of-five runs out of road. No score from the published table is changed.

15 min read115
Methodology

First-build stability: a v2 axis proposal (June 2026)

Proposing first-build stability as the fifth BuilderProof axis: the fraction of OQ-7 prompts that complete without manual intervention. Failure-mode taxonomy, measurement protocol, scoring rubric and open questions, dated June 20, 2026.

9 min read110
Methodology

How We Benchmark AI App Builders: The BuilderProof Methodology v1

The BuilderProof methodology v1, dated June 19, 2026, in full: four axes, the OQ-7 test brief, environment standards, scoring weights, reproducibility steps, the operator disclosure, and the v2 open questions. This is the rubric that produces every June 2026 BuilderProof score.

9 min read116