Bolt vs v0 (2026): Benchmarked Across 6 Axes
Bolt and v0 scored across BuilderProof's six axes in July 2026. v0 finishes 47/60, Bolt 45/60, and they split the axis wins. Here is the per-axis scorecard, with every score cited.
Versioned, reproducible benchmarks of AI app builders - methodology-first and re-tested monthly.
40 posts
Bolt and v0 scored across BuilderProof's six axes in July 2026. v0 finishes 47/60, Bolt 45/60, and they split the axis wins. Here is the per-axis scorecard, with every score cited.
Lovable, Bolt, and Replit scored on BuilderProof's six reproducible axes in 2026. The three pairwise results form a near-cycle, not a clean 1-2-3. See who wins each row.
On BuilderProof's six-axis rubric, Replit edges Lovable 45 to 44 in 2026. Replit wins first-build stability, deployment breadth, and default auth posture; Lovable wins code portability and front-end output quality. A reproducible, documentation-sourced head-to-head.
On BuilderProof's six-axis rubric, Replit and Bolt tie 44 to 44 in 2026. Replit wins first-build stability and deployment breadth; Bolt wins output coherence and default-on credential hygiene. A reproducible, documentation-sourced head-to-head.
A reproducible, documentation-based head-to-head: Lovable vs Bolt scored against BuilderProof's six published axes, with an honest per-axis verdict. July 2026, community-editable.
A proposed community-editable BuilderProof axis scoring how well the app an AI builder generates protects sign-in and per-row data access. Five 20-point sub-axes, a fixed protocol, and a provisional documentation-based cohort table for July 2026.
First-build scores rate one generation. Most real work is the follow-up edit. We propose iteration fidelity: a five-part rubric, a repeatable protocol, and a provisional July 2026 cohort table.
We are proposing a fifth BuilderProof axis to score whether an AI app builder ships code that can leave the platform. Five sub-axes, 0 to 100, provisional cohort scores included.
Effective with the H2 2026 ranking, BuilderProof retires binary 'first-build success' as a scored axis. The signal saturated across the seven-builder cohort. It becomes a precondition (must pass to be ranked) and the rank weight moves to time-to-first-functional-build, measured as p50 and p90 over six trials on the v1 prompt set.
Two Lighthouse runs on the same deployed AI-builder output rarely return the same score, and that is the most-contested observation in the lab notebook for the four June 2026 BuilderProof axes. This note documents the variance phenomenon, the reproducibility protocol the next iteration will adopt, and where median-of-five runs out of road. No score from the published table is changed.
Proposing first-build stability as the fifth BuilderProof axis: the fraction of OQ-7 prompts that complete without manual intervention. Failure-mode taxonomy, measurement protocol, scoring rubric and open questions, dated June 20, 2026.
The BuilderProof methodology v1, dated June 19, 2026, in full: four axes, the OQ-7 test brief, environment standards, scoring weights, reproducibility steps, the operator disclosure, and the v2 open questions. This is the rubric that produces every June 2026 BuilderProof score.