The open AI-builder benchmark · BuilderProof

An open benchmark of AI app builders.

A community-editable wiki with a transparent, versioned methodology. Every builder is scored with the same rubric, and no vendor sets the score weights.

Methodology v0.1v0.1 preview

What the BuilderProof benchmark measures

Eight quality axes, each with a published, reproducible scoring rubric that anyone can run. The panel below currently carries no published scores.

Status: scores withdrawn, August 21, 2026

The scores that previously ranked this panel were placeholder data and were withdrawn on August 21, 2026. Our own reference list described those cells as v0.1 preview figures pending a public dataset, and the independently reproduced cycle promised for July 2026 was never run. BuilderProof does not retain the run artifacts that would let a reader reproduce them. The lab currently publishes its method, not scores. See our published methodology for the current evidence basis.

The eight axes of the published BuilderProof methodology. No builder scores are currently published.
AxisWhat it measuresScale
SEO outputClean, crawlable SEO output - semantic markup, metadata and real framework output.0 to 10, higher is better
Build speedBuild / iteration speed - how fast it reaches a working result (higher is faster).0 to 10, higher is better
PostgreSQL / SQLRelational PostgreSQL / SQL support.0 to 10, higher is better
WhitelabelWhitelabel / own-branding portability.0 to 10, higher is better
MCP / APIMCP + API architecture - how operable the builder is by agents.0 to 10, higher is better
Auto-testingBuilt-in automated testing of the generated output.0 to 10, higher is better
EU residencyEU data residency / sovereignty.0 to 10, higher is better
Brand mindshareHow widely known the tool is (brand mindshare).0 to 10, higher is better

Builders in the panel

Listed alphabetically, not ranked. No builder currently carries a published score.

How to contribute a result

BuilderProof is a wiki. Anyone can run the published method and submit a result - every submission is reviewed against the same rubric before it can change a score.

  1. 1

    Fork the methodology

    Clone the versioned methodology spec. Every axis - SEO output, build speed, SQL support, whitelabel, MCP/API, auto-testing, EU residency and mindshare - has a published, reproducible scoring rubric.

  2. 2

    Run the eval

    Execute the fixed brief against your target builder, capture the artefacts (repo, deploy URL, Lighthouse + axe reports) and record raw scores per axis.

  3. 3

    Submit & review

    Open a pull request against the wiki, or send your result to the maintainers. Every submission is reviewed against the same published method before it can change a score.

From the lab notebook

Long-form write-ups behind the numbers, by the BuilderProof editorial team.