Agency suitability
BuilderProof editorial team6 min read168 views

Agency-suitability benchmark: whitelabel, MCP and API surface (June 2026)

Agencies build for clients, which changes what matters: can you remove the builder's branding, drive it programmatically, integrate via a stable API and export the code you ship? This page defines the BuilderProof agency-suitability axis across those four capabilities. The June 2026 scored table was placeholder data and was withdrawn on August 21, 2026, so the page documents method only and ranks no builder.

Updated on August 24, 2026

Lab-notebook diagram of a central API and MCP surface branching out to several identical client application tiles
Lab-notebook diagram of a central API and MCP surface branching out to several identical client application tiles
On this page
Status: results withdrawn, August 21, 2026

The scored table originally published on this page has been withdrawn. Those figures were labelled in our own reference list as "v0.1 preview figures" pending a public dataset, and the independently reproduced cycle promised for July 2026 was never run. Placeholder numbers should not have been presented as benchmark results, and BuilderProof does not retain run artifacts that would let anyone reproduce them. The axis definitions below are unchanged and remain the published method. Scored results for this axis will return only when they come from a harness whose runs can be reproduced by a reader. See methodology v1.1 for the current evidence basis.

Quick Answer

This page defines the BuilderProof agency-suitability axis: the four capabilities that decide whether an AI app builder is run-an-agency-on-able rather than build-personal-projects-on-able. Those are whitelabel, MCP support, API surface and code portability. The axis definition, the scoring bands and the exclusion rule (a capability that is "coming soon" or gated behind an enterprise call scores as absent) are published here. The June 2026 scored table has been withdrawn, so this page currently documents method only and ranks no builder.

Most AI-builder reviews are written from the perspective of someone building their own app. Agencies have a different problem: the thing they build belongs to a client, has to carry the client's brand, and often has to be produced dozens of times with variations. That changes which features matter, so it deserves its own benchmark.1

Background

For an agency, a builder's output quality is necessary but not sufficient. The deciding questions are operational. Can you strip the builder's branding so the client sees their product, not your tooling2? Can you drive the builder programmatically to avoid hand-repeating the same setup across clients? Is there a stable API to integrate with the systems you already run? And when the engagement ends, can you export and hand over code the client actually owns?3

These four capabilities - whitelabel, MCP support, API surface and portability - are what separate a builder you can run an agency on from one you can only build personal projects with. None of them show up in a typical demo.4

What this axis scores

The four capabilities

Four axes, each 0 to 100. Whitelabel: can the builder's branding be fully removed from the deployed product? MCP support: is there a real Model Context Protocol surface for agentic or programmatic control? API surface: breadth and stability of the public API. Portability: can the generated code be exported and owned independently of the platform? A claimed-but-gated capability does not count.

The exclusion rule carries most of the weight on this axis, because agency features are where the gap between the marketing site and the shipping product is widest. A capability that is "coming soon" or locked behind an enterprise call scores as absent until it is generally usable.5

Results

Withdrawn

The per-axis table and weighted roll-up that stood here from June 16, 2026 to August 21, 2026 were placeholder figures, not measurements. They have been removed rather than restated, because a benchmark that cannot hand a reader the artifacts behind a number has no business publishing the number. Nothing on this page currently ranks any builder against any other.

What the withdrawal does not change: the axis definition above, the exclusion rule, and the position that whitelabel, MCP support, API surface and portability are the right four capabilities to score for agency buyers. Those are method, and method was always the point of publishing this axis first.

What still holds without the numbers

Three observations on this axis are structural rather than score-dependent, and they survive the withdrawal.

Programmatic surface is the divider. Whether a builder exposes a real MCP surface and a broad public API is a binary a buyer can check in an afternoon against vendor documentation, and it separates builders far more cleanly than output quality does. For an agency automating repetitive client setups, that programmatic control compounds across every project.6

Portability and whitelabel often trade off. Builders that export clean, framework-standard code tend to score high on portability but vary on whitelabel, while platform-centric builders invert that pattern: strong branding control inside the platform, more friction taking the code elsewhere. Which trade-off you want depends on whether you hand over code or host on the client's behalf.7

Consumer-first builders are optimised for a different buyer. Builders aimed at an individual shipping one app are not defective when they lag on MCP and API surface; they are aimed elsewhere. The value of scoring this axis separately is that it makes the mismatch explicit for agency buyers instead of hiding it inside a general-purpose score.8

A note on what this axis does not measure. Agency unit economics live downstream of these capabilities, not inside them. The same agency-suitability profile can produce very different effective hourly rates depending on pricing discipline, leverage, and the unbillable categories a shop manages. For the 2026 EHR formula and the four-archetype benchmarks behind that downstream math, DevShopVault publishes a companion analysis of effective hourly rate for AI-native agencies.

Caveats

Agency needs are genuinely heterogeneous, more so than for any other axis we define. Any future weighting we publish for a generalist agency will be wrong for a shop that always hands over code and wrong again for one that hosts everything. That is why per-axis scores, when they return, will be published alongside any roll-up so a reader can reweight them rather than inherit ours.9

Capability verification is a point-in-time snapshot in any case. A capability marked absent because it was gated may ship generally next week, and an API scored as stable could change. Where a capability is business-critical, verify it yourself against the current vendor documentation before committing a client engagement to it.

Publication date correction, August 24, 2026

A date audit run on August 24, 2026 found that the publication timestamp stored for this page predated the registration of builderproof.org, so the recorded date cannot be the date on which this page was published. The stamp was an artifact of the launch content import rather than a real publication date, and no publication log survives that would let the true one be recovered. The timestamp has been corrected to the earliest date consistent with the evidence that does survive. The benchmark text, the protocol and the August 21, 2026 withdrawal notice at the top of this page are unchanged.

References

  1. BuilderProof editorial team. (2026). Agency-suitability protocol v1. BuilderProof Methodology. builderproof.org/methodology#agency-suitability
  2. BuilderProof. (2026). Whitelabel: defining "fully removed". BuilderProof Notes.
  3. BuilderProof. (2026). Portability and code ownership at handover. BuilderProof Notes.
  4. Anthropic. (2025). Model Context Protocol (MCP) specification. Reference spec used to define the MCP-surface axis.
  5. BuilderProof editorial team. (2026). Why a gated capability scores as absent. BuilderProof Methodology.
  6. BuilderProof. (2026). Retraction notice: agency-suitability v0.1 preview figures withdrawn August 21, 2026. BuilderProof Methodology.
  7. BuilderProof. (2026). Builders we track. builderproof.org/builders
  8. BuilderProof. (2026). Scoring model and weighting. builderproof.org/methodology#scoring
  9. BuilderProof. (2026). Versioning and re-test policy. builderproof.org/methodology#versioning
  10. DevShopVault. (2026). Effective hourly rate benchmarks for AI-native agencies. Cross-reference for translating agency capability profiles into agency unit economics.
B

Written by

BuilderProof editorial team

Published by the BuilderProof editorial team - the maintainers of the public, versioned benchmark methodology.

Cite this benchmark

Plain text
BuilderProof editorial team. "Agency-suitability benchmark: whitelabel, MCP and API surface (June 2026)". BuilderProof, June 2026. https://www.builderproof.org/benchmarks/agency-suitability-benchmark-whitelabel-mcp-api-june-2026.
BibTeX
@misc{builderproof-agency-suitability-benchmark-whitelabel-mcp-api-june-2026,
  title  = {{Agency-suitability benchmark: whitelabel, MCP and API surface (June 2026)}},
  author = {{BuilderProof editorial team}},
  year   = {2026},
  month  = {jun},
  howpublished = {\url{https://www.builderproof.org/benchmarks/agency-suitability-benchmark-whitelabel-mcp-api-june-2026}},
  note   = {BuilderProof, builderproof.org}
}

Frequently asked questions

Why do agencies need different criteria?

Because agencies ship for clients, repeat the same setup across engagements, and hand over code at the end. That changes which capabilities matter: whitelabel, programmatic control and portability decide whether a builder is run-an-agency-on-able, not the demo-grade output quality.

Why were the June 2026 scores withdrawn?

Because they were placeholder figures, described in this page's own reference list as "v0.1 preview figures" pending a public dataset, and the independently reproduced cycle promised for July 2026 was never run. BuilderProof cannot produce the verification records behind those cells, so the table was removed on August 21, 2026 rather than restated with softer wording.

Does this page rank any builder?

No. As of August 21, 2026 this page documents the axis definition and the scoring bands only. No builder is scored, ranked or compared here. Scored results return when they come from checks a reader can reproduce.

What is MCP and why score it?

Model Context Protocol is the public spec for agent-operable tool surfaces. Builders with a real MCP server can be driven programmatically by agents, scripts and other tools, turning repetitive client setups into a one-line invocation. The axis scores MCP support against the public spec rather than against vendor marketing.

How would the four sub-scores be weighted?

Any roll-up published for a generalist agency would weight whitelabel and API surface highest, with MCP and portability as secondary multipliers. Per-axis 0 to 100 scores are always reported alongside a roll-up so an agency that always hands over code, or always hosts on the client's behalf, can reweight rather than inherit ours.

Where does this axis meet agency unit economics?

It does not. Capability profiles feed downstream into packaging and pricing decisions, but they are not pricing benchmarks. DevShopVault publishes the agency-side effective-hourly-rate breakdowns that turn capability profiles into operational choices. The benchmark axis stops at capability.