Does the browser fill it in? An autofill posture axis proposal (October 2026)
The person testing a freshly generated app has an empty credential store on a brand new origin, so nothing is offered and every form looks the same. This proposal scores whether field purposes are declared, whether a credential pair is distinguishable, whether a save is ever offered, and whether identifiers survive a redeploy.
Updated on October 3, 2026
On this page
Quick answer (October 2026): this is a pre-registration, not a result. It proposes that AI app builders be scored on autofill posture: whether the forms a generated application emits can be filled, and saved into, by the credential and address store the person's own browser already carries. The failure is invisible to whoever builds and tests the app, because a fresh browser profile on a brand new origin has an empty store, so nothing is offered and every form behaves identically whether it declares a field purpose or not. Seven weighted signals, four postures and a ten-step protocol are proposed below. No builder is scored and no placement is published.
Why we are proposing this axis
A generated business application is mostly forms. Sign up, sign in, profile, billing address, shipping address, contact, checkout. Every one of those collects values a person has typed into some other website before, and every modern browser keeps a store of exactly those values so the person does not have to type them again.
Whether that store can reach your fields is not a matter of luck. It is decided by a small number of attributes in the markup the builder emits, and by the structure it wraps those fields in. The WHATWG HTML Living Standard puts the mechanism plainly in section 4.10.19.7, read 3 October 2026.
The specification, verbatim: "User agents sometimes have features for helping users fill forms in, for example prefilling the user's address based on earlier user input. The autocomplete content attribute can be used to hint to the user agent how to, or indeed whether to, provide such a feature."
A hint is a thing you can omit. Omitting it is the common case in generated output, and omitting it is not neutral.
The trap: the empty-profile illusion
Every axis we propose names the illusion that hides its defect. This one is the empty-profile illusion, and it is the thirty-fourth member of the one-actor family we have been tracking: a defect whose only cheap evidence at test time is generated by a single source.
The person testing a freshly generated application opens it in a browser, on an origin that did not exist an hour ago, and fills the sign-up form by hand because there is nothing else to do. They sign in. They type an address into the profile form. Everything works.
Nothing was offered to them, and nothing could have been. The credential store had no entry for this origin. The address store may have had an address, but if no field declared a purpose the browser had only its own guesswork to go on, and guesswork that quietly declines to fire looks exactly like a form with nothing to fill.
This family of defect has a particular shape here, and it is worth naming precisely because it is new to the set. In the previous members, the input needed to exercise the failure was either authored by the tester, manufactured by the system under test, or a string the application had deliberately stopped emitting. Here the input lives in a store that the application can neither read nor write, did not create, and cannot seed. It is filled by the user agent, over months, from the person's history on other sites. The only way to have the input is to have been a different person earlier.
There is no way to fake it from inside the work, and there is almost no way to observe it either. We come back to the observation problem in the limitations.
What the specification actually says, including the part that undercuts the obvious remedy
The obvious remedy, once somebody notices, is to sprinkle autocomplete tokens across the form. That is most of the work, and the specification is explicit about why the omission matters.
On the default, verbatim: "If the autocomplete attribute is omitted, the default value corresponding to the state of the element's form owner's autocomplete attribute is used instead (either "on" or "off"). If there is no form owner, then the value "on" is used."
So an undeclared field is not inert. It is in the on state, and on has a defined meaning.
On what on means, verbatim: "When the autofill field name is "on", the user agent should attempt to use heuristics to determine the most appropriate values to offer the user, e.g. based on the element's name value, the position of the element in its tree, what other fields exist in the form, and so forth."
This is the cleanest instance we have found of a pattern we keep meeting: the default is not a choice, and it is a choice that mostly works. The heuristics are good. For a field named email next to a field named password, a browser will usually get it right. That success is exactly what stops anyone discovering that a declaration mechanism exists, until the form contains a field the heuristics cannot read: a second address line, a country that wants a code rather than a name, a one time code, or a password field that the browser cannot tell is a new password rather than the existing one.
The second thing the specification does is constrain the user agent, and this is the part that undercuts the simple reading.
On what a filling user agent may not do, verbatim: "A user agent prefilling a form control's value must not cause that control to suffer from a type mismatch, suffer from being too long, suffer from being too short, suffer from an underflow, suffer from an overflow, or suffer from a step mismatch."
And immediately after, verbatim: "Where it's not possible for the canonical format to be used, user agents should use heuristics to attempt to convert values so that they can be used."
Read those two together. A conforming browser will not hand your field an over long value. What it will do instead is convert the value so that it fits. The specification gives its own worked example: a field with a length limit of one, declared as an additional name, is filled with the initial rather than the name. That is correct behaviour, and it is also a silent shortening performed by a party the application never spoke to.
On the floor under all of it, verbatim: "User agents must only prefill controls using values that the user could have entered."
And on timing, verbatim: "The autocompletion mechanism must be implemented by the user agent acting as if the user had modified the control's data, and must be done at a time where the element is mutable"
That last clause ties filling to the mutability concept the same specification defines in section 4.10.18.2, which says of a form control designated mutable that it determines "whether or not the user can modify the value or checkedness of a form control, or whether or not a control can be automatically prefilled". A field that is disabled or read only at the moment the document settles is not a field anything will fill.
There is one more requirement that matters for forms collecting two addresses at once.
On coherence, verbatim: "When the user agent autofills form controls, elements with the same form owner and the same autofill scope must use data relating to the same person, address, payment instrument, and contact details."
The specification provides section- prefixes and the shipping and billing tokens so that a single form can describe two subjects. A checkout page that collects a billing address and a shipping address with no scope tokens has told the browser that both groups concern the same subject, and a conforming browser will fill them accordingly.
Where the accessibility standard declines the question
The field purpose token is not only an autofill mechanism. It is also how WCAG 2.2 Success Criterion 1.3.5, Identify Input Purpose, is met in HTML. That overlap is real and we want to be exact about where it stops, because a reader could reasonably assume a build that conforms to WCAG at level AA has this axis covered.
It does not, and the criterion says so itself.
From the Understanding document for SC 1.3.5, read 3 October 2026 in the unpublished editor's draft at w3c.github.io, which carries its own notice that it might include content that is not finalized, verbatim: "Whether or not user agents actually autofill inputs is not relevant when evaluating this criterion. What matters is whether or not the inputs programmatically expose their purpose."
That is a stated decline, and it is the strongest kind of boundary evidence we look for. The accessibility criterion is satisfied by the declaration alone. Whether the fill lands, lands in the right field, survives validation, or is ever offered at all is outside what it evaluates.
The same document declines in the other direction too, verbatim: "user agents may use heuristics, rather than programmatically determined purposes, to autofill inputs", and it records that this is not sufficient to meet the criterion. So a form that fills perfectly by guesswork fails SC 1.3.5, and a form that declares every purpose and then suppresses filling entirely passes it. The criterion and this axis can move in opposite directions in both directions.
The scope note matters even more for a generated business application, verbatim: "This success criterion is specifically scoped to inputs collecting information about the user." And, in the same note, verbatim: "An input field for information that is not about the user does not need to programmatically expose its purpose, even if that purpose is included in the Input Purposes list."
Most fields in a generated CRUD application are not about the user. They are about a customer record, a supplier, an invoice line, a property, a patient. Those fields are entirely outside the criterion, and they are also the fields where a declared purpose would help a person who has to enter forty of them. A build can be at full conformance and have declared purposes on five fields out of ninety.
Finally, the criterion anticipates the suppression case, verbatim: "In some instances, authors may want to prevent user agents from autofilling form fields. However, in order to meet the requirements of this criterion, input purposes still need to be programmatically determinable." Declaring and suppressing are two separate decisions, and the standard expects both to be made deliberately.
What the browser side says about saving, not only filling
Filling is half of it. A credential store has nothing to offer on the second visit unless something was saved on the first, and saving is where generated applications have a structural problem.
MDN's reference for the autocomplete attribute, read 3 October 2026, on where the values come from, verbatim: "The source of the suggested values is generally up to the browser; typically values come from past values entered by the user, but they may also come from pre-configured values."
MDN also records that the suppression keyword does not mean what an author might assume it means.
Verbatim: "In most modern browsers, setting autocomplete to "off" will not prevent a password manager from asking the user if they would like to save username and password information, or from automatically filling in those values in a site's login form."
And in its guidance section, verbatim: "Browsers may also ignore autocomplete="off" on login fields to support password managers."
That is a documented tension rather than a contradiction. The specification defines off as a legitimate author instruction, listing among its meanings that the value "is particularly sensitive", or "that it is a value that will never be reused", or "that the document provides its own autocomplete mechanism and does not want the user agent to provide autocompletion values". Browsers honour that for most fields and deliberately decline to honour it for credential fields, because the population harmed by a site suppressing a password manager is larger than the population protected.
MDN is also explicit about the workaround authors reach for next, verbatim: "Using invalid or non-standard values (such as made-up strings to circumvent autofill) has a similar effect: the browser cannot match the field to any known purpose, so it cannot offer relevant suggestions." An invented token is not a stronger off. It is an undeclared field with extra steps, and it fails the accessibility criterion as well.
On the structural requirements for saving, Google's sign in form guidance on web.dev is the clearest published statement we could find. We note its own stated date: the page records that it was last updated 2020-06-29, and we read it 3 October 2026. It is cited here for the structural requirements it states, which the WHATWG form owner rules and MDN's own notes corroborate, and not for anything time sensitive.
From the web.dev checklist, verbatim: "Give input name and id attributes stable values that don't change between page loads or website deployments."
That single sentence is the sharpest generated-application finding in this proposal. A stored credential entry is matched against the field it was saved from. Builders that emit hashed, index suffixed or framework generated identifiers produce a form whose fields are not the same fields after the next deploy, and a user's saved entry silently stops being offered. Nothing errors. Nothing is logged. The person simply types their password again and assumes they imagined it working last time.
On submission, verbatim: "Help password managers understand that a form has been submitted." The page gives navigating to a different page and emulating navigation with the History API as the two ways to do it, and adds, verbatim: "With an XMLHttpRequest or fetch request, make sure that sign-in success is reported in the response and handled by taking the form out of the DOM as well as indicating success to the user."
A generated single page application that posts credentials with fetch, keeps the form mounted, and swaps the view with client side state has done none of those three things. It is the default shape of most generated sign in flows we have read, and it is the shape in which a credential is never offered for saving.
MDN's note on what browsers may require of a field before they will store it for later autofill names a stable name or id, descent from a form element, and ownership by a form with a submit button. A sign in built from loose inputs and a click handler satisfies none of them.
The rubric
Seven signals, weighted to 100. This is a proposal. The weights are the part most worth arguing about before anything is scored.
Scroll to see more
| Signal | Weight | What a failing case looks like |
|---|---|---|
| Every control that collects information about a person carries a valid field purpose token from the published list | 22 | No control declares a purpose, so the form is in the on state and whatever happens is the browser reading field names and layout |
| The credential pair is distinguishable: a username-class control is present and declared, and the password control declares whether it is the existing one or a new one | 20 | Sign in and sign up both emit a bare password control, so no store can tell which to offer, and a password change form offers the old password as the new one |
| The controls have a real form owner and the submission is legible as a submission | 16 | Inputs sit outside any form element, or the form has no submit control, or sign in posts with fetch and leaves the form mounted, so nothing is ever offered for saving |
| Field identifiers are stable across deploys | 14 | name and id are generated per build, so a saved entry stops matching after a redeploy that changed nothing a user can see |
| A filled value survives the application's own shape and validation rules unchanged | 12 | The server wants a two letter country code and the store supplies a country name, or a tight length limit makes a conforming browser shorten the value before the field ever sees it |
| Where filling is suppressed, the suppression is a recorded decision and the purpose is still declared | 9 | off is sprayed across the whole form, or invented tokens are used to defeat filling, with no reason written down and no purpose declared |
| Scope and grouping are declared wherever one form collects two subjects | 7 | Billing and shipping groups in one form with no scope tokens, so the coherence rule fills both from a single address |
Three weighting decisions worth arguing about now
Why the declaration signal carries the most. Every other failure is downstream of a browser not knowing what a field is for. An application that has declared postal-code can get the surrounding structure wrong and then fix it. An application that has declared nothing has handed the entire question to heuristics it does not control and cannot test, and there is nothing to argue with.
Why identifier stability carries 14 rather than something smaller. We expect this to be contested, and we are open to the argument that it belongs lower because it is a build system property rather than a markup property. Our reason for putting it this high is that it is the only signal on the list whose failure is caused by the thing AI builders do most, which is regenerate. A form that is correct on Monday and has new identifiers on Tuesday is a form that was never really correct, and the person who loses their saved entry has no way to attribute it. The shape is one we have measured one layer up, in our permalink durability axis: a string the application handed to somebody else, which it then stops honouring after a change it considers purely internal, and reports nowhere. There the string is an address sitting in a bookmark. Here it is a field identifier sitting in a credential store, and the person on the other end of both has no channel back to tell you.
Why suppression scores at all rather than being a penalty. Suppressing a fill is frequently right. A one time code should not be offered from a store, and the specification says so in its own definition of the keyword. The signal is not whether an application suppresses but whether the suppression was decided rather than inherited, and whether the purpose is still declared underneath it so the accessibility criterion is still met.
Four postures
To make a result legible at a glance, a score would map to one of four levels.
- Level 0, heuristic. No control declares a purpose. Whatever filling happens is the browser guessing from names, types and position, and the application has no idea whether it happens at all.
- Level 1, declared. Controls that collect information about a person carry valid tokens, but the credential pair is not distinguished, or the controls have no real form owner, or the submission path is invisible to a credential store. Filling may work; saving does not.
- Level 2, fillable and savable. Tokens are correct, the credential pair is distinguished, the controls sit in a form with a submit control, and a completed sign up produces an offer to save.
- Level 3, stable under regeneration. Everything above, plus identifiers that survive a redeploy, filled values that pass the application's own validation without alteration, and any suppression recorded as a decision with the purpose still declared.
The gap worth watching is between Level 1 and Level 2. It is where almost all generated output we have read would sit, and it is the gap that cannot be closed by adding attributes alone, because it is a question about structure and about how the application submits.
The measurement protocol
Ten steps. Steps 4 and 9 are the ones that make the rest mean anything.
- Generate an application from the standard brief, which must include a sign up form, a sign in form, a password change form, and a profile or checkout form carrying a postal address and a telephone number.
- Record the emitted markup for every form control: element, type,
name,id,autocompletevalue, computed form owner, and whether that form owner contains a submit control. - Classify every control as collecting information about a person or not, and compute the share of the first group carrying a valid token from the published list. Record invalid and invented tokens separately from absent ones.
- No information control. Open the application in a browser profile whose credential store and address store are both empty, and record what is offered on every form. Nothing should be offered. If anything is offered, the profile is not empty, the instrument is contaminated, and every later reading in this run is void rather than merely noisy.
- Seed the profile the way a person would: complete the sign up form by hand and accept the browser's offer to save, then save one postal address through the browser's own settings interface.
- Open the sign in form in a fresh session and record whether a credential is offered, which control receives the username, which receives the password, and whether both land.
- Open the profile or checkout form and record, per control, whether it filled, with what, and whether the value is one the application then accepts.
- Submit both forms unchanged and record every validation message the application raises against a value it did not receive from a human.
- Deploy separation. Redeploy the application with no source change at all, reload in the same profile, and repeat step 6. A saved entry that is no longer offered after a no change redeploy is the identifier stability failure, and it is detectable no other way.
- Record every suppression in the emitted source: each
off, each invented token, each disabled or read only control that carries a purpose, together with whatever reason the generated code gives, if any.
Steps 4 and 9 are the expensive ones. Step 4 requires a genuinely clean profile per run rather than a reused one, and step 9 requires a second deploy. A single deploy run can report everything except the identifier stability signal, and should say so rather than scoring that signal zero.
How this relates to axes already published here
Four axes already published here touch this territory. None of them covers it, and in two cases the relationship is sharper than a boundary.
Accessibility posture: a documented decline, in both directions
Our accessibility posture axis scores what a builder documents about the accessibility of its output, and its fifth sub criterion covers accessible forms, labelling and validation. It states its own scope plainly: it measures documentation, not behaviour.
The boundary runs both ways and it is unusually clean, because the accessibility standard itself draws it. SC 1.3.5 is met by the declaration and explicitly does not evaluate whether filling happens. So a form with perfect visible labels, correct programmatic label association and full conformance can carry zero field purpose tokens on the ninety fields that are not about the user, and a form with every token correct can be unlabelled and unusable with a screen reader. Neither axis needs amending. They need scoping, and the scope note is already written in the standard.
There is a second-order point worth recording. The accessibility axis's highest scoring remedy is standardising on an accessible component library, and a component library cannot know what a field is for. Our own write up quotes Radix stating that labelling is up to you. The more thoroughly a build abstracts its input element behind a generic wrapper, the further the purpose token sits from the place that knows the answer.
Length and truncation: the neighbour's remedy is this axis's triggering event
Our length and truncation axis rewards declaring a user perceived limit on a control, and separately quotes the rule that constraint validation applies only to a value changed by the user.
Put that beside the specification text above and the two meet at a point. A conforming user agent must not prefill a value that would make the control too long, and where the canonical format will not fit it should convert the value instead. So declaring the limit is what causes the conversion. A build that scores well on the length axis by putting a tight, well chosen limit on an initial field is the build in which a filled value is silently shortened by a party the application never spoke to, and the length axis's own instrument cannot see it, because its enforcement model is scoped to typing.
This is the strongest form of carve out we look for, where the neighbouring remedy does not merely fail to observe the event but causes it. Stated the other way: a field purpose token does nothing at all about byte budget, unit confusion or round tripping, which is the whole of the length axis. Neither needs amending.
Input validation: a genuine tension, resolved by scope
Our input validation axis scores server side boundary validation and treats rejecting malformed input with a 400 or 422 as a pass, and coercing or persisting it as a failure. That is right, and we are not softening it.
The tension is that the shapes a credential and address store supplies are not shapes the application chose. A store may hold a country as a name where the schema wants a two letter code, a telephone number with a country prefix where the schema wants digits, or a single street address where the schema wants three lines. A build at the top posture of the validation axis will reject all three, correctly by its own rubric, and the person who used autofill is refused while the person who typed by hand succeeds.
The resolution is scope rather than choice. Validate the value; do not require a shape that no store can produce. The validation axis asks whether the boundary is defended. This axis asks whether the defended boundary is reachable by the mechanism the user agent actually uses, and signal 5 exists precisely to measure that intersection rather than to argue with it.
Output encoding: a third provenance class, which is a composition and not a boundary
Our output encoding axis is about the sink, and it separates strings the application generated from strings a person typed.
An autofilled value is neither. It is a string the user agent supplied from a store that neither party wrote, inserted by a mechanism the specification describes as acting as if the user had modified the control's data. For the output encoding axis that is a distinction without a difference, and correctly so: the sink does not care where a string came from, and treating a filled value as trusted because the browser supplied it would be exactly the mistake that axis exists to catch.
So this is a composition rather than a boundary, and saying so is more useful than drawing a line that is not there. The two axes share a vocabulary of provenance and disagree about whether provenance matters, for good reasons on both sides.
Limitations and open questions
Every axis we propose names its own blind spots, because a scorecard that hides them is worse than no scorecard.
-
It measures one browser's store. The specification leaves the source of suggestions to the user agent, and MDN says the same in so many words. A result is a reading of one browser on one platform, and a cross browser matrix would multiply the protocol cost by the number of engines.
-
The observation channel is close to absent, and this is the deepest problem. There is no event for a fill. The only in page signal is a CSS pseudo class, and MDN records of it that "This feature is not Baseline because it does not work in some of the most widely-used browsers", that a vendor prefixed alias exists, and that "If the user edits a control, that control will no longer match :autofill, even if the value is the same as the autofilled value". A value that was filled and then touched is indistinguishable from a value that was typed. We can see the markup and we can watch a browser; the application itself cannot, which is why this is an external benchmark question rather than something a builder could instrument away.
-
It cannot see a third party password manager. A large share of real users have an extension doing this work, with its own heuristics and its own view of what counts as a login form. The protocol reads the browser's own store, and a build that works with one may not work with the other.
-
It rewards declaration, and a declared token is not a guarantee. The specification says a user agent should provide matching suggestions, not that it must offer anything at all. Steps 6 and 7 of the protocol exist to catch the gap between declaring and filling, but a negative there is not always the application's fault.
-
The fields outside the accessibility scope are the ones most likely to be missed, and the rubric may still under-weight them. Signal 1 counts controls collecting information about a person, which follows the standard's own framing. A reviewer working from a conformance checklist will see five such fields in a generated business app and ninety that are out of scope. Whether a benchmark should push declaration onto fields the standard excludes is a live question we would rather argue in public than settle quietly.
-
A build with nothing to declare should not be punished for it. An application with no authentication cannot distinguish a credential pair, and an application that collects no address cannot declare a scope. Signals 2, 3 and 7 are scored as a full pass in those cases rather than left blank, on the same principle we applied in our version skew proposal, where a build that ships "no lazily loaded routes, no client-side navigation and no client-side data fetching has essentially nothing to skew" and signals 1 to 3 are scored as a full pass rather than left blank.
-
Step 9 needs two deploys and step 4 needs a genuinely clean profile. Both are the kind of step that gets quietly skipped under time pressure, and a run that skips either should report the signal as not measured rather than scoring it. We would rather publish a result with one signal marked unmeasured than a result that silently assumes a passing value.
An open question we cannot resolve from documentation. The webauthn token exists in the published list, and the specification shows it combined with a password token on a single control so that a user agent may offer either a saved password or a public key credential through conditional mediation. We do not yet know whether any AI app builder emits it, and we do not know whether a passkey offer and a password offer on the same control would make this axis easier or harder to score. We have not written a signal for it, and we would rather say that than invent one.
References
Every claim above is from a document we fetched and read on 3 October 2026. Where a page states its own date, we have given it.
- WHATWG HTML Living Standard, section 4.10.19.7 Autofill, and section 4.10.18.2 Mutability: https://html.spec.whatwg.org/multipage/form-control-infrastructure.html
- MDN, HTML attribute: autocomplete: https://developer.mozilla.org/en-US/docs/Web/HTML/Reference/Attributes/autocomplete
- MDN, CSS pseudo class :autofill: https://developer.mozilla.org/en-US/docs/Web/CSS/:autofill
- W3C, Understanding Success Criterion 1.3.5 Identify Input Purpose, read in the unpublished editor's draft, which carries its own notice that it might include content that is not finalized: https://w3c.github.io/wcag/understanding/identify-input-purpose.html
- Google, web.dev, Sign in form best practices, which records its own last updated date of 2020-06-29: https://web.dev/articles/sign-in-form-best-practices
This is a pre registration. No builder is scored here, no placement is published, and the weights above are a starting point for argument rather than a settled formula. If you think the identifier stability signal is in the wrong place, that is the argument we most want to have.
Written by
BuilderProof editorial teamCite this benchmark
BuilderProof editorial team. "Does the browser fill it in? An autofill posture axis proposal (October 2026)". BuilderProof, October 2026. https://www.builderproof.org/benchmarks/does-the-browser-fill-it-in-autofill-posture-axis-october-2026.
@misc{builderproof-does-the-browser-fill-it-in-autofill-posture-axis-october-2026,
title = {{Does the browser fill it in? An autofill posture axis proposal (October 2026)}},
author = {{BuilderProof editorial team}},
year = {2026},
month = {oct},
howpublished = {\url{https://www.builderproof.org/benchmarks/does-the-browser-fill-it-in-autofill-posture-axis-october-2026}},
note = {BuilderProof, builderproof.org}
}Frequently asked questions
What is the autofill posture axis?
It is a proposed BuilderProof benchmark axis that scores whether the forms an AI app builder generates can be filled, and saved into, by the credential and address store the person's own browser already carries. It is a pre-registration, not a result, and no builder is scored on this page.
Does conforming to WCAG cover this?
No, and the accessibility standard says so itself. The Understanding document for Success Criterion 1.3.5 states that whether or not user agents actually autofill inputs is not relevant when evaluating that criterion, and that what matters is whether the inputs programmatically expose their purpose. It is also scoped to fields collecting information about the user, which in a generated business application is a small minority of the fields on screen.
Why does setting autocomplete to off not stop a password manager?
Because browsers deliberately decline to honour it for credential fields. MDN records that in most modern browsers, setting autocomplete to off will not prevent a password manager from asking whether to save username and password information, or from automatically filling those values in a login form, and that browsers may ignore the keyword on login fields to support password managers. Using an invented token instead does not help either, because the browser then cannot match the field to any known purpose.
Why does identifier stability matter for a generated app?
A saved credential entry is matched against the field it was saved from, and Google's sign in guidance asks authors to give input name and id attributes stable values that do not change between page loads or website deployments. A builder that emits hashed or per-build identifiers produces a form whose fields are not the same fields after the next deploy, so a stored entry silently stops being offered. Nothing errors and nothing is logged, which is why the protocol below includes a no-change redeploy as a separate step.
Related benchmarks
Accessibility (a11y) Posture: a proposed benchmark axis for AI app builders (August 2026)
A neutral, documentation-based benchmark axis for whether AI app builders commit to accessible output, scored across v0, Lovable, Replit, Base44, and Bolt.new (August 2026).
Does the Whole Name Get Saved? A Proposed Axis for Length and Truncation Correctness (September 2026)
Every length limit is checked with a value the developer typed themselves, and on that value all four units agree. A pre-registered seven-signal axis for whether the value submitted is the value stored.
Input-Validation and Data-Integrity Posture: A Proposed Axis for Whether AI App Builders Validate Untrusted Input at the Boundary (August 2026)
A candidate BuilderProof benchmark axis that scores whether the code AI app builders emit validates untrusted input at the server boundary, or trusts whatever the client sends. Rubric, four posture levels, an adversarial-payload reproduction protocol, and an open call for comment.