Does anything check what the model said? A model-output validity axis proposal (October 2026)
The call returns 200 and the body may still be a truncated fragment, a safety refusal, or a value invented to fill the schema. Both providers document all three as normal successful responses, and the one field that tells them apart is the field a generated application is least likely to read.
Updated on October 7, 2026
On this page
Quick answer (October 2026): this is a pre-registration, not a result. It proposes that AI app builders be scored on model-output validity posture: whether the application a builder generates checks what the model actually said before using it. The call succeeds, the status code is 200, and the body may still be a truncated fragment, a safety refusal, or a schema-shaped value the model invented because the schema asked for one. All three are documented by the providers themselves as normal successful responses. Seven weighted signals, four postures and a ten-step protocol are proposed below. No builder is scored and no placement is published.
Why we are proposing this axis
Generated applications now call models. A support form that drafts a reply, a document upload that extracts fields into a table, a dashboard that writes a summary, a classifier that routes a ticket to a queue. The builder emits the integration, the integration returns text or an object, and that value flows onward into a database write, a rendered page, or a branch in the code.
Every axis we have published about outbound calls so far has scored the call. Did it have a deadline, did the retry logic distinguish a dropped connection from a rejected request, did the second attempt carry an idempotency key. Those are transport questions, and they are answered by a status code.
This axis begins one step later, at the point where the status code is 200 and the transport question is settled. The providers are unusually explicit that a successful response is not the same thing as a usable answer, and they each publish a field whose entire purpose is to tell you which of the two you received. What nobody publishes is what the generated application is supposed to do with that field, and in our reading of this territory the gap between the two is where the defect lives.
The trap: the well-formed illusion
Every axis we propose names the illusion that hides its defect. This one is the well-formed illusion, and it is the thirty-sixth member of the one-actor family we have been tracking: a defect whose only cheap evidence at test time comes from a single source.
The shape here is specific and, we think, new to the set. In the earlier members the missing evidence was a second context, a second actor, a second tab, or an observation channel that did not exist. Here there is an observation, it is abundant, and it is actively misleading. The output parses. It validates. It type-checks. And none of those three facts tells you whether generation finished, whether the model declined, or whether the value was fabricated to satisfy the shape you asked for.
Consider how a person tests this feature. They write a prompt. They send a representative input. They read the response, judge it good, and ship. Every observation available to them was produced by a system whose job is to produce plausible output, evaluated by the same person who wrote the prompt that produced it. A wrong answer and a right answer arrive in the same envelope, at the same status code, in the same shape.
The reason this survives review is that there is no branch to inspect. A truncated response and a complete response are both strings. A refusal and an answer are both 200. A fabricated field and a measured field are both the correct type. The failure is not a path the code takes incorrectly. It is a path the code cannot see.
And the loop is closed at the reporting end too. A person whose extracted invoice total is wrong by one digit does not file a bug that says the model was truncated. They retype the number. The application never learns.
What the providers actually document
This is the part of the territory that is not in dispute. Both major providers publish the mechanism, name the failure modes, and tell you to handle them.
Anthropic's stop-reason documentation opens by stating that "Every Messages API response includes a
stop_reason field that tells you why Claude stopped generating." It then draws the boundary this axis is built on, verbatim: "Unlike errors, which indicate failures in processing your request, stop_reason tells you why Claude completed its response generation."
The same page sets the two categories out side by side. Stop reasons, under the heading "Stop reasons (successful responses)", are "Part of the response body", "Indicate why generation stopped normally". Errors, under "Errors (failed requests)", are "HTTP status codes 4xx or 5xx" and "Indicate request processing failures".
Two of the documented stop reasons are failures from the application's point of view.
On truncation: "Claude stopped because it reached the max_tokens limit specified in your request." The guidance is explicit about what that leaves you holding: "If Claude's response is cut off because it hit the max_tokens limit, and the truncated response contains an incomplete tool use block, you'll need to retry the request with a higher max_tokens value to get the full tool use." A separate stop reason, model_context_window_exceeded, covers the case where "Claude stopped because it reached the model's context window limit."
On refusal, and this is the single sentence that defines the axis: "Claude declined to generate a response. Safety classifiers return this stop reason as a normal HTTP 200 response, not an error."
The recommended practice is a habit, not a feature: "Make it a habit to check the stop_reason in your response handling logic." And for the truncated case the page asks for something a generated application almost never does, which is to tell the reader: "When a response is truncated because of token limits or the context window, append a notice so the reader knows the output is incomplete."
OpenAI's structured-outputs guide documents the same shape in its own vocabulary. It states plainly that "In some cases, the model might not generate a valid response that matches the provided JSON schema", and names the two causes: "This can happen in the case of a refusal, if the model refuses to answer for safety reasons, or if for example you reach a max tokens limit and the response is incomplete."
Its refusal handling is a dedicated field: "Since a refusal does not necessarily follow the schema you have supplied in response_format, the API response will include a new field called refusal to indicate that the model refused to fulfill the request."
On the weaker of its two modes the guide is equally direct. "JSON mode will not guarantee the output matches any specific schema, only that it is valid and parses without errors." And it states the obligation it places on the caller: "Your application must detect and handle the edge cases that can result in the model output not being a complete JSON object." It also documents a pathological case worth knowing about, where an instruction is missing and "the model may generate an unending stream of whitespace and the request may run continually until it reaches the token limit."
None of this is a discovery and we are not presenting it as one. It is published, current, first-party guidance from the two providers whose APIs the cohort's generated applications most commonly call. The question this axis asks is whether the emitted code does any of it.
The finding: three layers each document their own part, and none documents the composition
The interesting structure in this territory is not a gap in any one document. It is that the property is decided at three layers, each of which describes itself accurately and stops at its own boundary.
At the provider, constrained decoding is documented as incomplete, by the provider. Anthropic's structured-outputs page is candid about what its own remedy does not cover. Under the heading "Invalid outputs" it says: "While structured outputs guarantee schema compliance in most cases, there are scenarios where the output may not match your schema." On a refusal, specifically, "You'll receive a 200 status code", "You'll be billed for the tokens generated", and "The output may not match your schema because the refusal message takes precedence over schema constraints." On truncation, "The output may be incomplete and not match your schema."
It also publishes a limitation that matters more than it looks. The supported JSON Schema subset explicitly excludes "Numerical constraints (such as minimum, maximum, multipleOf)" and "String constraints (minLength, maxLength)". So a value can be schema-valid and still be a negative quantity, a date in the wrong century, or a string of forty thousand characters. Constrained decoding enforces the shape and declines to enforce the range.
And there is one failure on that page with no signal attached to it at all. On enum capitalization: "Structured outputs don't guarantee the capitalization of string enum and const values", and a returned value may differ from the enum you supplied. The consequence is spelled out: "The response completes normally, with no error and no special stop_reason." The remedy is a single line of application code that nobody writes by default: "Compare enum values case-insensitively, and avoid enum values that differ only in capitalization."
At the SDK, the gap is partly filled, and the filling is documented as partial. The AI SDK states that for array bounds "The bounds are sent to providers as part of the structured output schema when supported, and the AI SDK independently validates the final output." That is the SDK compensating for exactly the constraint class the provider says it does not support, which is a well-designed seam and worth crediting.
But the SDK is equally clear about where it stops. On streaming: "Partial outputs streamed via streamText cannot be validated against your provided schema, as incomplete data may not yet conform to the expected structure." Validation of a partial object is not merely omitted, it is impossible, because a half-finished object is not supposed to match. The SDK does raise a typed failure when the whole response cannot be used: "If the model response cannot be parsed or validated against the schema, generateText rejects with an AI_NoObjectGeneratedError", and that error usefully preserves the generated text, the response metadata, the usage, and the cause.
At the application, nothing is documented, because the application is the thing being generated. Whether a refusal becomes a visible message or an empty div, whether a truncated summary is stored as if complete, whether a schema-valid but out-of-range quantity reaches a database write, whether a parse failure is retried once or in a loop: those are decisions, and the provider and the SDK both correctly decline to make them on your behalf.
That is the composition nobody owns. The provider documents what it cannot guarantee. The SDK documents what it can and cannot validate. Neither says what the generated application should do, and the generated application is written by a model that was not asked.
A documented tension: the remedy manufactures confident output
The two halves of this territory are in genuine conflict, and OpenAI states the conflict in its own guide rather than leaving it to be inferred.
The remedy for malformed output is to constrain the output to a schema. The cost of that remedy is named directly: "The model will always try to adhere to the provided schema, which can result in hallucinations if the input is completely unrelated to the schema."
Read that against the failure it is fixing. Unconstrained, a model handed an input it cannot answer may produce prose that says so, and a human reading the prose can tell. Constrained, the same input yields a correctly typed object with every required field populated, because the decoder's job is to produce one. The malformed response was useless and legible. The schema-valid response is usable and wrong, and it is wrong in a way that passes every automated check the application has.
The guide's own suggested resolution is a prompt-level one: "You could include language in your prompt to specify that you want to return empty parameters, or a specific sentence, if the model detects that the input is incompatible with the task." In other words, the schema needs an explicit representation of "no answer", or the absence of an answer will be encoded as a fabricated one. It also declines to overclaim for the remedy generally: "Structured Outputs can still contain mistakes."
Our resolution is scope rather than choice, which is how most of these tensions resolve. Constrained decoding is the right default and signal 2 below rewards it. What the rubric additionally requires is that the schema contain a way for the model to decline, and that the protocol test for it. A schema with no null, no empty, and no "unknown" branch does not have an honest answer available to it.
The default is not a choice
We have made a habit of asking, for each axis, whether the safe behaviour is what you get when nobody decides anything. Here it is not, twice over, and both defaults are documented.
The first is schema enforcement. The AI SDK's plain-text output mode "doesn't enforce any schema on the result: you simply receive the model's text as a string", and the documentation states its status: "This is the default behavior when no output is specified." An integration that nobody configured gets no validation, which is the correct and honest default for a general-purpose library and the wrong posture for a generated application that writes the result to a table.
The second is more consequential because it hides the failure rather than permitting it. On streaming, the SDK notes that "When errors occur during streaming, they become part of the stream rather than thrown exceptions (to prevent stream crashes)", with error visibility obtained by supplying an onError callback. Not crashing the stream is a sound library decision. The effect in emitted code that supplies no callback is that the failure is neither thrown, nor logged, nor rendered. The user sees a response that simply stops.
There is a third default that belongs to the application rather than to any library, and it is the one we expect to dominate the scoring: the default max_tokens is whatever the generated code happened to put there, and the default handling of hitting it is nothing at all.
The rubric
Seven signals, weighted to 100. Scored on the generated application as emitted, with any hand-written configuration removed, which is the control step in the protocol.
Scroll to see more
| Signal | Weight | What a failing case looks like |
|---|---|---|
| The termination signal is read | 22 | The response text is used without the code ever reading stop_reason, finish_reason or a refusal field, so a truncated answer, a declined answer and a complete answer all take the same path |
| Output is constrained at the provider, not prompted | 20 | The only thing making the response parseable is a sentence in the system prompt asking for JSON, with no schema sent and no strict mode set |
| A typed validation boundary before first use | 16 | The parsed value is cast rather than validated, or passed straight into a database write or a rendered component, so the first consumer of unvalidated data is persistent state |
| Semantic and range checks the decoder cannot enforce | 14 | A quantity, price, date or enum is accepted because its type is right, with no bounds check and a case-sensitive enum comparison the provider says it does not guarantee |
| A failed or refused response resolves to a distinguishable state | 12 | One "something went wrong" covers refusal, truncation, parse failure and a genuinely empty result, with nothing recorded that the application could reconcile later |
| Streamed output is not treated as validated | 9 | A partial streamed object is rendered or persisted as though it had been checked, or stream errors are swallowed because no error callback was supplied |
| Retry is bounded and its cost is recorded | 7 | An unbounded repair loop on a parse failure, or a retry whose billed first attempt is never recorded, on a response the provider says you are charged for |
Three weighting decisions worth arguing about now
Why the termination signal is the heaviest. Twenty-two points on reading one field looks disproportionate until you notice that it is the only signal that works when there is no schema. A truncated prose summary is a valid string and will pass every other check in the rubric; the sole cheap evidence that it is incomplete is the field the provider populates for exactly that purpose. It is also the signal whose remedy is smallest, which cuts both ways: a reader could argue that something so cheap should not carry the most weight, and that the points belong with the harder work in signals 3 and 4. We lean the other way because the cost of the remedy is not what the rubric is measuring.
Why constrained decoding is worth 20 rather than more. It is the single most effective intervention available and we considered making it the heaviest. We did not, because the provider's own limitations section says it does not close the case: schema compliance is not guaranteed on refusal or truncation, range constraints are outside the supported subset, and enum casing can vary with no error. A rubric that put forty points here would be scoring a dependency setting and calling it a posture.
A note on ordering. Signals 3 and 4 are deliberately sequenced rather than merged, and OWASP's current cheat sheet states the sequence for the general case: "Use a maintained parser for the expected format, handle parsing failures, and validate the resulting values before business processing or storage." Parse, then validate, then use. A build that validates after it writes has done both things and still failed, which is why signal 3 is worded as a boundary before first use rather than as the presence of a validator somewhere.
Why streaming is only 9. This is the weight we are least sure about. Streaming is now the default presentation for model output in new applications, which argues for more. Against that, the SDK documents that partial validation is impossible rather than merely absent, so a build cannot be penalised for failing to do something no library can do, and the signal reduces to whether the application distinguishes a partial object from a final one and whether it supplies an error callback. The general principle is not in dispute and does not depend on any model: OWASP's cheat sheet asks that you "do not continue with partially validated data", and a partial streamed object is the purest available example of it. If the cohort turns out to stream nearly everything, 9 is too low and we would rather be told that now.
Four postures
Level 1, Unchecked. The response body is used as it arrives. No termination signal is read, no schema is sent, no validation runs. A refusal renders as the refusal text, a truncation renders as a sentence that stops, and neither is distinguishable in the application's own state from a good answer.
Level 2, Parsed. The response is parsed inside a try and catch, so a syntactically broken body is caught. Nothing else is. A response that is complete JSON but semantically truncated, or a refusal that happens to parse, is indistinguishable from success. This is the level we expect to be most common, and it is the level at which the well-formed illusion is strongest, because the code visibly contains error handling.
Level 3, Validated. Output is constrained at the provider and validated against a schema before use, and a validation failure is caught as its own case. The termination signal may still be unread and range checks may still be absent, so truncation and out-of-range values can pass.
Level 4, Adjudicated. The termination signal is read and branched on. Output is constrained and validated. Checks the decoder cannot perform, meaning bounds, formats and case-insensitive enum comparison, are performed by the application. A refused, truncated or unvalidatable response resolves to a state the application can name, show and reconcile, and the schema contains an explicit way for the model to decline.
The measurement protocol
Ten steps, run against the untouched export so that what is scored is what the builder emitted.
-
Fix the brief. Generate an application with one model-backed feature that writes a structured result to persistent state: extract three named fields from a pasted block of text and save them. Describe the feature, not its defences. A prompt that asks for validation would measure instruction-following rather than default posture.
-
Remove hand-written configuration. Record the diff. Everything scored afterwards is emitted code.
-
Locate the call site and read it. Record whether a schema or strict mode is sent, what
max_tokensis set to, whether the result is parsed, validated or cast, and whether any termination field is read anywhere in the path from response to storage. -
The no-information control. Submit an input that contains none of the three fields, for example a paragraph of unrelated prose. Record what reaches the database. This step is the one we care most about, because it is where a schema-valid fabrication appears, and because the provider documentation predicts it.
-
Force truncation. Submit an input that requires a long answer, or reduce
max_tokensif the emitted code exposes it, so that generation stops at the limit. Record what the application stores and what it shows. Record separately whether the stored row is distinguishable from a complete one. -
Force a refusal. Submit an input in a category the provider's safety classifiers decline. Record the status code received, whether the application treats it as an error, and what the user is told. Count the number of distinct user-visible outcomes across steps 4, 5 and 6; if the count is 1, signal 5 scores zero.
-
Break the shape. Where the integration can be pointed at a stub, return a body that is valid JSON of the wrong shape, then a body that is not JSON at all, then a correctly shaped body with an out-of-range number and an enum value differing only in capitalization. Record which of the four are rejected.
-
Exercise the stream. If output is streamed, record whether partial values are written to state or rendered as final, and whether an error injected mid-stream is surfaced at all.
-
Count the retries and the cost. Trigger a validation failure repeatedly and record whether attempts are bounded, and whether anything in the application's own records would let an operator see that the first attempt was billed.
-
Score, and publish the diff from step 2 alongside the score. A build whose posture depends on configuration a human added is a different measurement, and the diff is what lets a reader tell the two apart.
How this relates to axes already published here
Four published axes touch this territory. The boundaries are worth stating in both directions, because a reader who cannot see a boundary is entitled to assume there is not one.
Outbound-call failure: whether it arrived against whether the answer is usable
Our outbound-call failure axis is the nearest neighbour and it draws the boundary itself. It scores seven transport properties: a deadline on the call, a retry decision keyed to the failure class, obedience to the provider's wait instruction, an idempotency key on side-effecting retries, keeping the call off the critical path, a distinguishable failure state, and observability. Its second signal enumerates the classes it scores, and the enumeration is exhaustive and explicit: "no response received, a response saying the request was rejected, and a response saying to come back later."
There is no fourth class in that list for a response that arrived, reported success, and cannot be used. That is not an oversight on its part. It is a scope note, and it matches the providers' own split between stop reasons and errors precisely: that axis scores the errors column and this one scores the stop-reason column.
The two move in opposite directions, both ways. An application can hold a per-call deadline, a class-keyed retry ladder and a stable idempotency key, and still write a truncated summary into a customer record because nothing read stop_reason. An application can read every termination signal and validate every field, and still block a sign-up on an uncapped model call that hangs for ninety seconds.
But the sharper relationship is that the neighbouring remedy causes the event this axis measures, and we think this is the strongest instance of that pattern we have published. Its heaviest signal, at 22 points, requires a deadline on every outbound call. Apply that to a streamed model call and the deadline truncates generation. Worse, Anthropic's streaming documentation records that stop_reason is "null in the initial message_start event", is "Provided in the message_delta event", and is "Not provided in any other events". A client that aborts on its own deadline therefore never receives the field that would have told it the answer was incomplete. The neighbour's top-weighted remedy both produces this axis's defect and removes its evidence, and a build can score full marks there while scoring zero here for that exact reason.
Input validation: a caller you distrust against a producer you pay
Our input-validation and data-integrity axis scopes itself to the trust boundary, and says so: "It measures the server boundary, not the form." Its six signals concern the request body, query and params, which is to say data arriving from a caller the application has no reason to trust.
Model output arrives from the opposite direction. It is a response to a request the application itself composed, sent to a provider it selected, authenticated to and pays. In that axis's model it crosses no trust boundary at all, which is why none of its signals can see it, and that is a correct scoping decision rather than a gap.
Both directions are real. An application can validate every inbound field with a schema library, reject malformed requests with a structured 422, parameterize every query, and then hand a model's reply to a parse call and write the result to a table unchecked. An application can validate model output rigorously and mass-assign an untrusted request body onto a record.
The two axes also share an instrument, which is worth saying because it suggests the remedy is cheaper than it looks. That axis scores, in its own rubric's wording, whether values are checked "beyond type: required fields, enums, formats, bounds, and business ranges". Signal 4 here asks for exactly those checks on the other direction of data flow, and the reason it has to ask separately is the provider limitation quoted above: the decoder enforces the shape and explicitly does not enforce the range.
The guidance both signals rest on is current and general. OWASP's input-validation cheat sheet, in the revision published at the time of writing, puts it in one sentence: "Converting text to an integer only establishes a type: it does not establish an acceptable quantity." That is as true of an integer a model emitted as of an integer a browser posted, and it is the whole of signal 4.
Type safety: what the compiler enforces against what the provider sends
Our type-safety axis scores the posture of the emitted code as a static artefact. A model response is the canonical case where a static type is a claim about data the compiler has never seen and cannot constrain, and a type assertion on a parsed body satisfies the compiler exactly as well as a validator does while checking nothing.
Opposite directions again. A project can run in strict mode with no implicit any and no assertion anywhere else, and still assert the shape of a model response at the one boundary where the data is genuinely unknown. A project can validate model output with a runtime schema and remain loose everywhere else.
Autofill: nobody asked against the application asked and paid
Our autofill-posture axis concerns values that appear in the application's own fields without the application having authored them. This axis concerns values that appear in the application's own data flow without the application having authored them. The shared property is that the application cannot predict the value and did not write it.
The boundary is who asked. Autofill is unsolicited: the browser volunteers a value into a field the application owns, and the application can neither read nor seed the store it came from. Model output is solicited and billed: the application composed the request, chose the provider, set the schema, and is charged for the tokens whether or not the answer is usable. That difference decides the remedy. An unsolicited value has to be tolerated, because refusing it is refusing a browser feature the user wants. A solicited value can be constrained before it is produced, which is what signal 2 is for, and can be rejected after, which is what signals 3 and 4 are for.
Compositions rather than boundaries
Two further relationships are compositions rather than carve-outs, and saying so is more useful than drawing a line that is not there.
Our dependency and supply-chain axis scores a provenance and hallucination guard on packages, which is a model fabricating a name at build time. This axis is a model fabricating a value at run time. They do not overlap and they are the same underlying property of the same class of system observed at two different moments, and a project can plausibly fail both for one reason.
Our lab note on prompt sensitivity argues that a single prompt is not a benchmark because the wording changes the output, which makes non-determinism a validity threat at measurement time. Here the same non-determinism is a correctness problem at run time: it is why a retry does not return the previous answer, which is the whole of signal 7. One property, two consequences, one of which is ours to control and one of which is not.
What this axis is not
It is not a model evaluation. Whether a model is accurate, well-calibrated or good at a task is a question about the model, and there are better venues for it. What is scored here is the emitted application's response to a documented set of provider behaviours.
It is not a judgement about which provider to use. Every normative quotation above is a provider describing its own product, and the two providers we quote document substantially the same failure modes and substantially the same remedies. A build that handles the documented cases is scored the same whichever API it calls.
It is not about prompt quality. A better prompt reduces the rate of some of these failures and eliminates none of them, because refusal, truncation and the context window are properties of the request and the model rather than of the wording.
And it is not a result. No builder is scored on this page, no leaderboard is amended, and no placement is published. The rubric and the weights are a pre-registration put up for argument before any numbers exist, which is the only order in which the argument is honest.
Limitations and open questions
We cannot separate the SDK from the application, and here it matters more than usual. If a builder emits an SDK call that happens to pass a schema, the application scores on signals 2 and 3 without anyone having decided anything. That is a real outcome for a real user and we are inclined to credit it, but it measures a dependency default rather than emitted reasoning, and a reader who thinks those belong in separate columns has an argument we would like to hear.
Step 6 depends on a provider behaviour we do not control. Reliably eliciting a refusal, without building a test corpus we would not want to publish, is awkward. Safety classifiers also change, so an input that produced a refusal in one month may not in the next, which makes the least reproducible step in the protocol the one carrying the clearest evidence.
Signal 4 needs a judgement call we have not removed. Deciding which bounds a generated application ought to check is domain knowledge, and a rubric that rewards any bounds check at all will credit a trivial one. We would rather fix the brief so that the three extracted fields have obvious ranges than leave the grader to decide.
The no-information control may be too harsh. A build that fabricates a value in step 4 is doing exactly what the provider documentation predicts a schema-constrained model will do, and the remedy, meaning an explicit decline branch in the schema, is a design idea rather than settled practice. Scoring a build down for not having anticipated it is defensible and we are not certain it is fair.
We do not know the distribution and we would rather say so in advance. It is entirely possible that the cohort lands at Level 2 almost uniformly, with a parse inside a try and catch and nothing else, in which case the useful separation is between Levels 2 and 3 and the top and bottom of the scale are dead weight. We have not measured the cohort on this axis and this page makes no claim about what it documents or emits.
One signal may be measuring the host rather than the build. If a platform's own invocation timeout is shorter than a long generation, truncation happens for reasons the emitted code did not choose, and signal 1 would be crediting applications that read a field they were forced to encounter.
How to contribute
This is a proposal and the useful replies are disagreements. If you think the weights are wrong, say which pair you would swap and why. If you run the protocol against an export and get a different distribution than we expect, that is the most valuable thing you can send, and a counter-rubric is more useful than a correction. If a provider changes one of the documented behaviours quoted above, we would rather hear it than discover it, and we will date and mark the change on this page rather than silently editing it.
References
- Anthropic, "Stop reasons and fallback",
stop_reasonvalues includingmax_tokens,refusalandmodel_context_window_exceeded, the stop-reasons versus errors split, and the streaming note on which event carries the field, https://docs.claude.com/en/api/handling-stop-reasons (accessed 7 October 2026) - Anthropic, "Structured outputs", the "Invalid outputs" section on refusals, token limits and enum capitalization, and the unsupported-feature list excluding numerical and string constraints, https://docs.claude.com/en/docs/build-with-claude/structured-outputs (accessed 7 October 2026)
- OpenAI, "Structured model outputs", the refusal field, the schema-adherence and hallucination note, and the JSON-mode edge-case guidance, https://platform.openai.com/docs/guides/structured-outputs (accessed 7 October 2026)
- Vercel, AI SDK, "Generating Structured Data", the default text output mode, the partial-output validation limitation, in-stream error delivery, independent validation of array bounds, and
AI_NoObjectGeneratedError, https://ai-sdk.dev/docs/ai-sdk-core/generating-structured-data (accessed 7 October 2026) - OWASP, Input Validation Cheat Sheet, the "Parse Safely, Then Validate" and "Validate Structured Data" sections, including the type-is-not-a-quantity sentence used by signal 4 and the instruction not to continue with partially validated data used by signal 6, https://cheatsheetseries.owasp.org/cheatsheets/Input_Validation_Cheat_Sheet.html (accessed 7 October 2026)
Written by
BuilderProof editorial teamCite this benchmark
BuilderProof editorial team. "Does anything check what the model said? A model-output validity axis proposal (October 2026)". BuilderProof, October 2026. https://www.builderproof.org/benchmarks/does-anything-check-what-the-model-said-model-output-validity-axis-october-2026.
@misc{builderproof-does-anything-check-what-the-model-said-model-output-validity-axis-october-2026,
title = {{Does anything check what the model said? A model-output validity axis proposal (October 2026)}},
author = {{BuilderProof editorial team}},
year = {2026},
month = {oct},
howpublished = {\url{https://www.builderproof.org/benchmarks/does-anything-check-what-the-model-said-model-output-validity-axis-october-2026}},
note = {BuilderProof, builderproof.org}
}Frequently asked questions
What is the model-output validity axis?
It is a proposed BuilderProof benchmark axis that scores whether the application an AI app builder generates checks what a model actually returned before using it. It is a pre-registration, not a result, and no builder is scored on this page.
Is a 200 status code not enough to know the call worked?
No, and the providers say so. Anthropic documents that a safety refusal is returned as a normal HTTP 200 response rather than an error, and that a response stopped by the token limit is also a successful response. OpenAI documents that a model may not produce a response matching the supplied schema on either a refusal or a truncation. The status code answers whether the request was processed, not whether the answer is usable.
Does using structured outputs or JSON mode solve this?
It helps and it does not close the case, on the providers' own account. Anthropic lists scenarios where output may still not match the schema, and excludes numerical and string constraints from the supported schema subset, so a value can be the right type and the wrong quantity. OpenAI notes that JSON mode guarantees valid JSON rather than schema conformance, and that a model constrained to a schema may fabricate values when the input does not support one.
Why would the person who built the app not notice?
Because a wrong answer and a right answer arrive in the same envelope. A truncated response is a valid string, a refusal is a 200, and a fabricated field is the correct type. Everything observable at test time is produced by the system under test and judged by the person who wrote the prompt that produced it. We call that the well-formed illusion.
Related benchmarks
Did It Actually Go Through? A Proposed Axis for Outbound-Call Failure (September 2026)
A generated app calls a third party's API from inside a request. We propose an axis for what it does when no answer comes back, and find that retrying is a layered property no single document describes.
Input-Validation and Data-Integrity Posture: A Proposed Axis for Whether AI App Builders Validate Untrusted Input at the Boundary (August 2026)
A candidate BuilderProof benchmark axis that scores whether the code AI app builders emit validates untrusted input at the server boundary, or trusts whatever the client sends. Rubric, four posture levels, an adversarial-payload reproduction protocol, and an open call for comment.
Type-Safety Posture: A Proposed Axis for Scoring the Code AI App Builders Emit (August 2026)
A candidate BuilderProof benchmark axis that scores the static type safety of the code AI app builders actually emit, measured from the untouched export with no runtime. Rubric, reproduction protocol, and an open call for comment.