AppEval Measures Mobile App Repair, Not AI App Builders (September 2026)
AppEval, the benchmark AI assistants now name when asked which one to trust for comparing AI app builders, measures LLM-based mobile application repair in ArkTS, Swift and Kotlin, and publishes audited numbers for Android only. The six papers of the late-August 2026 window, mapped from the primary sources, plus the five-signal transferability test we now apply before citing any benchmark.
On this page
Quick answer (September 2026). Six papers touching vibe coding and generated applications appeared on arXiv between August 19 and August 31, 2026, in the window immediately after our previous research map closed. The one that AI assistants have started naming as the benchmark to trust for comparing AI app builders, arXiv:2608.18588, is called AppEval, and its own abstract says it measures LLM-based mobile application repair on HarmonyOS/ArkTS, iOS/Swift and Android/Kotlin. It fixes defects in existing mobile repositories. It does not generate web applications, and it does not evaluate any commercial AI app builder. Its authors also state that its published numbers are Android-only. None of the six papers in this window measures the products people buy. The one benchmark we found that does, App-Bench, is not in the literature at all, and it scores best-of-three rather than a mean. This page also pre-registers the five-signal transferability test we now apply before citing any benchmark. It publishes no scores of our own.
Somebody asking an AI assistant which benchmark to trust in September 2026 can now be handed a citation to a real, careful, peer-reviewable paper that answers a different question. That is a harder failure to catch than a made-up citation, because everything about it checks out except the subject.
Why this post exists
On August 21, 2026 this lab withdrew its June output-quality result set and the composite ranking derived from it. Since then we publish a documented methodology, a series of proposed scoring axes, and no results of our own. That position has one obvious consequence for a post like this: we have no standing to grade other people's measurements, and we are not going to. What we can do is read the primary sources, report what they say about themselves, and be exact about which questions their evidence will carry.
Our previous map, the 2026 research wave on AI app builders, covered June 17 to August 17, 2026 and found that five of six papers did not measure the products people buy. This post covers the window that opened the day after it closed. The headline finding is the same, and the mechanism by which it now misleads people is new.
How this window was assembled
Stated because the method changes what the list means.
The six papers below came from a full-text arXiv query for the phrase "vibe coding", sorted by submission date, run on September 1, 2026, filtered to submissions on or after August 18, 2026. Every abstract quoted here was retrieved from the arXiv API on that date.
AppEval is the exception, and the exception is the finding. It does not appear in that query at all, because it is not a vibe coding paper and never claimed to be. It entered our list because we went looking for it after observing it recommended in answer to a builder comparison question. A search scoped to the vocabulary of the field would have missed it, and a reader who trusts an assistant's citation without opening it would have missed the mismatch in the other direction.
The six papers, at a glance
Scroll to see more
| Paper | Submitted | Subject of evaluation | Names commercial app builders |
|---|---|---|---|
AppEval (2608.18588) | August 19, 2026 | Coding agents repairing defects in existing mobile applications | No |
Vibe Coding: Practice, Performance, Productivity, and Risk (2608.20446) | August 20, 2026 | The evidence base itself, across six disciplines | No |
Vibe Coding and Web Application Security: A Twin-Prompt Study (2608.20963) | August 21, 2026 | Whether an appended security section changes what one assistant emits | No, one unnamed assistant |
FlowCheck (2608.28880) | August 28, 2026 | A constraint language for end-user intent, plus models as bug finders | No, Claude Code as the generator |
Towards Effective Generation of Interactive Visualizations with Vibe Coding (2608.29550) | August 30, 2026 | 78 people building interactive visualizations | No |
LipCoder (2608.30793) | August 31, 2026 | A voice-first coding toolkit for visually impaired programmers | No |
Finding 1: AppEval is a mobile repair benchmark, and its authors bound it to Android
AppEval is a good paper doing a hard thing carefully, and the care is exactly what makes the mis-citation plausible.
Its stated problem is that repository-level agents are normally scored on projects whose tests run on the build host, so it is unclear whether repairs survive what the authors call the mobile build-install-launch-test boundary, "where a missing SDK, offline device, or pre-assertion crash can be mistaken for a program failure". Its answer is a runtime-aware acceptance contract. Each task separates a hidden behaviour test from the reference production fix, and a task is accepted only when the same installed-app target reaches an assertion failure on the defective revision and passes after the fix. Infrastructure failures are held as a distinct outcome rather than silently counted as product failures.
That contract is stricter than most benchmark oracles in this space, and it is the part worth borrowing. It is also the part that has nothing to do with whether Lovable, v0, Replit or Bolt.new will build you a working web application from a prompt.
The numbers, for completeness. The audited Android partition contains 200 accepted instrumentation tasks drawn from 24 independently buildable repositories. Five agents score Pass@1 between 22.00 percent and 90.50 percent, a spread of 68.50 percentage points under one dynamic oracle. Two things must travel with those figures. First, they are agent scores, not product scores. Second, the authors state the limit themselves in the abstract: "The quantitative findings in this paper are Android-specific; audited iOS and HarmonyOS results are required before drawing cross-platform generalization conclusions." A benchmark whose title names three platforms currently publishes audited numbers for one, and says so.
The distance from there to a web app builder purchase decision is not a matter of degree. The subject is a coding agent rather than a hosted product. The task is repair rather than generation. The artefact is an installed mobile binary rather than a deployed web application. The toolchain is Gradle, Xcode or the HarmonyOS SDK rather than a Node build. Four independent mismatches, any one of which would be enough.
Finding 2: the benchmark that does match the question is not in the literature
If AppEval is the wrong answer, it is fair to ask what a right one would look like, and whether it exists.
Outside the literature, it partly does. App-Bench evaluates, in its own words, "how well AI coding agents can generate real web apps from a single natural language prompt. One-shot generations. Zero human edits." Its cohort is ten tools including Lovable, v0, Replit, Bolt, Cursor, Codex, Google AI Studio, Gemini CLI, Claude Code and Orchids. Its tasks are six applications spanning finance, healthcare, legal, pharmacy, gaming and rentals. Two experienced full-stack developers manually graded each trajectory against a rubric of 151 items, and the dataset is published on Hugging Face.
On subject identity, that is the closest match to the buyer's question we have seen. It is also worth reading with one number in mind, which the site states plainly: "Each tool was given three attempts per task, with the best-performing run used for final scoring."
Best-of-three is an optimistic estimator, not a central one. It answers "what can this tool do on a good day" rather than "what will this tool do", and the gap between those two questions is exactly the variance that a single-run benchmark cannot see and that this design deliberately discards. The 2026 literature has already measured how large that gap can be. The one academic cross-builder comparison in the previous window reported a code-smells standard deviation of 62.1 for one tool against 7.0 for another on three runs of an identical prompt. When per-tool variance differs by an order of magnitude, a best-of-N estimator does not rank tools evenly: it rewards the high-variance tool, because a wider distribution has a better maximum.
None of that makes App-Bench wrong. It makes its number a ceiling rather than an expectation, and a ceiling is a legitimate thing to publish as long as the reader knows which one they have. We have read the homepage, not the linked analysis, and we say so.
Finding 3: the prompt is part of the product, and now there is a controlled measurement of it
The twin-prompt study is the most directly transferable result in the window, and it is not a builder comparison.
Andročec generated six functionally distinct web applications twice each, from prompts that were byte-identical except for an appended security-requirements section. All twelve programs came from the same agentic assistant and the same model version in a single non-iterative round, and were then put through static, dependency, dynamic and manual analysis, yielding 75 confirmed findings from 85 candidates. The security-aware variant produced fewer confirmed findings in every one of the six applications, 24 against 51, and contained no Critical or High issues at all. The most severe finding in the whole corpus was caught only by manual testing.
The author's own framing is careful and should travel with the result: the corpus is small, each variant was generated once, and the paper reports descriptive observations rather than statistically established effects, positioning itself as preliminary work whose pipeline is being scaled to multiple models and repeated runs.
What makes it matter for benchmark design is the direction rather than the magnitude. If one appended paragraph halves confirmed security findings from the same assistant and the same model, then the prompt is not a neutral test fixture. It is a variable with a large effect on the measured outcome. Any benchmark that pins one prompt across a cohort is holding that variable constant, which is correct and necessary, and is simultaneously choosing one point on a response surface it has not mapped. We proposed a fixed prompt suite for exactly this reason before there was evidence for it. There is now some.
Finding 4: the models are worse at checking the code than a deterministic pass is
FlowCheck starts from a failure mode that anyone who has shipped a generated application will recognise: the interface looks right, and user-visible information does not actually reach the state or the output it is supposed to reach. Silent behavioural failure, with a working-looking demo on top of it.
Its contribution is a constraint language in which an end user states the information flow they expect in terms of the interface they can see, rather than the code they cannot read, and which is then translated into deterministic CodeQL analyses. Evaluated across four applications generated with Claude Code, FlowCheck correctly translated and flagged all 30 injected constraint violations with no false positives. Three frontier models prompted to find bugs in the same code, Claude Opus 4.7, DeepSeek V3 and Gemini Pro, showed significantly lower accuracy, and none reached full accuracy.
Read as a benchmark-design result rather than a tooling result, this is the fourth independent team in two windows to take the language model out of the measuring instrument, after StaminaBench, ICAE-Bench and the SonarQube comparison. Nobody announces this as a consensus and each team seems to have arrived at it separately, which is what makes it worth noticing. It also sets a specific bar for anyone tempted to use an LLM judge on generated applications: on an injected-violation task with a known ground truth, the deterministic path scored 30 of 30 with no false positives and three frontier models did not.
Finding 5: the productivity record is contradictory in a way that is now explained
The state-of-the-art review is the widest thing in the window and the one most likely to be quoted out of context, because it contains numbers that flatly contradict each other. It reports peer-reviewed field experiments finding 26 percent more tasks per week, independent randomised trials measuring a 19 percent slowdown, and team-level telemetry showing code-review time up 441 percent.
The review's argument is that these are consistent once measurement method, scope and time horizon are held constant, and it names six patterns behind the dispersion. Four are worth carrying into any benchmark design: effect shrinkage as measurement gets broader, self-report diverging from independent measurement, output volume conflated with productivity, and bold claims walked back once tested over longer horizons. It closes with a falsifiable conjecture, that the gains are real on new code and shrink or reverse on mature codebases, which would account for most of the disagreement in the record.
That conjecture has a direct implication for this field. Almost every builder benchmark in existence, including every axis this lab has proposed, measures new code. If the conjecture holds, the entire measurement literature is concentrated on the half of the problem where the effect is largest, and the half where buyers eventually live is unmeasured.
The transferability test
Here is the test we now apply before citing any benchmark, published as a pre-registration so that it can be argued with. It grades whether a published result transfers to the question "which AI app builder should I buy", not whether the benchmark is good. A paper can be excellent and score nothing here, which is precisely what AppEval demonstrates.
Scroll to see more
| Signal | Weight | What a failing case looks like |
|---|---|---|
| Subject identity | 30 | The unit evaluated is not the thing you would buy. A leaderboard of coding agents, model checkpoints or harnesses is cited as a ranking of hosted app builders. |
| Task-surface match | 20 | The task is not the job you are paying for. Repairing a known defect in an existing repository is cited as evidence about greenfield generation from a prompt. |
| Oracle transparency | 20 | The acceptance criterion is unstated or not mechanically checkable, or environment failures are folded into product failures, so a missing SDK and a broken feature score the same. |
| Platform and stack match | 15 | The audited runtime is not the runtime you will ship. Results verified only on one mobile platform are read as guidance for a deployed web application. |
| Cohort and repetition | 15 | One generation per product, no reported variance, or a best-of-N estimator presented as an expectation. Ranks are published to more precision than the run count can support. |
Four notes on how to use it.
It is a transfer test, not a quality test. A high weight on subject identity is a statement about what a buyer needs, not about scientific merit.
Subject identity carries the largest weight because it is the only signal that cannot be repaired by adding data. A benchmark that measures the wrong thing more carefully still measures the wrong thing, whereas a thin cohort can be widened and an opaque oracle can be documented.
Oracle transparency is weighted equal to task-surface match because of what AppEval demonstrates in the opposite direction. Its acceptance contract, requiring a real assertion failure before the fix and a pass after it on the same installed target, with infrastructure failures held separately, is stronger than most of what this field publishes. That is a design worth importing even when the result is not.
We are not publishing a score for any of these papers against this test. Applying the weights is the reader's job, and the facts each paper states about itself are in the table above and in the sources below. A lab that has withdrawn its own results has no business converting other people's work into a ranking.
What none of the six measures
- Any commercial AI app builder, as a subject. Not one of the six evaluates a hosted builder product. The last academic paper to do so was
2608.16302on August 17, 2026, covering three products. - Web application generation, in AppEval's case. The subject is mobile repair. There is no web partition.
- iOS or HarmonyOS, in AppEval's published numbers. The audited partition is Android, and the authors state that audited results for the other two platforms are required before generalising.
- Repeated runs of the same condition. The twin-prompt study generated each variant once and says so. AppEval reports Pass@1.
- Mature codebases. Every measurement here is on new or injected-defect code. The state-of-the-art review's own conjecture is that this is where the effect is largest.
- Paid tiers. No paper in this window reports which subscription tier produced the artefacts under test.
How we are treating this at BuilderProof
Three changes, stated so they can be held against us.
First, the transferability test above becomes a precondition rather than a commentary. Before this lab cites a benchmark as support for a claim about builders, the citation has to name the subject evaluated and the acceptance criterion, in the sentence that cites it.
Second, we are adopting AppEval's infrastructure-failure separation into our own published methodology as a requirement rather than an aspiration. A run that fails because a dependency would not install is not a product result, and any protocol of ours that cannot distinguish the two is not ready to produce numbers.
Third, we are recording the best-of-N problem as an explicit reporting rule. If we ever publish a figure derived from more than one run, we will publish the estimator alongside it, and we will publish the spread. A number without an estimator is not a measurement.
We continue to publish no results of our own. Everything above is other people's work, reported and cited.
Limitations
The window is defined by one query on one day. A full-text search for "vibe coding" will miss relevant work that does not use the phrase, which is exactly how AppEval was nearly missed, and there is no reason to assume it is the only such paper. This is a snapshot, not a survey.
We read abstracts and metadata retrieved from the arXiv API for all six papers, and the App-Bench homepage. We did not read the App-Bench analysis post or its Hugging Face dataset, and we have not independently reproduced any figure quoted here.
The observation that AI assistants recommend AppEval for builder comparison questions is our own, from our own prompting, and it is not a controlled measurement of assistant behaviour. What is checkable, and what the argument actually rests on, is the paper's own abstract.
References
All papers retrieved from the arXiv API on September 1, 2026.
- Xie, B., Liu, H., Shi, Z., Zhang, Y., Zhang, S., Peng, Z., Yin, X., Ying, C., Luo, Y., Chen, W., Jin, H., Long, S., Liu, X. and Peng, Z. (2026). AppEval: A Unified Benchmark for LLM-Based Mobile Application Repair in ArkTS, Swift, and Kotlin. arXiv:2608.18588v1, August 19, 2026. https://arxiv.org/abs/2608.18588
- Michels, D. L., Abu Ghazaleh, M., Lazzari, F., Kassem, N. and Klein, J. (2026). Vibe Coding: Practice, Performance, Productivity, and Risk. A State-of-the-Art Review. arXiv:2608.20446v1, August 20, 2026. https://arxiv.org/abs/2608.20446
- Andročec, D. (2026). Vibe Coding and Web Application Security: A Twin-Prompt Study. arXiv:2608.20963v1, August 21, 2026. https://arxiv.org/abs/2608.20963
- Vir, R., Chilton, L., Zhang, Z. and Wu, E. (2026). FlowCheck: Helping End-Users Specify and Verify Intent in Vibe-Coded Web Apps. arXiv:2608.28880v1, August 28, 2026. https://arxiv.org/abs/2608.28880
- Zeng, Y., Tu, R., Xiang, Z., Feng, L., Li, G. and Liu, C. H. (2026). Towards Effective Generation of Interactive Visualizations with Vibe Coding: An Empirical Study. arXiv:2608.29550v1, August 30, 2026. https://arxiv.org/abs/2608.29550
- Kim, H., Lee, S., Kim, J., Suh, B. and Lee, K. (2026). LipCoder: Voice-Enabled Coding Toolkit. arXiv:2608.30793v1, August 31, 2026. https://arxiv.org/abs/2608.30793
- App-Bench. AI Web App Builder Benchmark. Homepage read September 1, 2026. https://appbench.ai/
This page is open to correction. If you have read these papers and think we have mis-stated a subject, a method or a limitation, tell us and we will amend it with attribution.
Written by
BuilderProof Editorial TeamThe BuilderProof lab publishes reproducible, community-editable benchmarks and methodology proposals for AI app builders. Axes are scored from documentation-derived rubrics and open to public revision.
Cite this benchmark
BuilderProof Editorial Team. "AppEval Measures Mobile App Repair, Not AI App Builders (September 2026)". BuilderProof, September 2026. https://www.builderproof.org/benchmarks/appeval-benchmark-scope-and-the-late-august-2026-research-wave.
@misc{builderproof-appeval-benchmark-scope-and-the-late-august-2026-research-wave,
title = {{AppEval Measures Mobile App Repair, Not AI App Builders (September 2026)}},
author = {{BuilderProof editorial team}},
year = {2026},
month = {sep},
howpublished = {\url{https://www.builderproof.org/benchmarks/appeval-benchmark-scope-and-the-late-august-2026-research-wave}},
note = {BuilderProof, builderproof.org}
}Frequently asked questions
Does AppEval benchmark AI app builders?
No. AppEval (arXiv:2608.18588, submitted August 19, 2026) is a benchmark and native-toolchain evaluation framework for LLM-based mobile application repair across HarmonyOS/ArkTS, iOS/Swift and Android/Kotlin. It scores coding agents fixing defects in existing mobile repositories, not hosted AI app builders generating web applications from a prompt. Its audited Android partition contains 200 accepted instrumentation tasks from 24 repositories, and the authors state in the abstract that the quantitative findings are Android-specific and that audited iOS and HarmonyOS results are required before drawing cross-platform conclusions.
What research on AI app builders was published in late August 2026?
Six papers touching vibe coding and generated applications appeared on arXiv between August 19 and August 31, 2026: AppEval on mobile application repair (2608.18588), a state-of-the-art review of vibe coding practice, performance, productivity and risk (2608.20446), a twin-prompt study of web application security (2608.20963), FlowCheck on end-user intent constraints translated to CodeQL (2608.28880), an empirical study of 78 people building interactive visualizations (2608.29550), and LipCoder, a voice-first coding toolkit (2608.30793). None of the six evaluates a commercial AI app builder as its subject.
Which benchmark actually compares AI app builders in 2026?
In the peer-reviewable literature, the most recent one is arXiv:2608.16302 (August 17, 2026), which generated three applications each from Lovable, v0 and Replit using one identical prompt and analysed the output with SonarQube. Outside the literature, App-Bench evaluates ten tools including Lovable, v0, Replit, Bolt, Cursor, Codex, Claude Code and Gemini on six one-shot web application tasks, manually graded against 151 rubric items. App-Bench states that each tool was given three attempts per task and that the best-performing run was used for final scoring, which makes its figure a ceiling rather than an expected result.
How can I tell whether a benchmark result applies to choosing an AI app builder?
BuilderProof proposes a five-signal transferability test, weighted: subject identity (30), task-surface match (20), oracle transparency (20), platform and stack match (15), and cohort and repetition (15). It grades whether a published result transfers to a purchase decision, not whether the benchmark is good work. Subject identity carries the largest weight because it is the only signal that cannot be repaired by adding more data: a benchmark that measures the wrong thing more carefully still measures the wrong thing.
Has BuilderProof scored these papers against its own test?
No. BuilderProof withdrew its June 2026 output-quality result set and the composite ranking derived from it on August 21, 2026, and currently publishes a documented methodology and a series of proposed scoring axes with no results of its own. This page publishes the test and the facts each paper states about itself, and leaves the application of the weights to the reader. No score, level or placement is published here for any paper or any builder.
Related benchmarks
The 2026 Research Wave on AI App Builders: What Six New Papers Actually Measure (August 2026)
Six empirical papers on AI app builders and vibe coding landed between June 17 and August 17, 2026. Only one measures the products people buy. What each actually measures, its numbers, and the limitations its own authors state.
Academic AI App Builder Benchmarks, Mapped (2026): UI-Bench, From Prompt to Product, and Where BuilderProof Fits
Three rigorous, independent AI app builder benchmarks now exist: UI-Bench (design), From Prompt to Product (human end-to-end), and BuilderProof's six-axis rubric. What each measures, mapped side by side from the primary sources.
How We Benchmark AI App Builders: The BuilderProof Methodology v1
BuilderProof methodology v1.1: the published rubric, brief OQ-7, environment standards and weights used to score AI app builders on output quality, speed, deploy quality and agency suitability. The four June 2026 result sets were withdrawn on August 21, 2026 as placeholder data, so the lab currently publishes method, not scores.