BuilderProof editorial team23 min read5 views

Does it fire before the word exists? A text-composition input axis proposal (October 2026)

While an input method is composing, the specification requires keydown, keyup, beforeinput and input to keep firing, and the key that confirms a candidate arrives before compositionend. The one attribute that separates those events from ordinary typing is the one attribute on this surface that is not Baseline.

Updated on October 9, 2026

Flat minimal diagram, off-white background: a row of five rounded squares, the first four thin slate-blue outlines and the last solid. A dashed arrow falls from the first to a small rounded rectangle; a solid arrow from the last crosses a long bar.
Flat minimal diagram, off-white background: a row of five rounded squares, the first four thin slate-blue outlines and the last solid. A dashed arrow falls from the first to a small rounded rectangle; a solid arrow from the last crosses a long bar.
On this page

Quick answer (October 2026): this is a pre-registration, not a result. It proposes that AI app builders be scored on text-composition posture: whether the application a builder generates can tell the difference between a keystroke that produced text and a keystroke that is still assembling it. While an input method is composing, the specification requires keydown, keyup, beforeinput and input to keep firing, and it requires the key that confirms a candidate to arrive as a keydown before compositionend. The one attribute that separates those events from ordinary typing is the one attribute on this surface that is not Baseline. Seven weighted signals, four postures and a ten-step protocol are proposed below. No builder is scored and no placement is published.

Why we are proposing this axis

Generated applications are full of handlers that run while somebody is typing. Submit on Enter. Search as you type. A character counter under a bio field. Inline validation that turns a border red. An autosave that fires on change. A builder emits these without being asked, because they are what a text field is expected to do.

Every one of them rests on an assumption that is never written down: that the value in the field is a thing the person meant to type.

For a developer working on a US or UK keyboard, that assumption is true on every keystroke, and it is true every single time they test. It is not true for a reader composing Japanese, Chinese or Korean, where a sequence of keystrokes is assembled into candidate text and then confirmed. And it is not true, as we will show from the specification's own worked example, for a reader on a French keyboard typing a circumflex.

The failure is not that composition is unsupported. The browser supports it perfectly. The failure is that the application's handlers fire during it, on text that does not exist yet, and nothing in the event stream looks wrong.

The trap: the finished-word illusion

Every axis we propose names the illusion that hides its defect. This one is the finished-word illusion.

The illusion is that the current value of a text field is a word somebody finished typing. On one keyboard, in one locale, that is true on every observation a developer will ever make while building. There is no error, no warning, no console message and no failed request. The handler receives a string, the string is well formed, and the handler does its job.

What makes this trap a member of the family we keep finding is its structure: every observation available at test time is generated by the same actor, and here that actor is the developer's own input method. A Latin keyboard cannot emit a composition session. The developer is not failing to look carefully; they are structurally unable to produce the condition without going into an operating system settings screen and installing a different input method. The reset lives outside the application entirely, on a surface the application cannot address and does not know exists.

There is a second closure on top of that one. When the defect does fire, the reader is not shown an error. They are shown a half-finished word in a submitted message, or a search that returned nothing for a fragment of a syllable, or a form that submitted when they pressed Enter to choose a character. Each of those reads as the reader's own mistake. There is no channel back.

What the specification actually documents

W3C

The W3C UI Events specification does not treat composition as an edge case. It gives it six numbered sections, including Composition Event Order, Canceling Composition Events, Key Events During Composition and Input Events During Composition. The behaviour below is quoted from it directly.

Key events keep firing, and they are flagged. Section 3.6.5 is explicit:

"During the composition session, keydown and keyup events MUST still be sent, and these events MUST have the isComposing attribute set to true."

The key that confirms the candidate is a keydown inside the session. The same section gives the event table. Step 1 is the keydown that initiates composition, with isComposing false. Steps 2 and 3 are compositionstart and compositionupdate. Step 5 is annotated, in the specification's own words, "This is the key event that exits the composition", and it carries isComposing true. Only at step 6 does compositionend arrive.

Read that ordering against a handler that submits a form when it sees Enter. The Enter that a reader presses to accept a candidate is a keydown, it arrives before compositionend, and to a handler that does not inspect isComposing it is indistinguishable from the Enter that means send. The specification documents the exact sequence in which the bug occurs.

The input event fires too, and this is the part that defeats the usual advice. The standard guidance for text fields is to stop listening to key events and listen to input instead. Section 3.6.6 removes that defence:

"The beforeinput and input events are sent along with the compositionupdate event whenever the DOM is updated as part of the composition."

So a React onChange, which is backed by input, fires on intermediate composition states. Every search-as-you-type request, every autosave, every inline validation pass hung off that event runs against text the reader has not committed. The specification is equally precise about the other end:

"Since there are no DOM updates associated with the compositionend event, beforeinput and input events should not be sent at that time."

That matters more than it first appears. The moment the value becomes final is the one moment in the whole sequence that does not produce an input event. An application that only listens to input is notified about every provisional value and is not notified about the committed one.

Intermediate values are not prefixes. The specification's handwriting-recognition example shows compositionupdate carrying "test", then "text", annotated "User rejects first word-match suggestion, selects different match". The data can also be empty: the specification says it "MAY be the empty string" where content has been deleted. Any code that treats successive values as a growing prefix, which is what a naive incremental search does, is working from a model the specification contradicts.

This is not only a CJK condition. The specification's dead-key section is the part most likely to surprise a reader who has filed this under internationalisation and moved on. Where the keyboard is handling a dead key:

"a key value of 'Dead' is reported. Instead, implementations generate composition events which contain the intermediate state of the dead key sequence reported via the data attribute."

Its worked example produces a circumflex e on a French keyboard, and the event table runs keydown with key Dead, compositionstart, compositionupdate carrying the bare combining accent, keyup with isComposing true, a second keydown with isComposing true, compositionupdate, compositionend, keyup. A reader typing an accented character in French, Portuguese or Spanish on a dead-key layout is running a composition session, with a provisional value that is a lone combining mark.

The observation channel is not guaranteed. The specification notes, twice, that "In some implementations or system configurations, some key events, or their values, might be suppressed by the IME in use." And for legacy code the signal is a sentinel rather than a flag: "If an Input Method Editor is processing key input and the event is keydown, return 229."

Cancelling is mostly not available. compositionstart is cancelable and its default action is to "Show a text composition system candidate window", but compositionupdate and compositionend are not cancelable, and the specification states plainly that "Most IMEs do not support canceling updates during a composition session." It also warns that cancelling the event is "distinct from canceling the text composition system itself", and that even where a session is terminated, "the compositionend event MUST be sent." An application cannot opt out of composition; it can only decide whether to notice it.

The finding: the events are Baseline and the attribute that reads them is not

MDN

The obvious remedy, once the behaviour above is understood, is one line: ignore the event when isComposing is true. We checked whether that remedy is portable, and it is not.

MDN publishes a Baseline availability status on each of these features. Measured on the live pages:

  • CompositionEvent: "Baseline Widely available. This feature is well established and works across many devices and browser versions. It's been available across browsers since July 2015."
  • Element: input event: "Baseline Widely available ... available across browsers since January 2020."
  • KeyboardEvent.isComposing: "Limited availability. This feature is not Baseline because it does not work in some of the most widely-used browsers."
  • InputEvent.isComposing: "Limited availability. This feature is not Baseline because it does not work in some of the most widely-used browsers."

The asymmetry runs the wrong way. The events you need to observe have been universally available for a decade. Both attributes that would let you discriminate a composing keystroke from an ordinary one, on the keyboard event and on the input event alike, are flagged as not Baseline on the vendor's own reference.

We could not find this contrast stated anywhere. Each MDN page carries its own status banner and neither draws the comparison. The UI Events specification offers no guidance at all on what an author should do: it defines isComposing and never recommends tracking session state as an alternative, and a search of the full specification text returns no remedy advice on this surface.

The consequence for a rubric is concrete and it is why signal 2 below is weighted as heavily as it is. Reading the flag is the wrong remedy to score as sufficient. The portable remedy is stateful: set a boolean on compositionstart, clear it on compositionend, and consult your own boolean, using isComposing only as a corroborating signal where it exists. An application that reads the attribute and nothing else has done real work and is still relying on a feature its own documentation marks as not widely available.

A documented tension: the correct listener is the one that fires at the wrong time

There is a genuine conflict here rather than a simple best practice, and it is worth stating rather than smoothing over.

The accumulated advice for text inputs is to prefer input over key events, and that advice is correct for almost every reason it is usually given: it covers paste, it covers drag and drop, it covers assistive technology, and it does not depend on key codes. MDN states the contrast with change plainly: "For text controls, the input event is fired as the user edits the value. This is unlike the change event, which only fires when the value is committed."

Firing as the user edits is exactly the property that makes input the right listener, and exactly the property that makes it fire on uncommitted composition states. The advice is not wrong. It is incomplete on a surface where "as the user edits" and "when the user has typed something" stop being the same thing.

The resolution is ordering rather than choice. Keep input as the listener. Gate the consequences of it, the network call, the counter, the validation pass, on composition state. The event is the right event; what needs a condition is the work it triggers.

The default is not a choice

A text field with a submit-on-Enter handler, a debounced search and a character counter is the default output of most generation prompts. None of those three handlers is wrong. None of them was chosen with composition in mind either, because the question never came up: the generated code is correct for the only keyboard anyone in the loop has used.

So the thing a rubric should score is not which decision was made. It is whether a decision was made at all. Signal 7 below carries that directly, and it is the signal most likely to be a flat zero across the field.

The rubric

Seven signals, one hundred points. Weights are proposed, not settled, and the three we expect to be argued about are set out underneath.

Scroll to see more

SignalWeightWhat a failing case looks like
Commit-triggering keys are gated on composition state22A keydown handler submits the form, sends the message or closes the dialog on Enter without consulting composition state, so the key a reader presses to accept a candidate submits a half-assembled word instead
Composition state is tracked from the events, not read from an attribute alone20The only guard is event.isComposing, an attribute both of whose host interfaces are documented as not Baseline, so the guard is absent on the browsers where it is unimplemented and the field silently reverts to the ungated behaviour
Work triggered by value changes is deferred past commit16A search request, autosave or inline validation pass fires from input on every compositionupdate, so the application issues requests for fragments the reader never typed and renders empty results under a partially formed word
Intermediate values are never treated as a prefix of the final one14Code accumulates or diffs successive values on the assumption they only grow, so a candidate the reader rejects and replaces, or a cleared buffer the specification allows to be the empty string, corrupts the accumulated state
Counts and limits are measured on committed text12A character counter or a maxlength-style guard reads the field mid-composition, so the number shown to the reader jumps around while a candidate window is open, or input is cut off at a boundary inside an unconfirmed sequence
The dead-key path is covered, not only the candidate-window path9Composition handling is treated as a CJK-only concern, so a reader typing an accented character on a dead-key layout still trips the ungated handlers that the specification's own worked example shows them running through
The decision is written down7Nothing in the emitted project states whether composing input was considered, so no reviewer can tell a deliberate choice from code that was never exercised by a composing keyboard

Three weighting decisions worth arguing about now

Why the commit key leads at 22. Of the failures in this family it is the only one that destroys data the reader has already produced. A premature search returns a poor result and recovers on the next keystroke. A premature submit sends a half-formed word to another person, or writes it to a record, and the reader cannot take it back from inside the application.

Why tracking state outscores a correct remedy at 20. This is the weight we expect to be challenged, since an application reading isComposing is doing the right thing conceptually. We weight it this high because the vendor's own availability data says the attribute does not work in some of the most widely used browsers, which means the flag-only remedy fails precisely where it cannot be observed failing. A guard that is absent on some engines and present on others is not a posture, it is a coin toss.

Why writing it down is only 7. It is the lightest signal on every axis we publish and we keep it light deliberately. A documented decision does not make a composing reader's form work. It is scored because its absence tells a reviewer that the question was never asked, and because it is the one signal an honest team can fix in an afternoon.

Four postures

Unaware. No composition handling anywhere. Key handlers fire on the confirming keystroke, value-change work fires on every intermediate state. The application behaves correctly for every keyboard that does not compose and incorrectly for every keyboard that does.

Flag-dependent. Handlers consult isComposing and nothing else. Correct on engines that implement the attribute, and silently back to Unaware on those that do not. This is the posture we expect to be most common among teams that have encountered the problem once.

Session-tracked. The application maintains its own composition state from compositionstart and compositionend, consults that state in commit-key handlers and in value-change consumers, and treats isComposing as corroboration rather than as the source of truth. Portable across engines.

Composition-complete. Session-tracked, plus the deferred-work discipline applied to every consumer rather than to the submit path alone, plus counts and limits measured on committed text, plus the dead-key path covered, plus the decision recorded in the project.

The measurement protocol

  1. Enumerate every text-entry control in the generated application, including search boxes, comment fields, chat composers and any contenteditable surface.
  2. For each, enumerate the handlers attached to keydown, keyup, beforeinput, input and change, and record which of them trigger a commit, a network request, a state write or a visible count.
  3. No-information control. Before any composing input, exercise each control with plain Latin typing and record the full event sequence and every consequence. This is the baseline against which a composing run is compared, and it is also the measurement that reproduces what the builder's own testing would have seen.
  4. Activate a candidate-window input method at the operating system level and repeat, recording for every event whether isComposing is present at all and what value it carries.
  5. Press the confirming key with a candidate window open and record whether the application commits, submits or sends. Record the result separately from any other failure, because this is signal 1.
  6. Count the value-change consequences fired between compositionstart and compositionend: the number of network requests, state writes and re-renders. Report the integer, not a pass or fail.
  7. Reject a candidate and select a different one, so the provisional value changes to something that is not an extension of the previous value, then clear the buffer so the provisional value becomes empty. Record whether accumulated state survives both.
  8. Repeat steps 4 to 7 on a dead-key layout producing an accented character, with no candidate-window input method active, and record the results separately. A build may pass one path and fail the other.
  9. Repeat the whole sequence on a second browser engine, chosen so that one engine implements isComposing and the other does not, and report the two results separately rather than merged.
  10. Search the emitted project for any recorded decision about composing input, and record its presence or absence verbatim.

How this relates to axes already published here

Text comparison: what the field eventually contains against whether the text exists yet

Our text-comparison axis is the nearest neighbour and the boundary is sharp. That axis scores normalisation, case folding, collation and the comparison rules applied to a stored string. Every one of its seven signals operates on text that already exists, and its remedy runs at the storage boundary, which is to say after compositionend. It is structurally unable to observe the event this axis measures, because the event is over before its rubric begins.

The boundary runs both ways, which is what makes it a boundary rather than a split hair. A build can store and compare composed text immaculately, normalising to NFC and declaring its collation, and still fire a search request on a lone combining accent. A build can gate every handler on composition state correctly and still store two spellings of the same name as two different people.

There is a pleasing detail here: that axis named its own trap the one-keyboard illusion, and wrote that every string a developer tests with is typed "on one machine, with one input method and one locale". It named the assumption and scored the consequences at rest. This axis scores the consequences in flight.

Autofill: a value that arrives with no event against a value that arrives with too many

Our autofill axis is the mirror image, and its own limitations section says so better than we could paraphrase: it records that "There is no event for a fill", and that the only in-page signal is a CSS pseudo-class.

So the two axes bracket the same field from opposite sides. Autofill scores a value that appears with no notification at all. This axis scores a value that announces itself many times before it means anything. One failure mode is silence and the other is noise, and a build can be excellent at one and oblivious to the other, because the remedies have nothing in common.

Browser support: the feature is Baseline and the discriminator is not

Our browser-support axis scores whether a build targets the engines its readers actually use, and it uses Baseline as its instrument. This axis applies that same instrument one level down, to a discriminator rather than to a feature, and gets an uncomfortable answer: the thing being observed is Baseline and the means of observing it is not.

That is a composition rather than a boundary dispute. A build can pass the browser-support axis outright, targeting every engine correctly, and still be Flag-dependent here, because targeting an engine is not the same as having an attribute available on it.

Model-output validity: a well-formed signal whose content is not yet usable

Our model-output-validity axis scores what an application does at the point where the status code is already 200 and the body may still be unusable. The structural parallel with this axis is close enough to state and not close enough to be a boundary: in both cases a well-formed signal arrives through the correct channel, the application's machinery reports success, and the payload is not yet something to act on.

The two are measured on entirely different surfaces, by different events, with different remedies, so they compose rather than compete. Naming the shared shape is more useful than drawing a line that is not there.

Compositions rather than boundaries

Two more relationships need a sentence and no link.

Our length-and-truncation work scores whether a whole name survives being saved, counting grapheme clusters rather than code units. Signal 5 above touches the same reader-visible number from the other end: not whether the count is computed correctly, but whether it is computed on text the reader has finished typing. A counter can be perfectly correct about graphemes and still be counting a provisional buffer.

Our output-encoding work separates strings the application generated from strings a user typed. A provisional composition value is the second kind, arriving earlier than that axis assumes any string arrives. Nothing in either rubric needs changing; the ordering is just worth noticing.

What this axis is not

It is not a verdict on any builder. No build is scored here, no placement is published, and this document produces no level, rank or number attached to a product. It is a pre-registration: the rubric, the postures and the protocol are published first so that the method can be argued with before any result exists.

It is not a claim that composition is poorly specified. The opposite is true. The UI Events specification handles this surface unusually well, with explicit event tables, stated ordering requirements and worked examples for candidate windows, handwriting and dead keys. The gap is between what the specification requires of a browser and what a generated application does with it.

It is not an accessibility rubric, although it overlaps with one. The population affected is readers using a particular class of input method, which includes a very large number of people typing their own first language.

Limitations and open questions

Every axis we propose names its own blind spots, because a scorecard that hides them is worse than no scorecard.

  1. It measures one input method at a time. Candidate-window behaviour differs between input methods and between operating systems, and the specification leaves much of it to the implementation. A result is a reading of the input methods tested, not a general claim.
  2. The observation channel is partly unreliable by specification. UI Events states that some key events or their values "might be suppressed by the IME in use". A measurement that sees no event cannot always distinguish an application that handled something from an environment that never delivered it. Step 3's no-information control limits this but does not remove it.
  3. The engine split in step 9 is a moving target. Which engines implement isComposing is exactly the kind of fact that changes, and a rubric that hard-codes it will rot. The protocol therefore asks for two engines chosen by current availability data and for the results reported separately, rather than naming engines in the rubric.
  4. Signal 3 is reported as an integer and scored as a judgement. How many intermediate requests is too many depends on what the request costs. We have not proposed a threshold and we are not confident one is meaningful across application types.
  5. Automated measurement is hard and we are not pretending otherwise. Synthesising a composition session in a test harness is not the same as driving a real input method, and a harness that dispatches composition events itself proves only that the handlers respond to synthetic events.
  6. Signal 7 may be unfairly easy to fail. Most generated projects document very little of anything. A near-universal zero on one signal carries no discriminating information, and if the first measurement round shows that, the signal should be re-scoped or dropped rather than kept for tidiness.
  7. A build with no live text handlers cannot fail most of this. An application whose forms commit only on an explicit button press, with no search-as-you-type and no inline counters, has nothing to gate. Signals 1, 3 and 5 should record a full pass in that case rather than being left blank, because declining to attach the handlers is a legitimate way to be correct.

How to contribute

This rubric is open to revision before any measurement round begins. The most useful contributions are a weighting you think is wrong and why, an input method whose behaviour diverges from the event tables above, a dead-key layout that behaves differently from the specification's worked example, and current availability data for isComposing on engines we should include in step 9. Corrections to anything quoted here are welcome and will be published as corrections rather than silently edited.

References

All sources were read in full on 9 October 2026 and all quotations were taken from the live pages on that date.

  • W3C UI Events, editor's draft: sections 3.6.2 Composition Event Order, 3.6.3 Handwriting Recognition Systems, 3.6.4 Canceling Composition Events, 3.6.5 Key Events During Composition, 3.6.6 Input Events During Composition, 3.6.7 Composition Event Types, the dead-key worked example in section 3.5, and section 7.3.1 on legacy keyCode determination.
  • MDN, CompositionEvent, including its Baseline status and the data and locale properties.
  • MDN, Element: compositionstart event.
  • MDN, KeyboardEvent.isComposing, including its Limited availability status.
  • MDN, InputEvent.isComposing, including its Limited availability status.
  • MDN, Element: input event, including its Baseline status and its statement of the contrast with change.

Cite this benchmark

Plain text
BuilderProof editorial team. "Does it fire before the word exists? A text-composition input axis proposal (October 2026)". BuilderProof, October 2026. https://www.builderproof.org/benchmarks/does-it-fire-before-the-word-exists-text-composition-axis-october-2026.
BibTeX
@misc{builderproof-does-it-fire-before-the-word-exists-text-composition-axis-october-2026,
  title  = {{Does it fire before the word exists? A text-composition input axis proposal (October 2026)}},
  author = {{BuilderProof editorial team}},
  year   = {2026},
  month  = {oct},
  howpublished = {\url{https://www.builderproof.org/benchmarks/does-it-fire-before-the-word-exists-text-composition-axis-october-2026}},
  note   = {BuilderProof, builderproof.org}
}

Frequently asked questions

What is the text-composition axis?

It is a proposed BuilderProof benchmark axis that scores whether the application an AI app builder generates can distinguish a keystroke that produced text from one that is still assembling it under an input method. It is a pre-registration, not a result, and no builder is scored on this page.

Is this only a problem for Japanese, Chinese and Korean input?

No. The W3C UI Events specification documents dead keys as generating composition events too, and its worked example produces a circumflex e on a French keyboard. A reader typing an accented character on a dead-key layout runs a composition session, so the affected population is far wider than candidate-window input methods alone.

Is listening to the input event instead of key events enough?

No. UI Events section 3.6.6 states that beforeinput and input are sent along with compositionupdate whenever the DOM is updated as part of the composition, so those events fire on uncommitted text. The specification also notes there are no DOM updates at compositionend, so the moment the value becomes final is the one moment that produces no input event.

Why not just check the isComposing attribute?

Because MDN marks both KeyboardEvent.isComposing and InputEvent.isComposing as Limited availability and not Baseline, while the composition events themselves are Baseline and widely available. A guard that depends only on the attribute is absent on the engines that do not implement it. The portable remedy is to track session state from compositionstart and compositionend.

Does the browser fill it in? An autofill posture axis proposal (October 2026)

The person testing a freshly generated app has an empty credential store on a brand new origin, so nothing is offered and every form looks the same. This proposal scores whether field purposes are declared, whether a credential pair is distinguishable, whether a save is ever offered, and whether identifiers survive a redeploy.

27 min read47
benchmarks

Does It Run in the Browser They Actually Have? A Proposed Browser-Support Axis for AI App Builder Output (September 2026)

A proposed benchmark axis for which browser engines a generated app can actually run in. Baseline defines widely available as 30 months of interoperability and names what it cannot see. Next.js floors at Firefox 111, Vite resolves its default to Firefox 114, and the preview pane reports a pass either way. Seven weighted signals, four postures, a ten-step protocol, and no scores.

19 min read126