• AI-assisted workflows
  • Modernise & Replatform

Accepting AI-generated front-end code: what review is for now

13 Min Read13 Min Read

Last updated on 5 Oct ‘26

Insights

AI assistants generate front-end components in minutes. What acceptance still has to decide: semantics, input, trust boundaries and reversal. The same report found that AI adoption "does continue to have a negative relationship with software delivery stability".

The short answer: a generator gives you a component that renders. Accepting it means deciding whether it can carry real use. That decision rests on four things a screenshot does not show. Does the markup mean what the interface says? Does the behavior hold for someone using a keyboard, zooming in, or leaving the path the demo followed? Is everything the server must decide actually decided on the server? Can the change ship in a small step and be undone? Where a tool can establish the answer, automate it. Human review then goes on the parts that need product context the code does not contain.

Last reviewed September 28, 2026.

Generation got cheaper. Acceptance did not.

The ground to concede first is real. AI assistants remove a large amount of front-end work: scaffolding, boilerplate, first drafts of components, and variations nobody wanted to type by hand. Google's 2025 DORA report (September 2025) found that 90% of respondents use AI at work. Unlike the previous year's report, it found AI adoption had a positive relationship with software delivery throughput and product performance. Nothing in this article argues for generating less.

The same report found that AI adoption "does continue to have a negative relationship with software delivery stability". DORA's explanation is that without "strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability". Sonar's State of Code developer survey (January 2026, over 1,100 developers) describes the same gap from the reviewer's side. 96% said they do not fully trust that AI-generated code is functionally correct. Only 48% said they always check it before committing. 38% said reviewing AI-generated code takes more effort than reviewing code written by colleagues.

Our reading of those findings is this. Producing a candidate got much cheaper. Deciding whether the candidate is fit to ship did not. When candidates arrive faster than decisions can be made, review shrinks to whatever can be checked quickly, and for a component the quickest check is whether it renders. That is an inference from survey and correlational data, not a measured cause. It does match the problem as teams describe it.

Why "does it render" passes so much

A front-end component has several consumers, and a visual review inspects only one of them.

  • The renderer draws it. This is what a screenshot or a quick look in the browser checks.
  • The accessibility tree is what screen readers and other assistive technology read. It is built from the markup's names, roles and states, not from its appearance. A `div` styled as a button and a real `button` can look identical and mean different things.
  • The keyboard and the viewport decide whether it works without a mouse and at the zoom levels people actually use.
  • The network carries every request the component makes. Anyone can send those requests, whether or not the interface shows them a button.
  • The release pipeline decides how the change reaches users and how it comes back out.

A generator works from what it is given: the prompt, the surrounding code and perhaps an image. When the request describes how something should look and behave on the main path, that is what the result can be checked against, and the other consumers go unexamined unless someone asks.

The gap between "it works" and "it is acceptable" shows up in measurement too, though not specifically for front-end components. Veracode's Spring 2026 GenAI code security update (March 2026), across more than 150 models, reports syntax pass rates above 95% and security pass rates of about 55%. The security figure has barely changed across model sizes and recent releases. For cross-site scripting (CWE-80), a front-end-relevant weakness, the security pass rate was 15%. This is a benchmark of coding tasks in four languages (Java, JavaScript, C# and Python), so treat it as neighboring evidence about the gap, not a measurement of your components.

Developers describe the same thing. In the 2025 Stack Overflow Developer Survey (July 2025), the frustration most respondents cited, at 66%, was "AI solutions that are almost right, but not quite". In Sonar's 2026 survey, 53% named "code that looks correct but isn't reliable" as a negative effect of AI on technical debt. These are perceptions, not defect counts. They explain why "it looks right" is weak evidence.

Four questions a render cannot answer

These four questions turn "what are we supposed to be checking?" into claims a change either supports or does not.

1. Does the markup mean what the interface claims?

Every control makes a claim: this is a button, it is called "Delete invoice", it is currently disabled. WCAG 2.2's Name, Role, Value criterion requires that claim to be available to assistive technology, not just drawn on screen.

A research paper at CHI 2026, Measuring the Semantic Accessibility Gap in LLM-Generated Web UIs (April 2026, extended abstract), analyzed 300 UIs from three commercial models. It found 541 semantic violations across six fault types. These are cases where an accessibility attribute is present but meaningless. The authors' example: an image with `alt="image"` passes every automated check yet tells a screen reader user nothing.

There is also a mechanism behind the pattern. The A11yn paper (October 2025) notes that models learn from a web full of accessibility faults. WebAIM's 2026 Million report (February 2026) detected WCAG 2 failures on 95.9% of the million home pages it tested. Pages using ARIA averaged 59.1 detected errors against 42 on pages without it. That is a correlation on existing websites, not a finding about generated code. It does show that adding accessibility attributes is not the same as making something accessible.

What to check: open the accessibility tree, or turn on a screen reader, for the new component. Does each control's name and role say what it visibly claims to do?

2. Does the behavior hold off the path it was built on?

WCAG 2.2 requires all functionality to be operable through a keyboard. The Reflow criterion requires content to work at a width equivalent to 320 CSS pixels, which W3C notes equals a 1,280-pixel viewport at 400% zoom. Neither shows up when the reviewer uses a mouse at default zoom.

The same applies to states. A generator builds the path it was described. States nobody described (empty, loading, error, partial permission, a name three times longer than the sample data) get whatever default the tool chose. The Prototype Hardening Checklist treats behavior a prototype does not establish as an unknown that belongs in the build plan, not as something to assume. Which states exist is a product decision, and nobody can check a state until someone has decided it exists.

What to check: tab through the component without touching the mouse. Zoom to 400%. Switch the data to empty, error and very long content, and look at what happens.

3. Where does the trust boundary sit?

The OWASP Top 10:2025 keeps Broken Access Control at number one. Its prevention guidance is plain: "Access control is only effective when implemented in trusted server-side code or serverless APIs, where the attacker cannot modify the access control check or metadata." Hiding a button in the interface decides nothing, because the request behind the button can be sent directly.

Generated front ends can make this line harder to see, because some stacks let the client talk to data directly. The disclosure of CVE-2025-48757 (May 2025) described Lovable-generated front ends that called the database from the browser with the public "anon" key, "relying exclusively on RLS" (row level security) to protect the data. Where those policies were missing or wrong, data was exposed. That was one platform's pattern, not evidence about every assistant. It does show the mechanism: the decision about who may read what lived in a place the generated interface could not enforce.

What to check: list every decision the component appears to make, such as hiding an action, filtering rows or showing a price. For each one, ask whether a direct request to the underlying endpoint would get past it.

4. Can it be released in a small step and reversed?

DORA's AI Capabilities Model (September 2025) names seven capabilities that amplify AI's positive effect. Two are about release. Working in small batches "amplifies the positive influence of AI on product performance". For version control, "the frequent use of rollback features boosts the performance of AI-assisted teams". A third, user-centric focus, carries a warning: without it, DORA found AI adoption can have a negative effect on team performance.

What to check: is the change small enough to accept as one decision? Can it be switched off or reverted without a data migration? Does anyone know what signal would tell you to reverse it?

Automate the facts, keep the decisions

A tempting way to divide the work is "what the generator can't know versus what it can". It does not hold for long. The CHI 2026 paper above found that LLM judges detected injected semantic faults, such as vague alt text and unclear link purpose, with recall of 80 to 92%, though they struggled with others, such as headings that do not match their content. A11yn trained a model against automated WCAG auditors and cut its inaccessibility rate by 87.5% and 58.1% under two auditors. Tools are reaching into ground that was recently manual, and any line drawn around tool capability is likely to keep moving.

A more durable line separates facts that can be established from the artifact from decisions that need context the code does not contain. Facts belong in automated gates so that no human re-checks them. Decisions belong in review, because nothing in the diff says which data is sensitive, which role may act, or which states your users will meet.

QuestionFacts a tool can establish (put them in CI)Decisions that need context (keep them in review)
MeaningMissing names, roles and labels; contrast; some semantic faults via LLM judgesWhether each name says what the control does in this product; whether structure matches the task
BehaviorKeyboard reachability in scripted tests; layout at 320 CSS pixels; regression of states already listedWhich states exist, what each should say, and who meets them
Trust boundaryKnown injection patterns; secrets in the bundle; tests that call endpoints without credentialsWhich data and actions are sensitive; which role may do what
ReleaseChange size limits; presence of a flag or revert pathWhat counts as failure, and who decides to reverse

One caution from WebAIM applies across the whole left column: "Absence of detected errors does not indicate that a page is accessible or conformant." A green gate only means the checkable facts passed. It does not mean the change has been accepted.

This split also answers the question in the title. Human review stops re-checking what a tool already checks and asks instead: which claims does this change make, and has each one been evidenced? A shared component library enforced in code, the work of Frontend Design Systems, moves more of the left column into the components themselves.

Running acceptance at generation speed

None of this requires slowing generation. It requires changing what a review is made of.

  1. Keep each change to one decision. One component or one behavior per review, in line with DORA's small-batch finding. Large generated diffs are hard to accept as a whole.
  2. Make the author state the claims. The person who prompted the change says which roles it affects, which states it handles, what data it touches and how it would be reversed. The generator produced the code, but a person stays accountable for it.
  3. Turn the facts into gates. Whatever the left-hand column can establish runs before a human looks, so review time goes on the right-hand column.
  4. Review the running thing, not a picture of it. Use a keyboard, 400% zoom, a screen reader or the accessibility tree, and realistic data, including empty and hostile data.
  5. Record what was accepted and on what evidence. Anyone who later changes the component can then see which claims they are about to break.

This is the position Tcules takes on its Implementation QA and Design Engineering pages: AI-generated code enters the same review, and coding agents change the speed of implementation, not the acceptance criteria. AI can help implement and verify, but it does not get to decide what "correct" means without product context. The same acceptance also has to cover the customer work that already exists, which is where interface regressions go unnoticed.

Where design engineering fits

Design engineering is still settling as a discipline in software, and teams draw its boundaries differently. Tcules uses it to mean carrying product and interaction decisions into coded prototypes, component systems and front-end implementation. It does not replace product design or software engineering.

Acceptance of generated front-end code falls into that overlap. Questions one and two are interaction decisions expressed in markup and behavior. Question three is a product decision about who may do what, which engineering then enforces. Question four is a delivery decision. All four need someone who can read both the product intent and the code.

Tcules uses design engineering selectively. A product may need it for one difficult interaction while its own engineers continue to own the wider application. Depending on where the doubt sits:

Common questions

Should we just prompt the generator to be accessible and secure? Yes, it helps, and research such as A11yn shows models can be trained to produce fewer detectable violations. The output still needs acceptance, because a prompt carries only the context you put in it, and the sensitive decisions depend on context that lives elsewhere.

Can an AI code reviewer do acceptance for us? It can take on more of the facts column, and that ground is growing. The decisions column depends on product context (which data is sensitive, which states matter, when to reverse) and on someone being accountable for the answer.

Does this mean we should generate less? No. The argument is about which work generation removes and which it leaves. Generation removed much of the writing. The accepting is still there, and it needs to be deliberate.

What is the smallest useful first step? Take one generated component that has already shipped and put it through the four questions. What that turns up will show whether the gap is in your gates, in your review, or in decisions nobody has made yet.

Tell us about the product problem you are working on.

Talk to Tcules fast and affordable

Start a project