LondonJS · 2026

How to build an
agentic software factory

The goal: don't write any code by hand.
The skills, rules and guardrails, not the theory.

where we are

This is how we write code now.

Weekly commits: flat through 2025, then a sharp climb from early 2026
Weekly commits, one repo · flat for a year, then the agents showed up

Same person. Same repo. The agents showed up in January.

the goal

Stop writing code. Not a single line, by hand.

a developer typing furiously
you today
→
lying back on a pile of cash
you, after building a software factory
the five levels

Five levels of using agents.

1The advisor. A chat tab on the side. You ask, it answers, you write every line.
2Pair programming. It drafts the boilerplate. You still rewrite the parts it got wrong by hand.
3Big PRs - you're the bottleneck. It writes 40 files. You review all of them. "No, not like that."
4Autonomous. You plan the architecture, it implements, you review. Few issues come back.
5The software factory. You name the feature. You never read the code.
Level 3 is the wall. A pull request touching 40 files, and you open them one at a time - wrong pattern, wrong architecture. Then you test the whole thing by hand. You can only run one agent, because you have to watch it. Tonight is about 3 → 4.
the gap

What's missing between
step 3 and step 4?

📚

Documentation

Nothing tells the agent how your codebase works, so it invents its own answer.

🛡️

Guardrails

Nothing stops it when it's wrong. No lint rule, no test, no review that fails.

the loop

Everything in this talk serves one loop.

all green agent writes code opens a PR CI runs red → agent fixes merge → prod
what's inside a well-defined codebase

Five pillars.

01

CLAUDE.md rules

The entry point. "Whenever you do X, do Y."

02

Architecture docs

How the data actually flows.

03

Deterministic guards

Linter rules. The law.

04

Tests

The most important pillar. Eyes on every change.

05

The review loop

Bots that catch what tests can't.

pillar 01 · CLAUDE.md

CLAUDE.md - the agent's entry point.

# CLAUDE.md
When you add a UI component:
  → check the component library first, reuse it
  → use our state hook, never raw useState
  → add a Storybook story for every new state

When you change a backend service:
  → add an integration test (service → DB → assert)
🌱
Start smallA handful of rules. Add one every time it gets something wrong.
🤖
Let AI write the rulesPoint it at your code: "look at how we do this, write the rule."
📈
It compoundsEvery rule is a mistake the agent never makes again.
pillar 02 · architecture docs

Architecture docs - how the system works.

Draw the main journey, not every table.

UI →API endpoint →business logic →DB/S3 →back to the UI

The spine of your app, so the agent knows where things live and why. Link it from CLAUDE.md and the agent reads two files instead of twenty to learn how your app works. That is context it gets to spend on your actual problem.

pillar 03 · deterministic guards

CLAUDE.md is a suggestion.
The linter is law.

Everything so far is a polite request: "please use this component, please name it that way." Some of it will not land. And mechanical rules don't belong in CLAUDE.md anyway - every line you put there is context the agent pays for on every task. The linter is the shortcut: it never has to learn the rule, it just finds out when it breaks one.

agent writes code →runs the check →❌ rule violated →sees the error, fixes it
pillar 03 · what kind of rules

Five kinds of guard.
We have 20 custom ones.

01

Use our component, not the raw one

The tooltip prop on Button, never a Tooltip wrapper. Durations from the animation enum.

02

React footguns types miss

Render <Foo/>, never call Foo(). AnimatePresence children need a key.

03

Effects & server components

Every useEffect carries a why-comment. An await in an RSC must be wrapped in try/catch.

04

Bundle & import boundaries

No barrel @repo/ui/icons imports. Heavy components load dynamically. No Node lib on the client.

05

Drift

CI regenerates the API client and fails if the diff isn't empty. No circular imports. No unused eslint-disable.

pillar 03 · why it matters

Every guard is a mistake
that already happened.

# three real rules from our frontend
@repo/prefer-tooltip-prop
  → use the Button tooltip prop, not a Tooltip wrapper

@repo/no-runtime-identifier-name
  → minifiers mangle function names in production

no-restricted-imports: @repo/ui/icons
  → import the icon directly, barrels bloat the bundle

Nobody could review their way to these. They're bugs we shipped once, and now they're impossible. And the agent fixes them alone - a red check needs no context, only the error.

pillar 04

Tests are the
most important
pillar

Without them, you can't let go. And letting go is the whole point.

pillar 04 · why it's the important one

You are the biggest bottleneck now.

Billing table redesignOpen
#128 · agent-a · ✓ checks pass · needs testing
Onboarding modalOpen
#131 · agent-b · ✓ checks pass · needs testing
Search filtersOpen
#134 · agent-c · ✓ checks pass · needs testing
Dark mode toggleOpen
#137 · agent-d · ✓ checks pass · needs testing
Export to CSVOpen
#140 · agent-e · ✓ checks pass · needs testing
Settings pageOpen
#143 · agent-f · ✓ checks pass · needs testing

Six agents finished at once - and every one still needs a human to click through it. Tests are how you stop being the thing that checks.

pillar 04 · the shape

The test mix that actually scales.

🗄️

Backend integration

Call the service, hit the real DB, assert the logic.

👆

Storybook interaction

Click the button, the modal opens, submit, it closes.

👁️

Visual

Does it still look right? A screenshot of every state.

🐌

A little end-to-end

Very little. Slow and flaky, so keep CI fast.

And the agent writes all of them.

pillar 04 · the objection

"We tried Storybook.
Nobody maintained it."

I know. That was true for me too, at two companies. But you don't write them any more - the agent does, because a CLAUDE.md rule says so. It's the same argument as types: "why write types when I could write plain JavaScript?" Because it's nearly free now.

Stories are your best line of defence on the UI.

uiverify.ai/blog · you don't write them by hand
10 agents.
10 PRs.
How do you know the UI didn't break?

Every one of them touched a shared component. Traditionally you'd open the app, find each page that uses it, and get each one into the exact state where you can see it. You are not doing that ten times.

uiverify.ai
UI Verify posting its verdict on the pull request: 35 intended, 1 likely regression
1 · on the pull request
UI Verify diff view: baseline vs new render side by side, flagged a likely regression with a written explanation
2 · in the dashboard
>the visual check failed on this PR, sort it out
uiverify - get_build(pr: 412)
2 changed · 1 likely regression
uiverify - get_diff("Builds table")
LIKELY REGRESSIONtable text rendered green, no intended change matches it
A debug colour left on the builds table. Reverting the token.
Update(dashboard/tokens.css)
- color: #16a34a;
+ color: var(--foreground);
uiverify - accept_build
check green · merge unblocked
3 · straight to the agent, over MCP
pillar 05 · the review loop

A second model, reviewing every PR.

you / your agent the PR reviewer model opens reviews

Same context as the coding agents - CLAUDE.md, the architecture docs - plus "go deep on the logic." If it's clean, the bot approves and it can merge.

the skills that make it real

/review-loop

Two models review the same diff. You triage and fix. It re-reviews - round and round until a clean pass finds nothing new.

Claude + Codex review the diff read · triage · fix apply the findings reviews fixes
the skills that make it real

/e2e-verify

your stack on unique ports →Playwright spins a real browser →clicks + screenshots the app
Aagent-a ·billing
localhost:5173
Bagent-b ·onboarding
localhost:5174
Cagent-c ·search
localhost:5175

Ten agents, ten isolated stacks - no port collisions. This is what catches the ones that only say they're done.

the skills that make it real

/babysit-pr

Opens the PR, then babysits CI - it polls, and fixes whatever goes red. Never approves, never merges.

Billing table redesignOpen
#128 · agent-a · opened just now
✓ lint
✓ unit tests
✕ visual - a diff to triage
⟳ e2e - still running
polls every ~10 min · reads the red · pushes a fix
the skills that make it real

The run makes the next run better.

/add-rule

A correction becomes a CLAUDE.md line and, when it's mechanical, a lint guard that goes red the next time.

/evaluate

Grades the finished run cold - a fresh model that never saw it asks how well the loop actually ran, and what to tighten.

the umbrella skill · one sentence in

/factory

/factory drives every station rules · docs · guards implement + tests /review-loop /e2e-verify /babysit-pr fixes PR findings green + approved merge → prod
afterwards, every correction →/add-rule ·/evaluate →a new rule + a new guard
the one thing to take home

There's always something
to improve.

1The advisor. A chat tab on the side. You ask, it answers, you write every line.
↓
2Pair programming. It drafts the boilerplate. You still rewrite the parts it got wrong by hand.
↓
3Big PRs - you're the bottleneck. It writes 40 files. You review all of them. "No, not like that."
↓
4Autonomous. You plan the architecture, it implements, you review. Few issues come back.
↓
5The software factory. You name the feature. You never read the code.

Wherever you are on this, there is a next step. Go take it.

scan it - take all three home

Your gift, LondonJS.

🛠️

The skills

Every skill from this talk - the deck-builder included. Free to install.

UI Verify

The whole platform. Free for early adopters.

📊

These slides

The deck itself, yours to keep and steal from.

Let's connect

Igor Luchenkov on LinkedIn - in/igor-luchenkov

uiverify.ai/londonjs

in/igor-luchenkov · Igor Luchenkov · LondonJS 2026 · thank you 🙏