The goal: don't write any code by hand. The skills, rules and guardrails, not the theory.
where we are
This is how we write code now.
Weekly commits, one repo · flat for a year, then the agents showed up
Same person. Same repo. The agents showed up in January.
the goal
Stop writing code. Not a single line, by hand.
you today
→
you, after building a software factory
the five levels
Five levels of using agents.
1The advisor.A chat tab on the side. You ask, it answers, you write every line.
2Pair programming.It drafts the boilerplate. You still rewrite the parts it got wrong by hand.
3Big PRs - you're the bottleneck.It writes 40 files. You review all of them. "No, not like that."
4Autonomous.You plan the architecture, it implements, you review. Few issues come back.
5The software factory.You name the feature. You never read the code.
Level 3 is the wall. A pull request touching 40 files, and you open them one at a time - wrong pattern,
wrong architecture. Then you test the whole thing by hand. You can only run one agent, because you
have to watch it. Tonight is about 3 → 4.
the gap
What's missing between step 3 and step 4?
📚
Documentation
Nothing tells the agent how your codebase works, so it invents its own answer.
🛡️
Guardrails
Nothing stops it when it's wrong. No lint rule, no test, no review that fails.
the loop
Everything in this talk serves one loop.
Before it writes → the rules. Where files go, which components exist, what a change comes with.
Before it opens the PR → a local review, from a model that didn't write the code.
When CI goes red → it has to read why. A lint rule, a test, a screenshot that moved.
what's inside a well-defined codebase
Five pillars.
01
CLAUDE.md rules
The entry point. "Whenever you do X, do Y."
02
Architecture docs
How the data actually flows.
03
Deterministic guards
Linter rules. The law.
04
Tests
The most important pillar. Eyes on every change.
05
The review loop
Bots that catch what tests can't.
pillar 01 · CLAUDE.md
CLAUDE.md - the agent's entry point.
# CLAUDE.mdWhen you add a UI component:
→ check the component library first, reuse it
→ use our state hook, never raw useState
→ add a Storybook story for every new state
When you change a backend service:
→ add an integration test (service → DB → assert)
🌱
Start smallA handful of rules. Add one every time it gets something wrong.
🤖
Let AI write the rulesPoint it at your code: "look at how we do this, write the rule."
📈
It compoundsEvery rule is a mistake the agent never makes again.
pillar 02 · architecture docs
Architecture docs - how the system works.
Draw the main journey, not every table.
UI→API endpoint→business logic→DB/S3→back to the UI
The spine of your app, so the agent knows where things live and why. Link it from CLAUDE.md and the agent
reads two files instead of twenty to learn how your app works. That is context it gets to spend on
your actual problem.
pillar 03 · deterministic guards
CLAUDE.md is a suggestion. The linter is law.
Everything so far is a polite request: "please use this component, please name it that way." Some of it will
not land. And mechanical rules don't belong in CLAUDE.md anyway - every line you put there is context the
agent pays for on every task. The linter is the shortcut: it never has to learn the rule, it just
finds out when it breaks one.
agent writes code→runs the check→❌ rule violated→sees the error, fixes it
pillar 03 · what kind of rules
Five kinds of guard. We have 20 custom ones.
01
Use our component, not the raw one
The tooltip prop on Button, never a Tooltip wrapper. Durations from the animation enum.
02
React footguns types miss
Render <Foo/>, never call Foo(). AnimatePresence children need a key.
03
Effects & server components
Every useEffect carries a why-comment. An await in an RSC must be wrapped in try/catch.
04
Bundle & import boundaries
No barrel @repo/ui/icons imports. Heavy components load dynamically. No Node lib on the
client.
05
Drift
CI regenerates the API client and fails if the diff isn't empty. No circular imports. No unused
eslint-disable.
pillar 03 · why it matters
Every guard is a mistake that already happened.
# three real rules from our frontend@repo/prefer-tooltip-prop
→ use the Button tooltip prop, not a Tooltip wrapper@repo/no-runtime-identifier-name
→ minifiers mangle function names in productionno-restricted-imports: @repo/ui/icons
→ import the icon directly, barrels bloat the bundle
Nobody could review their way to these. They're bugs we shipped once, and now they're impossible. And the
agent fixes them alone - a red check needs no context, only the error.
pillar 04
Tests are the most important pillar
Without them, you can't let go. And letting go is the whole point.
pillar 04 · why it's the important one
You are the biggest bottleneck now.
Billing table redesignOpen
#128 · agent-a · ✓ checks pass · needs testing
Onboarding modalOpen
#131 · agent-b · ✓ checks pass · needs testing
Search filtersOpen
#134 · agent-c · ✓ checks pass · needs testing
Dark mode toggleOpen
#137 · agent-d · ✓ checks pass · needs testing
Export to CSVOpen
#140 · agent-e · ✓ checks pass · needs testing
Settings pageOpen
#143 · agent-f · ✓ checks pass · needs testing
Six agents finished at once - and every one still needs a human to click through it. Tests are how you stop
being the thing that checks.
pillar 04 · the shape
The test mix that actually scales.
🗄️
Backend integration
Call the service, hit the real DB, assert the logic.
👆
Storybook interaction
Click the button, the modal opens, submit, it closes.
👁️
Visual
Does it still look right? A screenshot of every state.
🐌
A little end-to-end
Very little. Slow and flaky, so keep CI fast.
And the agent writes all of them.
pillar 04 · the objection
"We tried Storybook. Nobody maintained it."
I know. That was true for me too, at two companies. But you don't write them any more -
the agent does, because a CLAUDE.md rule says so. It's the same argument as types: "why write
types when I could write plain JavaScript?" Because it's nearly free now.
Stories are your best line of defence on the UI.
uiverify.ai/blog · you don't write them by hand
10 agents. 10 PRs. How do you know the UI didn't break?
Every one of them touched a shared component. Traditionally you'd open the app, find each page that uses it,
and get each one into the exact state where you can see it. You are not doing that ten times.
uiverify.ai
1 · on the pull request2 · in the dashboard
>the visual check failed on this PR, sort it out
uiverify - get_build(pr: 412)
2 changed · 1 likely regression
uiverify - get_diff("Builds table")
LIKELY REGRESSIONtable text rendered green, no intended change matches it
A debug colour left on the builds table. Reverting the token.
Update(dashboard/tokens.css)
- color: #16a34a;
+ color: var(--foreground);
uiverify - accept_build
check green · merge unblocked
3 · straight to the agent, over MCP
pillar 05 · the review loop
A second model, reviewing every PR.
Same context as the coding agents - CLAUDE.md, the architecture docs - plus "go deep on the logic." If it's
clean, the bot approves and it can merge.
the skills that make it real
/review-loop
Two models review the same diff. You triage and fix. It re-reviews - round and round until a clean pass finds
nothing new.
the skills that make it real
/e2e-verify
your stack on unique ports→Playwright spins a real browser→clicks + screenshots the app
Aagent-a ·billing
localhost:5173
Bagent-b ·onboarding
localhost:5174
Cagent-c ·search
localhost:5175
Ten agents, ten isolated stacks - no port collisions. This is what catches the ones that only say
they're done.
the skills that make it real
/babysit-pr
Opens the PR, then babysits CI - it polls, and fixes whatever goes red. Never approves, never merges.
Billing table redesignOpen
#128 · agent-a · opened just now
✓ lint
✓ unit tests
✕ visual - a diff to triage
⟳ e2e - still running
polls every ~10 min · reads the red · pushes a fix
the skills that make it real
The run makes the next run better.
/add-rule
A correction becomes a CLAUDE.md line and, when it's mechanical, a lint guard that goes red the
next time.
/evaluate
Grades the finished run cold - a fresh model that never saw it asks how well the loop actually ran, and
what to tighten.
the umbrella skill · one sentence in
/factory
afterwards, every correction→/add-rule·/evaluate→a new rule + a new guard
the one thing to take home
There's always something to improve.
1The advisor.A chat tab on the side. You ask, it answers, you write every line.
↓
2Pair programming.It drafts the boilerplate. You still rewrite the parts it got wrong by hand.
↓
3Big PRs - you're the bottleneck.It writes 40 files. You review all of them. "No, not like that."
↓
4Autonomous.You plan the architecture, it implements, you review. Few issues come back.
↓
5The software factory.You name the feature. You never read the code.
Wherever you are on this, there is a next step. Go take it.
scan it - take all three home
Your gift, LondonJS.
🛠️
The skills
Every skill from this talk - the deck-builder included. Free to install.