Give agents a way to check their work
Build repeatable controls, document user journeys, and keep verification working as the app changes.
These notes paraphrase the supplied source. Reported results and opinions belong to the speaker. Exercises are suggestions from this library.
Close the verification loop
Verification means trying the real behavior
An agent needs to run the app, perform the user action, and inspect the result. A successful build alone cannot tell you whether a button works or whether a page feels slow. Lauren treats this ability as shared team infrastructure. For a button, verification includes reaching the screen, clicking the button, and observing the expected change. For a performance fix, it includes a comparable measurement before and after. The proof must match the behavior claimed, not just the files edited.
Choose tools that expose the running app
A browser debugging connection or a simulator lets an agent inspect state, capture screenshots, and collect traces. If the runtime has poor debugging support, build the missing tools. Lauren even argues that verification quality can influence the choice of technology.
Turn repeated actions into commands
Save working automation instead of rebuilding it
A small command-line tool can open a session, send text, inspect the interface, and save evidence. Each agent then uses the same commands instead of writing a new browser script for every task. This preserves both the procedure and its entry point. The next agent can discover the command, run it against the app, and inspect its output instead of reconstructing the same setup from scratch. The command still needs maintenance when the app changes.
Design commands for recovery
Subcommands and useful help make capabilities discoverable. JSON output makes results easy to parse. Errors should explain the next action. Destructive commands should support a preview of their effects.
Make the development setup repeatable
Verification also depends on installing dependencies, starting the app, creating test data, and handling test accounts. Browser automation cannot repair an undocumented setup process by itself.
Separate execution environments
Remote agents move resource use off the laptop
Lauren favors cloud environments over many local worktrees when running large numbers of agents. Each remote environment can install dependencies and exercise the app independently. The benefit is isolation as well as extra capacity. Separate environments prevent one task's dependencies, working files, or test state from interfering with another. Remote execution still needs reproducible setup and a way to retrieve useful evidence.
Reuse a prepared environment
A snapshot after the initial setup can shorten later starts. The useful principle is repeatable isolation. The number of agents a laptop can handle depends on the workload, so her local capacity estimates are examples.
Write a map of user-visible features
Document how a person reaches each feature
A feature map lists capabilities, entry points, expected behavior, and useful automation commands. For settings, record how to open the dialog, move between tabs, search, and close it. The map starts with the user's action rather than the file that implements it. Someone checking a settings change should know which control opens settings, which state to inspect, and what result proves the change works.
Include states that change the route
Account permissions, feature flags, focus, and keyboard shortcuts can alter the same journey. Record those differences so an agent does not mistake an unavailable feature for a broken one.
Maintain the map alongside the app
A stale feature map sends every new agent down the wrong path. Lauren proposes a recurring maintenance routine. The map is useful only while its instructions match the running product.
Use evidence in everyday tasks
Connect feature work to proof
Ask for a feature and the verification needed to show it works. Screenshots, recordings, and runtime measurements make the result reviewable. The requested proof should reflect the change. A screenshot can show layout or a visible state; a recording can show the action and transition. For a performance claim, measurements need a defined baseline and comparable conditions.
Reproduce reports before changing code
A feedback-triggered agent can try the reported behavior on the current version, fix a reproduced defect, and return evidence. For performance work, establish the baseline before comparing a change.
Try it yourself
Choose one feature. Write its user steps, automate those steps with one command, and save the observed result. Run it again in a fresh session.
Check your understanding
Why is a working verification command more useful than a paragraph telling an agent to test carefully?
Show an answer
It makes the repeated operation executable and consistent. The agent still needs judgment about what to verify, but it no longer has to invent the mechanics each time.