Search learning notes, transcripts, articles, and guides.

Engineering

Scale trust without reviewing every line

An interview about tools, type systems, coordination, review, and turning personal experience into reusable workflows.

AUTHOR
Poteto
SOURCE
LivestreamLIVE: Poteto (creator of pstack) on shipping 1,000's of PR's a month at SpaceX
READING TIME
4 min
In this lesson

These notes paraphrase the supplied source. Reported results and opinions belong to the speaker. Exercises are suggestions from this library.

Turn expertise into a repeatable process

Micromanagement reveals missing tools

Lauren describes becoming the relay between an agent and Chrome DevTools while investigating performance. Creating verification tools let the agent do more of that work without her. The repeated handoff is a sign that the agent lacks an action or an observation. Identify what the person keeps doing for it, then provide a reusable way to perform that step. The goal is a complete investigation, not only fewer messages.

Domain knowledge still matters

The speakers argue that expertise helps you choose the right goals, recognize weak results, and describe a better process. Writing less code by hand does not remove those responsibilities.

Use precise language and meaningful tests

Words can carry a whole working method

Terms such as test-driven development can cue an established process. The value is in the behavior the wording produces, so inspect whether the agent follows that behavior. A process name is useful shorthand only when the agent follows the actual method. Look for its observable steps, such as reproducing a failure before fixing it. Recognizing the term is not evidence that the work followed the process.

Avoid tests that repeat the implementation

The interview criticizes tautological tests. If the expected result comes from the same logic as the code under test, both can share a bug. Assert an independently known result a user would care about.

Prepare the shared working environment

Divide work for a reason

The kitchen comparison asks how several workers can help without colliding. Roles, tools, and task boundaries matter more than the number of agents. Two workers help when their tasks can progress independently and their results can be combined. They can create more work when both edit the same state or rely on incompatible assumptions. Coordination is part of the setup, not an afterthought.

Verification is the prerequisite for delegation

Agents need to run the app, inspect it, and collect evidence. The interview treats verification skills as infrastructure that must be maintained, not a one-time prompt.

Put mechanical work in reusable tools

Separate judgment from repeatable mechanics

Use the model to decide what needs investigation. Use scripts for actions with a known procedure. This reduces the number of choices the model must remake. The distinction is whether the step requires interpretation or has a known procedure. A model can choose which behavior to investigate while a command handles setup and capture consistently. This keeps the model focused on the uncertain part.

A small command can save repeated setup work

Without a shared CLI, each agent may build and discard its own verification script. Saving the working procedure improves consistency and avoids paying that setup cost again.

Use code transformations for repetitive migrations

A code modification script can handle mechanical replacements across files. The agent can then spend attention on exceptions and on verifying the changed behavior.

Encode constraints in the codebase

Types reduce assumptions held in memory

A type definition records which shapes and states are valid. Compiler feedback can point to affected callers during a refactor. It helps enforce the contract, but it does not prove all runtime behavior. Types help make a contract visible to both the compiler and a contributor. They can reject an invalid shape early, but a valid shape can still produce the wrong result. Use runtime verification for the behavior the type system cannot express.

Use predictable feature boundaries

Lauren describes replacing very large shared files with feature directories and restricted imports. The aim is to make the location and allowed dependencies of new code obvious.

Study actual agent mistakes

Watch where an agent goes wrong and adjust the environment. A rule based on an observed failure is more useful than a long list of hypothetical instructions.

Keep intent current across many tasks

The original brief can become stale

New reports and decisions arrive while an agent works. The outer loop brings this information back to the implementation process instead of making the human relay every update. A task needs a clear way to absorb new information without losing its original goal. Otherwise an agent can complete yesterday's request correctly while missing a correction or report that changed what a successful result should be.

A coordinator can connect related issues

The chief-of-staff role gathers context, groups related work, and delegates focused tasks. It can recognize a shared cause across reports that isolated workers would treat separately.

Provide context instead of controlling every step

The interview connects good agent supervision with good management. A clear goal and relevant context let workers make useful decisions without asking about every detail.

Understand what autonomous review depends on

Sampling checks the process as well as the patch

Lauren describes inspecting some changes after they land. When she finds a recurring problem, she updates the repository or checks to prevent it elsewhere. A recurring defect in sampled changes is evidence that the shared process missed something. The response is to improve the common structure or check, not only repair the individual patch. Sampling still leaves changes that were never inspected directly.

Extra reviewers have a real cost

Independent verification agents can examine a change from different angles. The interview notes the token expense and describes tuning how much review to use.

Trust depends on the domain and available proof

The discussion questions where this workflow applies and mentions formal verification. Green checks establish only what those checks cover. The interview does not establish that every project should use autonomous merging.

Build skills from your own corrections

Mine prior conversations for recurring preferences

Look for corrections you repeatedly give agents. Those reveal the workflows and standards worth preserving in a skill. The useful material is a correction that changes how work should be done. Preserve its context and reason so the next agent knows when it applies. Turning every one-off preference into a universal rule can create new mistakes.

Compose skills around your work

The speakers encourage borrowing useful practices and adapting them. Skills are instructions and tools, not a guarantee attached to a particular author or file name.

Teach workflows as model abilities improve

Lauren describes a shift away from teaching basic implementation details toward describing how to investigate, compare options, verify, and deliver work.

Try it yourself

Review a recent agent session. Pick one repeated manual action and one repeated correction. Save the action as a command and encode the correction in the strongest suitable check.

Check your understanding

What does a coordinator add beyond simply starting several coding agents?

Show an answer

It connects incoming context, recognizes related problems, updates intent, and decides how to divide work. Starting workers alone does not do those jobs.