Turn workflows into prompts and verifiable systems
How natural language replaces rigid tooling, why type safety and CI gates enable autonomous agents, and how frictionless verification lets small teams scale.
These notes paraphrase the supplied source. Reported results and opinions belong to the speaker. Exercises are suggestions from this library.
Measure model reasoning against domain grammar
Evaluate reasoning through technical domain constraints
Theo created Skatebench to evaluate models by describing the physics of a skateboarding trick (board rotation axes, direction, and skater movement) and asking the model for the standard trick name. Google Gemini 3.1 Pro reached 95% on his benchmark while other frontier models scored 82% to 88%. He attributes this to a combination of 3D spatial reasoning and multilingual grammar translation.
Test new models with progressive autonomy
When evaluating a new model, Theo integrates it directly into everyday work rather than relying on synthetic benchmarks alone. He asks it to evaluate its available tools and codebase setup, before progressing from targeted bug fixes to full-stack architectural migrations.
Align tool choices with open source control
Open tools prevent misaligned incentives
Theo explains that closed development tools create organizational risk when the creator's product direction diverges from the user's workflow. Choosing open source software ensures users can inspect, fork, and adapt tools when upstream maintainers introduce breaking changes.
Compose small, reusable utilities
The Unix philosophy of combining small, focused tools informs both stack construction and agent ergonomics. Rather than building monolithic bespoke frameworks, composing well-maintained primitives with end-to-end type safety creates predictable, testable feedback loops.
Decouple the interface from background execution
Separate the user interface from agent compute
T3 Code was built after Theo and Julius experienced performance drops and rigidity in existing desktop apps. It runs a lightweight daemon and WebSocket server that binds to harness tools like Claude Code and Codex, allowing multiple clients (Electron desktop, web, or mobile) to connect over a private Tailscale network.
Offload long-running threads to dedicated hardware
Running agent sessions on remote desktop machines prevents laptops from thermal throttling, battery drain, or losing progress when closed. Developers can monitor and steer ongoing threads across multiple machines without hosting heavy builds locally.
Prioritize incomplete tasks in task lists
T3 Code structures its sidebar around unfinished work rather than chronological order, reducing the mental overhead of tracking active background tasks.
Move past skeuomorphic abstractions in agent tooling
Skeuomorphic interfaces ease initial transitions
Theo draws an analogy to iOS 6 skeuomorphism, where digital calendars resembled leather desk pads to help users trust touchscreens. Early AI tooling mimics human developer habits—rigid terminal interactions, manual Git steps, and plan modes—to make agent behavior familiar.
Advanced models outgrow artificial modes
Plan mode initially restricted models with modified system prompts to prevent premature file edits. As frontier models have grown more capable, rigid mode boundaries and complex orchestration pipelines can be replaced with direct prompts and Markdown instructions.
Structure instructions for agent amnesia
Manage the Groundhog Day effect with concise guidelines
Agents arrive at each task with no memory of prior sessions. Project instruction files (such as AGENTS.md) serve as a daily briefing letter, orienting the model to repository layout and constraints without filling its context window.
Transition repetitive skills into cloud services
While local skills offer reusable guidance, minor wording changes across models can produce inconsistent results. Theo suggests converting recurring multi-step audits into background services triggered on demand or on a schedule.
Reduce cognitive surface area with type safety
Structural guardrails protect limited context windows
Automatic transmission cars still rely on internal gears; similarly, full-stack type safety and automated linters remain essential with AI agents. Compilers catch regressions immediately, allowing agents to make changes safely without needing to load the entire repository into context.
Focus on how much an agent can succeed without knowing
In large codebases, complete global context is impractical for humans and agents alike. Strong type definitions and clear module boundaries allow agents to modify local components with high confidence.
Enforce organizational standards with natural language gates
Combine deterministic checks with natural language reviews
Linters catch syntax and structural issues, but subtle architectural rules—such as library conventions or design system compliance—often slip through. Theo uses Macroscope review agents in CI to evaluate pull requests against Markdown-defined criteria before human review.
Execute Markdown as software
Just as Node.js turned JavaScript into an executable language on servers, modern AI models make natural language and Markdown executable. Theo runs scheduled cron jobs that pipe Markdown instructions to agents to evaluate pull requests and overwrite a static HTML status dashboard.
Scale throughput by eliminating verification friction
Verification confidence determines autonomous scale
Raising confidence in automated checks from 40% to 85% enables teams to run dozens of background agents across backlogs without fear of silent regressions. Theo reports merging 28 pull requests in an evening because rigorous automated checks pre-screened every change.
Frictionless previews accelerate human decisions
Ephemeral preview environments (such as Vercel preview deployments) allow developers to inspect running changes immediately without pulling branches or clearing local ports. Removing local setup friction ensures reviewers spend time verifying user outcomes rather than troubleshooting environments.
Automate what you thoroughly understand
Theo cautions against using agents to avoid learning a domain. High-leverage automation requires deep understanding of failure modes and success criteria to establish effective guardrails.
Try it yourself
Identify a repetitive multi-step task in your project—such as auditing open pull requests or checking design system usage. Write a plain Markdown document outlining the pass/fail criteria and execution steps, then run an agent against that file to verify if natural language instructions alone can produce an accurate status summary without custom code.
Check your understanding
Why does full-stack type safety become more important, rather than obsolete, when delegating code edits to AI agents?
Show an answer
Agents operate with limited context windows and experience session amnesia on every new task, preventing them from memorizing the full codebase. Full-stack type safety provides immediate, deterministic compiler feedback when an interface changes, enabling agents to verify the ripple effects of their edits locally without needing to understand or load the entire system into context.