Essays on AI, craft, and the future of software.
Your company has 47 agents running in production. You know about 12 of them. Here's why agent sprawl is the new shadow IT, and the governance architecture that prevents it from becoming a security incident.
You can't unit test a system whose output is non-deterministic and whose execution path varies on every run. But you can build an evaluation pipeline that catches regressions before users do. Here's how.
MCP for tools, A2A for coordination, AGENTS.md for context. The three protocols becoming the TCP/IP of agentic AI, and how to build on them today.
If your agent is 85% accurate per step, a 10-step workflow succeeds just 20% of the time. Here's the math that kills agent demos, and the architecture that survives production.
The best interface is no interface. How we measure success by what users stop noticing, and the counterintuitive design principles that get you there.
How we design systems where multiple AI agents coordinate, delegate, and recover from failure without a human in the loop for every decision.
Everyone's building agents. Few understand why they break. We dig into the loop (plan, act, observe, repeat) and where the failure modes hide in production.