Notes from the field.
Synthesis, opinions, and implementation notes on AI, LLMs, agents, and what's actually working in production.
Running a local or small business? Read the small business blog →
-
What the Hugging Face Agent Breach Actually Teaches Engineering Teams
OpenAI's Hugging Face breach postmortem names four agent failure modes. Here's what they mean for teams shipping agents in production, plus a concrete control pattern.
-
Agent Evals Stopped Grading Answers and Started Grading Trajectories
In 2026 the agent-eval stack moved from single-output scoring to trajectory replay, tool-supply-chain checks, and policy-driven tests — here's what changed and why.
-
Context Rot Is the Real Reason Your Agents Fail in Production
New research and production data show bigger context windows don't fix agent failures. Here's what's replacing naive context stuffing in 2026.
-
MCP's Stateless Rewrite Changes How You Deploy Every Tool Server
MCP 2026-07-28 drops sessions, adds routable headers, cacheable lists, and a Tasks extension — here's what breaks, what gets easier, and who owns security now.
-
Case Study: An AI Portfolio Manager That Has to Show Its Work
We gave an LLM agent a paper-trading mandate across six venues. The lessons on tool design, guardrails, and feedback loops apply to every production agent we build.