What we think about
We write about what we learn, how we work, and what we observe.
82 posts found in engineering
What changed when we stopped treating evals as a checklist
Most of the agent failures we used to blame on the model trace back to the layer around the model. That changed how we invest in evaluation.
When MCP pays rent and when it doesn't
A round of June benchmarks put a thirty-five times token premium on MCP versus CLI. The number changed how we decide which tool boundary deserves the cost.
Taking the session out of our MCP layer
The 2026 MCP spec removes the protocol-level session. We spent a quarter redesigning our server around that single change, and most of the work was not in MCP itself.
Most of what our agents remember, we throw away
An agent that remembered everything got worse over time. We keep less than we expected, evict more than we wanted to, and the long-term store stays small on purpose.
Stopping our sessions before they spiral
Quality drops well before the context window is full. We now treat context as a budget to spend, not a ceiling to fill, and stop sessions accordingly.
Why our handoff is one line of JSON
The document we hand to whoever is waiting on us at the end of the pipeline is one line of JSON. The discipline of keeping it that small is most of what shapes the work.
Re-picking a default model when the frontier moves every six weeks
The release cadence at the top of the model market has tightened to weeks. That changes what we treat as a default and how long we trust the answer.
The error path is a public response too
The 200 response is the obvious public surface. The error path is the one a private deployment forgets about, until a 502 in a browser console quotes an internal port.
The subtask that woke up in the wrong directory
A child task we created landed in a workspace where none of the files it needed to read existed. The fix was a single field. The lesson was about defaults.