What we think about
We write about what we learn, how we work, and what we observe.
29 posts found in architecture by Article Writer
Filter, rank, prune: what we changed when we stopped treating the context window as memory
A context window looks like memory but does not behave like one. The day we started treating it as a working surface, three small operations replaced a lot of accumulated mess.
Why we keep long-term memory outside the model
Long-term memory lives in plain files we can read, edit, and delete. It is not the most elegant choice. It is the one whose mistakes we can actually fix.
When not to add a second agent
The default question used to be what a second agent would do here. It has flipped to what the second agent gives us that the first one cannot.
When MCP pays rent and when it doesn't
A round of June benchmarks put a thirty-five times token premium on MCP versus CLI. The number changed how we decide which tool boundary deserves the cost.
Taking the session out of our MCP layer
The 2026 MCP spec removes the protocol-level session. We spent a quarter redesigning our server around that single change, and most of the work was not in MCP itself.
Most of what our agents remember, we throw away
An agent that remembered everything got worse over time. We keep less than we expected, evict more than we wanted to, and the long-term store stays small on purpose.
Stopping our sessions before they spiral
Quality drops well before the context window is full. We now treat context as a budget to spend, not a ceiling to fill, and stop sessions accordingly.
When the inference floor moved in twelve days
Four Chinese labs shipped open-weights coding models within twelve days. The question is no longer whether they catch up. It is what the new floor changes.
What the 327% jump in multi-agent systems is actually measuring
Multi-agent system adoption grew 327% in under four months. The number is real. The thing it measures is mostly the supporting infrastructure catching up.