2026-08-22
Classify by the interaction edge, not the symptom, and trace back to the earliest
failure nobody recovered from. Plus the six checks worth running before you blame the
model — and where the new taxonomy is weaker than it looks.
2026-07-31
A status report generated by the same process being checked isn't evidence — whether
that process is a code review or a health check. What our own fleet's audit caught in
nine days, and what slipped through when the check was missing.
2026-07-28
An orchestrator started mixing up its own tasks mid-run. The instinct is the kill
switch — and it's the expensive move. What atomicity, sagas and SRE load-shedding say
about stopping a fleet safely, and the cases where you hit the switch anyway.
2026-07-24
A factorial study, a presence-effect RCT, and our own fleet's task history, argued
through to a specific design: what actually earns a place in an agent's instruction
file, and the two-variant template built from it.
2026-07-22
Sessions can now run in their own git worktree instead of a shared checkout. What a
worktree actually buys an agent, the real disk/complexity cost, and the internal
incident that decided the cleanup mechanism had to ship in the same release as the
isolation, not later.
2026-07-18
Running one Claude Code or Codex session is a solved problem. Somewhere around the
second or third at once, something shifts — and it's usually not the agents themselves
that start costing you time.
2026-07-18
A live benchmark of agent-memory systems against a bare-model floor and a flat-file
baseline — 6 live-tested, 7 tiered. One leader (narrowly), one real surprise, no
vendor claims taken at face value.
2026-07-18
Our standing rule for the harness equipment registry: a claim from a summary or a
README never gets cataloged as fact until it's re-confirmed against the primary
source. What that looks like in practice, and why we hold our own product to the same
bar.