AI architect. I build agent tooling and measure it: benchmarks with published losses, silent-failure hunts on real repos.
Latest: an A/B retrieval benchmark of memory stacks — harness-comparison, cited in claude-mem #3693, write-up at ai-architect.tools/notes.
Open source: DeusData/codebase-memory-mcp #1832 — open upstream PR proposing Markdown-to-file REFERENCES_FILE graph edges, with focused tests, after a documentation fan-in blind spot surfaced in my A/B retrieval work.
Each repository is having it's own traced implementation and explanation on how they work and what they do: Cortex · Cortex Viz · Session Optimizer · Codebase · Spec · Zetetic Agents.
The losses are in there too: one evaluation suite returns NOT PRODUCTION-GRADE against its own target, one commit gate can be configured off, one session banner states three wrong numbers out of four, these are making object of subject I'm treating now.
Tools: Cortex · Zetetic Agents · ai-architect.tools




