[ Blog_Posts ]

One harness treats MCP as core infrastructure and spends tens of thousands of lines on it. The other refuses it on principle and tells the agent to write its own tools instead. Both positions are coherent, and the disagreement is not about capability. It is about who owns the schema, what a stable boundary is worth, and what you pay for it in tokens and attack surface on every single request.

Read More

Codex made Item a wire type and thread/fork a protocol call. Pi gave every session entry a parentId. Claude Code compresses the transcript in five stages. Three harnesses, three answers to what the durable unit of agent state actually is, and the answer decides whether running out of context costs you a summary or costs you the work.

Read More

Pi is famous for a system prompt under a thousand tokens. I measured it: 550. Then I measured the four tool schemas that ship beside it in every single request: 588. The celebrated number is the smaller half. Meanwhile Codex ships a different prompt length for every model, from 1,436 to 5,070 tokens, and the reason why is the most useful thing in either codebase.

Read More

OpenAI changed two API settings on ARC-AGI-3 and GPT-5.6 Sol went from 13.3% to 38.3% while spending six times fewer output tokens. Same model, same weights, same benchmark. The official harness had been throwing away the model's reasoning after every move. Which means the score was never a measurement of the model, and neither is most of what we compare.

Read More