Long-form writing on runtime, distributed systems, kernel internals, Go concurrency, memory models, eBPF, and assembly walkthroughs.
Blog

The same lock-free C program runs clean on Intel for hours and fails on ARM64 in a minute. The source is identical. This is store→load reordering — why x86's TSO hides it, why ARM64 exposes it, and what a fence actually does — with a real litmus test and its numbers.
2026-09-06
8 min read
Blog

ja, jb, je, jl don't read your operands. They read bits in EFLAGS that some earlier instruction wrote — and it isn't always the cmp you think. Once you see that in GDB, reverse engineering and crash-dump triage stop being guesswork.
2026-09-06
6 min read
Blog

Every x86 machine boots in 16-bit real mode and switches to protected mode in its first moments of life. That switch is a GDT, an IDT, one bit in CR0, and a far jump — and those tables are where ring 0 vs ring 3 actually comes from. Walked through real bootloader assembly.
2026-09-06
7 min read
Blog

kprobe and fentry both hook kernel functions from eBPF, but they install differently, cost differently, and hand you arguments differently. A source-level comparison, grounded in a real eBPF project, plus why the choice is a production cost — not trivia.
2026-09-06
9 min read
Blog

top shows 100% CPU. perf shows the truth. Two threads doing identical work run in 841ms or 8118ms depending on one thing your profile never mentions: where their data sits in memory. Cache misses, TLB misses, and false sharing, with real code and real numbers.
2026-09-06
6 min read
Blog

Rebuild an ARM64 Linux kernel with Yocto, boot it in QEMU, and step through it in GDB — on a laptop, no hardware. The workflow is the easy part; the part nobody tells you is why GDB can't find your source and how to fix it.
2026-09-06
5 min read
Blog

Codex ships 43,591 lines of sandboxing across four platform backends, and treats enforcement and approval as two separate axes. Pi ships none and says so in its README. Both are defensible. What is not defensible is the assumption in between, where a dialog box gets mistaken for a boundary. A source-level look at what your coding agent can actually do to your machine.
2026-09-04
8 min read
Blog

Claude Code's agent loop is 1,729 lines. Codex's is 983. Pi's is 794. Three teams, three languages, no shared code, and the loop lands in the same place every time. What does not converge is where each one puts its boundary, and that turns out to be the whole design. A source-level read of two harnesses that are now open, against one that leaked.
2026-09-01
12 min read
Blog

Sixty-five bytes are enough to kill llama.cpp. Not by corrupting memory, by handing the GGUF loader a tensor dimension of zero and letting a bounds check divide by it. The check was already there and looked correct. Zero is a legal dimension the format is supposed to support, which is why rejecting it was never the fix. A note on two crashes I fuzzed out of the loader, what the review changed, and why a downloaded model deserves the suspicion you give a downloaded executable.
2026-08-26
8 min read
Blog

A bug fix can be written, reviewed, merged, tagged, and correct, and still never reach a single user. The fix for a real iOS connectivity bug had existed upstream for over a year while the people hitting the bug filed me too comments on an issue, because one git submodule pointer between the library and the client was never advanced. Merging changes the source of truth. It does not move the pointers, and in a dependency graph the last move is owned by no one.
2026-08-22
8 min read