Skip to content
tomo-labs

17

July 2026

Friday, 17 July

19:30 kata first live numbers The new kata engine ran its first real workloads against tomo-oi on hy3-free, like for like: same binary, same fence parsing, only the loop policy differs. Core-14 came out level on passes (13/14 each) with kata 28% leaner on total tokens and faster on wall clock. 18:05 M0 slice zero estate audit Spec 2105's M0 starts by reconciling what the June experiment journal says shipped against what the tomo tree actually carries. The audit checked one identifiable symbol per patch set across the committed tree and git log -S. 07:45 swebench-live, the six-wall ceiling The campaign to solve all fifteen swebench-live tasks with tomo-oi and be the cheapest tool in the lab lands at nine solved, and this is the write-up that proves the other six are not a tomo gap but a property of how those benchmark instances were cut. 06:30 hy3 gitingest-94 six tools The same free model on the same task through six tools. Two of tomo's engines pass, and both cost a fraction of codex and claude-code. 05:30 hy3 three-tool A/B, tomo-oi fixed to pass The reference column set up a fair fight on the ground the product cares about, so this slice runs it: one free model, hy3-free, through three tools, pi and opencode and tomo-oi, on one task, gitingest-94, in the same isolated harness. 03:00 codex-real reference column This slice steps away from the tomo-oi campaign to pin a reference column: real codex, the Rust CLI on a ChatGPT subscription, run against all fifteen swebench-live tasks on gpt-5.6 at medium effort, one graded pass each, in the same isolated harness. 02:00 briefcase-2085, a free model solves it The campaign's third slice runs the free roster on briefcase-2085, a well-framed git-config bug where the issue names the failing call and even proposes the fix. This one is neither the harness's fault, as the first task was, nor a diagnosis trap, as the second was. 00:30 sqllineage-661, the flagship also misses The free models could not solve sqllineage-661, and all of them patched the public entry point instead of the parser where the bug lives. The obvious next question is whether a stronger model closes it, so tomo-oi ran it on the three gpt-5.