21
July 2026
Tuesday, 21 July
Hy3 + current Tomo, the 15-task baseline
The post-LeetCode Tomo OI baseline runs all fifteen offline SWE-bench-Live tasks on hy3-free at pass@1. It solves five, including Kubernetes Python after a ten-minute first completion, and records 1.72 million provider-reported tokens. That number is a lower bound: 29.
Tomo OI reverses the Pi LeetCode cost gap
The first Luna LeetCode board found Tomo correct but expensive: twelve to fourteen model calls and up to 63 thousand tokens per problem.
TAOCP partial GPT matrix
A deliberately stopped TAOCP solver experiment shows that two full proof audits consumed more tokens and list-equivalent cost than solution generation. The completed paired cases also show no quality gain from slow mode despite 7.57 times the generation tokens.