Skip to content
tomo-labs

20

July 2026

Monday, 20 July

23:20 gpt-5.6-luna, LeetCode agent board The same gpt-5.6-luna model solves the same recent easy, medium, and hard LeetCode tasks through leetcode-solver, tomo, pi, opencode, Codex, and Claude Code. 01:30 qwen3-30b-a3b, local board Second model on the local roster board. qwen3-30b-a3b is a general MoE, not a coder tune, and it runs on the RTX 4090 behind the llmgw gateway driven through the uniform oi code-as-action harness. It scores 1 of 15. 01:15 qwen3-coder-30b-a3b, local board First model on the local roster board. qwen3-coder-30b-a3b runs on the RTX 4090 behind the llmgw gateway and is driven through the same uniform oi harness as the free zen models, over tailnet. It scores 1 of 15. 01:00 north-mini-code-free, fair board north-mini-code-free is the fourth and last free zen model on the abort-aware oi harness across all fifteen swebench-live tasks, and it is the one that breaks the pattern. It scores 0 of 15. 00:45 nemotron-3-ultra-free, fair board nemotron-3-ultra-free is the third free zen model on the abort-aware oi harness across all fifteen swebench-live tasks. It scores the same 3 of 15 as the other two free models, and it does it the hard way: it hits the thirty-round ceiling on every single task, spends 5. 00:30 mimo-v2.5-free, fair board mimo-v2.5-free is the second free zen model taken through the abort-aware oi harness across all fifteen swebench-live tasks. It scores the same 3 of 15 as deepseek-v4-flash-free, but the passes are a different set and the failure shape is the mirror image. 00:15 deepseek-v4-flash-free, fair board The earlier read on the free zen models was that deepseek-v4-flash-free never produced a clean multi-task pass. That read was an artifact of the free tier, not the model.