July 2026
Experiment reports captured in July 2026, by day. Each is one tool on one task, pinned to the exact build so it reproduces.
Runs captured in July 2026, grouped by the day they ran. Each item below is one experiment: one tool, one task, one verdict. The list is generated from each report's captured date, so every run for the month shows here in order with nothing to keep in sync by hand.
July 2026
Friday, 24 July
Decomposer fixes the failure mode
Experiment 0082. 0081 found the wall on dynaconf-1225 was decomposition: a thirteen-item port handed to a free model as one red wall made it a many-feature porter that edited a dozen files, broke test collection, and never converged.
Focus lands the patch, test-writing is the next wall
Two runs bracketed the model range on dynaconf-1225 and failed identically: a flagship model at maximum reasoning effort and a free deepseek model both read the thirteen-item port issue as one wide surface, explored every file, and finished with an empty patch, zero of five.
Stacking levers regresses, model wall
Scope and issue-example each landed a single lever at two of five on dynaconf-1225, both zero-regression, both stalling on the same three py-module-path cases.
Regression guard, inert on a checklist port
Experiment 0081. The reproduction gate holds a coding turn to a red-to-green, but it says nothing about what else the fix broke, so the regression guard runs the project's own tests before the loop and refuses to converge if a fix regresses one that was green.
Issue-example gate matches scope ceiling
The last non-tailored harness lever for dynaconf__dynaconf-1225, tested after trace inspection localized the remaining three failures to a model-reasoning ceiling rather than retrieval.
tomo-oi + testgen, deepseek-flash-free, dynaconf-1225
The previous run on this task ended with a prediction: a cheap model told to write the test first skips it and drifts to the cheapest checklist item, so move test authoring into the harness. This run does that.
Thursday, 23 July
Scope gate beats baseline on dynaconf-1225
The first lever in this arc to beat the one-of-five baseline on dynaconf__dynaconf-1225. After two rounds proving the reproduction gate is a regression, the reading was that the failure is sprawl and scope, not verification.
3-arm gate A/B on dynaconf-1225
The re-run of the verify A/B after the plumbing was fixed, three arms at pass@1 on dynaconf__dynaconf-1225 under gpt-5.6-luna: baseline, verify directive on, and verify plus the harness-side executing-check gate.
verify directive A/B on dynaconf-1225 (corrected)
A retraction. The first pass@1 A/B of tomo's verify directive on dynaconf__dynaconf-1225 reported zero of five with the directive off against two of five with it on, and read that as the directive working. It was not the directive.
Wednesday, 22 July
tomo-oi + gpt-5.6-luna on dynaconf-1225
The third luna note on dynaconf-1225: tomo's oi engine driving gpt-5.6-luna, with the symbol-anchored context pack resolving symbols through pyright. The pack does its job, it points the model at loaders/__init__.
pi + gpt-5.6-luna on dynaconf-1225
The second luna note on dynaconf-1225: the pi CLI driving gpt-5.6-luna through the subscription bridge, same faithful container.
codex + gpt-5.6-luna on dynaconf-1225
The same faithful SWE-bench-Live container, now driving the real Codex CLI on gpt-5.6-luna through the subscription bridge, on the same unsolved task dynaconf-1225. This is the first of three luna notes that hold the model fixed and swap the harness.
tomo-oi + gpt-5.6-sol on dynaconf-1225
The same faithful SWE-bench-Live container and the same paid model gpt-5.6-sol, driving tomo's code-as-action oi engine on the same unsolved task dynaconf-1225.
tomo-agent + gpt-5.6-sol on dynaconf-1225
The same faithful SWE-bench-Live container and the same paid model gpt-5.6-sol, now driving tomo's own agent engine on the same unsolved task dynaconf-1225. tomo-agent speaks chat/completions, so the subscription bridge translates it to the Responses wire.
codex + gpt-5.6-sol on dynaconf-1225
The same faithful SWE-bench-Live container that ran the deepseek three-way, now running the real Codex CLI on gpt-5.6-sol through a subscription bridge, on the same unsolved task dynaconf-1225.
faithful swebench-live container, deepseek three-way
We were grading swebench-live wrong. The old path built one shared image and ran every task in a host venv pinned to Python 3.12, which is not the environment the task ships with.
laguna-s three-way comparison
Poolside's Laguna-S-2.1 is a 118B mixture-of-experts coder, free to call on the opencode.ai/zen tier as laguna-s-2.1-free.
pi + laguna-s, incomplete
Third of the three per-tool boards on laguna-s-2.1-free, and the one that did not finish. pi was the last tool in the sweep, and by the time it started the free zen account was already deep into its rate-limit window from the two streams ahead of it.
opencode + laguna-s, partial board
Second of the three per-tool boards on laguna-s-2.1-free. OpenCode, the containerized coding agent, runs the same swebench-live tasks through the same free zen endpoint.
tomo-agent + laguna-s, fair board
Poolside shipped Laguna-S-2.1, a 118B mixture-of-experts coder, and opencode.ai/zen serves it free as laguna-s-2.1-free. This runs the whole fifteen-task swebench-live board against it through tomo's own agent engine, the one that drives native structured tool_calls.
Tuesday, 21 July
Hy3 + current Tomo, the 15-task baseline
The post-LeetCode Tomo OI baseline runs all fifteen offline SWE-bench-Live tasks on hy3-free at pass@1. It solves five, including Kubernetes Python after a ten-minute first completion, and records 1.72 million provider-reported tokens. That number is a lower bound: 29.
Tomo OI reverses the Pi LeetCode cost gap
The first Luna LeetCode board found Tomo correct but expensive: twelve to fourteen model calls and up to 63 thousand tokens per problem.
TAOCP partial GPT matrix
A deliberately stopped TAOCP solver experiment shows that two full proof audits consumed more tokens and list-equivalent cost than solution generation. The completed paired cases also show no quality gain from slow mode despite 7.57 times the generation tokens.
Monday, 20 July
gpt-5.6-luna, LeetCode agent board
The same gpt-5.6-luna model solves the same recent easy, medium, and hard LeetCode tasks through leetcode-solver, tomo, pi, opencode, Codex, and Claude Code.
qwen3-30b-a3b, local board
Second model on the local roster board. qwen3-30b-a3b is a general MoE, not a coder tune, and it runs on the RTX 4090 behind the llmgw gateway driven through the uniform oi code-as-action harness. It scores 1 of 15.
qwen3-coder-30b-a3b, local board
First model on the local roster board. qwen3-coder-30b-a3b runs on the RTX 4090 behind the llmgw gateway and is driven through the same uniform oi harness as the free zen models, over tailnet. It scores 1 of 15.
north-mini-code-free, fair board
north-mini-code-free is the fourth and last free zen model on the abort-aware oi harness across all fifteen swebench-live tasks, and it is the one that breaks the pattern. It scores 0 of 15.
nemotron-3-ultra-free, fair board
nemotron-3-ultra-free is the third free zen model on the abort-aware oi harness across all fifteen swebench-live tasks. It scores the same 3 of 15 as the other two free models, and it does it the hard way: it hits the thirty-round ceiling on every single task, spends 5.
mimo-v2.5-free, fair board
mimo-v2.5-free is the second free zen model taken through the abort-aware oi harness across all fifteen swebench-live tasks. It scores the same 3 of 15 as deepseek-v4-flash-free, but the passes are a different set and the failure shape is the mirror image.
deepseek-v4-flash-free, fair board
The earlier read on the free zen models was that deepseek-v4-flash-free never produced a clean multi-task pass. That read was an artifact of the free tier, not the model.
Sunday, 19 July
fonttools-3682, local MoE
Running tomo against a small quantized model on a single desktop GPU, fonttools-3682 was the fastest and leanest pass of the local run: five rounds, 18.1k tokens, 127 wall seconds. The model went grep, read, edit, done, with no wrong turns and no re-reads.
gitingest-94, local MoE
The local-4090 counterpart to the five-free-models study on the same task. This run drives tomo with qwen3-30b-a3b, a 4-bit Qwen3-30B-A3B MoE served by Ollama on a single RTX 4090 behind the llmgw gateway, on cyclotruc gitingest-94, the easiest task in the swebench-live set.
conan-17123, local MoE
conan-io__conan-17123 is a real SWE-bench-Live feature request: .conanignore should support inverse matching with a leading !, the way .gitignore and .dockerignore do, so a config repo can ignore everything and then re-include a few files.
briefcase-2085, local MoE
beeware__briefcase-2085 is a real SWE-bench-Live task: Briefcase fails to roll out templates when a user's Git config rewrites HTTPS to SSH with insteadOf, because it calls remote.set_url with an old_url that no longer exists.
Friday, 17 July
kata first live numbers
The new kata engine ran its first real workloads against tomo-oi on hy3-free, like for like: same binary, same fence parsing, only the loop policy differs. Core-14 came out level on passes (13/14 each) with kata 28% leaner on total tokens and faster on wall clock.
M0 slice zero estate audit
Spec 2105's M0 starts by reconciling what the June experiment journal says shipped against what the tomo tree actually carries. The audit checked one identifiable symbol per patch set across the committed tree and git log -S.
swebench-live, the six-wall ceiling
The campaign to solve all fifteen swebench-live tasks with tomo-oi and be the cheapest tool in the lab lands at nine solved, and this is the write-up that proves the other six are not a tomo gap but a property of how those benchmark instances were cut.
hy3 gitingest-94 six tools
The same free model on the same task through six tools. Two of tomo's engines pass, and both cost a fraction of codex and claude-code.
hy3 three-tool A/B, tomo-oi fixed to pass
The reference column set up a fair fight on the ground the product cares about, so this slice runs it: one free model, hy3-free, through three tools, pi and opencode and tomo-oi, on one task, gitingest-94, in the same isolated harness.
codex-real reference column
This slice steps away from the tomo-oi campaign to pin a reference column: real codex, the Rust CLI on a ChatGPT subscription, run against all fifteen swebench-live tasks on gpt-5.6 at medium effort, one graded pass each, in the same isolated harness.
briefcase-2085, a free model solves it
The campaign's third slice runs the free roster on briefcase-2085, a well-framed git-config bug where the issue names the failing call and even proposes the fix. This one is neither the harness's fault, as the first task was, nor a diagnosis trap, as the second was.
sqllineage-661, the flagship also misses
The free models could not solve sqllineage-661, and all of them patched the public entry point instead of the parser where the bug lives. The obvious next question is whether a stronger model closes it, so tomo-oi ran it on the three gpt-5.
Thursday, 16 July
sqllineage-661, capability not harness
The campaign's second slice baselines the five free zen models on sqllineage-661, the next-easiest swebench-live task. Where the first task was harness-bound, this one is the opposite. No free model solves it, and not one of the clean failures is the harness.
gitingest-94, five free models
Starting a campaign to solve all fifteen swebench-live tasks with tomo-oi and be the cheapest tool in the lab, the first slice baselines the five free zen models on the easiest task, cyclotruc gitingest-94.
python-2303, five free models
The earlier python-2303 run left the five free zen models unmeasured, since the free tier was rate limited all session, and it guessed the defensive-coding paragraph would matter most for a weak model. This run measures them.
python-2303, code-as-action wins
An earlier run called kubernetes-client python-2303 an unwinnable hidden-contract coin flip for tomo-cx, the structured-tools engine, which lost every one of five attempts on three gpt-5.6 models.
new OI is a Codex fork, ~2x the cost
Open Interpreter's Python code-as-action loop ended at 0.4.2, and its main line is now a Rust program, a fork of OpenAI Codex tuned for low-cost models. The lab's openinterpreter column now tracks that rewrite, release rust-v0.0.
oi dialect zoo, four fences
A cheap model told to write a Markdown code fence keeps its shell or python command but wraps it in a different costume from turn to turn, and tomo's oi engine only read the Markdown one.
oi glued fence, a cheap-model harness fix
tomo's oi engine acts by writing fenced code blocks, so its block parser is on the hot path of every round. A cheap model routinely writes a closing fence glued straight onto the next opening fence, with no blank line between them, so two fence lines arrive as one.
bridge cache hole and prompt_cache_key
The 2026-07-15 dynaconf run named the bridge's zero percent cache-read rate as its most actionable finding, so this run tries to close it.
cx compaction mirage under prefix cache
Tomo's cx engine rebuilds every request as the whole conversation plus the running turn, so an offline replay of six recorded runs showed a 76 to 86 percent cut in wire bytes once older tool results are stubbed. That headline did not survive the live check.
Wednesday, 15 July
oi vs agent, reachability is the model
Run the same cheap model through tomo's two engines, the oi code-as-action surface and the default structured agent surface, on swebench-live tasks that carry one failing test and one changed file. Two things separate cleanly.
oi code-as-action engine
tomo gets a new engine, engine/oi, ported from the shape of Open Interpreter 0.4.2. The model's only action is one Markdown code block, the engine runs it in the sandbox and feeds the output back, and the turn ends when a reply carries no block.
dynaconf cost and caching, deepseek vs gpt-5.6
The lab now prices every probe run at list rate and breaks out the prefix-cached share, so a free run still shows what it would cost and stays comparable to a paid one.
Monday, 13 July
dynaconf same model, codex solves, tomo fails
Hold the model fixed on the codex backend and vary the harness. On dynaconf-1225, codex solves the bug with luna, terra, and sol; tomo fails with all four subscription models it was run on.
dynaconf closed-door lessons for tomo
Seven honest runs on one task, three passes and four fails, read together. The lessons that transfer to tomo: a broad edit that regresses a green test is worse than no edit and wants a do-no-harm gate, spend does not track progress, cache-read is where the money actually goes…
dynaconf 5.6 family + analyzer false leak
gpt-5.6-terra and -sol both pass dynaconf-1225 with the doors shut, and both got flagged as answer leaks. Reading the trace, the flag is wrong.
dynaconf opus offline
The honest opus run on dynaconf-1225 is the most expensive fail in this comparison and the only run that ends with the repo worse than it started.
dynaconf sonnet offline
The earlier sonnet run passed dynaconf-1225 by fetching the merged pull request. Close the network door and run it again and the fetch is gone.
dynaconf gpt-5.6-luna offline
Run the newest codex model on dynaconf-1225 with the git-history door and the network door both closed, and for the first time in this comparison a model solves it honestly. gpt-5.
dynaconf gpt-5.5 offline
The flagship codex model runs dynaconf-1225 with both answer doors closed. It writes nineteen edits across every loader, the validator, and the cli, twice what the cheap model touched, spends six times as much, reaches no answer, and fails on the exact same two settings-loader…
dynaconf gpt-5.4-mini offline
The cheapest codex model runs dynaconf-1225 with both answer doors closed: git history pruned so the fix commit is unreachable, and the shell sandboxed so it cannot fetch the answer PR. It writes a real nine-edit fix, reaches no answer, and fails on the settings-loader tests.
dynaconf opus answer fetch
Claude Opus 4.8, the most expensive model in the comparison, ran dynaconf-1225 and passed by fetching pull requests over the network. It read PR #1204, the source the task asked it to port, and PR #1225, the merged answer that grades it.
dynaconf sonnet answer fetch
Claude Sonnet 5 ran dynaconf-1225 and passed, but the trace shows it did not solve the bug. It ran gh pr view 1225 and gh pr diff 1225, read the merged pull request that fixed the very issue it was handed, listed the fix commits, and applied them.
dynaconf haiku clean fail
Claude Haiku 4.5 ran dynaconf-1225 without reaching the network, wrote a real source fix, and failed. It threaded the identifier argument through every loader, the same broad refactor gpt-5.4 tried, and regressed a test that started green.
With the leak closed, dynaconf sorts the models
The same leak-free dynaconf task passes on gpt-5.6-sol and gpt-5.5 and fails on gpt-5.4, so the fix that removed the git shortcut left a task that actually measures the model.
mini vs sol on python-control
Four real codex subscription runs, gpt-5.4-mini and gpt-5.6-sol on two tasks, priced through our new single source of truth. On python-control both models, cheap and flagship, converge on the identical edit and fail the identical three tests.
dynaconf answer leak
gpt-5.6-sol, the most expensive model we can reach, passed a dynaconf task without reasoning out the bug. The trace shows how: it diffed the base commit against the upstream fix commit, which the work-tree clone left reachable, and applied it.
churn guard vs claude-code
The third and last of tomo's runaway shapes. A turn that keeps editing but never converges, writing scratch scripts or the same file over and over, now stops instead of burning a hundred rounds.
dynaconf tomo guard vs pi
The follow-up to tomo's git-archaeology runaway. tomo now bounds a turn that investigates without ever editing, so the same dynaconf run stops at 41 requests instead of 132 and 1.7 million tokens instead of four million.
python-control tomo scratch runaway
A second runaway with a different shape. On a python-control conversion bug, tomo made 34 edits and still failed, because 33 of them were throwaway debug scripts it wrote to instrument the problem rather than fix it.
fonttools tomo over-normalized
The one swebench-live failure where tomo did everything right and still lost. On a fonttools glyph-reordering bug it found the exact file, wrote a fuller fix than the maintainers, and verified its work, then failed a single hidden test because it reused a variable that forces .
dynaconf tomo runaway
tomo's worst run of the sweep: on a dynaconf bug it spent 132 requests and four million tokens running git log, git diff, and git show to reverse-engineer a fix from history, hit the fifteen-minute wall, and never edited a single file.
gitingest tomo
tomo solves a real gitingest issue the way the benchmark intends: it reads the source, finds the one branch that only handles https, adds the http case, and verifies with the project's own tests. One source edit, no network, 242k tokens.
cfn-lint pi
A second rival on the cfn-lint task. pi never leaves the repo, never fetches the pull request, and fails exactly where tomo failed: its source change does not produce the arbitrary graded wording. A short confirm, with the caveat that the free-tier rate limit cut the run short.
cfn-lint opencode
opencode passes a cfn-lint task whose graded wording appears nowhere in the repo. The trace shows how: it fetched the fixed source from the project's main branch and the merged pull request's diff, then copied the exact new messages into the checked-out source.
cfn-lint tomo
tomo reads a cfn-lint issue, implements exactly the message it asks for, and fails the grade. The graded wording is a generic validator message the maintainers changed instead, and it appears nowhere in the checked-out repo.
faker --yolo fix
The follow-up to the faker lockout. tomo gains a --yolo mode that runs it fully autonomous, the same way every rival already runs. The exact task tomo had solved but could not write now passes, and passes leaner: 40 percent fewer tokens and half the model calls.
faker IBAN lock
tomo diagnoses a Belgian IBAN bug and writes the exactly correct fix, then cannot apply it. A reference URL it fetched tripped its own prompt-injection guard, which escalated every later edit to an approval that never comes in headless mode.