2026-07-15
Experiment reports captured on 2026-07-15, one tool on one task each, newest first.
Runs captured on 2026-07-15, newest first. Each is one tool on one task with the full trace read closely.
July 2026
Wednesday, 15 July
oi vs agent, reachability is the model
Run the same cheap model through tomo's two engines, the oi code-as-action surface and the default structured agent surface, on swebench-live tasks that carry one failing test and one changed file. Two things separate cleanly.
oi code-as-action engine
tomo gets a new engine, engine/oi, ported from the shape of Open Interpreter 0.4.2. The model's only action is one Markdown code block, the engine runs it in the sandbox and feeds the output back, and the turn ends when a reply carries no block.
dynaconf cost and caching, deepseek vs gpt-5.6
The lab now prices every probe run at list rate and breaks out the prefix-cached share, so a free run still shows what it would cost and stays comparable to a paid one.