Testronaut Research / October 2026
Jev optimization experiments
Guardrails, model routing, and what actually makes an agentic test cheaper without making it less reliable.
Research question
Can a small reasoning layer make an agentic test smarter about how it spends intelligence?
Jev was tested first as a guardrail around a primary execution model, and then as a selector deciding which model should execute work. The experiments separate those effects rather than assuming model switching itself is the optimization.
Primary finding
The strongest evidence favored guardrails without model switching.
Guarded GPT-5.5 completed 15/15 OpenAI mission-file attempts while using 13.7% fewer model tokens than the unguarded Terra control.
Experiment 1
Guardrails and same-provider routing
Seventy aggregate runs invoked three mission files each. Conditions progressed from unguarded controls to guarded execution and then same-provider routing. Token deltas are relative to the provider-specific control.
| Provider | Condition | Strict passes | Mean model tokens | Token delta |
|---|---|---|---|---|
| OpenAI | Terra control | 5/15 | 300,073 | — |
| OpenAI | Guarded GPT-5.5 | 15/15 | 258,635 | -13.7% |
| OpenAI | Guarded Terra | 8/15 | 275,114 | -8.2% |
| OpenAI | Terra → Luna | 11/15 | 296,299 | -1.1% |
| Gemini | 3.8 control | 15/15 | 491,428 | — |
| Gemini | Guarded 3.5 Flash-Lite | 15/15 | 427,766 | -12.0% nominal |
| Gemini | 3.8 → 3.7 | 15/15 | 461,256 | -5.1% nominal |
Experiment 2
Cross-provider selection at mission boundaries
One hundred and twenty isolated mission executions compared fixed controls with cross-provider treatments, making the model-selection decision attributable to a specific execution.
| Condition | Strict passes | Mean model tokens | Selector tokens | Mean duration |
|---|---|---|---|---|
| GPT-5.5 fixed | 15/15 | 87,422 | 0 | 141.5s |
| Terra fixed | 7/15 | 91,698 | 0 | 106.2s |
| Gemini 3.8 fixed | 10/15 | 104,653 | 0 | 138.5s |
| Gemini 3.5 Flash-Lite fixed | 13/15 | 93,446 | 0 | 109.6s |
| GPT-5.5 ↔ Gemini 3.8 | 12/15 | 98,490 | 627 | 135.9s |
| GPT-5.5 ↔ Gemini 3.5 Flash-Lite | 15/15 | 104,897 | 633 | 132.1s |
| Terra ↔ Gemini 3.8 | 11/15 | 102,091 | 621 | 131.8s |
| Terra ↔ Gemini 3.5 Flash-Lite | 12/15 | 99,735 | 724 | 114.5s |
Treatment missions where Jev selected a Gemini candidate.
Mean selector tokens across cross-provider treatments.
GPT-5.5 passes on the discriminating file-transfer mission; Terra and Gemini 3.8 were 0/5.
Interpretation
Optimize the successful trajectory, not the individual call
1. Stabilize guardrails
Prioritize completion/evidence checks, fail-open behavior, primary-model continuity, and explicit telemetry.
2. Keep routing conservative
Treat execution-model routing as beta/opt-in. Prefer one primary and one compatible same-provider secondary while evidence grows.
3. Learn from history
Use versioned reliability, token, latency, price, and sample-size history with a reliability floor and cold-start fallback.
Methodology
How to read these results
- CLI baseline: Testronaut 1.11.0.
- Blocking: five randomized blocks in every condition.
- Strict success: every expected phase exists and passes.
- Token accounting: model and Jev tokens are separate; counts are not price-weighted spend.
- Intervals: Experiment 1 paired intervals are two-sided 95% Student-t intervals with 4 degrees of freedom.
Caveats
Six initial Gemini 3.8 requests in Experiment 2 returned provider HTTP 500 responses. The resumable benchmark reran incomplete units using persisted selections, producing the final 120 completed reports.
Results are specific to the tested repositories, missions, model versions, dates, and provider behavior. They are not universal model rankings.
Open data
Public research package
The publication-safe artifacts below come directly from the research package. They intentionally exclude raw browser events, DOM content, screenshots, HTML reports, environment files, credentials, absolute report paths, and verbose logs. Private reports remain the audit source.
Complete methodology, results, interpretation, and caveats.
Definitions for the normalized public fields and experiment tables.
Condition-level results, paired deltas, intervals, routing, and completion gates for Experiment 1.
Fixed-control and cross-provider condition summaries for Experiment 2.
The original package also contains normalized run-level and mission-level CSVs plus frozen experiment matrices. Those can be added to this public catalog once the complete publication bundle is committed as static artifacts.
The working conclusion
The cheapest model is not necessarily the cheapest test. The unit of optimization in an agentic system is the trajectory.