Testronaut Research / October 2026

Jev optimization experiments

Guardrails, model routing, and what actually makes an agentic test cheaper without making it less reliable.

70
aggregate runs
Experiment 1
210
mission-file attempts
Experiment 1
120
isolated missions
Experiment 2
5
randomized blocks
every condition

Research question

Can a small reasoning layer make an agentic test smarter about how it spends intelligence?

Jev was tested first as a guardrail around a primary execution model, and then as a selector deciding which model should execute work. The experiments separate those effects rather than assuming model switching itself is the optimization.

Primary finding

The strongest evidence favored guardrails without model switching.

Guarded GPT-5.5 completed 15/15 OpenAI mission-file attempts while using 13.7% fewer model tokens than the unguarded Terra control.

Experiment 1

Guardrails and same-provider routing

Seventy aggregate runs invoked three mission files each. Conditions progressed from unguarded controls to guarded execution and then same-provider routing. Token deltas are relative to the provider-specific control.

ProviderConditionStrict passesMean model tokensToken delta
OpenAITerra control5/15300,073—
OpenAIGuarded GPT-5.515/15258,635-13.7%
OpenAIGuarded Terra8/15275,114-8.2%
OpenAITerra → Luna11/15296,299-1.1%
Gemini3.8 control15/15491,428—
GeminiGuarded 3.5 Flash-Lite15/15427,766-12.0% nominal
Gemini3.8 → 3.715/15461,256-5.1% nominal
For guarded GPT-5.5, the paired 95% interval for token change was -19.8% to -7.5%. Gemini 3.8 → 3.7 retained 15/15 strict passes with a nominal 5.1% token reduction, making it a candidate for confirmation rather than a default recommendation.

Experiment 2

Cross-provider selection at mission boundaries

One hundred and twenty isolated mission executions compared fixed controls with cross-provider treatments, making the model-selection decision attributable to a specific execution.

ConditionStrict passesMean model tokensSelector tokensMean duration
GPT-5.5 fixed15/1587,4220141.5s
Terra fixed7/1591,6980106.2s
Gemini 3.8 fixed10/15104,6530138.5s
Gemini 3.5 Flash-Lite fixed13/1593,4460109.6s
GPT-5.5 ↔ Gemini 3.812/1598,490627135.9s
GPT-5.5 ↔ Gemini 3.5 Flash-Lite15/15104,897633132.1s
Terra ↔ Gemini 3.811/15102,091621131.8s
Terra ↔ Gemini 3.5 Flash-Lite12/1599,735724114.5s
50/60

Treatment missions where Jev selected a Gemini candidate.

621–724

Mean selector tokens across cross-provider treatments.

5/5

GPT-5.5 passes on the discriminating file-transfer mission; Terra and Gemini 3.8 were 0/5.

Interpretation

Optimize the successful trajectory, not the individual call

1. Stabilize guardrails

Prioritize completion/evidence checks, fail-open behavior, primary-model continuity, and explicit telemetry.

2. Keep routing conservative

Treat execution-model routing as beta/opt-in. Prefer one primary and one compatible same-provider secondary while evidence grows.

3. Learn from history

Use versioned reliability, token, latency, price, and sample-size history with a reliability floor and cold-start fallback.

Methodology

How to read these results

  • CLI baseline: Testronaut 1.11.0.
  • Blocking: five randomized blocks in every condition.
  • Strict success: every expected phase exists and passes.
  • Token accounting: model and Jev tokens are separate; counts are not price-weighted spend.
  • Intervals: Experiment 1 paired intervals are two-sided 95% Student-t intervals with 4 degrees of freedom.

Caveats

Six initial Gemini 3.8 requests in Experiment 2 returned provider HTTP 500 responses. The resumable benchmark reran incomplete units using persisted selections, producing the final 120 completed reports.

Results are specific to the tested repositories, missions, model versions, dates, and provider behavior. They are not universal model rankings.

Open data

Public research package

The publication-safe artifacts below come directly from the research package. They intentionally exclude raw browser events, DOM content, screenshots, HTML reports, environment files, credentials, absolute report paths, and verbose logs. Private reports remain the audit source.

The original package also contains normalized run-level and mission-level CSVs plus frozen experiment matrices. Those can be added to this public catalog once the complete publication bundle is committed as static artifacts.

The working conclusion

The cheapest model is not necessarily the cheapest test. The unit of optimization in an agentic system is the trajectory.