# Jev optimization experiments: guardrails and model routing

Publication package: `jev-optimization-experiments-2026-10`  
Generated: 2026-10-09T20:43:18.475Z  
CLI baseline: Testronaut 1.11.0

## Executive summary

Two blocked experiments evaluated Jev-assisted browser-testing optimization. The first contained **70 aggregate runs and 210 mission-file attempts** across OpenAI and Gemini workloads. The second contained **120 isolated mission executions** testing fixed-model controls and mission-boundary cross-provider selection. All conditions used five randomized blocks.

The most defensible result is that **Jev guardrails are useful without model switching**. Guarded GPT-5.5 completed 15/15 OpenAI mission-file attempts and used 13.7% fewer model tokens than the unguarded Terra control, with a paired 95% interval of -19.8% to -7.5%. Guarded Terra also improved strict completion and nominally reduced tokens. On Gemini, guardrails preserved the already-perfect reliability of Gemini 3.8 Flash with a smaller nominal token reduction.

Same-provider routing is promising but not yet broadly proven. Gemini 3.8 → 3.7 retained 15/15 strict passes with a nominal 5.1% token reduction and similar duration. OpenAI routing coverage was below one routed turn per aggregate run, making attribution underpowered.

Mission-boundary cross-provider routing did **not** reduce token volume. The strongest fixed control, guarded GPT-5.5, achieved 15/15 strict passes at 87,422 mean model tokens per isolated mission. GPT-5.5 ↔ Gemini 3.5 Flash-Lite also achieved 15/15, but required 104,897 mean model tokens. Jev's extra selector call averaged only about 621–724 tokens, so selector overhead was not the cause; the model choices themselves produced longer executions.

## Experiment 1: optimization ladder

| Provider | Condition | Strict pass | Mean model tokens | Paired token delta vs control | Mean duration |
|---|---|---:|---:|---:|---:|
| OpenAI | Terra control | 5/15 (33%) | 300,073 | baseline | 293s |
| OpenAI | Guarded GPT-5.5 | **15/15 (100%)** | **258,635** | **-13.7% [-19.8, -7.5]** | 416s |
| OpenAI | Guarded Terra | 8/15 (53%) | 275,114 | -8.2% [-18.0, 1.7] | 328s |
| OpenAI | Terra → Luna | 11/15 (73%) | 296,299 | -1.1% [-12.1, 9.9] | 376s |
| Gemini | 3.8 control | **15/15 (100%)** | 491,428 | baseline | 529s |
| Gemini | Guarded 3.5 Flash-Lite | **15/15 (100%)** | **427,766** | -12.0% [-34.7, 10.8] | 582s |
| Gemini | 3.8 → 3.7 | **15/15 (100%)** | 461,256 | -5.1% [-19.7, 9.5] | 538s |

The table highlights decision-relevant conditions; complete condition, run, and mission tables are included in `data/optimization/`.

## Experiment 2: cross-provider mission-boundary routing

| Condition | Design | Strict pass | Mean model tokens | Mean in-mission Jev tokens | Mean selector tokens | Mean duration |
|---|---|---:|---:|---:|---:|---:|
| c1_gpt55 | fixed-control | 15/15 (100%) | 87,422 | 58,180 | 0 | 141.5s |
| c2_terra | fixed-control | 7/15 (47%) | 91,698 | 64,868 | 0 | 106.2s |
| c3_gemini38 | fixed-control | 10/15 (67%) | 104,653 | 59,372 | 0 | 138.5s |
| c4_gemini35lite | fixed-control | 13/15 (87%) | 93,446 | 61,226 | 0 | 109.6s |
| x1_gpt55_gemini38 | cross-provider-treatment | 12/15 (80%) | 98,490 | 58,348 | 627 | 135.9s |
| x2_gpt55_gemini35lite | cross-provider-treatment | 15/15 (100%) | 104,897 | 67,577 | 633 | 132.1s |
| x3_terra_gemini38 | cross-provider-treatment | 11/15 (73%) | 102,091 | 59,736 | 621 | 131.8s |
| x4_terra_gemini35lite | cross-provider-treatment | 12/15 (80%) | 99,735 | 65,819 | 724 | 114.5s |

Jev chose a Gemini candidate for 50 of 60 treatment missions. That bias was not aligned with the observed workload: GPT-5.5 was the only fixed model with perfect strict completion, and it used the fewest mean model tokens. The file-transfer mission was particularly discriminating: GPT-5.5 passed 5/5, Terra and Gemini 3.8 passed 0/5, and Gemini 3.5 Flash-Lite passed 3/5 in their fixed controls.

## Product implications

1. Stabilize guardrail-only operation: completion/evidence checks, fail-open behavior, primary-model continuity, and explicit telemetry.
2. Keep execution-model routing beta and opt-in. Limit a run to one primary and one compatible same-provider secondary.
3. Treat Gemini 3.8 → 3.7 as a conservative candidate for a larger confirmation study, not a universal default.
4. Do not ship zero-shot cross-provider selection as an optimization default.
5. Next test history-informed routing using versioned per-model reliability, tokens, latency, and price, with a reliability floor and cold-start fallback.

## Methodology and interpretation

- Both experiments used five randomized blocks (`n=5` per condition design).
- Strict success required every expected phase in a mission file to exist and pass.
- Experiment 1 aggregate runs each invoked three mission files. Experiment 2 intentionally isolated each mission to make mission-boundary selection attributable.
- Model tokens and Jev tokens are reported separately. Token counts are not price-weighted spend.
- Experiment 1 paired intervals use a two-sided 95% Student-t interval with four degrees of freedom.
- Six initial Gemini 3.8 requests in experiment 2 returned provider HTTP 500 errors. The resumable benchmark reran those incomplete units using their persisted selections; the final dataset contains 120 completed reports.
- Results describe these repositories, missions, model versions, dates, and provider behavior. They are not universal model rankings.

## Data and privacy

This public package contains normalized measurements, matrices, methodology, and checksums. It intentionally excludes raw browser events, DOM content, screenshots, HTML reports, environment files, credentials, absolute report paths, and verbose logs. The private source reports remain the audit source for the normalized rows.
