In April 1970, Apollo 13 carried something that seems almost absurd in retrospect: a five-inch slide rule.
This was not because the Apollo program lacked computers. The spacecraft depended on its guidance computer, while much larger systems on Earth supported the calculations required to send people to the Moon and bring them home.
And yet the crew still carried a slide rule.
There is an appealing engineering lesson in that combination. The goal was not to use the most powerful computational tool available for every calculation. It was to use enough computation for the problem in front of you.
I have been thinking about that idea a lot while working on a different kind of optimization problem.
Agentic end-to-end testing gives us increasingly capable models that can inspect an interface, reason about what they see, recover when something unexpected happens, and continue toward an objective. Give the agent more intelligence and, generally, it becomes better at navigating uncertainty.
But that intelligence has a cost.
Every observation, decision, recovery, and detour consumes inference. Once a test becomes an autonomous trajectory rather than a predetermined sequence of commands, an apparently simple question becomes surprisingly difficult:
How much intelligence does this particular part of the mission actually need?
That question led to Jev.
A slide rule for the agent?
The slide-rule comparison needs one important qualification: Jev is not about replacing a capable model with something primitive.
The analogy is about allocation.
If a small calculation can be solved with a slide rule, there is little reason to monopolize a larger computer to solve it. Likewise, if a browser-testing decision can be made safely with a bounded check, we should ask whether it needs the full reasoning budget of the primary execution model.
That is an especially interesting question for agentic testing because the cost of a model decision is not isolated to a single API call. One decision changes what happens next. A weaker choice might require another observation, a retry, a recovery, or an entirely different path through the application.
So the real optimization target cannot simply be:
Which model call is cheapest?
It has to be:
Which sequence of decisions produces the cheapest successful trajectory?
That distinction became the most important result of these experiments.
From probes to mission control
This work follows naturally from the probe pattern I wrote about previously.
A probe gives an agentic mission a narrow way to inspect its current state before deciding whether a procedure is necessary. Instead of blindly repeating a login flow, for example, the mission can first determine whether it is already authenticated.
That lets a mission remain independently runnable without forcing every larger test journey to repeat setup that has already been completed.
Jev explores the same general idea at another layer.
A probe asks: Do I need to perform this procedure?
A guardrail can ask: Is the mission still proceeding in a useful and valid way?
And model routing eventually asks: What computational resource should perform this work?
Those sound like three versions of the same optimization problem. In practice, our experiments suggest they have very different risk profiles.
The hypothesis: spend intelligence where it matters
The basic hypothesis behind the Jev experiments was straightforward.
A strong primary model is useful because browser automation is messy. Interfaces change. Labels are ambiguous. Loading states appear. Authentication flows redirect unexpectedly. A human-readable mission has to be translated into concrete actions against whatever the browser is showing right now.
But not every moment of that process necessarily requires the same level of model capability.
We wanted to test two broad ideas: could Jev improve execution without switching models, by applying bounded guardrails around the primary agent? And could Jev reduce the cost of a mission by deciding when another model was sufficient for the work?
The second idea is intuitively attractive. If some portions of a mission are easy, why not hand them to a cheaper model?
That was also the idea that turned out to be much harder than it looked.
What we tested
We ran two blocked experiments using Testronaut 1.11.0.
The first experiment was an optimization ladder across OpenAI and Gemini workloads. It contained 70 aggregate runs and 210 mission-file attempts. Conditions included unguarded controls, guardrail-only configurations, and same-provider routing. Each aggregate run invoked three mission files.
The second experiment deliberately isolated missions so that model selection at mission boundaries could be attributed more cleanly. It contained 120 isolated mission executions, comparing fixed-model controls with cross-provider routing treatments.
Both experiments used five randomized blocks.
Strict success meant that every expected phase in a mission file existed and passed. We also kept model tokens and Jev tokens separate rather than collapsing everything into a single number.
These are workload-specific experiments, not universal model rankings. They describe these missions, model versions, repositories, dates, and provider behaviour. That limitation matters.
But the results were still useful enough to change how I think this feature should evolve.
Result one: guardrails were more useful than model switching
The strongest result was also the simplest.
On the OpenAI workload, the unguarded Terra control completed 5 of 15 mission-file attempts under the strict success definition and averaged 300,073 model tokens.
Guarded GPT-5.5 completed 15 of 15 and averaged 258,635 model tokens.
That is a 13.7% reduction in model tokens, with a paired 95% interval of -19.8% to -7.5%, while strict completion improved from 33% in the Terra control to 100%.
Guarded Terra also improved strict completion, from 5/15 in the Terra control to 8/15, while nominally reducing tokens by 8.2%.
On Gemini, guardrails preserved the already-perfect reliability of the Gemini 3.8 control. Guarded Gemini 3.5 Flash-Lite also completed 15/15 attempts and showed a nominal token reduction, although the uncertainty was much wider.
The important point is not that one model 'won.'
The important point is that Jev was useful even when it did not choose another execution model at all.
That changed the product question. Originally, model routing looked like the clever part. But the experiment suggested that making the existing agent more disciplined may be more valuable than frequently changing the agent doing the work.
Apollo had mission rules, too
This is where the Apollo comparison becomes more interesting than the slide rule alone.
Apollo was not simply a story of putting different computers next to each other and choosing whichever one was convenient. The program also depended on mission rules: explicit guidance for how crews and controllers should respond to conditions during a flight.
That is much closer to what the strongest Jev result looked like.
The agent did not necessarily need a different brain. It benefited from better boundaries around how that brain operated.
In software terms, this is less glamorous than dynamic model routing. It is also exactly the sort of thing that often makes a system dependable.
Before trying to optimize who performs the next piece of work, make sure the current system knows what useful work looks like, when enough evidence has been collected, how to fail open when supervision is unavailable, and when to preserve continuity with the primary model.
A five-inch slide rule is useful because it has a well-understood job. Mission rules are useful because they constrain decisions.
The lesson is not “use less intelligence.” It is give intelligence an appropriate role.
Result two: same-provider routing is interesting, but not proven
We did see one routing result worth following up.
Gemini 3.8 → 3.7 retained 15/15 strict passes, with a nominal 5.1% reduction in model tokens and similar duration to the Gemini 3.8 control.
That is promising. It is not enough evidence to turn the strategy into a universal default.
On the OpenAI side, routing coverage was below one routed turn per aggregate run, which left us with too little exposure to attribute the result confidently to routing.
This is an important distinction when building AI products. A result can be interesting enough to justify another experiment without being strong enough to justify a product promise.
For now, I think execution-model routing belongs in beta and should remain opt-in.
Result three: cross-provider routing backfired
The second experiment was designed to give routing a cleaner opportunity to demonstrate value. Instead of changing models during a mission, Jev selected between candidates at the mission boundary. This made it easier to compare the resulting execution against fixed-model controls.
The strongest fixed control was guarded GPT-5.5. It completed 15/15 missions at an average of 87,422 model tokens per isolated mission.
A cross-provider treatment selecting between GPT-5.5 and Gemini 3.5 Flash-Lite also completed 15/15. But it averaged 104,897 model tokens.
The routed version preserved reliability and consumed substantially more model tokens.
At first glance, it would be easy to blame the extra selection call. But Jev's selector averaged only about 621–724 tokens across the cross-provider treatments. The selector itself was not expensive enough to explain the difference.
The choices changed the trajectories.
Across the 60 treatment missions, Jev chose a Gemini candidate 50 times. That preference did not align well with what the fixed controls told us about this particular workload. GPT-5.5 was the only fixed model with perfect strict completion, and it also used the fewest mean model tokens.
One file-transfer mission made the difference especially visible. GPT-5.5 passed it 5/5 times in its fixed control. Terra and Gemini 3.8 passed 0/5. Gemini 3.5 Flash-Lite passed 3/5.
A selector that does not know that history is making its decision with an important piece of mission intelligence missing.
The cheapest model is not necessarily the cheapest test
This was the result I found most useful.
When thinking about model routing, it is tempting to treat models like interchangeable compute tiers. One is more capable and expensive; another is cheaper; therefore we should identify the easy work and move it downward.
Agentic systems complicate that model.
A browser agent is not evaluating a static function. It is participating in a feedback loop:
- Observe the application.
- Interpret the state.
- Choose an action.
- Change the application.
- Observe the result.
- Recover or continue.
A weaker decision at step three can change steps four through twenty. It might click the wrong control and recover. It might require more DOM context. It might need another screenshot. It might take a longer navigation path. It might reach an ambiguous state and spend several turns determining what happened.
A model that costs less per call can therefore produce a test that costs more to finish.
The unit of optimization in an agentic system is the trajectory.
The question is not simply what we paid for the next decision. It is what that decision caused us to pay before we obtained reliable evidence that the mission succeeded or failed.
What Jev should become next
These experiments produced a much clearer development path than the original hypothesis did.
The first priority is guardrail-only operation: completion and evidence checks, fail-open behaviour, continuity with the primary model, and telemetry that makes Jev's interventions visible.
Execution-model routing can remain available as an experimental capability, but it should be conservative. For now, a run should use one primary model and at most one compatible same-provider secondary rather than treating every available model as a candidate.
Gemini 3.8 → 3.7 is a reasonable candidate for a larger confirmation study. It is not a universal default.
And zero-shot cross-provider selection should not become an optimization default based on these results.
The more interesting next experiment is history-informed routing.
Instead of asking a selector to infer the best model from a mission description alone, give it evidence about what has happened before:
- Reliability by model and mission type.
- Model tokens consumed.
- Latency.
- Price.
- Model and provider version.
- Enough observations to distinguish evidence from a cold start.
Then put a reliability floor underneath the optimization objective.
In other words, do not ask:
Which model looks cheapest for this mission?
Ask:
Given what we have actually observed, which model is likely to complete this mission reliably at the lowest total cost?
That is a much more interesting routing problem.
From bigger models to better systems
There is a broader idea here that extends beyond testing.
The current AI ecosystem naturally focuses on model capability. Benchmarks compare models. Products advertise which models they support. Developers debate which model is smartest, fastest, or cheapest. Those questions matter.
But mature agentic systems will increasingly have another layer of intelligence around the model itself.
They will decide what context is necessary. They will recognize when work has already been completed. They will constrain unproductive behaviour. They will learn which tools and models perform well for particular jobs. And they will use observed outcomes rather than intuition alone to allocate computation.
That is where probes and Jev fit into the same picture for Testronaut™.
- A probe asks whether part of the mission needs to run.
- A guardrail helps determine whether the execution is producing useful evidence.
- A router may eventually decide which model should perform the work.
None of those mechanisms makes the primary agent less important. They make the system around the agent more deliberate.
Apollo did not reach the Moon because NASA found one computer powerful enough to make every decision.
It combined onboard computation, ground systems, human judgment, procedures, mission rules, and, yes, tools as simple as a slide rule.
The sophistication was in the system.
That is increasingly how I think about agentic testing as well.
The goal is not to spend the fewest tokens on every decision. It is not to route every possible action to the cheapest model. And it is certainly not to add AI simply because another decision can be delegated to AI.
The goal is to build a testing system that knows how much intelligence a decision needs, where that intelligence should come from, and whether the resulting trajectory is still producing trustworthy evidence.
Sometimes that may require the supercomputer.
Sometimes the slide rule is enough.
And knowing the difference is the interesting part.
Research note
The accompanying public research summary contains normalized measurements, methodology, and interpretation for these experiments. Raw browser events, DOM content, screenshots, HTML reports, environment files, credentials, absolute report paths, and verbose logs are intentionally excluded. Results should be interpreted as workload-specific rather than as universal model rankings.public research summary
