How We Build

Reusable Rockets: How Probes Cut the Cost of Agentic Testing

How agentic tests can remain atomic in isolation while reusing existing state to run efficiently as part of a larger suite.

September 29, 202613 min read

By Shane Fast

Reusable Rockets: How Probes Cut the Cost of Agentic Testing

Agentic tests can run atomically on their own and efficiently inside larger suites because they can determine which instructions are worth executing before they act.

Agentic testing is often described as a more flexible way to execute a test. Give an agent an objective and it can interpret the interface, find the right controls, and adapt when the path is not exactly as expected.

That is useful, but it is only the first layer of what makes agentic testing interesting.

The more significant shift happens when the agent can reason about the instructions themselves. Instead of asking only, “How do I perform this step?”, it can first ask, “Is this step worth performing at all?”

Consider a mission that includes instructions for logging in. When the mission runs by itself, those instructions are essential. But when it runs as part of a larger test suite, a previous mission may have already established an authenticated session. A conventional sequence will usually execute the login steps again unless that dependency has been handled explicitly elsewhere. A human tester would take one look at the dashboard, recognize that they are already logged in, and move on.

With a probe, an agentic test can make the same kind of practical judgment. It can inspect the current state, determine whether the user is already authenticated, and execute the login instructions only when they are actually needed. The agent is no longer just flexible in how it follows the test. It is applying bounded reasoning to decide which parts of the test still have value.

That is where the reusable-rocket comparison comes in. The mission still carries everything it needs for an independent launch, including the complete login flow. But when useful state from an earlier mission can be recovered and reused, it does not need to pay for the same launch twice.

The result is a different way to design end-to-end tests: missions that remain independently runnable, yet become faster, cheaper, and more seamless when assembled into a larger journey.

The maintenance cost of almost-identical tests

Earlier in my career, I worked with test suites that contained what were effectively several versions of the same test.

One version assumed the user was already authenticated. Another started from the login screen. A third ran as a different type of user or expected slightly different application state. Each test made sense in isolation, but together they became painful to maintain.

A small change to a shared workflow rarely meant updating just one test. It meant finding every variation, understanding why it was slightly different, and carefully applying the same change without erasing the contextual differences that justified the duplication in the first place.

This is where test maintenance can become a combinatorial problem. Every new role, starting state, feature flag, environment, or execution context creates another possible variation. Even when the underlying behaviour is nearly identical, the suite begins accumulating parallel versions of the same journey.

The problem is not unique to any one testing framework. Conventional test suites have plenty of ways to share fixtures, preserve authentication state, and arrange prerequisites. But those relationships usually need to be encoded explicitly through setup hooks, shared helpers, dependencies, or orchestration. As the number of contexts grows, understanding which test requires which setup can become a maintenance burden of its own.

I recently encountered a smaller version of this problem while authoring missions for Testronaut™.

Atomic missions versus seamless journeys

Imagine a mission that verifies an authenticated user can create a new project.

When I run that mission by itself, it needs to be capable of establishing an authenticated session. Otherwise, its success depends on whatever happened to run before it. That makes the mission fragile, difficult to reuse, and harder to troubleshoot independently.

But that same mission might also run immediately after another mission that already logged in. In a broader test journey, repeating the entire login process adds no new confidence. It simply consumes more time, more model tokens, and more money. If multifactor authentication or an external identity provider is involved, the repeated setup may be considerably more expensive than the assertions I actually care about.

This creates an awkward choice:

  • Keep the login instructions and preserve the mission's independence, but repeat unnecessary work during a larger run.
  • Remove the login instructions and make the larger run more efficient, but leave the mission dependent on an earlier step.
  • Maintain separate standalone and chained versions of the mission, bringing back the duplication problem.

What I wanted was the best of both approaches: missions that could run atomically when isolated, yet flow seamlessly and efficiently when assembled into a larger journey.

The answer was surprisingly simple. Before launching the login sequence, let the mission ask one narrow question:

Am I already logged in?

That question is a probe.

What is a probe?

A probe is a small, focused check that lets a mission inspect the application's current state before deciding what to do next.

In the authentication example, the intent looks something like this:

Probe whether the user is already authenticated.

If the user is authenticated:
  Continue to the project workflow.

If the user is not authenticated:
  Complete the login flow, then continue.

The probe does not remove the login capability from the mission. It prevents the mission from blindly executing that capability when the required state has already been established.

Run the mission alone and it can log in for itself. Run it after another authenticated mission and it can recognize that the prerequisite is already satisfied. The mission retains its independence without forcing every larger test run to start from zero.

I think of this pattern as adaptive isolation: a mission carries the setup it needs to operate independently, but can adapt when it becomes part of a larger journey.

Retro NASA safety-manual comic showing why blindly following every instruction can be inefficient

Why probes matter in agentic testing

In a conventional test, conditional behaviour can be a warning sign. Tests should be deterministic, and excessive branching can make it unclear which path was actually verified.

That concern still applies to agentic testing. A probe should not make expectations optional. What it can do is help the agent choose the most appropriate route to the same expected outcome.

The distinction is important:

  • The route can adapt to the current state.
  • The expectation should remain explicit.

If the goal is to verify that an authenticated user can create a project, it should not matter whether authentication was established by the current mission or preserved from the previous one. The assertion about project creation remains the same.

This gives probes several practical benefits.

Lower token usage

Agentic tests spend tokens interpreting the interface, selecting actions, responding to unexpected states, and confirming outcomes. Skipping a workflow that has already been completed avoids paying for the same reasoning twice.

Faster test runs

Authentication, navigation, data creation, uploads, and asynchronous workflows can take far longer than the probe that determines whether they are needed. A small state check can remove entire sections of repeated execution.

Fewer duplicated missions

Without a probe, it is tempting to create separate versions for different starting states. A context-aware mission can often represent those variations without duplicating the primary test intent.

Better composability

Missions become easier to combine in different orders. Each mission can bring its own prerequisites while cooperating with state established earlier in the run.

More resilient recovery

If a run is interrupted after completing an action but before recording its result, a probe can help determine what actually happened. The agent may be able to resume safely instead of repeating the action or failing because the application has moved ahead.

Probes beyond authentication

Authentication is an easy example, but it is only one place where a mission benefits from situational awareness.

1. Starting-location probes

Before navigating through several menus, a mission can determine whether the browser is already on the required page.

Are we already viewing the billing settings page?

If so, the mission can begin its assertions immediately. If not, it can navigate there. This is especially useful when several missions operate in the same area of an application.

2. Required-data probes

A mission may need a customer, project, workspace, invoice, or other record before it can test the intended behaviour.

Is there already a test project named Apollo Demo?

If the record exists, the mission can reuse it. If it does not, the mission can create it first. This keeps the test independently runnable without creating another duplicate record every time it participates in a larger run.

The record should still be uniquely identifiable and safe to reuse. A probe should not grab an arbitrary production-like record simply because it looks convenient.

3. Workflow-stage probes

Multi-step processes do not always begin from a clean slate. A user may have already completed onboarding, added an item to a cart, uploaded a document, or advanced an application to the next stage.

Has onboarding already been completed?
Does the current order already contain a test item?
Is this application awaiting review or still in draft?

The probe allows the mission to identify the current stage and take the shortest valid route to the state under test.

4. Role and permission probes

Different users may see different navigation, controls, and actions.

Does the current user have access to team administration?

A mission can use that information to confirm it has the intended test identity, switch to the correct account, or fail early with a useful prerequisite error. That is much clearer than blindly searching for an administrative control that will never appear.

5. Feature-availability probes

Features may vary by subscription tier, feature flag, region, staged rollout, or environment.

Is the new reporting experience enabled for this account?

This can help a mission select the correct interface path. However, the expected availability must remain explicit. If the test environment is supposed to have the feature, its absence should be a failure rather than a reason to silently skip the test.

6. Interface-variant probes

The same capability may appear differently across responsive layouts, experiments, or incremental redesigns.

Is navigation displayed in the sidebar or inside a collapsed menu?
Is the account selector rendered as tabs or as a dropdown?

A probe can identify the active interface variant and let the mission use the appropriate interaction without duplicating the entire test.

7. Interruption probes

Cookie banners, announcements, onboarding prompts, confirmation dialogs, and unsaved-change warnings can obstruct the intended workflow.

Is an expected dialog blocking the page?

The mission can dismiss or respond to a known interruption before continuing. Unexpected interruptions should still be reported, since they may reveal a real product or environment issue.

8. Asynchronous-readiness probes

Some actions do not complete immediately. Reports need to generate, documents need to process, messages need to arrive, and background jobs need to finish.

Has the uploaded document finished processing?
Has the verification email arrived?
Is the generated report ready to view?

Repeated probes can replace arbitrary fixed waits with state-aware progression. The mission continues when the application is ready, fails after a clear limit, and avoids waiting longer than necessary.

9. Cleanup probes

Cleanup should be safe even when an earlier step fails or another mission has already removed the test data.

Does the temporary workspace still exist?

If it exists, remove it. If it does not, cleanup is already complete. This makes teardown more resilient and prevents a missing resource from obscuring the result of the test that actually mattered.

10. Partial-completion probes

An action can succeed even when the test runner misses the confirmation. A network delay might hide the response after an account was created or a payment was submitted.

Was the account created even though the confirmation step timed out?

Before retrying a potentially destructive or duplicate-producing action, the mission can inspect the application and determine whether the first attempt actually succeeded. It can then resume, clean up, or report the ambiguous state accurately.

Probe the route, not the result

Probes are powerful precisely because they introduce conditional behaviour. That also makes them something to use deliberately.

Consider these two instructions:

If the user is already logged in, skip the login steps and continue testing invoice creation.
If invoice creation is unavailable, skip that part of the test.

The first adapts the route while preserving the test's purpose. The second removes the very expectation the mission was supposed to verify.

A useful probe should follow a few rules:

  • Ask a narrow, observable question. “Is the user authenticated?” is better than “Does everything look ready?”
  • Be cheaper than the work it may skip. There is little value in spending more time and tokens proving that setup is unnecessary than simply performing the setup.
  • Select a route without weakening the expectation. A probe can change how the mission reaches the assertion, not whether the assertion matters.
  • Treat unexpected states as information. If the probe discovers something outside its known outcomes, the mission should report it clearly rather than improvising indefinitely.
  • Make the chosen path visible in the report. A reviewer should be able to tell whether setup ran, was skipped, or failed its prerequisite check.
  • Keep probe scope bounded. A simple state check should not turn into a second open-ended agentic mission.

These constraints keep probes from becoming a way to excuse flaky tests. The goal is not to make a mission permissive. It is to make the mission aware of what has already happened.

Reusable missions for context-aware testing

The most interesting part of probes is not the login time they save. It is what they suggest about the structure of agentic test suites.

A traditional scripted test is primarily a sequence: do this, then this, then this. Its starting conditions are usually established outside the sequence or assumed to be true.

An agentic mission can carry both an objective and enough awareness to determine how to reach it from the application's present state. It does not need to abandon determinism or invent a new goal. It can simply avoid repeating work that the environment proves has already been completed.

That creates a useful middle ground:

  • More independent than a tightly coupled sequence of tests
  • More efficient than resetting and rebuilding every prerequisite
  • More maintainable than keeping parallel versions for every context
  • More transparent than hiding complex dependencies in orchestration

For Testronaut™, probes make it possible to author missions that are atomic when they need to stand alone and cooperative when they join a larger flight plan. Like a reusable rocket, the mission keeps everything it needs for an independent launch without requiring every run to rebuild and repeat the same journey from the ground up.

A good mission does not need to relaunch every stage of the journey. Sometimes it just needs to look through the window, determine where it already is, and continue from there.