---
title: "Agent Readiness Evals — test the tasks agents need to complete"
description: "Turn a real agent customer journey into a controlled evaluation. Compare docs, interfaces, prompts, models, and workflows under the same success criteria."
date: 2026-09-05
---

# Test the customer tasks agents need to complete.

Agent Readiness Evals turns a real buying or product workflow into a repeatable
evaluation. Change one thing, run the same task again, and see whether the agent
experience improved.

- [Join the waitlist](https://tansohq.com/#waitlist)
- [Talk to a founder](https://cal.com/katrina-laszlo/30-minute-meeting)

## Start with the job

Define the goal, starting state, allowed tools, and the evidence required for
success. Tanso then runs that task against the product experience being tested.

**Self-serve task example:** Choose the right plan, create an account, get
scoped access, and complete a purchase.

**Sales-led task example:** Qualify the product, book the right meeting, then
evaluate and integrate with granted access.

## What every run measures

Every run uses the same seven checkpoint names: **Discover, Understand, Try,
Sign up, Access, Pay, and Use**.

- **Completion:** Did the agent finish the customer task?
- **Friction:** Where did it backtrack, ask for help, or hit a human-only gate?
- **Latency:** How long did the task take?
- **Retries:** How often did it repeat an action or switch approaches?
- **Recovery:** Did the product provide a safe, machine-readable next step?
- **Failure mode:** What actually prevented success, with evidence attached?

## Compare treatments

Compare a baseline against a documentation change, a new API, a different
checkout, or another agent/model. The task and success criteria stay fixed.

Each result remains attached to its task definition, environment, agent/model,
timestamp, and evidence.

Each report row uses this schema:

- `checkpoint`: `Discover`, `Understand`, `Try`, `Sign up`, `Access`, `Pay`, or `Use`
- `status`: `ready`, `partial`, or `blocked`
- `evidence`: a source URL, request and response, or direct observation
- `blocker`: a specific blocker or `null`
- `recommended_next_step`: the highest-priority fix or verification step

## Controlled results are not production proof

Agent Readiness Evals reports how an agent performed in a defined run. It does not claim
to reconstruct every journey happening in production.

[Join the Tanso Agent Readiness waitlist](https://tansohq.com/#waitlist)

Agent Readiness pricing has not yet been published. The sanctioned human
handoff is [Talk to a founder](https://cal.com/katrina-laszlo/30-minute-meeting).
