ModulesLocatorsAgentCacheLabQuiz

Core Concepts: Goals, Assertions and Locators

Every e2e test is built from three kinds of step. Learn when each one is the right tool, what it costs, and how the replay cache makes agent steps nearly free on the second run.

Module 2 of 5 · AI-Powered E2E Testing

Module 02 ~40 min read + lab TypeScript

What you will learn

Prerequisites: Module 1 and the TaskFlow project from its lab, with one green run.

1. Three kinds of step

A test is a list of steps. Each step is one of three kinds, and the kind decides whether a model is called.

KindAPICalls a model?Use it when
Goalagent.actYes, unless the replay cache replays itYou care about the outcome, not the exact clicks
Assertion (judgement)agent.assert, agent.waitFor, agent.extractYes, every timeThe check is about meaning ("greets Ada by name") or the data is hard to locate
Locatorscreen and expectNoYou know the exact control or the exact value
Goal agent.act('sign up…') agent plans the clicks model: yes, cached later Judgement agent.assert('greets…') model reads the screen model: yes, every run Locator expect(status)… exact, retries until true model: never The locator check after the goal is also what lets the cache record the goal.
The sign-up test from Module 1, read as three steps. Figure: drawn for this course.

Await everything

Await every agent call, every screen action and every expect on a locator. If a test returns while a step is still running, it fails with STEP_NOT_AWAITED.

2. Anatomy of a test file

A test receives fixtures: ready-made objects the runner prepares for each attempt. You ask for the ones you need by name.

FixtureWhat it gives you
appThe app lifecycle on every platform: open(path), restart(), clearState(), back(), screenshot(label).
screenLocator queries and gestures on what is shown now.
agentact, assert, waitFor, extract.
browser (web) / device (mobile)Engine-specific controls: routes, cookies and dialogs on the web; permissions and location on devices. Import test from @e2e-dev/web or @e2e-dev/mobile to have them typed.

Group tests with describe, share set-up with hooks, and pass options as the second argument:

import { beforeEach, describe, test } from 'e2e'; describe('billing', { tags: ['billing'] }, () => { beforeEach(async ({ app }) => { await app.open('/billing'); }); test('upgrades', { retries: 2, timeout: 60_000 }, async ({ agent }) => { await agent.act('upgrade to Pro'); }); });
npx e2e run tests/signup.e2e.ts:12 # the test at line 12 npx e2e run --grep checkout # tests whose title matches npx e2e run --target web # one target only

0.18 and the docs

The docs show import { test, expect } from 'e2e'. That works in the 0.18.0 release too; the generated example imports test (plus describe and the hooks) from @e2e-dev/web instead, which adds the typed browser fixture. Both forms ran in our tests. Every API shown in this module exists in 0.18.0; we checked the package's type definitions.

3. Locators: screen and expect

A locator finds a known control without a model. Prefer queries a user would recognise: role and name first, label second, test id last. These are the same accessible names a screen reader announces, so a test that relies on them also nudges the app towards accessibility.

GroupMethods
QueriesgetByRole(role, name?), getByLabel, getByText, getByPlaceholder, getByDisplayValue, getByTestId
Refinement.filter({ hasText, has }), .first(), .last(), .nth(i), and queries chained inside a locator
Actionstap/click, fill, clear, press, check, selectOption, hover, dragTo, setInputFiles, longPress, swipe
Reads (no retry)textContent(), inputValue(), getAttribute(), isVisible(), count(), allTextContents()
// Role and name first, label second, test ID last. await screen.getByRole('button', 'Sign in').tap(); await screen.getByLabel('Email').fill('ada@example.test'); // Scope into one row. const row = screen.getByRole('listitem').filter({ hasText: 'Design review' }); await row.getByRole('button', 'Archive').tap(); // Assertions retry until they pass or time out. await expect(screen.getByRole('status')).toHaveText('2 remaining');

expect has two families of matcher:

Locator matchers (retry)

toBeVisible, toBeHidden, toBeEnabled, toBeChecked, toBeFocused, toHaveText, toContainText, toHaveValue, toHaveAttribute, toHaveCount, toHaveAccessibleName. Always await them.

Value matchers (immediate)

toBe, toEqual, toMatchObject, toContain, toHaveLength, toMatch, toBeGreaterThan and friends, for plain values such as the result of agent.extract. Also expect.poll and expect.soft.

One match, please

An action, a read, or a single-node assertion fails when its query finds more than one node (LOCATOR_AMBIGUOUS). Narrow it with .filter(), .first() or .nth().

4. The agent: goals and judgements

Goals: agent.act

One call, one goal. The agent chooses and performs the actions, returns when the goal is complete, and throws when it fails. You control the order of the goals; the agent controls the flow inside each one.

await agent.act('complete checkout with the test card'); await agent.act('invite {email} as an editor', { params: { email: 'ada@example.test' } }); // A value that changes every run: wrap it in unique() so the cache can still replay the step import { unique } from 'e2e'; const email = `ada+${Date.now()}@example.test`; await agent.act('sign up with email {email}', { params: { email: unique(email) } });

Judgements: assert, waitFor, extract

A judgement looks at the screen now. The judging model sees the statement, the current screen and the agent's context, but not the earlier steps or the acting agent's summaries. So a judgement cannot be talked into passing by an over-confident act.

import { expect } from 'e2e'; import { z } from 'zod'; await agent.assert('the dashboard shows a trial badge'); await agent.waitFor('a download link appears', { timeout: 120_000 }); const data = await agent.extract('every todo title', { schema: z.object({ titles: z.array(z.string()) }), }); expect(data.titles).toContain('Buy milk');
CallBehaviourFails with
assertOne judgement (plus one repair call if the answer is malformed). Attaches a screenshot to the report by default.ASSERTION_FAILED if false, ASSERTION_INCONCLUSIVE if the evidence is not on screen
waitForJudges again at most every interval (default 3 s) until timeout, and skips model calls while the screen is unchanged.STEP_TIMEOUT
extractReads data into any Standard Schema type (Zod, Valibot, ArkType) and validates it.ASSERTION_INCONCLUSIVE rather than a made-up placeholder

Judgements read a redacted text snapshot of the screen. For something only pixels show, such as a chart's direction, pass { vision: true } (snapshot plus screenshot) or { vision: 'only' }. Images add input tokens.

Judgement or locator?

For text a model wrote, check the meaning with agent.assert. When the exact text matters ("Please enter a valid email."), use a locator and toHaveText: it is free, instant and never ambiguous.

5. How an agent step works

  1. What the model sees. The instruction, the params, the steps done so far, and a redacted text snapshot of the screen: roles, names, text and states. Never raw HTML, cookies, headers or environment values. Password fields are masked; configured secrets appear as <secret:name>.
  2. What it may do. Tools the engine supports: tap, type, press, select, scroll, navigate, and pixel tools when the snapshot is not enough. The runner checks every action and counts it against the budget.
  3. Guard rails. Repeated calls and cycles are detected. After three failed actions in a row the model is told to change approach; after five, it must give a verdict.
  4. How it ends. With a verdict: passed, failed, or blocked (a prerequisite, the environment or a limit stopped it, such as STEP_BUDGET_EXHAUSTED or STEP_TIMEOUT).

Cost comes from model calls and tokens. assert and extract make one judgement; screenshots add tokens; act asks for a screenshot only when the text snapshot is not enough. The cheapest model call is the one the cache makes unnecessary.

6. The replay cache

When an agent.act step is followed by a check that passes (a locator assertion or an agent.assert), the runner records the actions. On the next run it replays them with no model call. If a control or the expected result no longer matches, because the app changed, the agent takes over from the current screen.

Second run of the sign-up test: agent.act finished in 615 ms, Cache 1 replayed, 782 tokens
The same sign-up test, run a second time. agent.act took 615 ms instead of 10.78 s and made no model call (Cache 1 replayed). The agent.assert still made its one call. Total: 782 tokens, against 15.9k on the first run. Screenshot: course run, e2e 0.18.0.
Summary saysMeaning
replayedThe recording completed without a model call.
handed offReplay started, then the agent took over (the app changed).
missedNo usable recording; the agent ran the step from the start.

7. Worked example: the invalid-email path

The sign-up form shows an error when the email is malformed. We know the exact controls and the exact message, so this test needs no model at all. A beforeEach hook opens the form for every test in the group.

TaskFlow sign-up form with Please enter a valid email shown in red
The state the test checks: the error is shown, and the form is still on screen. Screenshot: course demo app.
// tests/concepts.e2e.ts import { describe, beforeEach, test } from '@e2e-dev/web'; import { expect } from 'e2e'; import { z } from 'zod'; describe('sign-up form', () => { beforeEach(async ({ app, screen }) => { await app.open('/'); await screen.getByRole('button', 'Start free trial').click(); }); test('rejects an invalid email', async ({ screen }) => { await screen.getByLabel('Full name').fill('Ada Lovelace'); await screen.getByLabel('Work email').fill('not-an-email'); await screen.getByRole('button', 'Start my trial').click(); await expect(screen.getByRole('alert')).toHaveText('Please enter a valid email.'); await expect(screen.getByRole('heading', 'Create your account')).toBeVisible(); }); test('the agent reads the form fields', async ({ agent }) => { const form = await agent.extract('the labels of every text field on the form', { schema: z.object({ labels: z.array(z.string()) }), }); expect(form.labels).toContain('Work email'); expect(form.labels).toHaveLength(2); }); });
Terminal: both sign-up form tests pass, 818 tokens and 1 model call in total
Our run. The locator-only test passed in 1.2 s with no tokens. The extract made exactly one model call (818 tokens), and plain value matchers checked its result. Screenshot: course run, e2e 0.18.0.

Lab: three kinds of step on TaskFlow

About 20 minutes, in the TaskFlow project from Module 1.

1

Run the worked example

Save the file from section 7 as tests/concepts.e2e.ts and run npx e2e run tests/concepts.e2e.ts. Both tests should pass. Note the token count: only the extract test spends any.

2

Watch the cache work

Run npx e2e run tests/signup.e2e.ts twice. Compare the agent.act time, the Cache line and the token total between the two runs. Then run it once more with --no-cache and compare again.

3

Write a test in each style

Add three tests to concepts.e2e.ts for the empty-name case (submit with only an email):

  • (a) locators only;
  • (b) one agent.act goal followed by an expect;
  • (c) one agent.act followed by an agent.assert.

Run each one twice. Which is cheapest? Which one survives if you rename the button to "Begin trial"?

4

Break a judgement on purpose

Change the assertion in signup.e2e.ts to 'the welcome screen greets Grace by name' and run it. Read the error code, then put it back.

Model answer for step 3

The required attribute on the name field means the browser itself blocks submission and shows its own bubble, which is not in the page's accessible tree. A good locator check is therefore that the form is still there and the welcome screen did not appear:

test('(a) empty name, locators', async ({ screen }) => { await screen.getByLabel('Work email').fill('ada@example.test'); await screen.getByRole('button', 'Start my trial').click(); await expect(screen.getByRole('heading', 'Create your account')).toBeVisible(); await expect(screen.getByRole('status')).toBeHidden(); }); test('(b) empty name, goal + expect', async ({ agent, screen }) => { await agent.act('submit the sign-up form with only the email {email} and no name', { params: { email: 'ada@example.test' }, }); await expect(screen.getByRole('heading', 'Create your account')).toBeVisible(); }); test('(c) empty name, goal + judgement', async ({ agent }) => { await agent.act('submit the sign-up form with only the email {email} and no name', { params: { email: 'ada@example.test' }, }); await agent.assert('the sign-up form is still shown and no welcome message appears'); });

(a) is free and fastest but breaks on a renamed button. (b) costs model calls once, then replays from the cache, and the agent adapts to a renamed button. (c) also adapts but pays for a judgement on every run. In step 4 the judgement fails with ASSERTION_FAILED, because the screen greets Ada.

Knowledge check

Pick one answer per question, then check your score.

1. Which step is never cached and calls the model on every run?

Why: Judgements (assert, waitFor, extract) always run live. Only act is replayed, and locator steps never use a model.

2. Your agent.act step is never saved to the cache. What is the most likely reason?

Why: The cache records a goal only after a later check (a locator assertion or agent.assert) verifies the result.

3. What does the judging model see during agent.assert?

Why: Judgements are independent of the steps before them, and never see raw HTML, cookies, headers or environment values.

4. screen.getByRole('button').tap() fails because the page has three buttons. What is the fix?

Why: Actions need exactly one match. Add the accessible name, or refine the locator.

5. You need the list of plan prices from a pricing page as numbers. Which call fits?

Why: extract returns structured data validated against a schema; you then check it with value matchers.

Self-check

Answer in your own words first, then open the model answer.

1. Why should every agent.act be followed by an exact check?

Two reasons. The check proves the outcome precisely, instead of trusting the agent's own account. And the cache only records a goal once a later check passes, so without it every run pays for the goal again.

2. When is agent.assert better than toHaveText?

When the check is about meaning rather than exact characters: text written by a model, a greeting that may vary in wording, or "the cart shows two items" across a complex layout. When the exact text matters, toHaveText is cheaper and stricter.

3. Why wrap a timestamped email in unique()?

A cache entry is keyed on the instruction and its params. A value that changes every run would never match, so the step would always miss. unique() tells the cache to record a slot for the value instead of the value itself.

References

  1. e2e documentation: Core concepts, How agent steps work, Caching. Accessed 9 October 2026.
  2. e2e reference: test, expect, screen and Locator, agent, app. Accessed 9 October 2026.
  3. Standard Schema: standardschema.dev; Zod: zod.dev.
  4. W3C. WAI-ARIA 1.2: roles and accessible names. w3.org/TR/wai-aria-1.2.

Image credits

All screenshots are from runs made for this course with e2e 0.18.0. The diagram was drawn for this course.

Summary

Key takeaways

  • Three kinds of step: goals (agent.act), judgements (assert, waitFor, extract) and locators (screen + expect). Only locators never call a model.
  • Tests get fixtures (app, screen, agent, browser/device) and are organised with describe, hooks and options.
  • Find controls by role and name first; locator matchers retry, value matchers do not; actions need exactly one match.
  • Judgements see only the current screen and the statement, never raw HTML or secrets.
  • Follow every goal with an exact check: it proves the outcome and lets the cache replay the goal for free next time.