ModulesExploreDebugCICapstoneAll courses

Explore, Debug and CI

Send an agent to hunt for bugs, prove what it finds with failing tests, read the evidence a failed run leaves behind, and make the suite run on every pull request.

Module 5 of 5 · AI-Powered E2E Testing

Module 05 ~45 min read + capstone TypeScript + YAML

What you will learn

Prerequisites: the TaskFlow project from Module 1 and the test vocabulary from Module 2. A GitHub account for the CI section.

1. Exploring without a test

So far you have told the agent exactly what to do. e2e explore turns that around: you give a goal, not a test file. The agent plans a step, operates the app, records findings, and repeats until the goal is covered or a budget runs out. Then it writes an assessment. It uses the same model, tools and action rules as agent.act(), and a failed step does not stop the run.

npx e2e explore 'Explore the trial sign-up like a first-time visitor and report anything off. Stop after the welcome screen.' npx e2e explore # default goal: "Explore the app and find bugs" npx e2e explore --max-steps 4 --headed # fewer steps (1-12, default 8), visible browser npx e2e explore --agent ux 'Review onboarding as a first-time user' npx e2e explore --session admin 'Explore the admin settings and find bugs'
Terminal output of e2e explore reporting a trivial warning about the misspelled word Recieve
Our explore run on TaskFlow. The agent walked the sign-up flow in 4 actions and flagged the typo we planted in Module 1 as a trivial warning, with expected and actual text, the steps to reproduce, and a screenshot as evidence. Screenshot: course run, e2e 0.18.0.
The evidence screenshot: the sign-up form with the Recieve product updates checkbox
The finding's evidence image, saved under .e2e/artifacts/web/explore-…/attempt-0/finding-1.png. Always compare a finding with its evidence before acting on it. Screenshot: course run.

Write a good goal

Name the flow and the user: "Explore checkout as a first-time buyer and check that totals match the cart." Add limits such as "stop before placing the order".

Issue or warning

An issue fails the run with exit code 1. A warning (like our typo) does not. The report lists the most severe issues first.

Budgets

--max-steps (1–12) and --timeout (3–15 minutes) cap the cost. Our run used 6 model calls and about 30k tokens.

From finding to regression test

A finding is a claim; a test is proof. Turn each real finding into a test that asserts the expected behaviour. It fails now, and passes once the bug is fixed, so the bug cannot quietly come back:

// tests/copy.e2e.ts import { test } from '@e2e-dev/web'; import { expect } from 'e2e'; test('the updates checkbox is spelled correctly', async ({ app, screen }) => { await app.open('/'); await screen.getByRole('button', 'Start free trial').click(); await expect(screen.getByRole('checkbox', 'Receive product updates')).toBeVisible(); });

Bug bashes and personas

A bug bash runs many explore sessions at once. You give your coding agent one prompt, such as "Bug bash the checkout and account settings on this branch." Following the e2e skill, it writes 5–10 one-sentence charters (first-time user, numbers and copy, edge input, state after reload, error paths) and runs one e2e explore per charter, four at a time. Then it merges the findings and rejects those with a known cause. For each finding left, it writes a repro test. A finding counts as a bug only when its test fails with ASSERTION_FAILED. Repro tests are tagged bugbash; keep them out of the merge gate with npx e2e run --exclude-tag bugbash. The full procedure is in npx e2e guide bug-bash.

Each charter can use its own persona: a named entry under agents in the config with its own model, system prompt (how the agent should behave) and context (what your app calls things). Select one with --agent:

agents: { default: { model: google('gemini-3.8-flash') }, buyer: { model: google('gemini-3.8-flash'), system: 'You are a first-time buyer.' }, admin: { model: google('gemini-3.8-flash'), system: 'You manage orders and approve refunds.' }, },

Agents do not inherit from each other: give every entry the model and context it needs. npx e2e run --agent buyer,admin tests/checkout.e2e.ts runs the same test once per persona.

2. Debugging a failed run

In Module 2 you met this test. It expects the wrong error text, so it fails. It is all locators, so it needs no model, which makes it a cheap way to practise reading failures:

// tests/broken.e2e.ts import { test } from '@e2e-dev/web'; import { expect } from 'e2e'; test('rejects an invalid email', async ({ app, screen }) => { await app.open('/'); await screen.getByRole('button', 'Start free trial').click(); await screen.getByRole('textbox', 'Full name').fill('Ada Lovelace'); await screen.getByRole('textbox', 'Work email').fill('not-an-email'); await screen.getByRole('button', 'Start my trial').click(); await expect(screen.getByRole('alert')).toHaveText('Email is invalid'); });
Terminal output showing ASSERTION_FAILED with expected text Email is invalid and observed text Please enter a valid email.
The failure block: the error code, the locator, what was expected and what was observed, the URL, and the path to the screen dump. Exit code 1. Screenshot: course run.

Read it top to bottom. ASSERTION_FAILED means the app was reachable and the locator matched (match count 1), but the text differs. So the app is fine; the test is wrong. If the locator had matched nothing you would see observed: no node, which points to a wrong role or name, or to a step that never happened. We hit exactly that while writing this course: with type="email" on the input, the browser's own validation blocked the submit, so no alert ever appeared.

Exit codes

ExitMeaningRetry the CI job?
0Every selected test passed, was flaky, or was skipped—
1A test failed or timed out (for explore: an issue was reported)Only if the test is flaky
2CLI, config, or agent policy errorFix the reported problem first
3Engine, app process, model provider, or artifact failureIf the cause is temporary
4Internal runner errorReport it
130Interrupted (Ctrl-C or a CI signal)—

The evidence

Screen at failure. This is the accessibility tree the agent reads, one node per line (#id role "name" [states]). It is the fastest way to see what the page really offered: here the alert is present, with different text.

# Screen at failure url: http://127.0.0.1:4173/ revision: b1 viewport: 1280x720 nodes: 15 #root document "TaskFlow" #n53 banner #n54 text="TaskFlow" #n55 navigation #n56 link "Pricing" href="/" #n57 main #n58 heading "Create your account" #n59 text="Full name" #n60 textbox "Full name" value="Ada Lovelace" #n61 text="Work email" #n62 textbox "Work email" value="not-an-email" #n63 text="Recieve product updates" #n64 checkbox "Recieve product updates" #n65 alert "Please enter a valid email." #n66 button "Start my trial" [focused]
Failure screenshot of the sign-up form showing Please enter a valid email.
The screenshot taken when the assertion failed: the red message reads "Please enter a valid email.", not "Email is invalid". Screenshot: course run (screenshots/001-failure.png).

Playwright trace. On the web engine, a failed attempt also keeps a Playwright trace, trace/trace.zip. Open it with npx playwright-core show-trace <path>/trace.zip (or drop it on trace.playwright.dev) to scrub through a filmstrip and see every action with before and after snapshots, console output and network calls.

Playwright trace viewer showing the action list, filmstrip and a DOM snapshot of the sign-up form with the error message
The trace of the failed test in the Playwright trace viewer: the filmstrip on top, the actions on the left, the page snapshot in the middle. Screenshot: course run.

Where the evidence lives: 0.18 vs the docs

In e2e 0.18.0 (stable, used here) each failed attempt writes to .e2e/artifacts/<target>/<test>/default/attempt-0/: failure/screen.txt, screenshots/001-failure.png and trace/trace.zip, and the failure block prints a screen path. The official docs describe the 0.19 nightly, which writes .e2e/results/<test>/trace.md: one readable Markdown page per failed test, with the error, every step, the app log, the agent's last turns, the screen at failure and links to the screenshots. The failure block then prints a trace path instead. The habit is the same in both: open the path the failure block prints.

Debugging flags

FlagUse it when
--headedYou want to watch the browser (or device) while it runs.
--video, --video=retain-on-failureYou need a recording, for example from CI. Videos are not masked; check them before sharing.
--trace [mode]Which attempts keep a trace: off, on, retain-on-failure, on-first-retry, on-all-retries.
--debugPrints phase timings and the agent step table to stderr: what the agent did, and how long it took.
--ai-traceWrites every model request and response to .e2e/ai-trace.json; summarise it with npx unbox-ai summary .e2e/ai-trace.json --run 1.
--no-cacheA replayed step behaves oddly; force the agent to run live. Replayed steps make no model call and leave no AI trace.
--last-failedRerun only the tests the previous run did not pass, read from .e2e/report.json.
--reporter jsonPrint the report to stdout for scripts, or let your coding agent read .e2e/report.json.

Let a coding agent fix it

e2e init installed the e2e skill (.agents/skills/e2e/, linked into .claude/skills/) and the e2e mcp server. A coding agent such as Claude Code, Codex or Cursor can therefore run the suite, read report.json and the screen dumps, open a live app session over MCP to check locators, and propose the fix. Prompts like "why did this run fail?" or "add a test for checkout" work without more setup. For agents that need a pointer, add this line to AGENTS.md:

End-to-end tests use e2e; read .agents/skills/e2e/SKILL.md before writing or running one.

3. Secrets and the model

Agent steps send what is on screen to a model provider. e2e draws a clear boundary:

Redaction has limits

It does not cover transformed values, images or app logs, and videos mask nothing. Use test accounts, never real customer credentials.

4. Continuous integration

A CI job does the same steps on any service that runs Node: install dependencies, install browsers (or boot a device), and run npx e2e run. This is the web workflow from the e2e docs, verbatim. It assumes the config has an app.command that starts the app, and that the Vercel AI Gateway makes the model calls:

# .github/workflows/e2e.yml name: e2e on: pull_request: push: branches: [main] permissions: contents: read concurrency: group: ${{ github.workflow }}-${{ github.ref }} cancel-in-progress: ${{ github.event_name == 'pull_request' }} jobs: e2e: runs-on: ubuntu-latest steps: - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 - uses: pnpm/action-setup@9fd676a19091d4595eefd76e4bd31c97133911f1 # v4.2.0 - uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0 with: node-version: 26 cache: pnpm - run: pnpm install --frozen-lockfile - name: Install browsers run: pnpm exec e2e-web install chromium --with-deps - name: Run e2e run: npx e2e run --reporter list,junit env: APP_URL: http://127.0.0.1:3000 AI_GATEWAY_API_KEY: ${{ secrets.AI_GATEWAY_API_KEY }} E2E_USER_ADMIN_USERNAME: ${{ secrets.E2E_USER_ADMIN_USERNAME }} E2E_USER_ADMIN_PASSWORD: ${{ secrets.E2E_USER_ADMIN_PASSWORD }} - name: Upload report if: ${{ !cancelled() }} uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 with: name: e2e-report path: | .e2e/report.json .e2e/junit.xml if-no-files-found: warn - name: Upload artifacts if: ${{ !cancelled() }} uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1 with: name: e2e-artifacts path: .e2e/results if-no-files-found: warn retention-days: 7

Pinned actions

Each action is pinned to a commit SHA, and dependencies come from the lockfile. e2e-web install chromium --with-deps installs the browser that the engine's Playwright version expects, with its Linux libraries.

Secrets

The model key and test-account credentials come from repository secrets. With npm and a Google key, use npm ci, npx e2e-web install chromium --with-deps and GOOGLE_GENERATIVE_AI_API_KEY instead.

Always upload

if: ${{ !cancelled() }} uploads evidence even when tests fail, so a test that failed and then passed on retry keeps its evidence.

On 0.18, upload .e2e/artifacts

The workflow uploads .e2e/results, the 0.19 layout. On 0.18.0 the evidence is under .e2e/artifacts, so set path: .e2e/artifacts, or simply path: .e2e as the docs' mobile workflow does. Also note that in CI the runner ignores reuseExisting; it starts your app itself, or reports APP_ALREADY_RUNNING.

Results on the pull request

The @e2e-dev/github reporter posts one comment per run with failures, source links and links to the workflow artifacts. Reruns update the same comment, and it also writes the job summary.

npm install --save-dev @e2e-dev/github
// e2e.config.ts import type { E2EConfig } from 'e2e'; import { web } from '@e2e-dev/web'; import { github } from '@e2e-dev/github'; export default { targets: [{ engine: web(), app: { url: 'http://localhost:3000' } }], reporters: ['list', github()], } satisfies E2EConfig;

Then give the workflow permission to comment, and pass the token to the test step:

permissions: contents: read pull-requests: write # ...in the "Run e2e" step: env: GITHUB_TOKEN: ${{ github.token }}

Mobile jobs follow the same shape: build a Release app for the simulator or emulator, boot the device, install the build, run npx e2e run --target ios (or android). iOS needs a macOS runner; Android needs Linux with KVM. The docs also have ready workflows for EAS Workflows, Bitrise and Codemagic.

Lab: find it, prove it, fix it

About 25 minutes, in your TaskFlow project.

1

Explore

Run the explore command from section 1. Confirm the typo finding and open its evidence image. Note the exit code with echo $?: it is 0, because the finding is a warning, not an issue.

2

Prove it

Add tests/copy.e2e.ts from section 1 and run it. It must fail. Read the failure block and the screen dump: which node did the locator miss?

3

Debug the broken test

Add tests/broken.e2e.ts, run it, and open the three pieces of evidence: the screen dump, the failure screenshot and the trace (npx playwright-core show-trace). Fix the expected text.

4

Fix the app and rerun

Correct "Recieve" in app/index.html, then run npx e2e run --last-failed. Both tests should now pass.

Expected results
  • Step 2 fails with ASSERTION_FAILED: expect.toBeVisible failed and observed: no node (match count 0), exit code 1: the tree shows checkbox "Recieve product updates", not "Receive".
  • Step 3: change the expected text to 'Please enter a valid email.' (it matches exactly, including the full stop).
  • Step 4: --last-failed reruns only the two tests from report.json; exit code 0.

Capstone: test your own app

Two to three hours. Use an app you are building, a course project, or one of the e2e example projects (Vite, Next.js, Astro, Expo, SwiftUI, Jetpack Compose, Kotlin Multiplatform, Flutter).

1

Set up

Install e2e, configure a model, and let the runner start your app with app.command (web) or point a mobile target at your build.

2

One mixed test for your most important flow

A goal (agent.act) for the flow, at least one agent.assert, and at least one exact expect. Run it twice and show the cache replay on the second run.

3

One explore session, one regression test

Run e2e explore with a written goal. Pick one finding, prove it with a failing test (or explain, with evidence, why it is not a bug).

4

CI

Add .github/workflows/e2e.yml, store the model key as a repository secret, open a pull request, and show a green run with the report uploaded as an artifact. Bonus: the @e2e-dev/github comment.

Deliverable

A repository link and a one-page write-up: the flows covered, the explore finding and its test, one failure you debugged (with the evidence you used), and the token cost of a full run with and without cache.

Knowledge check

Pick one answer per question, then check your score.

1. An explore run reports one trivial warning and no issues. What is its exit code?

Why: Only an issue (or nothing explorable) gives exit code 1. Warnings are reported but do not fail.

2. In a bug bash, when does a finding count as a confirmed bug?

Why: Findings are claims. A failing repro test is the proof, and it doubles as the regression test after the fix.

3. A failure says observed: text "Please enter a valid email." (match count 1). What does this tell you?

Why: A match count of 1 means the locator resolved. Compare expected and observed text; here the test, not the app, was wrong.

4. A replayed agent.act step behaves strangely and you want the agent to run it live. Which flag?

Why: --no-cache turns replay off for the run, so the model drives the step again.

5. Why do the CI upload steps use if: ${{ !cancelled() }}?

Why: A failing step normally stops later steps. This condition keeps the upload running for every run that was not cancelled.

Self-check

Answer in your own words first, then open the model answer.

1. Why is a finding from e2e explore not yet a bug?

The explorer can be wrong: a limit of the explorer, the local environment, intended design or seed data can explain it. Compare it with its evidence, then write a test that asserts the expected behaviour. Only a failing test proves the bug.

2. You see observed: no node. Name three likely causes.

A wrong role or accessible name, text that is not exact, the element sits in an iframe, or an earlier step never happened (for example, native form validation blocked the submit). Read the screen dump to see what was there.

3. Why are screenshots withheld after a secret fill?

Once a secret has been typed, the app could display it anywhere on the screen, outside the masked field. Withholding pixels for the rest of the attempt keeps the value out of model input and reports.

4. Exit code 3 in CI: retry or fix?

Code 3 is an engine, app process, model provider or artifact failure. Retry if the cause is temporary (for example a provider outage); otherwise fix the environment, such as a missing browser or a wrong key.

References

  1. e2e documentation: Exploring without a test, Bug bashes, Agents and personas, Debugging a run. Accessed 9 October 2026.
  2. e2e documentation: Security model, Signing in, Coding agents. Accessed 9 October 2026.
  3. e2e documentation: Continuous integration, GitHub Actions, Pull request comments, CLI reference. Accessed 9 October 2026.
  4. Playwright: Trace viewer.
  5. GitHub Docs: Using secrets in GitHub Actions.

Image credits

All screenshots are from runs made for this course with e2e 0.18.0 and the TaskFlow demo app.

Summary

Key takeaways

  • e2e explore takes a goal, drives the app and reports findings; issues fail the run, warnings do not.
  • A finding becomes a bug only when a repro test fails; bug bashes run many explorers and prove each finding.
  • Read failures top-down: error code, expected vs observed, then the screen dump, screenshot and trace.
  • Exit codes tell CI what to do: 1 test failure, 2 config, 3 environment or provider, 4 runner bug.
  • Passwords are opaque Secret handles; secret values are redacted and screenshots stop after a secret fill.
  • On GitHub Actions: install, install browsers, run with --reporter list,junit, always upload the evidence, and optionally comment on the PR.

You have finished the course

Mark every module complete and claim your certificate. Then do the capstone; it is the best proof of what you learned.

Related on this site: CI/CD (DevOps Lab) · Testing AI-generated code (Vibe Coding)