ModulesViewportsLabQuizAll courses

Web Testing

Everything the web engine adds on top of the core API: getting the app running, testing several screen sizes, controlling the network, signing in once, checking your API, and catching visual regressions.

Module 3 of 5 · AI-Powered E2E Testing

Module 03 ~40 min read + lab Web engine (Playwright)

What you will learn

Prerequisites: Module 2: Core Concepts and the TaskFlow project from the Module 1 lab.

Everything here was run

Every test in this module ran on TaskFlow for the course. The viewport, mock and API tests below passed on e2e 0.18.0 without any model calls. Visual comparison needed the 0.19 nightly; section 7 explains why.

1. Getting the app running

A web target needs an app.url. You can start the app yourself, or add app.command and let the runner do it.

Point at a running app

app: { url: process.env.APP_URL ?? 'http://localhost:3000' }. Reading APP_URL lets you aim the same suite at a preview deployment without editing the file.

Let the runner start it

Add command. The runner spawns it, waits until app.url answers (60 s by default, startupTimeout), and stops it when the run ends, fails or is interrupted.

app: { url: process.env.APP_URL ?? 'http://localhost:3000', command: { executable: 'npm', args: ['run', 'dev'], reuseExisting: true, // use a server you already started (ignored in CI) log: '.e2e/logs/app.log', // keep the app's stdout and stderr }, },

Let the runner pick the port. With port 0 the runner chooses a free port and substitutes it for {port} in the command's args and env. Use 127.0.0.1 (or [::1]); localhost:0 is rejected. Tests read the real address from app.baseUrl.

app: { url: 'http://127.0.0.1:0', command: { executable: 'pnpm', args: ['dev', '--port', '{port}'], env: { PORT: '{port}' } }, },

Two gotchas

The command runs without a shell and inherits only PATH, HOME and temp variables, so pass anything else in env. And with Next.js 16 dev servers, add allowedDevOrigins: ['127.0.0.1'] to next.config.ts when the target opens 127.0.0.1, or the page renders but never hydrates.

2. Several viewports and browsers

The browser and the initial viewport are options of the engine; the app belongs to the target. To test several sizes, add one target per size. Every test runs once per target, and targets that declare the same app.command share one server.

// e2e.config.ts (the config used for the run below) import type { E2EConfig } from 'e2e'; import { web } from '@e2e-dev/web'; const app = { url: 'http://127.0.0.1:4174', command: { executable: 'python3', args: ['-m', 'http.server', '4174', '--bind', '127.0.0.1', '--directory', 'app'], reuseExisting: true, }, }; export default { targets: [ { name: 'desktop', engine: web(), app }, // 1280 x 720 default { name: 'phone', engine: web({ viewport: { width: 390, height: 844 } }), app }, ], } satisfies E2EConfig;
Terminal output: two targets, desktop and phone, four test files and six tests passed in 2.9 seconds
Three tests, two targets, six results: each file ran on desktop and phone in parallel. No line mentions a model, because none of these tests uses agent. Screenshot: course run, e2e 0.18.0.
TaskFlow landing page at phone width, 390 pixels
TaskFlow at the phone target's 390 px width. Screenshot: course demo app.

3. The browser fixture

screen, app and expect work on every platform. For browser-only operations, import test from @e2e-dev/web and take the browser fixture. In a suite that also has device targets, add { requires: ['browser'] } to these tests so they are skipped on phones.

TaskCall
Mock an APIbrowser.route('**/api/quote', route => route.fulfill({ json: {...} })), registered before app.open. Each handler calls exactly one of fulfill, continue, fallback or abort.
Wait for a real responsePromise.all([browser.waitForResponse('**/api/orders'), button.tap()])
Handle a dialogconst off = await browser.onDialog('accept'); … await off(); An unhandled alert/confirm fails the next step.
Download a fileconst file = await browser.waitForDownload(() => link.tap()); The file is stored as a test artifact.
Reach into an iframebrowser.frameLocator('#payment-frame').getByLabel('Card number').fill(...)
Read page statebrowser.evaluate(() => localStorage.getItem('flag')); the function is serialized, so pass data as the second argument.
URL and titleexpect(browser).toHaveURL('/dashboard'), expect(browser).toHaveTitle(/Dashboard/)
Cookies, init scriptsbrowser.setCookies([...]), browser.addInitScript(fn, arg)

Worked example: a mocked price list

TaskFlow's lab adds a static app/api/plans.json. This test replaces the real response before the page opens, then asks the page to fetch it. The page sees only the mock.

// tests/browser.e2e.ts import { test } from '@e2e-dev/web'; import { expect } from 'e2e'; test('the page sees a mocked price list', async ({ app, browser }) => { await browser.route('**/api/plans.json', async (route) => { await route.fulfill({ json: [{ id: 9, name: 'Mocked', cents: 4200 }] }); }); await app.open('/'); const plans = await browser.evaluate(() => fetch('/api/plans.json').then((r) => r.json())); expect(plans).toEqual([{ id: 9, name: 'Mocked', cents: 4200 }]); }); test('the call to action fits every screen', async ({ app, screen, browser }) => { await app.open('/'); await expect(browser).toHaveTitle('TaskFlow'); await expect(screen.getByRole('button', 'Start free trial')).toBeVisible(); const width = await browser.evaluate(() => window.innerWidth); expect([1280, 390]).toContain(width); });

Mocks make agent tests reliable too: an agent.act goal against a mocked checkout always sees the same prices, so its cached replay stays valid.

4. Signing in once

Logging in at the start of every test is slow and spends model calls if an agent does it. Instead, a setup test signs in once and saves a named session (cookies, local storage and IndexedDB). Other tests declare the session and start signed in.

// tests/auth.setup.e2e.ts import { test } from '@e2e-dev/web'; import { expect, credentials } from 'e2e'; test.setup('authenticate as admin', { sessions: ['admin'] }, async ({ app, screen, session, browser }) => { const admin = credentials.user('admin'); await app.open('/login'); await screen.getByLabel('Username').fill(admin.username); await screen.getByLabel('Password').fill(admin.password); await screen.getByRole('button', 'Sign in').tap(); await expect(browser).toHaveURL('/dashboard'); // prove it worked before saving await session.save('admin'); }); // tests/dashboard.e2e.ts test('the dashboard opens directly', { session: 'admin' }, async ({ app, browser }) => { await app.open('/dashboard'); await expect(browser).toHaveURL('/dashboard'); });
// e2e.config.ts: declare the credential; the password comes from the environment credentials: { admin: { username: 'admin@example.test', password: process.env.ADMIN_PASSWORD ?? '' }, },

The password is a handle, not a string

admin.password is an opaque Secret with no .value. You can pass it to fill() or as an agent.act param: the model sees only its name and purpose, the runner types the value, and the value is redacted from model input and reports. After a secret is filled, screenshots are withheld for the rest of the attempt. Sessions last for one run and are encrypted on disk.

The runner always runs the setup a test depends on, even if you select only tests/dashboard.e2e.ts. If the setup fails, its dependants are skipped with cause setup-failed.

5. API tests in the same suite

API checks live in the same files and run in the same command. A test that takes only app calls no model. Use plain fetch against app.baseUrl and the value matchers; toMatchSchema accepts any Standard Schema (Zod, Valibot, ArkType) and returns the value already typed.

// tests/plans-api.e2e.ts import { test, expect } from 'e2e'; import { z } from 'zod'; const Plan = z.object({ id: z.number(), name: z.string(), cents: z.number() }); test('GET /api/plans.json lists three plans', async ({ app }) => { const response = await fetch(new URL('/api/plans.json', app.baseUrl)); expect(response.status).toBe(200); expect(response.headers.get('content-type')).toContain('application/json'); const plans = expect(await response.json()).toMatchSchema(z.array(Plan)); expect(plans).toHaveLength(3); expect(plans.map((plan) => plan.name)).toEqual(['Free', 'Team', expect.any(String)]); });

For state the API writes after it answers, expect.poll re-reads until a matcher passes. To call the API as a signed-in user, write a fixture that forwards browser.cookies() with each request (see the API page of the docs).

6. Visual comparison

Version note: needs 0.19 (nightly)

toHaveScreenshot is documented on the e2e site but is not in the 0.18.0 stable release: there the test fails with expect(...).toHaveScreenshot is not a function. It works in the nightly build we used, 0.19.0-nightly-20261008205824. Install it in a separate branch or project until 0.19 ships:

npm i -D e2e@nightly @e2e-dev/web@nightly

expect(screen).toHaveScreenshot(name) compares the whole screen with a PNG stored next to the test; expect(locator).toHaveScreenshot(name) compares one element. Each target and operating system gets its own file, because fonts render differently.

// tests/visual.e2e.ts import { test, expect } from 'e2e'; test('landing page looks right', async ({ app, screen }) => { await app.open('/'); await expect(screen).toHaveScreenshot('landing.png'); await expect(screen.getByRole('button', { name: 'Start free trial' })).toHaveScreenshot('cta.png'); });

Run 1: no baseline. The first run has nothing to compare against, so it writes the screenshot and fails on purpose. Look at the image, commit it, run again. (Our test has two screenshots, so the second run wrote cta the same way; the third run passed.)

Terminal: toHaveScreenshot failed, no stored screenshot at tests/visual.e2e.ts-snapshots/landing-web-darwin.png; wrote this run's there
The first run writes tests/visual.e2e.ts-snapshots/landing-web-darwin.png and asks you to check and commit it. Screenshot: course run, e2e 0.19 nightly.

Then we changed one CSS value: the button colour from green to blue. Every functional test still passed. The visual test did not.

Terminal: 4299 pixels, 0.47 percent of the image, differ; diff, actual and expected image paths
4,299 pixels (0.47% of the image) differ. The failure names three images in the test's results folder. Screenshot: course run, e2e 0.19 nightly.
Expected: green Start free trial button
landing-expected.png (baseline)
Actual: blue Start free trial button
landing-actual.png (this run)
Diff: changed pixels of the button in red
landing-diff.png: changes in red

If the change was intended, accept it with npx e2e run --update-snapshots, review the new PNGs and commit them. To tolerate noise or hide content that changes every run:

await expect(screen).toHaveScreenshot('dashboard.png', { maxDiffPixelRatio: 0.01, // up to 1% of pixels may differ mask: [screen.getByTestId('last-updated')], // painted over before comparing });
OptionDefaultMeaning
threshold0.2How far one pixel's colour may drift and still count as the same (0 to 1)
maxDiffPixels0How many pixels may differ
maxDiffPixelRatio0What share of pixels may differ (0 to 1)
mask / maskColornone / #ff00ffLocators painted over before comparing

On the web, animations are stopped and the text caret hidden before the capture. In CI, baselines must come from the machine that renders them (a Linux runner for web), so generate them there with --update-snapshots and commit the result.

7. Coming from Playwright

You do not have to switch in one go. @e2e-dev/web ships its own pinned playwright-core, so your @playwright/test stays installed and both runners live in one project. e2e only picks up tests/**/*.e2e.ts, so keep the file patterns apart and move one spec at a time.

Playwrighte2e
use.baseURL, webServerapp: { url, command } on the target
projectstargets
use.viewport, use.browserNameweb({ viewport, browser })
use.httpCredentialsweb({ basicAuth })
project dependencies for sign-intest.setup + { session }
page.getByRole(...) etc.the same on screen; names match the whole string, so add { exact: false } where Playwright matched a fragment
locator.click()locator.tap() (click() is an alias)
page.locator('css'), page.frameLocatorbrowser.locator('css'), browser.frameLocator
reporter: 'html'no equivalent yet; .e2e/report.json every run, markdown reporter for a summary

A good first candidate is a flow whose steps change often, such as a wizard or checkout: keep navigation and final assertions deterministic and replace the brittle middle with one agent.act goal. The docs' migration page also offers a ready-made prompt that has a coding agent do this one spec at a time.

Lab: TaskFlow on two screens

About 20 minutes, using your Module 1 project. No model key is needed for steps 1 to 4.

1

Add a tiny API

Create app/api/plans.json:

[{"id":1,"name":"Free","cents":0},{"id":2,"name":"Team","cents":900},{"id":3,"name":"Business","cents":2400}]
2

Two targets

Change targets in e2e.config.ts to the desktop + phone pair from section 2 (keep your agents block and port). Run npx e2e list and check every test now appears twice.

3

Add the tests

Add tests/browser.e2e.ts (section 3) and tests/plans-api.e2e.ts (section 5). Run npx e2e run tests/browser.e2e.ts tests/plans-api.e2e.ts: six passes.

4

Break it on purpose

Change "Team" to "Teams" in plans.json and rerun. Read the failure, then put it back. Run only the phone with --target phone.

5

Optional: visual baseline (nightly)

In a copy of the project, install the nightly, add tests/visual.e2e.ts from section 6, run until it passes, change the button colour in app/index.html, and open the -diff.png. Then accept or revert.

What you should see in step 4

Both targets fail the API test at the last assertion, with ASSERTION_FAILED: expected ["Free","Teams","Business"] to equal ["Free","Team","Any<String>"] (that is the exact message from our run). The schema check still passes, because the shape is fine; only the content changed. That is the point of having both a schema assertion and a value assertion.

Knowledge check

Pick one answer per question, then check your score.

1. You want every test to run on a desktop size and a phone size. What do you configure?

Why: One target per size is the documented way; each test then runs and reports once per target. setViewport is for one test that checks a resize, and the agent cannot resize.

2. When must browser.route be registered to mock the first page load's API call?

Why: A route intercepts requests made after it is registered, so register it before the page opens.

3. In a setup test, what does the model see when agent.act receives admin.password as a param?

Why: A Secret has no plaintext accessor. The runner authorizes the fill and redacts the value from model input and reports.

4. The first run of a new toHaveScreenshot test fails. Why?

Why: With nothing to compare against, the run writes the PNG and fails so that a person looks at it before it becomes the reference.

5. Your Playwright test uses getByRole('button', { name: 'Save' }) and the button reads "Save draft". What changes in e2e?

Why: e2e string names match the whole text, case-sensitively. { exact: false } gives Playwright's substring behaviour.

Self-check

Answer in your own words first, then open the model answer.

1. Why use port 0 instead of a fixed port?

Two checkouts or two CI jobs on one machine would fight over a fixed port. With http://127.0.0.1:0 the runner picks a free port, substitutes it for {port}, and tests read it from app.baseUrl. Cache entries stay valid when the port changes.

2. Why should a setup test assert that sign-in worked before session.save()?

Because save() stores whatever state exists, signed in or not. Without the check, a broken login would be saved and every dependent test would fail later with a confusing error.

3. When is a visual test better than a locator assertion?

When the bug is in how something looks rather than what it says: a colour, a layout shift, an overlapping element. Locator assertions would still pass in those cases, as our blue-button run showed.

References

  1. e2e documentation: Web, Working with the browser, Signing in, API, Visual comparison, From Playwright, Browser engine reference. Accessed 9 October 2026.
  2. e2e 0.18.0 bundled documentation (node_modules/e2e/docs), used to confirm which features exist in the stable release.
  3. Standard Schema: standardschema.dev; Zod: zod.dev.
  4. Playwright documentation: playwright.dev.

Image credits

All screenshots are from runs made for this course on 9 October 2026.

Summary

Key takeaways

  • app.command starts and stops the app; port 0 with {port} avoids conflicts.
  • One target per viewport or browser; every test runs once per target.
  • The browser fixture handles mocks, dialogs, downloads, iframes, cookies and page state.
  • A setup test signs in once and saves a session; passwords are Secret handles the model never sees.
  • API tests use fetch + toMatchSchema in the same run, with no model.
  • toHaveScreenshot (0.19 nightly) catches visual changes; accept intended ones with --update-snapshots.