History & Walking Skeleton

Tests as Executable Evidence


Learning Objectives

  • You can run and read the supplied Deno, Vitest, and Playwright tests.
  • You can extend a test with a claim about observable behavior.
  • You can distinguish passing, failing, and ignored tests from untested properties.
  • You can explain what a successful local client build establishes.

In the previous chapter, we checked a response and followed its data into the page. Tests let us write down expectations and check them repeatedly after changes. The practice application’s tests cover the behavior we have been studying. The tools may be unfamiliar; begin with what each test does and expects.

Run one suite at a time in an isolated project

We will start by running the supplied suites, then look inside them. From the root of the practice application (practice/), run these commands one at a time. wsd-practice-test is an example Compose project name: choose a distinct name for your copy, different from its development Compose project, and use that same Compose test project name throughout.

Terminal window
docker compose -p wsd-practice-test -f compose.yaml -f compose.test.yaml run --build --rm api deno task test
docker compose -p wsd-practice-test -f compose.yaml -f compose.test.yaml run --build --rm --no-deps client deno task test
docker compose -p wsd-practice-test -f compose.yaml -f compose.test.yaml run --build --rm e2e-tests

Always retain both -f files and the separate -p name for database and browser tests. The override replaces database storage with temporary storage, removes published ports, and supplies the internal client/API addresses. The separate project keeps these services distinct from development.

Once a suite finishes, read its result before starting the next command. You can stop here and return to the individual test examples with the output beside you. See Docker’s documentation for merging Compose files and one-off services.

run creates a one-off runner, --build builds its image when needed, and --rm removes that runner afterward. Dependencies such as the database can remain running. The API and browser suites share this Compose test project’s state; separate suites are not automatically independent.

The client command uses --no-deps because component tests supply controlled responses rather than contacting the API. If a runner exits with a nonzero status or an assertion fails, start with the first failure. The sections below will help you connect that output to its test.

If you already have Deno on the host, the optional root deno task test shortcut runs the same three suites sequentially. Its helper derives a Compose test project name from the development Compose project name with -test appended. Use one naming workflow consistently when inspecting or cleaning up services; that derived Compose project name may differ from the explicit example above.

Read the direct Hono tests

Open api/src/app_test.ts. The test imports app through the API’s @src/ mapping without starting app-run.ts:

import { assertEquals } from "@std/assert";
import { app } from "@src/app.ts";
Deno.test("GET /health returns an OK response", async () => {
const response = await app.request("/health");
assertEquals(response.status, 200);
assertEquals(await response.json(), { status: "ok" });
});

Deno.test registers a named test. Its async callback contains the work to perform and the checks to make. The test first sends a request to /health, then uses assertEquals(actual, expected) to check the status and parsed JSON. await lets each asynchronous operation finish before its result is used.

Hono’s app.request() sends the request through the application without a listening TCP server. We can check the route here; checking access through a published host port would need a different test. The second supplied test checks an unknown route’s 404 response. See the Hono testing guide.

These tests make no database query, although importing the application still requires DATABASE_URL because it imports the shared database module. To run just these tests with the supplied configuration and no dependent services:

Terminal window
docker compose run --build --rm --no-deps api deno task test:unit

The API’s test:unit task discovers tests under src/; its full test task also discovers tests/.

Now try writing a test for a library-status route. Return only the test file: the activity supplies the application so you can focus on what to check. Your test must accept the correct implementation before it can earn credit for detecting the supplied faulty versions.

Make the test catch the regression

0 / 10 points

Task

Complete app_test.ts for this small Hono application:

import { Hono } from "hono";
const app = new Hono();
app.get(
"/library-status",
(c) => c.json({ branch: "Harbor", availableSeats: 7 }, 200),
);
export { app };

Requirements

The exercise starter supplies the imports, test callback, request, and JSON parsing. Complete its assertion TODOs with assertEquals. Your non-skipped test must:

  • pass for the application shown above,
  • require response status 200,
  • require branch "Harbor", and
  • require availableSeats to equal 7.

Submit

Submit only app_test.ts at the submission root. Keep its import of ./app.ts; the grader supplies the application separately. Grading first runs your test against the stated implementation and then checks that the same test detects single-behavior regressions. A test that always fails, is skipped, or does not execute cannot establish the required behavior.

Passing the reference earns 1 point. Detecting each of the wrong status, wrong branch, and wrong seat count earns another 3 points, independently, after the reference passes. Syntax/import errors, skipped tests, ordinary runtime errors, and timeouts do not count as detecting a regression.

Understand a failing test

An assertion compares an expected value with an actual value. Try making the health test fail on purpose: temporarily change the expected status to 201. The application returns 200, so the assertion should report a difference. Read the test name, assertion location, expected value, and actual value.

Here we know the cause because we changed the expectation ourselves. In other failures, the assertion location shows where the difference was detected; the source defect may be elsewhere.

Restore the correct expectation and rerun. When the test passes again, you have seen it detect a mismatch and accept the intended response. Here we repaired a wrong expectation. If the implementation were wrong instead, we would fix it while preserving the expectation of the intended behavior.

A passing health test tells us that this request returned the status and JSON we checked. Other behavior still needs its own assertions. An ignored test has not run its checks, and a syntax or import error may stop a test before it reaches an assertion. Deno’s testing documentation explains discovery, steps, and results.

A mixed test report

0 / 5 points

A room-search application has a database test that explicitly allows skipping when its connection setting is absent. A test run reports:

  • a route-status test completed successfully;
  • a database test was ignored because its connection setting was absent;
  • a browser test failed while submitting a search; and
  • no test mentions whether a search survives a page reload.

Which classification is accurate?

Read the PostgreSQL integration test

Next, open api/tests/todos_test.ts. This test repeats the kind of database investigation we did by hand. It needs the real migrated test database, so missing configuration or migrations will cause it to fail. Here is the first case, with the second case omitted to focus on inserting and reading a row:

import { assertEquals } from "@std/assert";
import { app } from "@src/app.ts";
import { sql } from "@src/database.ts";
import type { Todo } from "@src/todoRepository.ts";
Deno.test("GET /api/todos reads the current PostgreSQL data", async (t) => {
try {
await sql`SELECT 1`;
await t.step("includes a stored todo", async () => {
const name = `Test todo ${crypto.randomUUID()}`;
try {
const [saved] = await sql<Todo[]>`
INSERT INTO todos (name) VALUES (${name}) RETURNING id, name
`;
const response = await app.request("/api/todos");
assertEquals(response.status, 200);
const todos: Todo[] = await response.json();
assertEquals(todos.find((todo) => todo.id === saved.id), saved);
} finally {
await sql`DELETE FROM todos WHERE name = ${name}`;
}
});
} finally {
await sql.end();
}
});

Read the case as four steps:

  1. Prepare data. Insert a uniquely named todo and keep the ID and name returned by PostgreSQL. A row created for a test is a test fixture.
  2. Perform the action. Request /api/todos through Hono. The route reads the real database.
  3. Check the result. Require status 200, then find the returned row with the saved ID and compare it with the inserted row. If that row is missing, find returns undefined and the equality assertion fails.
  4. Clean up. Remove the uniquely named row in finally, including when an assertion fails.

The full file awaits two sequential t.step cases. Its second case deletes a record and checks that the API no longer returns it. Each case cleans up its rows, but the outer finally closes the shared SQL client only after both cases finish. Closing it after the first case would prevent the second case’s queries.

The opening SELECT 1 establishes the shared connection in the parent test. Deno can then track that connection across both cases and check that the group closes it. Keep Deno’s resource checks enabled so they can report connections or other resources left open.

You can relate the setup, request, assertions, and cleanup to adding, reading, and removing your practice row in Chapter 1.7. Here we are reading the supplied integration test; there is no additional integration-test file to submit in this chapter.

Read the component tests

The component tests check behavior we can see on the page without starting the whole application. Open client/src/lib/Counter.test.ts:

import { render, screen } from "@testing-library/svelte";
import { userEvent } from "@testing-library/user-event";
import { describe, expect, it } from "vitest";
import Counter from "./Counter.svelte";
describe("Counter", () => {
it("increments when the user clicks the button", async () => {
const user = userEvent.setup();
render(Counter);
const button = screen.getByRole("button", { name: "Count: 0" });
await user.click(button);
expect(screen.getByRole("button", { name: "Count: 1" }))
.toBeInTheDocument();
});
});

describe groups related tests and it defines one case. render(Counter) places the component in the test page. getByRole finds the button by its accessible name, the label a user of assistive technology would encounter. After the awaited click, the assertion checks that a button named Count: 1 is present.

This test performs one click. What would it leave unchecked about subsequent clicks?

The exercise starter below gives you the incrementing component and a test scaffold so you can focus on extending the test. This is a separate exercise: the earlier decrement activity left the practice application’s counter unchanged.

Test repeated counter clicks

0 / 10 points

Task

Complete client/src/lib/Counter.test.ts to test the incrementing counter. The supplied component starts with a button labelled Count: 0. Each click increases its count by one. The existing test checks the first click; extend it so it also establishes what happens after the second click.

Requirements

  • Keep the import of ./Counter.svelte and render that supplied component.
  • Retain a check of the first click and add a check of the second click. Await the simulated clicks before checking their results.
  • The test must pass the incrementing component and fail with an assertion if the component stops incrementing after its first click.
  • Use static imports from vitest, @testing-library/svelte, @testing-library/user-event, and ./Counter.svelte only. DOM matchers such as toBeVisible are already configured; do not import extra setup.

Submit

Submit only client/src/lib/Counter.test.ts, with that directory structure. Do not submit or change Counter.svelte. The grader supplies the component, dependencies, and test setup; this exercise does not require a running API.

Your test must pass the correct component (2 points). The same test must then fail with an assertion when a faulty component stops incrementing after its first click (8 points). Fault-detection points require a passing reference first. Skipped tests, syntax/import errors, ordinary runtime errors, and runner/process timeouts do not count as evidence of detecting the fault.

Then open client/src/lib/TodoList.test.ts. Its cases cover the states we considered earlier: initial and pending state, successful rows, an empty response, and a failed request followed by a retry. The following excerpt keeps the imports, cleanup, and empty-response case:

import { render, screen } from "@testing-library/svelte";
import { userEvent } from "@testing-library/user-event";
import { afterEach, describe, expect, it, vi } from "vitest";
import TodoList from "./TodoList.svelte";
afterEach(() => vi.unstubAllGlobals());
describe("TodoList", () => {
it("shows an empty state when there are no todos", async () => {
const user = userEvent.setup();
vi.stubGlobal("fetch", vi.fn().mockResolvedValue(Response.json([])));
render(TodoList);
await user.click(screen.getByRole("button", { name: "Load todos" }));
expect(await screen.findByText("No todos yet.")).toBeInTheDocument();
});
});

vi.stubGlobal temporarily replaces fetch with a function that returns a successful response containing an empty array. The component still performs its load interaction, but the test supplies the response instead of contacting Hono or PostgreSQL.

The click starts asynchronous work. findByText waits for No todos yet. to appear; getByRole, used for the already visible button, looks for a match immediately. afterEach restores the replaced globals after each case so a response chosen for one test does not carry into the next.

These tests render the component on its own. To check that it also appears and works as part of the complete page, we will turn to a browser test next. Svelte’s testing guidance discusses these different scopes.

Read the Playwright tests

Open e2e-tests/tests/todos.spec.ts. Its first test checks the counter again, this time through the complete page in a running browser:

import { expect, test } from "@playwright/test";
test("the counter updates local state", async ({ page }) => {
await page.goto("/");
await page.getByRole("button", { name: "Count: 0" }).click();
await expect(page.getByRole("button", { name: "Count: 1" })).toBeVisible();
});

An asynchronous status update

0 / 5 points

A Playwright test clicks Load rooms. The page eventually updates its status element from Loading rooms to Rooms loaded:

const status = page.getByRole("status");
await page.getByRole("button", { name: "Load rooms" }).click();

Which next line checks for the completed update without assuming a fixed speed?

Playwright supplies the page object. page.goto("/") opens the application using the base URL in playwright.config.ts. The test finds a button, clicks it, and waits for the updated button to be visible. It does not import or render Counter itself: the running client serves the page containing that component.

The other supplied test finds the Todos region, clicks Load todos, and asserts that Finish walking skeleton appears in that region. This second local path reaches the actual Hono API and migrated database. A locator obtained from another locator searches within that part of the page: for example, todos.getByRole("button", { name: "Load todos" }) looks inside the region stored in todos.

The todo assertion asks whether the expected row appears in Todos, allowing other valid todos to appear there too. Keep that focus when extending the test: adding another valid row should not break a check for the expected one.

The e2e-tests/tests directory is bind-mounted into the runner, so test-source edits are visible without rebuilding that image. Dependency or runner-configuration changes can still require a rebuild. Failure traces are saved under e2e-tests/test-results/ on the host.

Playwright’s web assertions wait for matching browser state. This lets us wait for the result we need instead of guessing how long a request will take with an arbitrary sleep. Await actions and asynchronous assertions so the test finishes only after its checks complete.

In the next activity, write a test for the same named region and load interaction. A controlled browser target is supplied, so return only your test, not the application. The grader checks whether your test accepts the working display and detects a version that omits the expected row.

Test loading a saved todo

0 / 15 points

Task

Complete e2e-tests/tests/student_todos.spec.ts. The controlled page has a region named Todos, containing a Load todos button. Initially the saved rows are not visible. Clicking the button makes the saved todos appear, including a row whose exact text is Finish walking skeleton.

Requirements

  • Use the supplied page and Todos region locators. Activate Load todos inside that region.
  • Use an awaited Playwright assertion to check that a row with the exact text Finish walking skeleton becomes visible inside the same region. Loading may finish after a short delay.
  • Allow other rows to appear too. The same text may also appear elsewhere on the page, so an assertion outside the Todos region is not enough.
  • Keep the static import from @playwright/test; do not add other imports.

Submit

The grader supplies the page, server, browser, and test configuration. page.goto("/") opens that controlled page. Submit only e2e-tests/tests/student_todos.spec.ts, with that directory structure. Do not submit application code or start your own server. This exercise focuses on the browser test, so it does not require a running API or database.

Your test must pass the correct page (3 points). The same test must then fail with an assertion when the expected saved row is omitted after loading (12 points). Fault-detection points require a passing reference first. Skipped tests, syntax/import errors, ordinary runtime errors, and runner/process timeouts do not count as evidence of detecting the fault. An awaited assertion that fails because the expected row never becomes visible is the intended evidence for the missing-row defect.

What each test checks

We have seen two browser checks with different purposes: the local test reaches the running client, API, and database, while the submitted exercise uses a controlled target to focus on display behavior. To understand what a test covers, look beyond the tool’s name or a green result: what runs for real, what is supplied, and what do the assertions check?

TestWhat runs for real?What does the example check?
Direct Hono health testHono application, without a listening serverResponse status and JSON
PostgreSQL integration testHono application and migrated test databaseStored and deleted rows in the API response
Vitest component testSvelte component in a simulated browser environmentRendered state after a click or controlled response
Local Playwright testRunning client in a browser; API and database for the todo caseVisible behavior on the assembled page

Before continuing, choose one test from the table and describe its setup, action, assertion, and cleanup. Then name one behavior it does not check.

Three test runners

0 / 5 points

A study-room application has Deno tests that call the Hono application through app.request(), Vitest tests that render individual Svelte components, and Playwright tests that use a real browser with the running client and API. Select every accurate statement about the boundaries these checks exercise.

Reset only the Compose test project when needed

Removing a one-off runner does not remove its dependencies. To discard the temporary test database and stop the isolated application:

Terminal window
docker compose -p wsd-practice-test -f compose.yaml -f compose.test.yaml down

This targets the chosen Compose test project and needs no -v; removing its database container discards the memory-backed data. The next test run recreates that database and applies migrations. The development Compose project’s named volume remains available.

Without this cleanup, the temporary database lasts while its container runs. Removing a test’s fixture rows lets us repeat that test; recreating the project lets us start again from the migrations. We can choose the cleanup that matches what we need to check.

Check the local client build

From the root of the practice application (practice/), run:

Terminal window
docker compose run --build --rm --no-deps client deno task build

If the build succeeds, you have checked that the current source and toolchain can produce the client build output. The supplied static adapter writes it inside the one-off container, which --rm removes afterward. This command checks the build locally; it leaves no permanent output on the host and does not deploy the application.

Save your practice changes after restoring any deliberate defects. Once the tests and build pass, you have a repeatable way to check these parts of the application after changes. You do not need to remember every testing API yet. Explaining what a test does, why it passes or fails, and what it leaves unchecked prepares you to write more tests as the application grows.

Check Your Understanding

  1. How does a direct Hono test send a request without starting a listening server, and what does the health example assert?
  2. Why does the database test create its own row, remove it in finally, and close the shared client only after all cases?
  3. What does replacing fetch let a component test control, and which parts of the application does that leave untested?
  4. How does the Playwright counter test differ from the Vitest counter test, even though both check a button click?
  5. What is the difference between a failed assertion, an ignored test, and an import error, and why are awaited assertions important?
  6. What does a successful local client build establish, and what would still need testing before deployment?