Skip to content

Latest commit

 

History

History
199 lines (139 loc) · 14.3 KB

File metadata and controls

199 lines (139 loc) · 14.3 KB

Playtestr

End-to-end testing for the terminal.

Press the keys. Check the screen. Catch the regression. Playtestr drives interactive CLIs and TUIs through a real pseudoterminal, compares rendered text snapshots, and produces readable diffs when something changes.

Download stable v0.1.0 · First test · Documentation · Examples · Support

  • Keyboard-driven tests: describe an interaction in a small JSON spec.
  • Rendered screen assertions: test text after cursor movement, redraws, and resizing.
  • Reviewable regressions: compare snapshots and deliberately update a selected baseline.
  • Bounded execution: set deadlines and output limits, with cancellation and process cleanup.

The standalone runner is written in Go but does not require Go to run. The MVP has a versioned test contract and machine-readable reports for trusted local applications.

Stable v0.1.0 publishes native archives for Linux x86-64 (amd64), Apple silicon macOS (arm64), and Windows x86-64 (amd64). Support is limited to the exact documented targets and evidence; it is not a claim for every terminal application, OS release, distribution, or architecture.

Playtestr starts a real pseudoterminal, sends keyboard input, and feeds output into a VT terminal emulator. Assertions inspect the rendered screen, including cursor movement and redraws.

The newest verified prerelease is v0.4.0-rc.3, adding the selected two-column wide-character repair while retaining suites, offline reports, workspaces and installer/macOS output repairs. See its downloads and installation. Stable v0.1.0 remains available below.

For a prerelease trial, see the captured pass/failure/recovery demo and the complete PR workflow. The workflow builds an actual example target and retains failure evidence; hosted pass, deliberate failure and recovery verified its summary and artifact steps.

Install and run a first test

Choose the archive that matches your host exactly. Each archive has an adjacent SHA-256 integrity file:

Host Archive Checksum
Linux x86-64 (amd64) playtestr_0.1.0_linux_amd64.tar.gz .sha256
Apple silicon macOS (arm64) playtestr_0.1.0_darwin_arm64.tar.gz .sha256
Windows x86-64 (amd64) playtestr_0.1.0_windows_amd64.zip .sha256

Follow the binary-only installation and first-test walkthrough. It verifies the checksum, creates a small host-native greeting target, produces a passing result, proves a deliberate failure with evidence, and restores the passing test. Release archives include one native binary, this README, the Apache 2.0 license, and third-party notices.

The target application and its runtime remain your responsibility. Playtestr runs trusted targets with your user permissions and is not a sandbox. See Platform support for exact release evidence and Terminal compatibility for rendered-screen limits.

To build the current development version with Go 1.25 or newer, use the source development guide.

Write a test

Specs are JSON. command is an executable followed by arguments and is executed directly without a shell. Executable lookup uses the environment that launches Playtestr. The cwd behavior is described below.

{
  "version": 1,
  "name": "My CLI",
  "command": ["my-cli", "configure"],
  "width": 80,
  "height": 24,
  "timeout_ms": 3000,
  "run_timeout_ms": 30000,
  "max_output_bytes": 2000000,
  "steps": [
    {"expect": "Project name"},
    {"text": "hello"},
    {"key": "Enter"},
    {"expect": "Created hello"},
    {"exit": 0},
    {"snapshot": "created.txt"}
  ]
}

version is required. Playtestr rejects missing or unsupported versions before it launches a target. The complete defaults, limits, normalization rules, and JSON Schema are in Test specification version 1.

Stateful local flows can opt into specification version 2. It copies a bounded reviewed fixture for every run and can provide managed home and temporary directories:

{
  "version": 2,
  "command": ["../bin/fixture", "workspace"],
  "workspace": {
    "fixture": "workspace-fixture",
    "cwd": ".",
    "home": "temporary",
    "temp": "temporary"
  },
  "steps": [
    {"expect": "fresh workspace seed=reviewed home=true temp=true"},
    {"exit": 0}
  ]
}

The command resolves before the copied working directory is selected. Failed runs normally clean up; --keep-workspace-on-failure prints and records an explicitly retained local path. Successful workspaces are always removed. See Repeatable workspaces, spec v2, and report v2. A temporary workspace controls selected local paths but is not a security sandbox.

Each step has exactly one action. Supported keys: Enter, ArrowDown, ArrowUp, ArrowLeft, ArrowRight, Escape, Tab, Backspace, CtrlC, CtrlE, CtrlF, CtrlO, CtrlQ, CtrlS, CtrlSpace, and CtrlZ. An exit step waits for the process and requires the exact exit code; intentionally nonzero expected codes are supported. Long-running TUIs do not need an exit step.

expect polls the current screen until the text appears or the per-step timeout expires. expect_not waits for text that an earlier expect observed to disappear after input or resize, which is useful for closing modals without arbitrary sleeps. If the process exits first, an unmatched assertion reports the exit code instead of waiting for a timeout. Snapshots compare the rendered screen after at least 150 ms without output.

After a resize, {"wait_for_redraw": true} requires target output after the resize and allows a one-second redraw window before Playtestr sends the next input. It works with continuously repainting dashboards, must immediately follow the resize, and does not replace a content assertion.

A snapshot must follow a successful expect or expect_not since the most recent input or resize, or a successful exit assertion. This makes application readiness explicit; quiet output alone does not prove that an app has finished rendering. Choose expected text that identifies the new state rather than text left over from the previous screen.

Screen snapshots preserve leading spaces and internal blank lines while trimming trailing spaces and unused rows. They compare text rather than colors or styles. A mismatch prints a unified expected/actual diff and saves both <spec>.actual.txt and <spec>.diff.txt from the same captured screen.

Resize a running terminal with one action:

{"resize": {"width": 100, "height": 30}}

The next snapshot requires a new successful expect, because resizing can trigger an asynchronous redraw.

Session limits

timeout_ms limits each wait step and defaults to 3 seconds. run_timeout_ms limits the whole run and defaults to 30 seconds. max_output_bytes counts raw PTY output before terminal parsing and defaults to 2 MB. An optional startup_timeout_ms requires the target to render visible text within that interval; leave it unset for programs that legitimately wait for input before rendering.

Ctrl+C cancels the active spec, performs bounded cleanup, and prevents later specs from starting. Playtestr exits with status 130 for this interruption.

Write an ordered machine report with --report results.json. It includes stable status and failure categories, step metadata, target exit, cleanup evidence, and artifact paths. It deliberately excludes command arguments, environment data, typed text, and terminal-screen contents. See Machine report version 1.

Starting with the verified v0.3.0-rc.1 prerelease, render that captured report and its admitted evidence as one portable offline diagnosis:

playtestr report --input artifacts/results.json --evidence-root artifacts --output artifacts/report.html

The renderer never launches the target or reads specs. It escapes and embeds bounded screen/diff text, labels missing report-v1 information, works without JavaScript or a server, and preserves an existing output if rendering fails. See Offline failure reports for its path, privacy, resource, and release boundaries.

Run a directory as a deterministic serial suite and keep each invocation's evidence together:

playtestr test --list tests/terminal
playtestr test --artifacts-dir artifacts/playtestr --report artifacts/results.json tests/terminal

Directory matches are recursive, sorted, deduplicated, bounded, and resolved before any target launches. The final summary accounts for every selected spec and points directly to each failed screen or diff. See Test suites and CI evidence for selection rules, limits, output layout, and exit statuses.

Working directory and environment

When omitted, cwd remains the directory where Playtestr was invoked. A relative cwd is resolved from the test file's directory.

Targets receive a small operational environment including executable lookup, temporary-directory, home-directory, and locale variables appropriate to the operating system. Add literal values with env, or explicitly select more host variables with inherit_env:

{
  "version": 1,
  "cwd": "../fixture-project",
  "env": {"APP_MODE": "test"},
  "inherit_env": ["CI"]
}

Environment values are never written to failure artifacts by Playtestr itself, although a target can still print them to its terminal.

go run ./cmd/playtestr test --update examples/menu.json
go run ./cmd/playtestr test --update --snapshot diagnostics.txt examples/menu.json
go run ./cmd/playtestr test --list examples/suite
go run ./cmd/playtestr test --artifacts-dir artifacts/playtestr --report artifacts/results.json examples/suite
go build -o bin/fixture ./cmd/fixture
go run ./cmd/playtestr test examples/workspace.json
go test ./...

--update explicitly writes baselines in a snapshots folder beside the spec. Updates stay in memory until the entire spec and process cleanup succeed. --snapshot updates one named baseline, requires --update, and accepts exactly one spec. Other snapshots still compare normally. Review baseline changes before committing.

A failure returns exit code 1. General failures save the final visible screen to <spec>.actual.txt; snapshot mismatches also save <spec>.diff.txt. Multiple spec paths can be passed to one invocation when no snapshot selector is used.

Sprint 4 also exercises the independently maintained Charm Gum TUI at a pinned version. See the external Gum trial for its install and test commands.

The repository includes deliberately failing fixtures for manual verification:

go run ./cmd/playtestr test examples/hang.json
go run ./cmd/playtestr test examples/output-flood.json
go run ./cmd/playtestr test examples/child-cleanup.json
go run ./cmd/playtestr test examples/cancel.json
go run ./cmd/playtestr test examples/snapshot-mismatch.json

The first three commands fail promptly, report why they stopped, confirm their cleanup mechanism, and save the final screen. The cancellation example waits until you press Ctrl+C, then exits with status 130 after cleanup. The snapshot-mismatch example prints an intentional unified diff and saves matching screen and diff artifacts.

Current scope

The tested terminal behavior and known emulator gaps are recorded in Terminal compatibility. Compatibility with one application does not certify every terminal application.

Windows uses a Job Object and Unix uses a dedicated process group to terminate managed descendants. The current xpty API starts a Windows target immediately before Playtestr can attach it to the Job Object, leaving a small launch-to-attachment window in which a very early child could escape management. Unix descendants can deliberately detach into another session. Only test trusted applications; local PTY execution is not a sandbox.

The GitHub Actions matrix runs native tests on Linux, macOS, and Windows and retains machine reports, screens, and diffs from its deliberate-failure check. Native and published-release results are recorded in Platform support. Releases are packaged by a separate workflow; the process is documented in Releasing.

The repository also contains a setup-only GitHub Action for installing one exact checksum-verified release. See CI installation for the verified immutable revision, pinning contract, supported hosts, adopter workflow, archive fallback, and maintenance boundary.

Recording, replay, exact-failure minimization, and styled snapshots remain post-MVP work.

Built on Charm's xpty and a locally maintained vt10x renderer. The selected wide-character repair is in verified v0.4.0-rc.3; release evidence records the exact scope. Combining clusters, emoji/ZWJ sequences and target-visible terminal queries remain excluded.

For completed releases, engineering status and the next implementation steps, see the project roadmap. Plans describe future work separately from the released capabilities above.