# Module 9 — Prove the code

Four local builds: author a regression, select testing boundaries, debug a
synchronous pipeline, and run an actual fail-fast quality command. Nothing here
calls a paid provider, hosts an API, authenticates a user or awards a portfolio.
Async comes later in Module 11; this module does not require it.

Prerequisites: functions, lists/dictionaries, exceptions, files, Modules 4/5's
test and fake-provider patterns, and Module 8's distinction between machine
metadata and separately supplied authored findings. This is advanced test design,
not your first introduction to `assert`.

## Start here

Extract `module-9-practice-v1.zip` into a NEW folder. Open a terminal inside its
`module-9` folder. Use Python 3.11 or later; the checked environment is CPython
3.12.10. Create a separate environment for this kit, not the website or an older
course environment. The dependencies are local development tools, not website
runtime dependencies. Installation downloads packages; the fixture exercises do
not use external network.

Windows PowerShell:

```powershell
py -3.12 -m venv .venv
.venv\Scripts\python.exe -m pip install -e ".[test]"
.venv\Scripts\python.exe -m quality_tools.gate
```

macOS/Linux (use an installed Python 3.11+):

```sh
python3 -m venv .venv
.venv/bin/python -m pip install -e '.[test]'
.venv/bin/python -m quality_tools.gate
```

Expected: format, lint, types and tests all succeed, followed by a JSON report
with `status: passed`, four zero exit codes and `exit_status: 0`. This checks
the provided **reference**, not your independent work. Direct tool versions are
pytest 9.1.1, Ruff 0.16.10 and mypy 2.4.0. Transitive/build dependencies are not
fully locked. Keep using the same environment's Python for every command.

Below, `python` means that exact environment executable: on Windows replace it
with `.venv\Scripts\python.exe`, on macOS/Linux with `.venv/bin/python`. No
activation or global installation is necessary. Run from the extracted kit root.

## Work one build at a time

Read the contract, predict the example, run it, then implement only the matching
function in `practice.py`. Do not replace learner functions with calls to the
reference's completed target functions. Reviewed validation/file/command/verifier
assistance is permitted and must be disclosed; using it does not show that you
implemented those helpers independently. All four starter functions deliberately
raise `NotImplementedError`; a blank regression file deliberately fails.

### Build 1 — Turn a bug into a test

```sh
python demo.py versions --broken
python demo.py versions
python -m quality_tools.regression --target reference --test examples/test_regression.py
python -m pytest -q -m build1 --target practice
```

The first output is `["v10", "v9"]`, the second `["v9", "v10"]`.
Write `sort_versions` and your own test in `learner_tests/test_regression.py`.
Import `sort_versions` from `candidate` in that test: the verifier supplies the
candidate in two private runs. Use a changed pair/triple (for example v99/v100),
not only the worked example. Your assertion must pass your corrected function
(pytest exit 0) and fail the specific lexical-sort bug (exit 1). Blank, print-only,
skipped-only and syntax-error tests do not prove this regression. Keep the red and
green transcripts and explain Arrange → Act → Assert in your own words.

Contract: an exact list, at most 10,000 entries, canonical ASCII v1..v999999,
no leading zeros, whitespace or Unicode digits. Empty is valid; duplicates stay;
numeric ascending output is a fresh list; input is unchanged. `_versions` is
reviewed schema assistance; you own ordering and the authored test.

### Build 2 — Test at the right boundary

```sh
python demo.py snapshot --output new-snapshot.json
python -m pytest -q -m build2 --target practice
```

Use an explicit NEW filename each time. Read its actual JSON and compare the
ordered typed values, byte count and SHA-256 receipt—not merely file existence.
Complete `collect_jobs` and your test map in [BOUNDARIES.md](BOUNDARIES.md).
IDs: an exact list of at most 1,000 distinct ASCII letters/digits/underscore/hyphen
strings, length 1..64. Validate all IDs before any provider call. Each response
has exactly the requested `id` and a `status` of queued/running/done/failed.
Return fresh rows in request order; never silently drop unknown fields. Empty
requests return an empty list. Provider errors propagate; there is no retry.

The fake verifies our declared boundary, not a live service. `save_snapshot` is
reviewed NEW-file assistance. It rejects invalid rows, known links/reparse points,
traversal, reserved filenames, missing parents and overwrite. It creates no parent
folders. File failures can leave a partial NEW file; they never return a verified
receipt or delete/roll back output. This is not hostile-filesystem race isolation.

### Build 3 — Debug a synchronous pipeline

```sh
python demo.py debug --broken
python demo.py debug
python -m pytest -q -m build3 --target practice
```

The deliberately broken example exits nonzero with `ZeroDivisionError`. The
correct example leaves two missing authored reviews **unreviewed**, with a null
mean; it does not invent two zero ratings. Write three hypotheses, disprove two,
reduce the input, repair the cause and retain a regression in
[DEBUGGING.md](DEBUGGING.md). Implement `review_pipeline`.

Supplied reviews must be an exact mapping of requested IDs to None or exact
integers 1..5 (bool is not a rating). Foreign keys/nonmapping reviews fail before
provider calls. Missing stays None. The mean includes only supplied authored
values, with an independent Decimal context and two places/HALF_UP. No supplied
values means no mean. No winner, automated defect or quality verdict is created.
Keep integer input/reviewed/unreviewed counts and `findings_present`/`unreviewed`.

Log controlled event/stage/job_id fields; do not copy prompts, reviews, secrets
or raw exception text to logs. Fetch, validate and review failures raise
`PipelineFailure` with that stage/id and preserve `__cause__`. The last failure
event names the failed stage; the last success event is completed/summary/batch.
Logger failures remain visible. Controlled events do not make a full traceback
safe to publish; inspect fixture tracebacks locally and redact real data.

### Build 4 — Create a quality command

```sh
python -m pytest -q -m build4 --target practice
python -m quality_tools.gate --target practice
```

Implement `run_quality`: require four (name, argv) tuples in format/lint/types/tests
order, nonempty tuple argv containing nonempty strings. Validate the whole plan
before running commands. Execute each once; stop on the first nonzero exit
(negative exits also fail). Reject bool/noninteger codes. Return status,
failed_stage, exit_status and only executed steps. Unexpected runner failures
propagate. Read [QUALITY.md](QUALITY.md) for the actual tool scope and failure lab.

The provided command runner uses the same Python, no shell, and a 90-second
per-command timeout. An execution audit refuses an invented green report, changed
command order, extra commands or continuing after failure. It actually runs
Ruff/mypy/pytest; a timeout is not a green
result. Format/lint scan the kit's Python files; strict typing checks only
`quality_tools` and `practice.py`; pytest runs the declared public suite. Your
authored regression is executed by its verifier, not guessed from source text.
The installed `quality-check` command is equivalent; editable setup ties it to
this kit folder. Portable wheel behavior is not claimed.

## Finish with evidence, not a green badge alone

Run all practice tests and the actual practice gate. Keep your changed-input
regression, test-boundary map, investigation and one failure from each quality
stage, then show that later stages did not run after the failure. Disclose which
reviewed helpers/examples/tutor suggestions you used. Explain a new unseen
boundary without copying the reference. Passing checks is local behavioral
evidence, not proof of authorship, independence, live-provider compatibility,
deployment readiness or course/portfolio completion.

The tutor remains optional and uses the main Training Hub login. This kit never
sends code, files, logs or findings to it. Reference reading, practice and checks
do not require a tutor account.

## Troubleshooting

- `No module named pytest/ruff/mypy`: use the kit environment's exact Python and
  run the editable test installation above. Do not install into the website.
- `NotImplementedError`: finish only the selected learner function; this is not
  an installation failure or a reason to weaken tests.
- `candidate` missing when running a regression directly: use the verifier;
  it creates the two candidate modules privately.
- Formatting/lint errors: read the filename and rule. Format deliberately with
  `python -m ruff format .`; the gate itself never rewrites your work.
- Existing snapshot: choose a NEW filename; never delete an old artifact to
  force a green result. Inspect any partial output before choosing a new name.

Tool references: [Ruff configuration](https://docs.astral.sh/ruff/configuration/),
[mypy command line](https://mypy.readthedocs.io/en/stable/command_line.html),
[pytest exit meanings](https://docs.pytest.org/en/stable/reference/exit-codes.html).
