# Build 4 — One command, four real tools

The gate runs these commands in order in this extracted editable kit:

| Stage | Actual command (same environment Python) | Scope / meaning |
| --- | --- | --- |
| format | `python -m ruff format --check .` | Formatting only; no rewrite |
| lint | `python -m ruff check .` | Configured E4/E7/E9/F/I rules, not every possible flaw |
| types | `python -m mypy --strict quality_tools practice.py` | Declared code types, not all tests/examples or runtime data validity |
| tests | `python -m pytest -q tests --target practice` | Observed public behaviors + authored lexical regression verifier |

Reference mode selects `--target reference` for the last command. It is a setup
check, not an independent learner result. Exit 0 for each stage is required;
any nonzero (including a signal-style negative return) stops every later command.
No shell interpolation or automatic formatter rewrite is used by the gate.
The reviewed entry audits actual runner calls and typed JSON report parity;
returning invented zero steps/codes or changing the command plan cannot earn a
printed passed report. This audit is assistance, not a hostile-code sandbox.
Missing tools and subprocess timeouts are visible failures, not excuses to skip
a stage. The runner has a per-command timeout, not a global wall-clock budget.

## Failure lab — work on a separate disposable copy

Use a NEW duplicate of your kit for each experiment. Keep your real work intact.
Introduce one fault at a time and retain the final JSON steps to prove fail-fast.

1. Format: change spacing in a Python file; expect `failed_stage: format`, one
   executed step. Repair deliberately with the formatter and inspect the diff.
2. Lint: add an unused import to a new file in `quality_tools`, format that file,
   then run the gate. Expect `failed_stage: lint`; types/tests must not run.
3. Types: add `def wrong() -> int: return "not an integer"` in a new
   `quality_tools` file and format it. Expect `failed_stage: types`; tests do not run.
4. Tests: add a clearly labeled deliberate failing assertion in a new `tests/`
   file, then format it. Expect `failed_stage: tests`, not `passed`.

Use invented faults only. Do not import real private credentials or modify a
provider/account to generate a failure. After the experiment, keep the faulty
copies labeled; do not weaken the public tests or delete evidence to get green.
For the final implementation, record all four successful stages and the changed
authored regression. Name reviewed orchestration/schema/verifier assistance.

## Coverage is a question, not a certificate

Ask which executable decisions your tests exercise and which failures remain
untested. Optional local teacher/development measurement needs an extra tool:

```sh
python -m pip install pytest-cov==7.0.0
python -m pytest -q --cov=quality_tools --cov-branch --cov-report=term-missing
```

This covers the in-process reference package, not learner independence, every
subprocess command-entry line, live providers or performance. The authored
regression executes in separate child processes; no claim of child-process
coverage is made by this command. A high percentage cannot rescue tests with
the wrong assertions. Run the actual gate too. This kit has no production CI,
deployment, security or course-completion certificate.
