One useful idea
A quality command combines checks with different jobs. Ruff format --check observes formatting without rewriting. Ruff check applies the declared E4/E7/E9/F/I lint rules, not every possible defect. mypy --strict checks declared types in quality_tools and practice.py, not all tests/examples or runtime input values. pytest runs the public suite and the actual authored-regression verifier. None of these alone establishes live-provider compatibility, security, speed or independence.
The command plan is exactly four (name, argv) tuples in format/lint/types/tests order. argv is a nonempty tuple of nonempty strings, not a shell string. Preflight the whole plan before any process. Use the same virtual-environment Python for every tool, execute each once and stop at the first nonzero exit. A negative signal-style exit also fails. Bool or other noninteger codes are not valid process results.
run_quality returns status, failed_stage, exit_status and only executed steps. The reviewed local entry provides actual subprocess execution with no shell and a 90-second timeout per command; it never formats automatically. Timeout or unexpected runner faults remain visible. Its execution audit compares an immutable declared plan, actual ordered runner observations and typed JSON report parity. Returning invented zero codes without running the tools cannot print a passed result. This audit is assistance, not a hostile-code sandbox.
Build your own orchestration, then demonstrate real faults in a separate disposable kit copy: spacing fails format; a formatted unused import fails lint; a formatted wrong typed return fails types; a formatted deliberate test assertion fails tests. For each, retain the last executed stage and show that later tools did not run. Do not weaken the checks or delete old evidence to force green. Complete the quality-failure lab and scope notes.
The worked example below actually runs the reference gate and parses its real final report. Your own final command selects --target practice and executes your completed functions plus authored regression. A reference pass is setup evidence only. Keep local test tools separate from website/production dependencies. High coverage means executable paths were exercised, not that assertions are meaningful; retain changed-input tests, helper disclosure and a separate observed defense for portfolio work.
Refresh first: Authored red/green regression, Observed testing boundaries, Visible causes and no false green result.
Trace a finished example
import json
import subprocess
import sys
# Actual reference tools, not a fake runner or browser tool simulation.
process = subprocess.run([sys.executable, "-X", "utf8", "-B", "-m", "quality_tools.gate"],
capture_output=True, text=True, encoding="utf-8", timeout=120)
if process.returncode != 0:
raise RuntimeError("Reference quality gate failed; inspect captured diagnostics locally")
report = json.loads(process.stdout.strip().splitlines()[-1])
print(" / ".join(step["name"] for step in report["steps"]))
print(report["status"], report["exit_status"])
print(len(report["steps"]), all(step["exit_code"] == 0 for step in report["steps"]))Expected output
format / lint / types / tests
passed 0
4 TrueThe child executes real Ruff formatting/lint, strict mypy and the 97 public reference cases. Its last line is the audited JSON report. This example captures rather than fabricates the tool output; failure produces no success-shaped report and diagnostics remain local.
The finished implementation is in quality_tools/core.py and quality_tools/gate.py. Reading it is guided practice, not independent evidence.
Predict fail-fast
Formatting passes but lint exits 7. Which later checks should run?
Compare your answer · self-reviewed
None. Retain the successful format step and failed lint step, return failure and stop before typing/tests. Nonzero is a failure even if it is not 1.
Find the forged green
A learner returns four zero-code steps without invoking the runner. Is a printed passed result permitted?
Compare your answer · self-reviewed
No. The reviewed entry compares actual runner calls to the immutable declared plan and typed report. Invented codes are not evidence of tool execution.
Recall coverage limits
Does 100% measured coverage prove a meaningful regression or independent course completion?
Compare your answer · self-reviewed
No. Paths may be exercised with weak assertions. Keep the actual red/green distinction, changed-input evidence, observation scope and assistance disclosure; portfolio defense remains separate.
Change it, then build your own
One controlled change
In a NEW disposable kit copy, introduce one fault for each of format/lint/types/tests. Format deliberately where needed to isolate later stages. Predict the report and retain actual last-stage evidence before repairing.
Your independent task
Implement run_quality in practice.py with complete plan preflight, exact integer codes, ordered single execution, fail-fast and truthful report steps. Reuse only the disclosed runner/entry assistance, not the finished reference orchestration. Finish earlier practice functions and your authored regression, run all practice tests, then the actual practice gate. Demonstrate each real tool-stage failure in a separate disposable copy.
What success looks like
Build 4 behavioral checks and the full practice gate pass only with actual ordered tool execution and completed earlier practice/regression. Four real injected faults stop at their declared stages. Retain changed tests, boundary map, investigation and actual transcripts; a quiz or reference report does not replace these artifacts or observed independent defense.
Hint 1 · a question
List the four roles and their order. Which evidence exists if execution stops at lint, and which cannot legitimately be claimed?
Hint 2 · a concept cue
Validate every (name, argv) pair before calling the runner. Append a step after its actual exit; nonzero returns failure without invoking later commands.
Hint 3 · a localized example
The final JSON contains only actual steps. A passed status requires all four real zero exits; the entry audit refuses invented results, reordered/extra calls and continuation after failure.
Need the complete worked solution?
Open quality_tools/core.py and quality_tools/gate.py from the kit. Trace it, close it, then try fresh inputs in your own files. Treat the attempt as guided; seeing the solution does not award a practical pass.
Course help is guidance, not independent evidence. With JavaScript, opening help records guidance locally; otherwise note it in your README. Reset does not erase that history.
Repair a failed check
If later tools run after failure, return immediately with only observed steps. If True passes as exit code, require the exact integer type. If formatting silently changes work, use --check in the gate and format deliberately outside it. If a forged report is refused, fix orchestration rather than bypassing the audit. A missing tool means use the correct kit environment and installation.
NotImplementedError means a practice stub is still unfinished. Read the failing test name and the last error line. Change one behavior, rerun that build, then rerun all implemented builds.
Show it works on new inputs
Use fresh invented fixtures and your completed functions/test file, retain the actual practice gate, boundary map and investigation, then demonstrate four real isolated tool-stage faults with no later execution. Explain one new untested boundary and helper assistance. Keep portfolio evidence and observed defense distinct from reference/tool/quiz success.
Self-review: name the input, result, refused case and reason. Your local test output and explanation are separate from a quiz score; this page does not certify a pass.
Keep the idea
A trustworthy quality command reports what ran, what failed and what remains untested. It does not turn a transcript into a certificate.
Module 9 checkpoint
Five questions, followed by the separate practical task above. JavaScript loads the scored questions; the build, files and hints remain available without it.