One useful idea
Liveness asks whether the process responds. Readiness asks whether its required dependencies can currently serve the declared work. A static 200 cannot establish database readiness. repository_probe uses a read-only actual SQLite child: declared application identity, schema version, quick_check and required columns. inspector_probe starts the actual trusted FFprobe version command; it does not analyze a clip or contact a provider.
Observe both ordered repository/inspector probes, measuring real monotonic durations and a fresh generated run ID. Both pass is ready; fail, unavailable or timeout is not ready. An unexpected programming fault raises ObservationFault with safe run correlation and retains its original local cause. Never turn that fault into a healthy result. Probe subprocesses are owned and awaited during cleanup, which can exceed the cooperative deadline.
Telemetry has fourteen fixed category/status cells and a bounded event ring, default 64. IDs belong in events, not metric labels; URLs, filenames, keys and raw exception strings belong in neither controlled output. Counters are process-local, limited and non-durable; a restart resets them. Do not claim monitoring history or all-user analytics. Keep local original diagnosis private until sharing review.
The optional HTTP adapter runs in your separate completed Module 10 environment with explicit Module 11 PYTHONPATH, not by installing overlapping practice.py modules together. Fixture auth is invented, unrelated to the Training Hub tutor. Actual /health/ready returns 200 or 503, a controlled fault 500, and events/metrics are no-store. The incident exercise changes only a NEW owned fixture database. A real failed/recovered observation is still assistance, not independent Portfolio III or observed human defense.
Refresh first: Async ownership and failures, SQLite identity and persistence.
Trace a finished example
import asyncio, os, sqlite3
from contextlib import closing
from pathlib import Path
from tempfile import TemporaryDirectory
from incident import FIXTURE_SCHEMA
from operation_tools import Settings, Telemetry, observe
from operation_tools.probes import repository_probe, inspector_probe
async def main(database):
settings = Settings("fixture-only-not-a-provider-key",
("media.example.invalid",))
telemetry = Telemetry()
probes = (repository_probe(database),
inspector_probe(Path(os.environ["DVP_FFPROBE"])))
for version in (2, 1, 2):
with closing(sqlite3.connect(database)) as connection, connection:
connection.execute(f"PRAGMA user_version={version}")
result = await observe(probes, settings, telemetry)
print(result.ready, [item.status for item in result.checks])
events = telemetry.events()
print(len({item["run_id"] for item in events}),
any(item["run_id"] == result.run_id for item in events))
# FIXTURE_SCHEMA is disclosed classroom assistance, not your migration.
# Only this example-owned temporary directory is removed on exit.
with TemporaryDirectory(prefix="dvp-readiness-example-") as owned:
database = Path(owned) / "fixture.db"
with closing(sqlite3.connect(database)) as connection, connection:
connection.executescript(FIXTURE_SCHEMA)
asyncio.run(main(database))Expected output
True ['pass', 'pass']
False ['fail', 'pass']
True ['pass', 'pass']
3 TrueThe actual owned database changes version 2 → 1 → 2; read-only probes observe pass → fail → pass without repairing it themselves. FFprobe is launched each observation. Generated IDs/timings vary; three distinct IDs and membership in actual events are compared, not hardcoded. No TCP server is started by this example.
The finished implementation is in operation_tools/observability.py and incident.py. Reading it is guided practice, not independent evidence.
Predict the state
The process returns liveness 200 but the database has the wrong schema version. Is it ready?
Compare your answer · self-reviewed
No. Liveness and readiness measure different facts. The actual readiness dependency check fails and the adapter responds 503 rather than a static successful health claim.
Find the label bug
A metric label includes each new run ID and filename. Why is this wrong?
Compare your answer · self-reviewed
It creates unbounded dimensions and can expose data. Keep the fixed category/status metric cells; use bounded events for generated run correlation, never filenames or keys.
Recall the limit
Do a reference incident receipt and quiz score complete the capstone?
Compare your answer · self-reviewed
No. They observe fixture behavior and knowledge. Independent artifacts, fresh transfer, assistance disclosure and observed defense remain separate evidence.
Change it, then build your own
One controlled change
Change the database to the wrong application identity, then restore your own fixture’s declared identity. Predict the readiness result without changing the inspector. Run incident.py into a fresh owned folder to observe a declared delayed child timeout and programming fault; those are controlled fixtures, not claims of an actual provider outage.
Your independent task
Implement observe in practice.py using the supplied validated models, traced-plan/outcome helpers and disclosed scheduler. Own ordered probe orchestration, real duration/run ID, readiness derivation and completed/fault/cancelled telemetry paths. Preserve original causes while exposing only safe correlation. Finish earlier boundaries, then run the practice incident and optional adapter with explicit target; no silent reference fallback.
What success looks like
The build4 group passes actual-process probes, bounded metrics/events, faults and cancellation. The practice incident retains actual healthy, schema failure, recovery, owned-child timeout and safe fault correlation in a NEW receipt. The optional localhost adapter is a separate HTTP observation, not implied by in-process calls.
Hint 1 · a question
Write a state table for completed-ready, completed-not-ready, unexpected fault and caller cancellation. Which paths return an Observation?
Hint 2 · a concept cue
Use the supplied trace/outcome contracts so completed outcomes are recorded even before a later sibling fault. Finish owned cleanup before propagating cancellation.
Hint 3 · a localized example
Observation.ready derives from two ordered pass statuses. Record one run event with the measured total and generated ID; do not place that ID in a metric category or log raw exceptions.
Need the complete worked solution?
Open operation_tools/observability.py and incident.py from the kit. Trace it, close it, then try fresh inputs in your own files. Treat the attempt as guided; seeing the solution does not award a practical pass.
Course help is guidance, not independent evidence. With JavaScript, opening help records guidance locally; otherwise note it in your README. Reset does not erase that history.
Repair a failed check
If readiness stays True during a schema failure, derive it from both actual statuses. If metrics grow a cell per run, remove variable labels. If faults become 200/empty results, preserve the controlled 500 and chained local cause. If practice imports reference accidentally, inspect selected target and environment/PYTHONPATH rather than installing both kits into one environment.
NotImplementedError means a practice stub is still unfinished. Read the failing test name and the last error line. Change one behavior, rerun that build, then rerun all implemented builds.
Show it works on new inputs
Keep a fresh practice incident receipt and your own changed-input failure/recovery demonstration. Compare safe run correlation with retained local diagnosis, restart the optional server to explain telemetry reset and disclose all supplied helpers. Defend what each actual process/HTTP observation proves and what remains outside its scope.
Self-review: name the input, result, refused case and reason. Your local test output and explanation are separate from a quiz score; this page does not certify a pass.
Keep the idea
Useful operations evidence describes observed dependencies and honest failure states. A process being alive, a dependency being ready and a learner being independently competent are different claims.
Module 11 checkpoint
Five questions, followed by the separate practical task above. JavaScript loads the scored questions; the build, files and hints remain available without it.