# Build 4 — diagnose from observations

This is a local invented fixture exercise, not a public deployment runbook. Never diagnose
by exposing credentials, private paths, media or raw exception strings to the optional tutor.
The Academy login is unrelated to this service's public invented fixture authorization.

## Start with the difference

| Actual local response | Meaning | Next observation |
| --- | --- | --- |
| `/health/live` 200 | This server handled the liveness request | Check readiness; dependency health is not established |
| `/health/ready` 200 | Both configured probes passed in this run | Retain the generated run ID and durations; not provider/media safety |
| `/health/ready` 503, repository fail | Owned repository identity/schema/integrity policy failed | Inspect schema/configuration locally; do not overwrite/delete the database |
| `/health/ready` 503, unavailable | Required dependency could not be checked | Diagnose missing/malformed data or bounded output locally; not a pass |
| `/health/ready` 503, timeout | The admitted cooperative probe exceeded its deadline | Find the matching bounded event; cleanup can outlast the nominal timeout |
| `/health/ready` 500, observation_fault | Unexpected programming/child/telemetry fault | Match run ID to fault events; inspect the original cause locally |

1. Confirm you launched the owned adapter, not another listener. Requests stay on `127.0.0.1`.
2. Record readiness's `run_id` / `X-Operation-Run-ID`, status and measured durations. Module10's
   `X-Request-ID` is a different request identifier. Neither belongs in metric labels.
3. GET `/observability/events`. Match that run ID; events contain only fixed category/status,
   run ID and duration. Older entries are dropped at capacity (64 by default), not archived.
4. GET `/observability/metrics`. Compare the fixed 14 category/status cells. Counts/times are
   process-local and reset on restart; no persistent audit or external collector is implied.
5. Separate liveness, readiness failure, elapsed timeout and unexpected fault. No failing
   dependency justifies restarting a healthy process automatically or declaring every asset safe.
6. Repair only the owned fixture configuration/data. Run a fresh readiness check and retain its
   new run ID. A historical 200, unchanged canned response or printed replay is not recovery.
7. Stop the owned service (Ctrl+C). Verify its listener/process exited before launching another.

## Repeatable incident exercise

Run `incident.py` with an absolute NEW output directory and your reviewed FFprobe executable.
It creates only its owned database and JSON receipt. It actually probes healthy schema, changes
its owned schema version, probes failure, repairs and observes recovery. Then it launches declared
owned child-delay and child-bug fixtures to distinguish timeout from programming fault. Those last
two are not actual FFprobe failures, public outages or provider requests. Run IDs/times vary.
No file is automatically deleted/overwritten; a failed run may leave partial NEW output.

Read `incident.json`, not this checklist, for your actual observations. An accepted fixture/run
does not establish learner independence, production security or completion of Portfolio III.

## Learner defense and changed-input transfer

In your own extracted kit, implement `practice.observe`. Disclose scheduler/model/trace/adapter
assistance. Explain each of the five observed phases using the matching actual events and metrics.
Explain why a timeout is not an internal TimeoutError and why liveness can remain healthy.
Then use a different NEW repository: omit a required table, change its identity, repair it, and
show a fresh failure/recovery pair. Do not edit an old portfolio database merely for this exercise.
Use a changed probe that returns False, raises expected CheckUnavailable and raises a real bug;
retain actual outputs/causes locally and prove no false-ready result. Explain queue versus admitted
duration, capped trace retention, fixed metric dimensions and process restarts/cleanup limitations.
Reference code, supplied fixtures, passing tests and copied alternatives are assistance, not
independent artifacts or an observed learner defense.
