DVPPython Studio
Module 4: Build a resilient pipeline / Build 3 of 4

Log an auditable run

Create bounded UTC audit events with useful context, honest levels and explicit limits on credential redaction.

Runs on your computer · 55–80 minutes · no paid services

Download practice filesFiles, commands & notes

Without JavaScript, use the step links and keep your files on your computer.

One useful idea

An audit event answers what happened, when and to which run/asset. The function returns exact at, level, event, run_id, asset_id and context fields. run_started, asset_inspected and run_finished are INFO; asset_failed is WARNING. A finish event means the run completed, not that every asset passed.

Use the reviewed safe single-component run ID and canonical relative asset path (or None). IDs remain visible: choose non-secret labels. A timestamp must be an aware UTC datetime; default to actual UTC when none is supplied. Import datetime, timedelta and timezone from datetime. datetime(2026, 10, 7, 12, 0, tzinfo=timezone.utc) constructs a fixed year/month/day/hour/minute value with explicit UTC. datetime.now(timezone.utc) supplies the actual moment. The * in the signature makes when keyword-only: call when=fixed, not an extra positional argument. Tests inject a fixed moment. Format whole seconds with isoformat, not a machine-dependent display. Event times are not media capture times or precise latency measurements.

Allow only declared verdict, controlled code and reconciled total/accepted/failed native-integer counts from 0 to one million. Import CODES from asset_intelligence.errors for the reviewed failure-code set. Copy nested counts. Omit unknown keys, raw messages and nested sidecars. Recognized credential keys token, access_token, api_key, password, authorization and secret are redacted case-insensitively, treating hyphens as underscores; retain their original key spelling. This bounded policy is not universal secret detection.

The reviewed emit_event(logger, event) sends one JSON message at the declared logging level. The caller owns Logger handlers: the helper neither configures global logging nor secretly opens a file. The pipeline collects structured events; explicit export writes JSONL.

Refresh first: Validation without mutation, Structured JSON records, Useful error categories.

Trace a finished example

from datetime import datetime, timezone
from asset_intelligence.core import audit_event

event = audit_event("asset_failed", "review-run", "missing.mov",
                    {"code": "missing_metadata", "API-Key": "invented-secret",
                     "raw_message": "do not retain this"},
                    when=datetime(2026, 10, 7, 12, 0, tzinfo=timezone.utc))
print(event["at"], event["level"])
print(event["context"]["API-Key"])
print("raw_message" in event["context"])
# 2026-10-07T12:00:00Z WARNING
# [REDACTED]
# False

The fixed aware UTC moment produces the same whole-second timestamp. The event selects WARNING. Controlled code survives, the recognized credential value is replaced and unknown raw_message is absent. No file is written.

The finished implementation is in asset_intelligence/core.py. Reading it is guided practice, not independent evidence.

Predict context

What survives from {"code": "corrupt_metadata", "raw_message": "..."}?

Compare your answer · self-reviewed

Only the controlled code. Unknown raw context is omitted rather than copied into the audit.

Find the timestamp ambiguity

Should a datetime without timezone information be assumed UTC?

Compare your answer · self-reviewed

No. Refuse it. A naive local time is ambiguous; ask the caller for an aware UTC value.

Recall output ownership

Why copy counts rather than reuse the caller’s dictionary?

Compare your answer · self-reviewed

Otherwise a later caller change could silently rewrite the recorded event. A fresh nested counts dictionary preserves the event’s intended evidence.

Try the idea in this browser

Runs in this browser · optional preparation · local project checks remain separate

Try a small function before opening your local files. Python downloads when you choose Run; if it cannot load, your code stays here and the local kit still works. The worker executes on your device, not on a DVP server. Only run code you trust: this is not a hostile-code security sandbox.

JavaScript loads the practice controls. Python starts only after Run.

Read the browser task briefs without running Python

Trace the declared failure boundary

Read the function, predict a changed normal/refused case, then Run. Implement the declared events/levels, safe non-secret IDs, aware UTC whole-second timestamp, exact fields and bounded context/count rules. Redact recognized credential values and omit unknown context; copy nested counts. Return data only, not a persisted log or universal security guarantee.

Build the boundary on changed inputs

Implement the declared events/levels, safe non-secret IDs, aware UTC whole-second timestamp, exact fields and bounded context/count rules. Redact recognized credential values and omit unknown context; copy nested counts. Return data only, not a persisted log or universal security guarantee.

Change it, then build your own

One controlled change

Use run_finished with counts total=4, accepted=3, failed=1 and a fixed UTC timestamp. Change a recognized credential key to uppercase/hyphen spelling, then add an unknown nested metadata key. Predict retained keys before running.

Your independent task

Implement audit_event in practice.py with its keyword-only when parameter. Validate declared events/IDs/time/context and exact reconciled counts; create fresh outputs, redact recognized credential values and omit unknown keys. Do not log raw sidecars or configure global handlers. See the README for exact allowed types/ranges and errors.

What success looks like

The build3 group passes deterministic timestamps, levels, context/redaction, nested ownership, IDs, invalid times and count/type refusal. Successful event construction does not prove universal credential safety or a persisted log file.

Hint 1 · a question

List allowed event names and context fields before coding. Which data is safe to retain?

Hint 2 · a concept cue

Build a new context field by field. Validate nested counts and their sum. Validate UTC before formatting the moment.

Hint 3 · a localized example

key.casefold().replace("-", "_") can identify the declared credential names. It does not detect secrets in IDs or arbitrary text; omit unknown context rather than claiming universal redaction.

Need the complete worked solution?

Open asset_intelligence/core.py from the kit. Trace it, close it, then try fresh inputs in your own files. Treat the attempt as guided; seeing the solution does not award a practical pass.

Course help is guidance, not independent evidence. With JavaScript, opening help records guidance locally; otherwise note it in your README. Reset does not erase that history.

Repair a failed check

If raw context survives, allowlist rather than copy the whole dictionary. If bool counts pass, use native integer checks. If time depends on the machine’s zone, require aware UTC. If credentials remain, compare a normalized key but preserve its original spelling in output.

NotImplementedError means a practice stub is still unfinished. Read the failing test name and the last error line. Change one behavior, rerun that build, then rerun all implemented builds.

Show it works on new inputs

Log a new mixed-run narrative with a fixed UTC time, missing/corrupt outcomes and reconciled final counts. Show known-key redaction, unknown-field omission and naive-time refusal. Explain how a caller-owned logger differs from writing audit.jsonl.

Self-review: name the input, result, refused case and reason. Your local test output and explanation are separate from a quiz score; this page does not certify a pass.

Keep the idea

Useful logs retain bounded evidence and acknowledge their limits. Do not turn sensitive raw inputs into permanent audit text.