DVPPython Studio
Module 4: Build a resilient pipeline / Build 1 of 4

Survive broken metadata

Continue past expected metadata failures while exposing unexpected software and filesystem errors.

Runs on your computer · 50–75 minutes · no paid services

Download practice filesFiles, commands & notes

Without JavaScript, use the step links and keep your files on your computer.

One useful idea

A resilient batch does not pretend that every item succeeded. An expected missing or corrupt sidecar can become a failure record while later items continue. An unexpected bug or permission error must interrupt the run, not disappear into a broad catch.

The kit defines MetadataError with four controlled codes: missing_metadata, corrupt_metadata, invalid_metadata and unsafe_metadata. ValidationError is its invalid_metadata subclass. Catch only this family around the inspector and record validation. Keep the requested path and code; never copy raw exception text into a report.

Before the first callback, require a list of unique canonical relative POSIX paths and a callable inspector. The reviewed relative_asset_path helper validates path text without reading files. For each path, call inspect, validate the returned record and require its source to match that path. Preserve accepted and failure order; return fresh total/accepted/failures data.

The worked example below uses an injected in-memory inspector, not filesystem evidence. The local tests additionally exercise the reviewed sidecar reader: bounded UTF-8 JSON, missing/corrupt/invalid/link refusals and unexpected I/O errors. Tiny media-suffix fixtures are text, not playable clips; declared metadata is not measured video quality.

Refresh first: Read-only inventories, Specific exception handling, Virtual environment and pytest setup.

Trace a finished example

from asset_intelligence.core import inspect_batch
from asset_intelligence.errors import MetadataError

def inspect(path):
    if path == "missing.mov":
        raise MetadataError("missing_metadata")
    return {"source": path, "width": 1920, "codec": "h264",
            "frames": 240, "fps": 24, "model": "Google Veo"}

result = inspect_batch(["first.mp4", "missing.mov", "last.mp4"], inspect)
print(result["total"], len(result["accepted"]), len(result["failures"]))
print(result["accepted"][-1]["source"])
print(result["failures"][0]["code"])
# 3 2 1
# last.mp4
# missing_metadata

All three path values are checked first. first.mp4 produces accepted data, missing.mov becomes a bounded failure and last.mp4 is still inspected. Total remains three: missing evidence is not removed from the denominator.

The finished implementation is in asset_intelligence/core.py. Reading it is guided practice, not independent evidence.

Predict continuation

Will last.mp4 be inspected after an expected missing_metadata failure?

Compare your answer · self-reviewed

Yes. Only that item becomes a failure row. The batch keeps running and total includes the failed path.

Find the swallowed bug

Should an unexpected RuntimeError become invalid_metadata?

Compare your answer · self-reviewed

No. Catch only MetadataError. Unexpected programming and filesystem errors must remain visible; changing their category invents evidence.

Recall validation order

An invalid path is last in the list. Should the inspector already have processed the first path?

Compare your answer · self-reviewed

No. Validate the entire caller list and uniqueness before invoking the inspector. This avoids partial callback work for an invalid request.

Try the idea in this browser

Runs in this browser · optional preparation · local project checks remain separate

Try a small function before opening your local files. Python downloads when you choose Run; if it cannot load, your code stays here and the local kit still works. The worker executes on your device, not on a DVP server. Only run code you trust: this is not a hostile-code security sandbox.

JavaScript loads the practice controls. Python starts only after Run.

Read the browser task briefs without running Python

Trace the declared failure boundary

Read the function, predict a changed normal/refused case, then Run. Implement inspect_batch with complete unique canonical path prevalidation, source identity, fresh accepted records, full total and bounded expected MetadataError failures in order. Expose unrelated errors. The supplied reviewed validator is guided support for the later Build 2; this in-memory callback is not local sidecar/file evidence.

Build the boundary on changed inputs

Implement inspect_batch with complete unique canonical path prevalidation, source identity, fresh accepted records, full total and bounded expected MetadataError failures in order. Expose unrelated errors. The supplied reviewed validator is guided support for the later Build 2; this in-memory callback is not local sidecar/file evidence.

Change it, then build your own

One controlled change

Change missing.mov to corrupt.webm and raise corrupt_metadata for it. Add another good path after the failure. Predict all three counts and the last accepted source before running.

Your independent task

Implement inspect_batch(paths, inspect) in practice.py. Use the provided relative_asset_path and your validate_ingest. While the validator stub is unfinished, you may import validate_ingest as reviewed_validate INSIDE inspect_batch and call that alias. A top-level import with the stub’s name would be overwritten by the later def. Label this assistance; in Build 2 replace the alias call with your own validator. Prevalidate unique paths and the callback, check source identity, catch only MetadataError and return exact fresh total/accepted/failures data. Preserve order; never store raw messages.

What success looks like

The build1 pytest group passes for good/missing/corrupt/later-good inputs, empty/all-failed batches, prevalidation and unexpected-error propagation. The untouched starter raises NotImplementedError. Passing with a reviewed validator is guided integration, not independent Build 2 evidence.

Hint 1 · a question

Separate caller errors, expected metadata failures and unexpected errors. Which category may become a report row?

Hint 2 · a concept cue

Validate all paths and uniqueness first. Put inspector, record validation and identity checking inside a small try block; append either accepted data or one controlled failure.

Hint 3 · a localized example

except MetadataError as error: failures.append({"path": path, "code": error.code}). Do not catch Exception or return from that branch.

Need the complete worked solution?

Open asset_intelligence/core.py from the kit. Trace it, close it, then try fresh inputs in your own files. Treat the attempt as guided; seeing the solution does not award a practical pass.

Course help is guidance, not independent evidence. With JavaScript, opening help records guidance locally; otherwise note it in your README. Reset does not erase that history.

Repair a failed check

If later assets disappear, remove break/return from the expected-failure branch. If a RuntimeError becomes a row, narrow the catch. If a source mismatch passes, compare normalized source with the requested path. If callback work occurs before a bad later path is rejected, move list validation before the loop.

NotImplementedError means a practice stub is still unfinished. Read the failing test name and the last error line. Change one behavior, rerun that build, then rerun all implemented builds.

Show it works on new inputs

Create a new four-path in-memory fixture with good, missing, corrupt and later-good results. Add an unexpected PermissionError case and show it propagates. Record counts, order, assistance and the difference between this simulation and actual sidecar tests.

Self-review: name the input, result, refused case and reason. Your local test output and explanation are separate from a quiz score; this page does not certify a pass.

Keep the idea

Resilience means continuing through declared recoverable cases while preserving failures—not suppressing every exception.