Module 8 Rubric#

Artifact#

reproducible training pipeline readiness review with profiling, cost, and deployment risks

Criterion

Excellent

Satisfactory

Needs Revision

Technical correctness

Neural-network concepts, code, metrics, and terminology are accurate for gpu workflows, scale, and deployment.

Most technical claims are accurate, with minor gaps or imprecision.

Claims are incorrect, unsupported, or disconnected from the module lab.

Experimental evidence

Uses notebook results, plots, and comparisons to support a clear argument; identifies what the toy setup proves and does not prove.

Uses notebook evidence but interpretation or limits are incomplete.

Reports outputs without analysis, comparison, or limitations.

Design tradeoffs

Explains architecture, optimization, data, compute, or deployment tradeoffs in concrete terms.

Names tradeoffs but does not fully connect them to design decisions.

Treats model choices as arbitrary or self-evident.

Risk and failure analysis

Identifies technical and operational failure modes, including overfitting, data mismatch, compute limits, or misuse where relevant.

Identifies some failure modes but mitigations are generic.

Omits meaningful failure analysis.

Communication

Recommendation is concise, reproducible, and understandable to an AI engineering reviewer.

Recommendation is understandable but partially supported.

Recommendation overclaims or cannot be reproduced from the submitted evidence.

Minimum Completion Standard#

A passing submission must include runnable notebook evidence, at least one baseline or parameter comparison, one documented limitation of the toy setup, and a recommendation for the next experiment or deployment gate.

Because this is the Populi final portfolio assessment, the readiness review must reuse, revise, or reject evidence from at least three earlier submitted artifacts and trace it through architecture, training, evaluation, compute, monitoring, and rollback. This integration is evaluated under Experimental evidence, Design tradeoffs, and Communication; this does not create a second Module 8 grade.

100-Point Scoring Guide#

Use this 100-point guide to plan and self-check your submission. The qualitative rubric above defines the performance levels; the table below shows how each criterion contributes to the total.

Criterion

Points

How to self-check

Technical correctness

20

Full criterion credit requires evidence that meets the Excellent description; use the other descriptions to identify what to revise before submission.

Experimental evidence

25

Full criterion credit requires evidence that meets the Excellent description; use the other descriptions to identify what to revise before submission.

Design tradeoffs

20

Full criterion credit requires evidence that meets the Excellent description; use the other descriptions to identify what to revise before submission.

Risk and failure analysis

20

Full criterion credit requires evidence that meets the Excellent description; use the other descriptions to identify what to revise before submission.

Communication

15

Full criterion credit requires evidence that meets the Excellent description; use the other descriptions to identify what to revise before submission.

Performance Bands#

Band

Points

Interpretation

Excellent

90-100

Graduate-level command; evidence is accurate, contextualized, and professionally defensible.

Proficient

80-89

Meets graduate expectations with minor gaps in depth, precision, or integration.

Developing

70-79

Demonstrates partial outcome achievement but needs revision for rigor, evidence, or communication.

Not Yet Demonstrated

Below 70

Does not yet provide adequate evidence of the aligned course outcomes.

🧑‍🌾 SAMWISE — Pre-Submission Self-Check#

🧑‍🌾 SAMWISE — Student note

Before submitting, point to specific artifact evidence for every criterion. Confirm that another reader can reproduce or inspect it, that your recommendation follows from it, and that you name at least one limitation and one responsible next step.

Use the performance bands honestly: revise a criterion when your evidence matches Satisfactory, Developing, or Not Yet Demonstrated rather than relying on polished prose to cover a missing result.

This is prewritten guidance; it does not provide answers or predict a grade.