Module 6 Rubric#

Artifact#

attention-pattern analysis explaining token interactions, context limits, and transformer risks

Criterion

Excellent

Satisfactory

Needs Revision

Technical correctness

Neural-network concepts, code, metrics, and terminology are accurate for attention and transformers.

Most technical claims are accurate, with minor gaps or imprecision.

Claims are incorrect, unsupported, or disconnected from the module lab.

Experimental evidence

Uses notebook results, plots, and comparisons to support a clear argument; identifies what the toy setup proves and does not prove.

Uses notebook evidence but interpretation or limits are incomplete.

Reports outputs without analysis, comparison, or limitations.

Design tradeoffs

Explains architecture, optimization, data, compute, or deployment tradeoffs in concrete terms.

Names tradeoffs but does not fully connect them to design decisions.

Treats model choices as arbitrary or self-evident.

Risk and failure analysis

Identifies technical and operational failure modes, including overfitting, data mismatch, compute limits, or misuse where relevant.

Identifies some failure modes but mitigations are generic.

Omits meaningful failure analysis.

Communication

Recommendation is concise, reproducible, and understandable to an AI engineering reviewer.

Recommendation is understandable but partially supported.

Recommendation overclaims or cannot be reproduced from the submitted evidence.

Minimum Completion Standard#

A passing submission must include runnable notebook evidence, at least one baseline or parameter comparison, one documented limitation of the toy setup, and a recommendation for the next experiment or deployment gate.

Point-Weighted Scoring Guide#

Use this 100-point guide for direct assessment and gradebook entry. The qualitative rubric above explains performance levels; this table converts those levels into auditable scoring evidence.

Criterion

Points

Scoring Evidence

Technical correctness

20

Score the highest level fully met by the submitted evidence; partial credit requires specific instructor comment.

Experimental evidence

25

Score the highest level fully met by the submitted evidence; partial credit requires specific instructor comment.

Design tradeoffs

20

Score the highest level fully met by the submitted evidence; partial credit requires specific instructor comment.

Risk and failure analysis

20

Score the highest level fully met by the submitted evidence; partial credit requires specific instructor comment.

Communication

15

Score the highest level fully met by the submitted evidence; partial credit requires specific instructor comment.

Performance Bands#

Band

Points

Interpretation

Excellent

90-100

Graduate-level command; evidence is accurate, contextualized, and professionally defensible.

Proficient

80-89

Meets graduate expectations with minor gaps in depth, precision, or integration.

Developing

70-79

Demonstrates partial outcome achievement but needs revision for rigor, evidence, or communication.

Not Yet Demonstrated

Below 70

Does not yet provide adequate evidence of the aligned course outcomes.

Direct Assessment and Evidence Retention#

For accreditation sampling, retain the submitted artifact, rubric score, instructor feedback, and any revision notes. At minimum, archive one high-performing, one satisfactory, and one needs-revision artifact per offering when available. Remove or redact student identifiers before using artifacts for program assessment review.

Calibration Guidance#

Before grading a live cohort, instructors should score one sample artifact together or compare notes against this rubric. Calibration should focus on whether evidence supports the recommendation, whether limitations are explicit, and whether the submission demonstrates the aligned module outcomes rather than surface polish alone.