Assessment and Feedback

Performance Rubrics: Criteria, Scoring, and Validity Checks

A performance rubric is a scoring guide for observable features of work on a defined task. It should represent the intended knowledge or skill, distinguish meaningful levels, and support the decision the score will inform. A polished grid does not make a task authentic, a criterion valid, or two scorers consistent.

Choose analytic or holistic scoring deliberately

Comparison of analytic and holistic rubrics
FeatureAnalytic rubricHolistic rubric
How it worksScores separate criteria, then reports the profile or combines scores.Assigns one description and score to the work as a whole.
Useful whenStudents need criterion-specific feedback or the decision depends on distinct components.Overall performance is the construct and raters can judge it as a coherent whole.
Main strengthShows where evidence differs across criteria.Can be faster and preserves the interaction among features.
Main riskToo many criteria can fragment the performance or reward repeated evidence.One strong feature can mask a serious weakness; feedback is less specific.
Scoring checkApply each criterion independently before calculating a total.Use anchor responses that illustrate each whole-performance level.

Start with the performance task

Write the task before writing level labels. The task must elicit the evidence named in the objective. Authentic here means the work uses knowledge or skill in a purposeful performance; resemblance to adult work is not enough. Indiana University’s authentic-assessment guidance emphasizes applying learning through meaningful tasks and defined criteria.

Example performance task

Objective: Use a supplied data set to recommend one of two plans, explain the pattern accurately, address a relevant counterargument, and communicate the recommendation to a named audience.

Prompt: The school is considering two plans for a shaded courtyard. Using the provided tables for cost, shaded area, maintenance time, and student-use counts, write a 250-350 word recommendation to the facilities committee. Select Plan A or Plan B, cite at least three relevant values, explain how the evidence supports the recommendation, address one reasonable counterargument, and identify one limitation in the data.

Conditions: Students may use the annotated data packet and calculator. A written response, dictated response with transcript, or approved communication format is permitted when it preserves the same evidence and reasoning criteria. Decorative design is not scored.

Constructed task data: These values were created for rubric practice and are not actual school findings.

Constructed courtyard plan data for the performance task
PlanCostShaded areaMonthly maintenanceMean midday users across three clear days
A$18,000300 m²6 hours68
B$24,000420 m²10 hours80

A four-by-four observable rubric

Score each criterion separately. The descriptions identify evidence in the response, not traits such as creativity, effort, or sophistication.

Four-criterion four-level rubric for the courtyard recommendation
Criterion4: Complete3: Adequate2: Partial1: Insufficient
Data accuracy and selectionCites at least three accurate, relevant values from two or more supplied tables and makes valid comparisons.Cites three accurate values; most are relevant and comparisons are valid, with one minor omission.Cites one or two accurate values, or includes a comparison error that weakens the recommendation.Uses no accurate supplied values or substantially misreads the data.
Evidence-to-recommendation reasoningStates a clear recommendation and explains how the selected evidence supports it under the stated priorities.States a recommendation and connects most evidence to it; one reasoning step is implied.States a recommendation but mainly lists data or relies on an unsupported preference.Gives no recommendation or provides reasoning unrelated to the supplied evidence.
Counterargument and limitationAddresses a relevant counterargument with evidence and identifies a specific data limitation that affects confidence or scope.Addresses a relevant counterargument and names a relevant limitation; one explanation is brief.Mentions a counterargument or limitation but not both, or treats either generally.Omits both or gives unrelated cautions.
Communication for audienceOrganizes the recommendation so the committee can identify the decision, reasons, cited values, and qualification without inference.Organization and language make the decision and reasons clear, with one local lapse.The decision can be found, but unclear references or organization obscure part of the reasoning.The response is too incomplete or unclear to identify the proposed decision and its basis.

Replace weak criteria and level language

Weak and improved rubric language
Weak versionProblemImproved version
Criterion: QualityCombines accuracy, reasoning, organization, and presentation.Use separate criteria such as data accuracy and evidence-to-recommendation reasoning.
Criterion: EffortIntent and time are not reliably visible in the final performance.If process is part of the objective, score a defined action such as submits a revision that addresses one identified criterion.
Level: ExcellentNames a judgment but not the evidence that earns it.Cites three accurate values from two tables and makes valid comparisons.
Levels: Always / usually / sometimes / neverFrequency is hard to infer from a single product.Describe the amount, accuracy, relevance, or completeness visible in that product.
Level: No errorsTreats all errors as equally important and makes perfection the top condition.Specify which errors change the interpretation and what accurate evidence must be present.

Score two constructed samples independently

The short samples below are written for rubric practice; they are not student records or reported classroom results.

Sample A

I recommend Plan A. It costs $18,000 instead of $24,000 and needs 6 maintenance hours each month instead of 10. Plan B shades 420 square metres, 120 more than Plan A, but the student-use table shows only 12 more users at midday. If maximum shade is the committee’s first priority, Plan B is stronger. Under a limited budget, Plan A adds substantial shade at lower cost and maintenance. The counts cover only three clear days, so they may not represent other weather or seasons.

Sample B

I choose Plan B because it has more shade and students will like it. The table says 420 square metres. It is the better plan even though it costs more. The school should choose quality.

Have scorers assign all criterion scores before seeing another score. The table shows a defensible independent application, including one disagreement to calibrate.

Independent scores for two constructed sample responses
Sample and criterionScorer 1Scorer 2Evidence and calibration decision
A: Data44Four accurate values from the cost, maintenance, shade, and use tables; comparisons are valid.
A: Reasoning44Recommendation is tied to the stated budget priority and evidence.
A: Counterargument and limit44Addresses maximum shade and explains why three clear days limit the use count.
A: Communication44Committee can identify the decision, reasons, values, and qualification.
B: Data22One accurate value; more is not quantified and no second table is used.
B: Reasoning22Recommendation is present, but better and quality are unsupported preferences.
B: Counterargument and limit12Costs more acknowledges contrary information but does not address it; no limitation appears. Apply level 2 only when mentioning one relevant counterargument satisfies the counterargument or limitation descriptor. Clarify this during calibration.
B: Communication22Decision is identifiable, but the evidence and basis are incomplete.

Calibrate before consequential scoring

  1. Select two to four responses that show different patterns, remove identifying information when authorized, and have every scorer score them independently.
  2. Compare criterion scores, not only totals. Each scorer points to the response feature and rubric phrase used.
  3. Resolve whether the rubric, task, scorer interpretation, or sample caused each disagreement. Revise ambiguous descriptions before operational scoring.
  4. Record annotated anchor responses. Recheck agreement after a scoring break and when a new response pattern appears.
  5. Use a defined adjudication route for unresolved differences. Do not simply average scores when the difference reveals a criterion interpretation problem.

Agreement is evidence about consistency under those scoring conditions, not proof that the task measures the intended construct. Review alignment and access separately.

Check access and language demands

Identify language, reading, speaking, handwriting, motor, sensory, technology, and time demands in the task and rubric. Keep a demand when it is part of the objective; reduce or provide an approved route around it when it is incidental. Apply required accommodations and accessibility procedures. Student-friendly wording can improve usability, but do not replace necessary disciplinary terms; define them and provide examples.

Rubric access and language review
Review questionRevision when needed
Does a response format add a skill that is not assessed?Offer an equivalent approved format and score the same observable criterion.
Do clarity or professional hide language conventions?Name the organization or language feature needed for the audience and objective.
Does one level count grammar errors unrelated to the construct?Remove that count or state the particular error that prevents meaning from being understood.
Can students and scorers distinguish adjacent levels?Add a contrasting example and revise the boundary using observable evidence.
Can assistive technology access the task and rubric?Test the actual files, reading order, table structure, labels, and response route before use.

Set weights after defining importance

Weight a criterion only when its importance to the objective and decision justifies the influence. In the example, data and reasoning might each receive 30%, counterargument and limitation 25%, and communication 15%. Publish the calculation before students begin. Check the effect with hypothetical profiles: a high communication score should not compensate for absent evidence if evidence is central.

Do not add points for the same evidence in several criteria. Report the criterion profile alongside any total when a minimum performance in one criterion matters. If a criterion is a threshold, such as a safety procedure, state the threshold rather than hiding it inside an average.

Copyable blank analytic rubric

Performance task:
Learning objective or construct:
Decision the scores will inform:
Allowed response routes and required accommodations:

Criterion 1:
4:
3:
2:
1:
Weight or threshold:

Criterion 2:
4:
3:
2:
1:
Weight or threshold:

Criterion 3:
4:
3:
2:
1:
Weight or threshold:

Criterion 4:
4:
3:
2:
1:
Weight or threshold:

Sample responses used for calibration:
Features marking adjacent levels:
Scorer disagreement and resolution rule:
Access or language barrier found:
Revision before use:

Review the score before using it

After scoring, ask whether the task elicited every criterion, the descriptions fit actual response patterns, scorers applied them consistently, and access conditions allowed students to show the intended performance. Record revisions. A rubric score describes the sampled work under stated conditions; it should not be stretched into a judgment about a learner’s general ability or potential.

Sources

What changed: This revision adds a complete performance task and rubric, independent sample scoring, calibration and weighting guidance, access checks, and a copyable blank rubric.