NR-722 · Week 7 of 8 · Assessment, feedback and evaluation instruments

NR-722 Week 7 Assessment and Feedback Design: How to Write It

The short answer

The seventh stage of a facilitation course usually turns to assessment: what formative and summative assessment are actually for, how an instrument earns validity evidence, how a scoring guide is built so two raters agree, and what makes feedback change performance rather than merely inform a learner they were wrong. Assessment is the part of education with the most technical vocabulary and the most direct consequences for learners, and doctoral rubrics grade the precision closely. Your section may print this as NR 722 or NR722; it is the same course. Chamberlain publishes no syllabi outside Canvas. The placement here is our teaching judgment from the course's catalog arc; your section's rubric decides what your week actually asks.

NR-722 Week 7 grading scale at Chamberlain, the criterion levels this assessment is scored on, from Chamberlain Tutors
How Chamberlain grades NR-722 Week 7, visualized by Chamberlain Tutors.

What NR-722 Week 7 asks for

What does a score actually tell you? A clinic running an annual skills day used a checklist that had grown by accretion over six years. It contained forty-one items, several of which described the same action in different words, three of which no longer matched current practice, and one of which was scored differently by every one of the four observers because it asked whether the staff member had communicated effectively with no description of what that looked like. Everyone passed. Nobody learned anything from the result, and when a genuine performance concern arose two months later, the skills day record was useless as evidence because it could not distinguish between a person who was competent and a person who had been observed by a lenient rater.

That is what this stage exists to prevent. Assessment design has a small number of ideas that do most of the work. Formative assessment exists to change learning while it is still in progress and is wasted if it arrives too late to act on. Summative assessment makes a judgment for a decision and needs enough evidence to bear the weight of that decision. Validity is not a property of an instrument but an argument about whether the interpretation of a score is justified for a stated purpose. Reliability is about consistency, and a scoring guide with vague criteria destroys it faster than any other single factor.

Deliverables here are typically an assessment plan, a scoring guide or rubric you construct, an analysis of an existing instrument against published criteria, or a feedback plan for a specific teaching situation. Where a discussion accompanies it, expect to be asked what your criteria actually mean, and remember posts do not reopen once submitted in Canvas.

Feedback is the other half of the stage and is where most learners in an educator course have the most to gain. Feedback that changes performance is specific to the task rather than to the person, arrives close enough in time that the learner can still remember their reasoning, includes what to do next rather than only what went wrong, and is followed by an opportunity to try again. Each of those is supported in the literature and each is easy to specify in a written plan.

One boundary is worth restating at this stage in particular. Building a scoring guide or an assessment plan as a scholarly product is coursework. Completing an actual evaluation of a real learner, signing a competency record, or recording a judgment a school or employer relies on is your own professional documentation and is never drafted, reconstructed or estimated with outside help. Keep the design work and the record entirely separate, and say so if your paper draws on a real evaluation process.

The NR-722 Week 7 method, step by step

Six moves for designing assessment and feedback that hold up to scrutiny.

  1. State the purpose and the decision the assessment serves

    Learning during instruction, or a judgment with consequences. The purpose determines everything else, and instruments fail most often because they were built for one purpose and used for the other.

  2. Return to the outcome verb and assess the same performance

    If the outcome says the learner will prioritize, the assessment must require prioritizing. Written tests of knowledge about a skill are evidence about knowledge, not about the skill.

  3. Write criteria that describe observable performance at each level

    Not adequate and excellent, but what the difference actually looks like in behavior. A criterion two raters would score differently is a criterion that has not been written yet.

  4. Build the validity argument explicitly

    Say what interpretation you want the score to support, what evidence supports it, and what would undermine it. Validity as an argument rather than a property is the doctoral framing and it earns marks directly.

  5. Design the feedback event, not just the feedback content

    When it arrives, who delivers it, how long it takes, whether it is written or spoken, and what the learner is asked to do with it. Feedback with no subsequent attempt attached rarely changes anything.

  6. Plan for rater agreement before the instrument is used

    Shared examples of each level, a brief calibration conversation, and one anchor performance everyone scores together. This is cheap, it is rarely done, and it is what makes the resulting scores mean something.

A layout and word budget for an assessment and feedback plan

Our frame for one instrument plus its feedback design, sized for roughly 1,500 to 1,800 words plus the scoring guide itself. It is our own outline rather than anything the university issues, and your week's rubric outranks it wherever they disagree.

SectionWhat belongs in itWord target
Purpose and decisionWhat the assessment is for, what decision follows from it, and who acts on the result.160 to 200
Outcome to evidenceThe outcome verb and the performance that would count as evidence of it, matched deliberately.200 to 250
The instrumentThe scoring guide with observable descriptors at each level, plus the reasoning behind the number of levels.340 to 400
Validity and reliability argumentThe interpretation being claimed, the evidence for it, threats to it, and the calibration plan for raters.300 to 360
Feedback designTiming, format, content structure, and the opportunity to act on it before the summative point.300 to 360
Fairness and limitsBias risks, accommodations, what the instrument cannot judge, and where the boundary of your role sits.200 to 250

Evidence craft for assessment writing

Use the field's technical terms exactly. Validity, reliability, formative, summative and criterion-referenced all have defined meanings, and misusing one in a course about assessment is more costly than misusing it anywhere else. Define each on first use and hold the definition.

Ground criteria in published performance descriptions where they exist. Competency statements and published assessment tools give you language that has already been through review. Adapting from them, with attribution, is both faster and more defensible than writing descriptors from scratch.

Say what the assessment cannot do. Every instrument has a boundary: a single observation cannot establish consistency, a checklist cannot capture judgment, a knowledge test cannot demonstrate transfer. Naming those limits is the mark of an educator who understands measurement.

Address bias concretely. Rater familiarity, language, communication style differences and the halo effect from an earlier strong performance are all documented threats. Name the two most likely in your setting and the specific procedure that reduces each.

Five mistakes that cost points in this week's territory

  • Criteria made of adjectives. Good, adequate and needs improvement describe nothing a second rater could apply consistently.
  • Formative in name only. Feedback delivered after the final judgment has been made is summative regardless of what it is called.
  • Validity treated as a property. Calling an instrument valid without saying valid for what interpretation misses the whole construct.
  • Knowledge tests standing in for performance. A written check about a skill measures knowledge about the skill and should be described that way.
  • Feedback with no second attempt. Telling a learner what was wrong at the last possible moment is information, not education.

Before you submit

  • Purpose and the decision the score supports are stated first
  • The assessed performance uses the same verb as the outcome
  • Every criterion describes observable behavior at each level
  • A validity argument names the interpretation, its evidence and its threats
  • Feedback has a timing, a format and a subsequent attempt attached
  • Bias risks are named with a specific procedure reducing each

Designing assessment for NR-722?

Send the rubric and the outcome you are assessing out of Canvas. A premium original draft comes back in 24 to 48 hours with observable criteria, a real validity argument and a feedback design that includes a second attempt, and revisions run until the grade lands.

Questions students ask about this stage

How many performance levels should a scoring guide have?
Three or four for most purposes, and the reasoning matters more than the number. Every additional level requires you to write a descriptor that is genuinely distinguishable from its neighbours, and raters cannot reliably discriminate between six shades of a performance they observed once. Three levels work well when the decision is essentially about whether a performance meets a standard, with one level below and one above. Four are useful when you want to separate developing from competent without collapsing them. What matters is that you can write an observable descriptor for each level that two people would apply the same way, and if you cannot write the descriptor, the level does not exist. Say in your paper why you chose the number you did, because that reasoning is exactly the judgment being assessed.
How do I give difficult feedback without damaging the climate I built earlier in the session?
Separate the performance from the person and separate the observation from the judgment. The structure that works reliably is to state what you observed in specific behavioral terms, state the standard, ask the learner for their own reading of the gap before offering yours, then agree one thing to change and the occasion on which they will try it again. Asking first is the step most often skipped and the one that does the most work, because a learner who identifies the gap themselves has already done the cognitive work that feedback is meant to trigger. Difficult feedback delivered that way is entirely compatible with a safe climate; what damages climate is feedback that is vague, that arrives long after the event, or that is delivered to some learners and not others for the same performance.
Can I use an existing instrument from my workplace as the basis of the assignment?
Usually yes, and analyzing a real instrument often produces a stronger paper than building one from nothing, because real instruments carry real flaws worth diagnosing. Attribute it properly, describe it accurately, and be careful about reproducing a proprietary tool in full if its terms do not allow that; summarizing its structure and quoting a criterion or two for analysis is normally sufficient and safer. Then do the work the assignment wants: assess it against published criteria, identify which criteria are unobservable, say where two raters would diverge and why, and propose a revision with your reasoning. Keep any actual completed evaluations of real staff out of the submission entirely. The instrument is a document you may analyze; the records made with it about identifiable people are not coursework material.

Keep going

Online now