| Course | NUR 640 Assessment and Evaluation in Nursing Education |
|---|---|
| Module | Module 3 |
| Paper type | milestone paper on redesigning a clinical performance evaluation tool |
| Length | About 1,030 words, 6 pages |
| Format | APA 7 student paper |
| School | Southern New Hampshire University |
| Program | MSN |
| Updated | September 2026 |
Free sample paper for NUR 640 Module 3
Milestone One: Redesigning a Medical-Surgical Clinical Evaluation Tool for Validity, Consistency and Fairness
[Student Name]
Southern New Hampshire University
NUR 640: Assessment and Evaluation in Nursing Education
Module Three Milestone One
[Instructor Name]
[Date]
Milestone One: Redesigning a Medical-Surgical Clinical Evaluation Tool for Validity, Consistency and Fairness
Clinical evaluation carries high stakes: a student who fails a clinical course may be dismissed from the program, and a student who passes carries that judgment into practice. Yet clinical performance is harder to judge than a written exam because every patient is different, instructors observe only part of each day and the criteria are often abstract. At the composite Lakemont College of Nursing, students in the second-semester medical-surgical course filed complaints that clinical grades depended on which instructor they had, and two failing grades were appealed in one term. This milestone revises the course's clinical evaluation tool. It argues that observable behaviors, anchored performance levels, documented evidence and rater calibration make clinical evaluation more valid, more consistent and fairer to students.
Problems With the Current Tool
The existing tool has twelve criteria, each rated satisfactory or unsatisfactory at the end of the rotation. Several criteria cannot be observed directly. "Demonstrates professionalism" could mean punctuality, appearance, communication or accepting feedback, and each instructor reads it differently. "Uses critical thinking" names a mental process rather than a behavior. The two-point scale hides the difference between a student who barely meets expectations and one who excels, so it offers no guidance for improvement. And because ratings are completed only at the end, a student can learn of a serious concern for the first time in the final week, too late to improve. These are threats to validity, because the ratings may not reflect the intended competencies, and to reliability, because different instructors may rate the same performance differently.
Rewriting Criteria as Observable Behaviors
Each criterion was rewritten so that an instructor can point to evidence. "Demonstrates professionalism" became three behaviors: arrives prepared with a completed patient data sheet; communicates changes in patient condition to the assigned nurse using SBAR within the shift; and responds to feedback by describing what they will change. "Uses critical thinking" was replaced by clinical judgment behaviors taken from the rubric Lasater (2007) built by watching nursing students work through simulations; it describes four dimensions: noticing, interpreting, responding and reflecting. For example, under noticing, the student identifies a change from baseline, such as rising heart rate or new confusion, and seeks related data without prompting. Each behavior is linked to a course outcome so that the tool measures what the course claims to teach.
Anchored Performance Levels
The two-point scale was replaced by four levels, shown for one criterion in Table 1, following the structure of the Lasater rubric. Each level describes what a student at that level typically does. The passing standard is set at Developing for the midterm and Accomplished for the final evaluation on every criterion marked essential, such as patient safety, which reflects the expectation that students progress during the rotation. A student who is rated Beginning on any safety behavior at any point receives a written learning plan within 48 hours.
Table 1. Example of Anchored Levels for One Criterion, Adapted From Lasater (2007)
| Level | Noticing: recognizes change in patient status |
|---|---|
| Beginning | Needs cues to notice abnormal findings; focuses on tasks rather than the patient's condition |
| Developing | Notices obvious changes but may miss subtle ones; seeks additional data only when prompted |
| Accomplished | Notices most changes from baseline and gathers related data independently |
| Exemplary | Monitors a variety of data, anticipates change and recognizes subtle patterns early |
Note. Each level describes what the student does, so instructors compare observed behavior with a description rather than with an impression.
Documenting Evidence and the Midterm Conference
Instructors will keep brief weekly anecdotal notes for each student, recording specific observed behaviors, both strengths and concerns, with the date and patient context. These notes are the evidence behind ratings and are shared with students weekly so there are no surprises. At midterm, each student completes a self-evaluation on the same tool and meets with the instructor to compare ratings. This conference is formative, reflecting evidence that feedback during learning raises achievement more than judgment at the end (Black & Wiliam, 1998): its purpose is to agree on targets the student will work on over the remaining weeks, which are recorded and reviewed at the final evaluation. Documented evidence also protects both students and faculty in an appeal, because ratings rest on recorded observations rather than memory.
Calibrating Raters
A better tool will not produce consistent ratings unless instructors apply it the same way. Jonsson and Svingby (2007) reviewed studies of scoring rubrics and found that reliability improves when rubrics are analytic and topic-specific and when raters are trained with examples, and that rubrics can support valid judgment of complex performance when criteria are clearly linked to what is being assessed. Before the next term, all six clinical instructors will watch three recorded simulation scenarios of students at different levels and rate each independently. Ratings will be compared criterion by criterion, and any criterion where two raters sit two or more levels apart goes to group discussion until the group agrees on what the anchors mean. The exercise will be repeated at midterm with a new scenario to check whether agreement holds.
Evaluating the Revision
The revised tool will be judged by evidence rather than impression. Measures will include agreement among instructors on the calibration scenarios, the number of grade appeals and complaints compared with the previous two terms, the proportion of students with learning plans who reach the passing standard by the final evaluation and student ratings of how clear the expectations were. Instructors will also be asked how long the tool takes to complete, since a tool that is too burdensome will be completed hastily.
Fairness receives its own check. Ratings will be reviewed at the end of the term for patterns by student group, such as whether multilingual students or students of color are more often rated Beginning on communication behaviors, since vague criteria leave room for bias. Any pattern will be discussed at the next calibration session, and the related anchors will be reworded if they reward a particular speaking style rather than safe, clear communication.
Conclusion
Clinical evaluation will always involve judgment, but that judgment can be structured. Rewriting criteria as observable behaviors, anchoring four levels with examples adapted from the Lasater rubric, documenting evidence weekly, holding a formative midterm conference and calibrating instructors should give Lakemont students a clinical grade that reflects their performance rather than their instructor assignment.
References
Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7-74. https://doi.org/10.1080/0969595980050102
Jonsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity and educational consequences. Educational Research Review, 2(2), 130-144. https://doi.org/10.1016/j.edurev.2007.05.002
Lasater, K. (2007). Clinical judgment development: Using simulation to create an assessment rubric. Journal of Nursing Education, 46(11), 496-503. https://doi.org/10.3928/01484834-20071101-04
What the NUR 640 Module 3 instructions ask for
Milestone One in NUR 640 often asks you to analyze an existing assessment from your setting, such as a clinical evaluation tool, and propose revisions grounded in measurement principles. Expect to describe the tool's purpose, identify weaknesses in validity, reliability and fairness, present revised criteria and scoring and plan how you will check that the revision works. Plan on four to six pages in APA 7. Quote or paraphrase actual criteria, rewrite them as observable behaviors, show anchored levels for at least one criterion, set a clear passing standard and describe how raters will be trained and compared. Address how students are told of concerns and appeal ratings. A short evaluation plan for the revision should close the paper.
How this NUR 640 Module 3 milestone one example is built
This milestone revises a medical-surgical clinical evaluation tool at a composite college after complaints that grades depended on the instructor. It critiques criteria such as "uses critical thinking" as unobservable and the two-point scale as uninformative. Criteria are rewritten as behaviors, with clinical judgment items drawn from the Lasater rubric's four dimensions, and a table shows anchored levels for noticing. Weekly anecdotal notes and a formative midterm conference are added. A calibration plan, supported by Jonsson and Svingby, has six instructors rate recorded scenarios, and evaluation measures include appeals and rater agreement. A fairness review looks for rating patterns by student group. Instructor workload is also tracked.
Where the NUR 640 Module 3 rubric puts the points
Grading of this milestone usually weighs the analysis of the current assessment, the quality of revisions, application of validity and reliability concepts, attention to fairness, a plan for evaluating the revision and APA 7 writing. Top-band papers quote specific criteria and explain exactly why each threatens validity or reliability. Graders reward revised criteria that are observable and anchored, a clear passing standard with a defined response to safety concerns and a calibration method with named steps. Including formative feedback, such as a midterm conference, and measurable evaluation outcomes shows the reviewer that the tool will be used to improve learning. Due process for students is often scored.
NUR 640 Module 3 help: the mistakes that cost points
Clinical evaluation papers lose points when criteria are revised but remain abstract, when levels are labeled without descriptors, when no passing standard is set or when rater consistency is assumed rather than tested. Another gap is ignoring due process, such as how and when students learn of concerns. Rewrite criteria as behaviors, anchor each level, set the standard, document evidence, give formative feedback and calibrate raters. If your milestone focuses on a different tool, such as a preceptor evaluation, a skills checklist or a simulation rubric, share it with your NUR 640 prompt so the revision fits your setting. Show how students hear about concerns early. Then test whether raters agree.
Get NUR 640 Module 3 written to your instructions
Share the NUR 640 milestone prompt, the tool you are revising and the rubric. The paper you receive will name specific threats to validity and reliability, rewrite criteria as observable behaviors, anchor performance levels and plan rater calibration, within 24 to 48 hours, free the first time. The paper above is an original model document written by our desk, not a submitted student paper and not an official Southern New Hampshire University document.
More NUR 640 papers and related MSN samples
- NUR 640 Module 1 Discussion: Formative Checks That Change Tomorrow's Class
- NUR 640 Module 2 Item Writing Paper: Clinical Judgment Items on Sepsis
- NUR 520 Module 5 Statistical Inference Paper: Confidence Intervals Around County Estimates
- NUR 631 Module 3 Milestone One: Framing the Staffing Issue With Data
- NUR 550 Module 3 Milestone One: A Documented, Repeatable Literature Search on Delirium Prevention
- NUR 557 Module 6 Case Paper: A Duodenal Ulcer, H. pylori and Bismuth Quadruple Therapy
NUR 640 Module 3 questions, answered
Where can I find a free NUR 640 Module 3 Milestone One sample?
The full paper is on this page: a medical-surgical clinical evaluation tool rebuilt with observable behaviors, anchored levels adapted from Lasater, a midterm conference and rater calibration.
What is the Lasater Clinical Judgment Rubric?
A rubric developed from observations of nursing students in simulation that rates noticing, interpreting, responding and reflecting on a four-step scale from beginning to exemplary.
Why are satisfactory or unsatisfactory clinical ratings a problem?
They hide differences in performance, give little guidance for improvement and, with vague criteria, are applied inconsistently by different instructors.
How do you make clinical grading more consistent across instructors?
Use observable, anchored criteria and have instructors rate the same recorded performances, compare scores and discuss disagreements until they agree on what the anchors mean.
What are anecdotal notes in clinical evaluation?
Brief dated records of specific behaviors an instructor observed, both strengths and concerns, that serve as the evidence behind clinical ratings.