| Course | HIM 400 Communication and Technologies II |
|---|---|
| Module | Module 5 |
| Paper type | undergraduate paper applying data mining to predict missed appointments |
| Length | About 1,060 words, 6 pages |
| Format | APA 7 student paper |
| School | Southern New Hampshire University |
| Program | BS Health Information Management |
| Updated | September 2026 |
Free sample paper for HIM 400 Module 5
Who Will Miss the Visit, and What Should We Do About It? Mining Appointment Data at Cold Brook Health
[Student Name]
Southern New Hampshire University
HIM 400: Communication and Technologies II
Module Five Short Paper
[Instructor Name]
[Date]
The organization, setting and figures below are a composite written as a model document. No real employer, client, colleague or patient is described.
Who Will Miss the Visit, and What Should We Do About It? Mining Appointment Data at Cold Brook Health
About one primary care appointment in eight at Cold Brook Health ends as a no-show. Each empty slot is a patient whose diabetes, blood pressure or depression went unchecked and a clinician's time that could have gone to someone on the waiting list. Operations leaders asked whether data mining could predict which appointments are likely to be missed. This paper describes how the analytics team built and tested such a model and why the most important decision was not which algorithm to use but what to do with a patient's risk score.
Data Mining as a Process
Data mining finds patterns in large data sets that can describe or predict behavior. Its main tasks include classification, which assigns records to categories such as attended or missed; clustering, which groups similar records without predefined labels; and association rules, which find items that occur together. Predicting no-shows is a classification task. The team followed the widely used CRISP-DM sequence: understand the business question, understand and prepare the data, build models, evaluate them and deploy the result. Writing down the business question first mattered, because it shaped every later choice, including which errors were acceptable.
Data and Predictors
The data set contained 148,300 scheduled primary care appointments from January 2024 through December 2025, of which 12.1% were no-shows. Pooling the published literature on missed appointments, Dantas et al. (2018) reported that rates varied widely across settings and that a patient's history of missed visits and the lead time between booking and the appointment were among the most frequently reported predictors. Guided by that review, the team built 14 candidate predictors from the scheduling system and the registry: prior no-shows in the past year, lead time in days, age, appointment type, day and hour, whether a reminder was confirmed, portal enrollment, insurance type and estimated travel distance. Clinical diagnoses were left out because the model's purpose was operational.
Building and Testing the Models
Rajkomar et al. (2019) caution that a model must be judged on data it has not seen, ideally from a later period, because patterns shift over time and a model tested on its own training data will look better than it is. The team therefore trained on 2024 appointments and tested on 2025. Two models were compared: logistic regression, which estimates how each predictor changes the odds of a no-show and is easy to explain, and a gradient-boosted tree model, which combines many small decision trees and can capture interactions. Discrimination on the 2025 test set, summarized as the area under the curve (AUC), reached 0.74 for logistic regression and 0.77 for the tree model. An AUC of 0.5 would mean the model ranks patients no better than a coin toss, and 1.0 would mean perfect separation. The small gain did not justify a model that staff could not interpret, so the team chose logistic regression.
What the Numbers Mean in Practice
A score becomes useful only when paired with a threshold and an action. The clinics can make about 15% of appointments the focus of extra attention each week, so the team flagged the highest-risk 15%. Table 1 shows what that means. Of flagged appointments, 31% were actually missed, compared with 12.1% overall, so outreach effort is concentrated where it is most needed. But the flags caught only 38% of all no-shows, so most missed appointments will still come as a surprise.
Table 1. Model Performance on 2025 Appointments at a 15% Flag Threshold
| Measure | Result | Plain meaning |
|---|---|---|
| AUC | 0.74 | Ranks a missed visit above an attended one 74% of the time |
| Share of appointments flagged | 15.0% | Set by clinic capacity |
| Positive predictive value | 31% | Share of flagged appointments that were missed |
| Sensitivity | 38% | Share of all no-shows that were flagged |
| Overall no-show rate | 12.1% | Baseline for comparison |
Note. Composite results from the logistic regression model.
Checking the Model by Subgroup
Obermeyer et al. (2019) examined a commercial risk tool used to enroll patients in care management and found that it assigned Black patients lower scores than white patients who were equally ill. The tool had been trained to forecast spending, and because unequal access had kept spending lower for Black patients, cost stood in poorly for need. The lesson is that a model can reproduce the conditions behind its training data. A no-show reflects barriers as much as choices, and at Cold Brook those barriers are concentrated among Medicaid patients and patients living more than 30 miles away. As expected, these groups were flagged far more often: 29% of appointments for patients on Medicaid and 34% for distant patients, against 15% overall. Accuracy was similar across groups, with areas under the curve between 0.72 and 0.75, so the model was not simply wrong for them. The question was what the flags would lead to.
Choosing the Action
Many clinics use no-show predictions to overbook. At Cold Brook, overbooking flagged slots would mean that patients on Medicaid and patients from distant towns, who already face the most barriers, would more often arrive to find a double-booked clinician and a long wait. Char et al. (2018) argue that the ethics of machine learning in health care depend on the intentions and incentives of those who design and deploy it, not only on the model's accuracy. The team recommended that flags trigger help instead: a reminder call from a scheduler who can offer a video visit, arrange a ride through the patient's insurance benefit or move the appointment to a nearer practice. Overbooking would not be used for flagged appointments.
Monitoring After Deployment
Once deployed, the model will be reviewed each quarter for three things: whether its accuracy holds on new appointments, whether flag rates by subgroup change and whether outreach actually lowers no-shows among flagged patients compared with the year before. Because the model predicts behavior that the outreach is designed to change, a successful program will make the model look less accurate over time, and the team will retrain it on new data each year.
Conclusion
Data mining gave Cold Brook a reasonably accurate way to identify appointments at risk of being missed. The more important work was deciding what a flag should mean. By testing the model on a later year, reporting results in plain terms, checking subgroups and linking predictions to help rather than overbooking, the team turned a technical result into a fair operational tool.
References
Char, D. S., Shah, N. H., & Magnus, D. (2018). Implementing machine learning in health care: Addressing ethical challenges. New England Journal of Medicine, 378(11), 981-983. https://doi.org/10.1056/NEJMp1714229
Dantas, L. F., Fleck, J. L., Cyrino Oliveira, F. L., & Hamacher, S. (2018). No-shows in appointment scheduling: A systematic literature review. Health Policy, 122(4), 412-421. https://doi.org/10.1016/j.healthpol.2018.02.002
Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447-453. https://doi.org/10.1126/science.aax2342
Rajkomar, A., Dean, J., & Kohane, I. (2019). Machine learning in medicine. New England Journal of Medicine, 380(14), 1347-1358. https://doi.org/10.1056/NEJMra1814259
What the HIM 400 Module 5 instructions ask for
The HIM 400 data mining assignment usually asks you to explain data mining methods and apply one to a health care question, often prediction or pattern discovery. Plan four to five pages in APA 7 with three or more scholarly sources and at least one table. State the business question first, name the data mining task, describe the data and predictors, and explain how the model was built and tested, ideally on data from a later period. Report results in terms a manager can use, such as positive predictive value and sensitivity at a chosen threshold, then examine how the model performs and is used for different patient groups before recommending an action and a monitoring plan.
How this HIM 400 Module 5 data mining short paper example is built
Cold Brook Health loses about 12% of primary care appointments to no-shows. Guided by Dantas and colleagues, the team builds 14 predictors from 148,300 appointments, trains on 2024 and tests on 2025 as Rajkomar and colleagues advise and chooses logistic regression over a slightly stronger tree model for interpretability. A table explains an area under the curve of 0.74, a 31% positive predictive value and 38% sensitivity at a 15% threshold. Obermeyer and colleagues frame a subgroup check showing Medicaid and distant patients flagged far more often, and Char and colleagues support using flags for rides and video visits instead of overbooking.
Where the HIM 400 Module 5 rubric puts the points
Data mining papers in HIM 400 are usually graded on correct explanation of methods, a clear link between the business question and the task, sound validation, accurate interpretation of performance measures, attention to bias and ethics and APA 7 mechanics. The strongest papers explain why an interpretable model might be preferred over a slightly more accurate one and translate statistics into plain meaning. Graders reward writers who separate a model's accuracy from the fairness of the decision it supports, and who plan monitoring after deployment. Noting that success can make a model look worse over time shows a thoughtful grasp of how predictions interact with the actions they trigger.
HIM 400 Module 5 help: the mistakes that cost points
Data mining papers lose points when they report accuracy without a baseline, test a model on its training data, describe algorithms in jargon without explaining what they do or ignore how predictions will be used. Another frequent gap is treating bias as a question about the algorithm alone. If your course provides a data set or asks for a specific method, such as clustering or association rules, send it with the prompt so the sample applies that method to your data. Say which software you use. A custom paper can study readmissions, length of stay or claim denials while following the same path from question to model, evaluation, subgroup check, action and monitoring.
Get HIM 400 Module 5 written to your instructions
Share the HIM 400 Module 5 prompt and any data set or method your instructor requires. The paper will frame the business question, describe the data and model, test it on unseen data, explain results in plain terms, check subgroups and recommend a fair use, delivered within 24 to 48 hours with the first sample free. The paper above is an original model document written by our desk, not a submitted student paper and not an official Southern New Hampshire University document.
More HIM 400 papers and related BS Health Information Management samples
- HIM 400 Module 1 Discussion: Why Health Technology Projects Fail
- HIM 400 Module 2 Database Structures Short Paper: A Relational Design for a Diabetes Registry
- HIM 400 Module 3 Data Extraction Short Paper: Queries, Data Requests and the Minimum Necessary
- HIM 400 Module 4 Project One: Trends and Patterns in Diabetes Control
- HIM 360 Module 7 Project Two: A Documentation Integrity Improvement Plan
- HIM 350 Module 1 Discussion: How Technology Changed Care Conversations After 2020
- HIM 220 Module 6 Data Governance Short Paper: Stewardship and Patient Matching
- HIM 200 Module 7 Project Two: A Plan to Improve Portal Use and Record Exchange
HIM 400 Module 5 questions, answered
Where can I find a free HIM 400 Module 5 Data Mining Short Paper sample?
This page carries the complete HIM 400 Module 5 paper: a no-show prediction model built, tested on a later year, checked by subgroup and used for outreach.
What is the difference between classification and clustering?
Classification assigns records to known categories using labeled examples; clustering groups similar records without predefined labels.
Why test a model on a later time period?
Patterns change over time, so testing on later data shows how the model will perform once it is actually used.
What do sensitivity and positive predictive value mean for a no-show model?
Sensitivity is the share of all no-shows the model flags; positive predictive value tells you how many of the flagged visits really go unattended.
Can an accurate model still be unfair?
Yes. If its predictions lead to actions that burden groups already facing barriers, the use can be unfair even when accuracy is similar across groups.