HIM 400 Module 5 Data Mining Short Paper Example

Reviewed by Delia Ravenscroft, MSN, RN

This HIM 400 Module 5 Data Mining Short Paper sample follows a predictive model from question to deployment and shows why a model can be accurate and still unfair in how it is used. It is written for SNHU HIM 400 (HIM-400), and it helps BS Health Information Management majors explain classification, validation and evaluation in terms a clinic manager can act on. The composite rural health system loses about 12% of primary care appointments to no-shows. Using two years of scheduling data, the paper describes the mining process, chooses predictors, compares two models tested on a later year and reports sensitivity and positive predictive value. It then checks results for patients on Medicaid and those living far away, and argues that flags should trigger help such as rides or video visits, not overbooking.

CourseHIM 400 Communication and Technologies II
ModuleModule 5
Paper typeundergraduate paper applying data mining to predict missed appointments
LengthAbout 1,060 words, 6 pages
FormatAPA 7 student paper
SchoolSouthern New Hampshire University
ProgramBS Health Information Management
UpdatedSeptember 2026

Free sample paper for HIM 400 Module 5

1

Who Will Miss the Visit, and What Should We Do About It? Mining Appointment Data at Cold Brook Health

[Student Name]

Southern New Hampshire University

HIM 400: Communication and Technologies II

Module Five Short Paper

[Instructor Name]

[Date]

The organization, setting and figures below are a composite written as a model document. No real employer, client, colleague or patient is described.

What this page is doingThe title pairs the prediction with the decision it is meant to support.
2

Who Will Miss the Visit, and What Should We Do About It? Mining Appointment Data at Cold Brook Health

About one primary care appointment in eight at Cold Brook Health ends as a no-show. Each empty slot is a patient whose diabetes, blood pressure or depression went unchecked and a clinician's time that could have gone to someone on the waiting list. Operations leaders asked whether data mining could predict which appointments are likely to be missed. This paper describes how the analytics team built and tested such a model and why the most important decision was not which algorithm to use but what to do with a patient's risk score.

What this page is doingThe introduction frames the problem and the paper's argument.
3

Data Mining as a Process

Data mining finds patterns in large data sets that can describe or predict behavior. Its main tasks include classification, which assigns records to categories such as attended or missed; clustering, which groups similar records without predefined labels; and association rules, which find items that occur together. Predicting no-shows is a classification task. The team followed the widely used CRISP-DM sequence: understand the business question, understand and prepare the data, build models, evaluate them and deploy the result. Writing down the business question first mattered, because it shaped every later choice, including which errors were acceptable.

What this page is doingThe paper defines data mining tasks and the process used.
4

Data and Predictors

The data set contained 148,300 scheduled primary care appointments from January 2024 through December 2025, of which 12.1% were no-shows. Pooling the published literature on missed appointments, Dantas et al. (2018) reported that rates varied widely across settings and that a patient's history of missed visits and the lead time between booking and the appointment were among the most frequently reported predictors. Guided by that review, the team built 14 candidate predictors from the scheduling system and the registry: prior no-shows in the past year, lead time in days, age, appointment type, day and hour, whether a reminder was confirmed, portal enrollment, insurance type and estimated travel distance. Clinical diagnoses were left out because the model's purpose was operational.

What this page is doingData and predictors are described with research support.
5

Building and Testing the Models

Rajkomar et al. (2019) caution that a model must be judged on data it has not seen, ideally from a later period, because patterns shift over time and a model tested on its own training data will look better than it is. The team therefore trained on 2024 appointments and tested on 2025. Two models were compared: logistic regression, which estimates how each predictor changes the odds of a no-show and is easy to explain, and a gradient-boosted tree model, which combines many small decision trees and can capture interactions. Discrimination on the 2025 test set, summarized as the area under the curve (AUC), reached 0.74 for logistic regression and 0.77 for the tree model. An AUC of 0.5 would mean the model ranks patients no better than a coin toss, and 1.0 would mean perfect separation. The small gain did not justify a model that staff could not interpret, so the team chose logistic regression.

What this page is doingTemporal validation and model comparison are explained.
6

What the Numbers Mean in Practice

A score becomes useful only when paired with a threshold and an action. The clinics can make about 15% of appointments the focus of extra attention each week, so the team flagged the highest-risk 15%. Table 1 shows what that means. Of flagged appointments, 31% were actually missed, compared with 12.1% overall, so outreach effort is concentrated where it is most needed. But the flags caught only 38% of all no-shows, so most missed appointments will still come as a surprise.

Table 1. Model Performance on 2025 Appointments at a 15% Flag Threshold

MeasureResultPlain meaning
AUC0.74Ranks a missed visit above an attended one 74% of the time
Share of appointments flagged15.0%Set by clinic capacity
Positive predictive value31%Share of flagged appointments that were missed
Sensitivity38%Share of all no-shows that were flagged
Overall no-show rate12.1%Baseline for comparison

Note. Composite results from the logistic regression model.

What this page is doingTable 1 translates performance measures into plain terms.
7

Checking the Model by Subgroup

Obermeyer et al. (2019) examined a commercial risk tool used to enroll patients in care management and found that it assigned Black patients lower scores than white patients who were equally ill. The tool had been trained to forecast spending, and because unequal access had kept spending lower for Black patients, cost stood in poorly for need. The lesson is that a model can reproduce the conditions behind its training data. A no-show reflects barriers as much as choices, and at Cold Brook those barriers are concentrated among Medicaid patients and patients living more than 30 miles away. As expected, these groups were flagged far more often: 29% of appointments for patients on Medicaid and 34% for distant patients, against 15% overall. Accuracy was similar across groups, with areas under the curve between 0.72 and 0.75, so the model was not simply wrong for them. The question was what the flags would lead to.

What this page is doingA subgroup check draws on research on algorithmic bias.
8

Choosing the Action

Many clinics use no-show predictions to overbook. At Cold Brook, overbooking flagged slots would mean that patients on Medicaid and patients from distant towns, who already face the most barriers, would more often arrive to find a double-booked clinician and a long wait. Char et al. (2018) argue that the ethics of machine learning in health care depend on the intentions and incentives of those who design and deploy it, not only on the model's accuracy. The team recommended that flags trigger help instead: a reminder call from a scheduler who can offer a video visit, arrange a ride through the patient's insurance benefit or move the appointment to a nearer practice. Overbooking would not be used for flagged appointments.

What this page is doingThe paper recommends a supportive use of predictions.
9

Monitoring After Deployment

Once deployed, the model will be reviewed each quarter for three things: whether its accuracy holds on new appointments, whether flag rates by subgroup change and whether outreach actually lowers no-shows among flagged patients compared with the year before. Because the model predicts behavior that the outreach is designed to change, a successful program will make the model look less accurate over time, and the team will retrain it on new data each year.

What this page is doingA monitoring plan closes the deployment step.
10

Conclusion

Data mining gave Cold Brook a reasonably accurate way to identify appointments at risk of being missed. The more important work was deciding what a flag should mean. By testing the model on a later year, reporting results in plain terms, checking subgroups and linking predictions to help rather than overbooking, the team turned a technical result into a fair operational tool.

What this page is doingThe conclusion restates the paper's main argument.
11

References

Char, D. S., Shah, N. H., & Magnus, D. (2018). Implementing machine learning in health care: Addressing ethical challenges. New England Journal of Medicine, 378(11), 981-983. https://doi.org/10.1056/NEJMp1714229

Dantas, L. F., Fleck, J. L., Cyrino Oliveira, F. L., & Hamacher, S. (2018). No-shows in appointment scheduling: A systematic literature review. Health Policy, 122(4), 412-421. https://doi.org/10.1016/j.healthpol.2018.02.002

Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447-453. https://doi.org/10.1126/science.aax2342

Rajkomar, A., Dean, J., & Kohane, I. (2019). Machine learning in medicine. New England Journal of Medicine, 380(14), 1347-1358. https://doi.org/10.1056/NEJMra1814259

What the HIM 400 Module 5 instructions ask for

The HIM 400 data mining assignment usually asks you to explain data mining methods and apply one to a health care question, often prediction or pattern discovery. Plan four to five pages in APA 7 with three or more scholarly sources and at least one table. State the business question first, name the data mining task, describe the data and predictors, and explain how the model was built and tested, ideally on data from a later period. Report results in terms a manager can use, such as positive predictive value and sensitivity at a chosen threshold, then examine how the model performs and is used for different patient groups before recommending an action and a monitoring plan.

How this HIM 400 Module 5 data mining short paper example is built

Cold Brook Health loses about 12% of primary care appointments to no-shows. Guided by Dantas and colleagues, the team builds 14 predictors from 148,300 appointments, trains on 2024 and tests on 2025 as Rajkomar and colleagues advise and chooses logistic regression over a slightly stronger tree model for interpretability. A table explains an area under the curve of 0.74, a 31% positive predictive value and 38% sensitivity at a 15% threshold. Obermeyer and colleagues frame a subgroup check showing Medicaid and distant patients flagged far more often, and Char and colleagues support using flags for rides and video visits instead of overbooking.

Where the HIM 400 Module 5 rubric puts the points

Data mining papers in HIM 400 are usually graded on correct explanation of methods, a clear link between the business question and the task, sound validation, accurate interpretation of performance measures, attention to bias and ethics and APA 7 mechanics. The strongest papers explain why an interpretable model might be preferred over a slightly more accurate one and translate statistics into plain meaning. Graders reward writers who separate a model's accuracy from the fairness of the decision it supports, and who plan monitoring after deployment. Noting that success can make a model look worse over time shows a thoughtful grasp of how predictions interact with the actions they trigger.

HIM 400 Module 5 help: the mistakes that cost points

Data mining papers lose points when they report accuracy without a baseline, test a model on its training data, describe algorithms in jargon without explaining what they do or ignore how predictions will be used. Another frequent gap is treating bias as a question about the algorithm alone. If your course provides a data set or asks for a specific method, such as clustering or association rules, send it with the prompt so the sample applies that method to your data. Say which software you use. A custom paper can study readmissions, length of stay or claim denials while following the same path from question to model, evaluation, subgroup check, action and monitoring.

Get HIM 400 Module 5 written to your instructions

Share the HIM 400 Module 5 prompt and any data set or method your instructor requires. The paper will frame the business question, describe the data and model, test it on unseen data, explain results in plain terms, check subgroups and recommend a fair use, delivered within 24 to 48 hours with the first sample free. The paper above is an original model document written by our desk, not a submitted student paper and not an official Southern New Hampshire University document.

More HIM 400 papers and related BS Health Information Management samples

HIM 400 Module 5 questions, answered

Where can I find a free HIM 400 Module 5 Data Mining Short Paper sample?

This page carries the complete HIM 400 Module 5 paper: a no-show prediction model built, tested on a later year, checked by subgroup and used for outreach.

What is the difference between classification and clustering?

Classification assigns records to known categories using labeled examples; clustering groups similar records without predefined labels.

Why test a model on a later time period?

Patterns change over time, so testing on later data shows how the model will perform once it is actually used.

What do sensitivity and positive predictive value mean for a no-show model?

Sensitivity is the share of all no-shows the model flags; positive predictive value tells you how many of the flagged visits really go unattended.

Can an accurate model still be unfair?

Yes. If its predictions lead to actions that burden groups already facing barriers, the use can be unfair even when accuracy is similar across groups.