| Course | ACC 430 Data Analytics for Financial Professionals |
|---|---|
| Module | Module 2 |
| Paper type | undergraduate data preparation and cleaning assignment |
| Length | About 1,000 words, 6 pages |
| Format | APA 7 student paper |
| School | Southern New Hampshire University |
| Program | BS Accounting |
| Updated | October 2026 |
Free sample paper for ACC 430 Module 2
Before the Margin Analysis: Cleaning and Documenting 1.26 Million Invoice Lines
[Student Name]
Southern New Hampshire University
ACC 430: Data Analytics for Financial Professionals
Module Two Assignment
[Instructor Name]
[Date]
The organization, setting and figures below are a composite written as a model document. No real employer, client, colleague or patient is described.
Before the Margin Analysis: Cleaning and Documenting 1.26 Million Invoice Lines
Introduction
The margin analysis requested by the CFO needs two years of invoice-line detail: what was sold, to whom, at what price and at what cost. That detail lives in the distributor's ERP system, with customer and product information in separate master files. Before any analysis, the data must be complete, consistent and reconciled to the company's books. Richardson et al. (2021) place this mastering of data at the heart of the analytics cycle, because errors introduced here pass silently into every chart and conclusion. This assignment documents the extraction, profiling, cleaning and reconciliation.
Extraction
The IT manager ran a query for all invoice lines dated within the last 24 months, returning 1,262,400 rows with invoice number, line number, date, customer ID, product ID, quantity, unit of measure, unit price, extended price and standard cost. The customer master supplied segment codes, such as school, health care, office and restaurant, and the product master supplied categories. The extract was saved untouched, and all cleaning was performed on a copy through a documented query so the steps can be rerun. A hash value of the original file was recorded, so anyone can confirm later that the starting point was not changed.
Profiling and Problems Found
Table 1. Data Quality Issues and Rules
| Issue | Rows affected | Cause | Rule applied |
|---|---|---|---|
| Duplicate invoice lines | 39,100 | Invoices reloaded during ERP migration | Keep one row per invoice and line number; remove exact duplicates |
| Inconsistent customer names | 140 variants of one district; similar for 31 other customers | Free-text entry before migration | Group by customer ID, not name; correct 12 IDs split across two records |
| Returns mixed with sales | 27,800 | Normal returns entered as negative quantities | Keep; flag as returns for separate analysis |
| Missing segment code | 208 customers, 2.4% of revenue | New accounts not coded | Code from billing address and sales rep notes; approved by sales manager |
| Unit of measure mismatch | 4,600 | Cases entered as eaches on six products | Convert to cases using product master pack size |
The duplicates were the largest problem. During the ERP migration 14 months ago, two weeks of invoices were loaded twice, and the duplicate lines inflate revenue in those weeks. Removing exact duplicates on invoice and line number eliminated them, and a check of those two weeks against the bank's deposit totals confirmed the corrected figures. The unit of measure problem was smaller but would have distorted price analysis badly: a case of 1,000 nitrile gloves recorded as one glove makes the unit price look a thousand times too high.
Decisions That Needed Business Approval
Two fixes involved judgment rather than mechanics. Assigning segments to 208 uncoded customers required knowing who they were, so the analyst proposed codes from billing addresses and the sales manager approved them. Deciding whether to keep returns was an analytical choice: removing them would make margins look better than they are, so they were kept and flagged. Brown-Liburd et al. (2015) caution that large data sets tempt analysts to filter aggressively to reduce noise, and that such filtering can introduce bias into the judgments made from the data. Recording who approved each judgment call addresses that risk.
Reconciliation
Table 2. Revenue Reconciliation to the General Ledger (millions of dollars)
| Item | Amount |
|---|---|
| Revenue in raw extract | 238.1 |
| Less duplicate lines removed | (3.4) |
| Revenue after cleaning | 234.7 |
| General ledger revenue, 24 months | 234.66 |
| Difference | 0.04 |
The remaining $41,000 difference traces to freight charges billed on invoices but recorded in a separate ledger account, which is acceptable for a product margin analysis and documented in the log so that no later reader mistakes it for an error. Cao et al. (2015) note that the value of large-scale analytics in financial work depends on whether the data can be reconciled to the financial records the analysis is meant to explain, and this reconciliation provides that link.
Validating the Joins
Joining invoice lines to the customer and product masters is where data most often disappear silently. A line whose customer ID does not exist in the master would be dropped by an inner join and its revenue would vanish from the analysis. To prevent that, the joins were run as left joins and unmatched lines counted: 312 lines referenced six customer IDs deleted from the master after accounts closed, and 95 lines referenced discontinued product IDs. Both groups were restored from archived master files rather than dropped, and the log records that step. After the joins, the line count and revenue total were rechecked against Table 2 and matched.
What Cleaning Changed in the Story
Cleaning was not cosmetic. Before duplicates were removed, the two weeks loaded twice showed a spike in revenue that would have looked like a successful promotion. Before units were standardized, gloves appeared to sell at absurd prices, which would have pulled average margins for the health care segment up by almost two points. And before missing segments were coded, about $5.6 million of revenue sat in an unknown category, much of it small restaurants, the very segment the CFO suspects is unprofitable. Each fix moved the eventual answer closer to the truth, which is why the preparation deserves as much care as the analysis.
The Data Quality Log
Every rule is recorded in a data quality log with the issue, the rule, the number of rows affected, the effect on revenue and cost, the date and the approver, in the order the rules were applied. The log, the original extract and the cleaning query together allow anyone to reproduce the cleaned data set or challenge a choice.
Conclusion
The cleaned data set contains 1,223,300 invoice lines whose revenue reconciles to the general ledger within $41,000, with duplicates removed, customers grouped correctly, returns retained and flagged, segments assigned with approval and units standardized. The margin analysis can now proceed on data that match the books. When the next month's invoices are added, the same query will apply the same rules, and the reconciliation will be rerun before any figure is reported, so the cleaning becomes a routine control rather than a one-time project that slowly goes stale.
References
Brown-Liburd, H., Issa, H., & Lombardi, D. (2015). Behavioral implications of big data's impact on audit judgment and decision making and future research directions. Accounting Horizons, 29(2), 451-468. https://doi.org/10.2308/acch-51023
Cao, M., Chychyla, R., & Stewart, T. (2015). Big data analytics in financial statement audits. Accounting Horizons, 29(2), 423-429. https://doi.org/10.2308/acch-51068
Richardson, V. J., Teeter, R. A., & Terrell, K. L. (2021). Data analytics for accounting (2nd ed.). McGraw Hill.
What the ACC 430 Module 2 instructions ask for
The Module Two assignment in ACC 430 usually asks you to prepare a data set for analysis: extract it from source systems, assess its quality, clean and transform it and load it into an analysis tool. Expect to profile the data for problems such as duplicates, missing values, inconsistent formats and outliers, decide how to handle each and document every decision. Many versions require joining tables, such as transactions with customer and product masters, and reconciling the result to a control total such as general ledger revenue. Explain each cleaning rule and its effect on record counts and totals, since undocumented changes make results impossible to verify. A log of transformations is usually expected.
How this ACC 430 Module 2 data preparation assignment example is built
The sample extracts 1,262,400 invoice lines for 24 months from the ERP and joins them to the customer and product master files. Profiling finds 39,100 duplicate lines created when invoices were reloaded during an ERP migration, 140 spellings of one school district, 27,800 return lines with negative quantities mixed with sales, 208 customers missing a segment code and 4,600 lines where cases were recorded as single units. Each is fixed with a written rule, and the effect of every rule on rows and dollars is recorded. Revenue in the extract exceeds the ledger by $3.4 million before cleaning and matches within $41,000 after. A data quality log records every rule, the rows it affected and who approved it, so the work can be reviewed line by line.
Where the ACC 430 Module 2 rubric puts the points
Rubrics for the ACC 430 data preparation assignment typically score profiling and identification of data quality issues, the appropriateness of cleaning rules, documentation, joins, reconciliation to control totals and the explanation. Top papers quantify each issue, explain why each rule is reasonable for the analysis to come, show record counts and totals before and after, and reconcile to an independent source such as the general ledger. Graders reward recognizing which fixes are judgment calls that someone in the business should approve, and a log that lets the reader trace each change. Common deductions include deleting rows without explanation, treating returns as errors, ignoring units of measure and skipping reconciliation.
ACC 430 Module 2 help: the mistakes that cost points
Data preparation papers most often go wrong by cleaning silently, so that a reader cannot tell what changed, or by applying rules that distort the analysis, such as deleting all returns, which would overstate margins. Another frequent gap is reconciling only record counts and not dollar totals. If your data come from a different source, such as payroll, purchasing or a public data set, the same steps of profile, rule, document and reconcile apply. Keep the original extract untouched and apply every change through a script or query you can rerun; that turns cleaning from a one-time effort into a repeatable process the next analyst can trust.
Get ACC 430 Module 2 written to your instructions
Send the ACC 430 Module 2 data description and instructions. The paper will describe extraction, identify each data quality problem, apply documented cleaning rules, reconcile to a control total and present a data quality log. No fee applies to a first request, and delivery takes about two days. The paper above is an original model document written by our desk, not a submitted student paper and not an official Southern New Hampshire University document.
More ACC 430 papers and related BS Accounting samples
- ACC 430 Module 1 Discussion: Turning "Margins Are Down" Into a Question
- OL 320 Module 7 Final Opportunity Analysis Plan
- ACC 330 Module 2 Filing Status and Dependents Assignment: Four People and One Return
- FIN 320 Module 3 Project Milestone One
- PSY 365 Module 8 Final Project Motivation Plan
ACC 430 Module 2 questions, answered
Where can I find a free ACC 430 Module 2 data preparation sample?
This page includes a full ACC 430 Module 2 assignment cleaning 1.26 million invoice lines with documented rules and a ledger reconciliation.
What is data profiling?
Examining a data set's structure, values and patterns to find problems such as duplicates, missing values, inconsistent formats and outliers before analysis.
Why reconcile analytics data to the general ledger?
To confirm the extract is complete and accurate. If revenue in the data does not match the ledger, any analysis of margins or customers will be off.
Should returns be removed from sales data?
Usually not. Returns are real transactions that reduce revenue and margin; they should be classified and kept, not deleted.
What is a data quality log?
A record of each data issue found, the rule applied to fix it, the records affected and who approved the change, allowing the work to be reviewed and repeated.