The Most Misunderstood Stage of Analytics
- Aug 3
- 8 min read
Why data cleaning isn't the stage before analytics. It is where analytics begins.

Data Cleaning Has Always Been the Overlooked Step-Sibling of Analytics
If you ask people which part of analytics they enjoy the least, data cleaning will probably appear somewhere near the top of the list. It is repetitive, time-consuming and often feels like the obstacle standing between us and the work we actually want to do. We look forward to building dashboards, discovering insights and helping organisations make better decisions. Data cleaning, on the other hand, is usually treated as the necessary chore that has to be completed before any of that can begin.
Today, AI promises something even more attractive. It can identify duplicates, standardise formats, detect anomalies and automate many of the repetitive tasks that have traditionally consumed hours of an analyst's time. It is no surprise that many organisations are asking whether data cleaning should simply become another task we hand over to technology. After all, if software can perform the work faster and more consistently, surely that allows us to focus on more valuable activities.
I certainly welcome that progress. In fact, I believe much of the repetitive work involved in preparing data should eventually become automated. However, I also think we may have misunderstood the role that data preparation plays within analytics itself.
We often think of data preparation as a single activity. It isn't. Some parts are about understanding the business behind the data. Others are simply about repeating work that has already been done. Confusing the two has caused us to overlook where analytics really begins.
The First Conversation with Your Data
Whenever I receive a new dataset, my first instinct is not to build charts or dashboards. My first objective is much simpler. I want to understand what the data is trying to tell me.
That understanding rarely comes from looking at summary statistics alone. It develops through hundreds of small observations that gradually build a picture of how the business really works. Why doesn't this report reconcile with another? Why are there three different spellings for the same customer? Why are certain fields consistently missing? Why is one department recording information differently from everyone else?
At first glance, these appear to be data quality problems. More often than we realise, they are business problems revealing themselves through the data. A duplicate customer record may not simply be careless data entry. It could be the result of different departments maintaining separate customer lists for years without agreeing on who owns the customer relationship. Missing values may not indicate poor data quality at all. They may reveal that one business unit follows a different process from everyone else. Inconsistent product categories may reflect business definitions that have quietly evolved over time without anyone noticing.
The longer I spend with a new dataset, the more I realise that messy data rarely becomes messy by accident. Every inconsistency has a history. Every anomaly has a reason. Every exception has a story waiting to be understood. Long before I begin thinking about dashboards or statistical tests, I am already learning how the organisation works.
By the time I create my first visualisation, I usually have a much clearer picture of the business behind the numbers. I understand where business rules are well established, where different teams interpret information differently and where operational processes have gradually drifted apart. None of those discoveries appear in a dashboard. They are uncovered while preparing the data.
Perhaps that is why I have never been comfortable describing data preparation as merely housekeeping. Housekeeping suggests that our job is simply to tidy the data before the interesting work begins. My experience has been very different. Preparing data is often the first meaningful conversation we have with the business through its data, and it is during that conversation that some of the most valuable insights first begin to emerge.

Data Cleaning Is Already Analytics
We often describe analytics as though it unfolds in a neat sequence of stages. First we clean the data. Then we analyse it. Finally, we communicate our findings. It sounds logical enough, but I have gradually come to realise that this isn't how analytics actually works.
The moment I begin asking why two reports don't reconcile, whether two records really refer to the same customer, or why one business unit records information differently from another, analytical thinking has already begun. I may not be building visualisations or testing hypotheses yet, but I am already investigating patterns, questioning assumptions and trying to understand the business behind the numbers.
That distinction matters because it changes how we think about data preparation. Instead of viewing it as the hurdle we must overcome before the interesting work begins, we begin to recognise it as the stage where many of the most important questions first emerge. Every answer improves our understanding of the data, and that understanding influences every analysis that follows.
Perhaps this is why experienced analysts approach a new dataset differently from beginners. Beginners often ask, "How do I clean this data?" Experienced analysts ask a different question. "Why is the data like this in the first place?" One question focuses on fixing the data. The other focuses on understanding the business that created it.
That shift in perspective changes everything. Once we stop seeing inconsistencies as obstacles and start seeing them as clues, data preparation becomes much more than a technical exercise. It becomes an investigation. Every duplicate record, unexpected value and conflicting definition is another piece of evidence helping us understand how the organisation really works.
Every Messy Dataset Has Something to Teach Us
One of the reasons I still enjoy working with new datasets is that they nearly always surprise me. The surprises rarely come from sophisticated statistical analysis. More often, they appear while I am still trying to understand the data itself.
A customer might appear several times under slightly different names. At first, it looks like poor data quality. A little investigation later, it becomes clear that different business units have been maintaining their own customer lists for years. What looked like duplicate records was actually telling the story of how the organisation had grown.
The same thing happens with missing values. It is easy to assume that someone simply forgot to enter the information, but missing data often turns out to be far more interesting than that. One department may have stopped collecting a particular field years ago because their process changed. Another may never have needed that information in the first place. Before we have even built our first chart, we are already learning something about how the organisation operates.
Even inconsistent product categories can become valuable clues. What appears to be poor discipline may actually reflect different interpretations of the same business concept. Marketing may classify products one way, Operations another and Finance a third. Cleaning the data forces those conversations to happen, and those conversations often improve the business just as much as they improve the data.
This is why I have become increasingly convinced that data preparation deserves far more credit than it usually receives. It doesn't simply improve the quality of our data. It improves the quality of our understanding. By the time we begin analysing trends or building dashboards, we are no longer looking at anonymous rows and columns. We are looking at a business whose processes, assumptions and definitions have gradually become much clearer.

The Real Problem Isn't Data Cleaning
If data preparation creates so much value, why do so many analysts dislike it? I don't think the answer is data cleaning itself. I think the answer is repetition.
There is a world of difference between working with a dataset for the first time and working with it for the twelfth month in a row. The first encounter is an opportunity to explore, question and understand. Each inconsistency teaches us something about the business, and every answer helps us build confidence in the data we are about to analyse.
The second, third and fourth time are very different. The questions have already been answered. The business rules have already been clarified. We know why certain fields need to be transformed, which records should be excluded and how different datasets relate to one another. At that point, repeating exactly the same preparation process month after month creates very little additional understanding. It simply repeats what we have already learnt.
Perhaps this is where we have unintentionally confused two very different kinds of work. One creates understanding. The other simply reproduces it. Both are called data cleaning, yet they contribute value in completely different ways.
That distinction matters because it changes what we should be trying to improve. The objective shouldn't be to eliminate data preparation altogether. It should be to preserve the learning while reducing the repetition.
Don't Skip It. Shorten It.
This is why I find today's conversations about AI both exciting and slightly worrying.
Exciting, because AI is becoming remarkably good at many of the repetitive tasks involved in preparing data. Standardising formats, combining datasets, identifying duplicates and applying established business rules are exactly the kinds of activities that technology should help us perform more efficiently.
Worrying, because we sometimes talk about automating data preparation as though every part of it creates value in the same way.
It doesn't.
The first encounter with a dataset is fundamentally different from every encounter that follows. The first is where we begin understanding the business. It is where we discover assumptions that have quietly drifted apart, definitions that are no longer shared and processes that no longer work as intended. Those discoveries are difficult to automate because they require curiosity, judgment and conversations that often extend far beyond the data itself.
Once those lessons have been learnt, however, the role of technology changes completely. AI no longer needs to help us understand the data. Instead, it helps us preserve that understanding by repeating the preparation process consistently, accurately and at a speed that would be difficult to achieve manually.
That, to me, is where AI creates its greatest value.
Not by replacing discovery.
By removing repetition.

A Different Way to Think About Data Preparation
Perhaps data cleaning has been misunderstood for so long because we have judged it by the effort it requires rather than the understanding it creates.
We celebrate dashboards because they make insights visible. We celebrate statistical models because they help us test ideas. We celebrate AI because it promises to make us more productive. Yet long before any of those things happen, something equally important has already begun.
We are learning.
We are learning how the organisation works, where its processes are strong and where they quietly begin to break down. We are learning which definitions are shared and which exist only within individual teams. We are learning which numbers deserve immediate confidence and which require another conversation before we can trust them.
Those lessons rarely appear in a dashboard.
They are earned while preparing the data.
Once that understanding exists, I want technology to do everything it can to remove unnecessary repetition. I want reports to refresh automatically. I want routine transformations to happen consistently. I want analysts to spend less time rebuilding yesterday's preparation process and more time extending yesterday's understanding.
Perhaps that is the balance we should be aiming for. The opportunity isn't to skip data preparation. It is to shorten the repetitive parts so that we have more time for the parts that only people can do.
Maybe data cleaning has never been the stage before analytics.
Maybe it has been the beginning of analytics all along.































Comments