The Numbers Are Different. But Does the Difference Really Matter?
How to tell whether a difference in your data is meaningful — or just normal variation.

Imagine your organisation has just completed a customer service improvement program. Before the program, the average customer satisfaction score was 78. Three months later, it is 81.
On the surface, that's good news. A three-point increase certainly looks like an improvement. The program team may be pleased, management may want to know what worked, and someone may already be preparing the green arrow for the dashboard. But before we celebrate, there is another question worth asking: is 81 genuinely different from 78, or are we simply seeing the kind of variation that naturally occurs whenever we measure something?
It's a question that comes up constantly in business. Sales increase after a campaign, employees score higher after training, one branch appears to outperform another, or customers using a new process report better satisfaction. In each case, we can see that the numbers are different. The harder question is whether that difference gives us enough evidence to conclude that something has really changed.
Numbers naturally move around
Suppose we hadn't introduced the customer service programme at all. Would we expect next month's satisfaction score to be exactly 78 again? Probably not. It might be 77, 79 or 80 simply because the customers surveyed were different, their experiences varied, or circumstances changed slightly.
This kind of natural movement is everywhere in data. Flip a fair coin ten times and you won't necessarily get exactly five heads. Survey one group of customers this month and another group next month and you probably won't get identical results either. Even when nothing fundamental has changed, the numbers we observe will move around.
So when satisfaction rises from 78 to 81, we have at least two plausible explanations. The programme may genuinely have made a difference, or the increase may simply be part of the variation we could have seen anyway.

This is where statistical testing becomes useful. Not because we need statistics to tell us that 81 is bigger than 78, but because we want to understand how much confidence we should place in that difference.
What are we actually trying to find out?
Terms such as hypothesis testing can make this sound much more intimidating than it really is. At its heart, the question is quite practical: is the difference we're seeing convincing enough that it would be difficult to explain as normal variation alone?
In our customer satisfaction example, we might begin by assuming that the program hasn't actually changed satisfaction. Statisticians call this the null hypothesis. From there, we look at the evidence and ask how unusual a result like ours would be if that starting assumption were true.
This is the thinking behind statistical tests and p-values. We aren't trying to prove that 81 is different from 78; that part is obvious. We're trying to judge whether the difference is strong enough, given the amount of variation in the data and the evidence available, to make us question the idea that nothing has really changed.
That is a much more useful way to think about hypothesis testing than simply treating it as a procedure for producing a number.
So what does a p-value actually tell us?
Suppose our analysis produces a p-value of 0.03. You may have heard this interpreted as, “There is only a 3% chance that the result happened by accident.” It's an understandable shortcut, but it isn't quite what the p-value tells us.
A better way to think about it is this: if there really were no underlying difference, how unusual would results like ours be? A small p-value tells us that what we observed would be relatively unusual under that assumption, which gives us evidence to question the idea that nothing has changed.
What it doesn't tell us is equally important. A p-value of 0.03 does not mean there is a 97% probability that our program worked. It doesn't tell us what caused the improvement, and it doesn't tell us whether the improvement is large enough to matter to the business.
That last point is particularly important because statistical significance and business importance are not the same thing.
A real difference isn't necessarily an important one
Imagine a large company introduces a new system that reduces the average time needed to complete a task from 10.0 minutes to 9.9 minutes. With enough observations, that tiny difference might be statistically significant. But whether saving six seconds makes the new system worthwhile depends entirely on the context.
If the task is performed millions of times, six seconds could add up to a substantial saving. If it happens only occasionally, the difference might be practically irrelevant. The statistical result hasn't changed; what has changed is what that result means to the business.
Now consider the opposite situation. A small pilot programme improves productivity by 12%, but only 15 employees participated. A 12% improvement could be very important to the organisation, yet with such a small group we may not have enough evidence to be confident that the same improvement would appear across the wider workforce.
These examples illustrate two questions that are related but should not be confused: is the difference statistically convincing, and is the difference large enough to matter?

Statistics can help us assess the evidence behind the first question. The second requires business context, judgement and an understanding of what the organisation is actually trying to achieve.
The amount of evidence changes the story
This is also why sample size matters. Suppose Branch A has a customer satisfaction score of 78 while Branch B scores 84. A six-point difference certainly catches the eye, but if each branch surveyed only ten customers, we'd probably be cautious about reading too much into it. If the same difference appeared across thousands of customer responses, we'd have considerably more evidence behind what we're seeing.
Larger samples generally allow us to estimate what is happening more precisely and make smaller differences easier to detect. That sounds entirely positive, but it creates an interesting issue of its own: with very large datasets, even tiny differences can become statistically significant.
This brings us back to the earlier example of reducing task time from 10.0 minutes to 9.9 minutes. With enough observations, we might become very confident that the difference is real while still being unsure whether anybody should care about it.
So more data doesn't eliminate the need for judgement. In some situations, it makes the distinction between statistical significance and practical importance even more important.
Let the business question lead the analysis
Once people become familiar with hypothesis testing, there is a temptation to jump quickly to the technique: should I use a t-test? Do I need ANOVA? Which test should I run?
Those are legitimate questions, but I wouldn't begin there. I'd first want to understand what we're actually comparing. Are these the same employees before and after training, or two different groups of employees? Are we comparing two branches or several? Are we interested in average scores, percentages or some other measure?
Those differences matter because they determine what kind of analysis is appropriate. There isn't one statistical test that we simply apply whenever two numbers look different.
For someone working with data in a business setting, I think the more useful principle is much simpler: let the question determine the technique, rather than allowing the technique to determine the question.
Once we are clear about what we're trying to learn, choosing the appropriate statistical approach becomes much easier — whether we make that choice ourselves, with the help of an analyst, or increasingly with the help of AI.
AI makes the calculation easier, not the judgement
AI changes this conversation in an interesting way. Give an AI tool a dataset and ask whether two groups are significantly different, and it may be able to recommend a statistical test, perform the calculation, report the p-value and explain the result within seconds.
That's genuinely useful. Statistical techniques that once required specialised knowledge are becoming far more accessible to ordinary business users.
But easier calculation doesn't necessarily mean easier analysis. If we don't understand the question we're asking, we can obtain a perfectly calculated answer to the wrong question.
We still need to consider whether the right groups were compared, whether the sample is appropriate, whether the measure itself makes sense and whether something else might have changed at the same time. And even if the analysis tells us that a difference is statistically significant, someone still has to decide whether that difference is important enough to act on.
This is a theme I keep returning to as AI becomes more capable. The value of human judgement doesn't disappear when the calculation becomes easier. It shifts towards knowing what to ask, what to question and what the result actually means.
Back to our customer satisfaction score
Let's return to where we started. Customer satisfaction has increased from 78 to 81, and we can now see why the three-point improvement is worth noticing but not necessarily worth celebrating immediately.
We would want to understand how much the measure normally fluctuates, how many customers were surveyed and whether the customers measured before and after the programme were reasonably comparable. Statistical analysis may then help us judge whether the three-point increase stands out sufficiently from normal variation to give us confidence that something has changed.
But even if it does, there is still another question to answer: is a three-point increase meaningful enough to matter to the organisation?
Perhaps three points represents a substantial improvement in an area where satisfaction has barely moved for years. Perhaps it takes the organisation above an important service benchmark. Or perhaps customers would barely notice the difference and the programme cost far more than the improvement was worth.
The statistical analysis cannot make that judgement for us.

This doesn't mean we should become suspicious of every improvement or reluctant to act until we have perfect evidence. It simply means understanding what the numbers can tell us, what they cannot tell us, and how much confidence we should place in the conclusion.
Perhaps “Is it significant?” isn't the final question
When someone asks whether a result is statistically significant, it's tempting to think statistics will give us a simple yes-or-no verdict. But significance is really only one part of a more useful analytical conversation.
First, is there a difference? We can often see that directly in the data. Then, how convincing is the evidence that the difference reflects something beyond normal variation? That's where statistical inference can help. Finally, is the difference large enough to matter? That requires judgement.
The three questions belong together because any one of them on its own can mislead us. A difference may look impressive but rest on very little evidence. Another may be statistically convincing but far too small to justify changing anything. And occasionally a modest-looking improvement may have enormous value when repeated thousands or millions of times.
Statistics can help us understand how much confidence to place in a difference, but they cannot decide what that difference is worth to us. That requires context, judgement and an understanding of the decision we're trying to make.
So the next time two numbers look different, don't rush to celebrate the result. Look at the variation around them and how much evidence sits behind the difference, then ask what that difference actually means for the business.
Numbers being different is an observation. Knowing whether the difference matters is analysis.
New to Statistical Significance? A Quick Reference
If some of the statistical terms in this article are unfamiliar, here’s the plain-English version.
What is variation?
Variation simply means that numbers naturally move around. Customer satisfaction might be 78 this month, 80 next month and 77 the month after, even when nothing important has changed.
That natural movement is one reason we shouldn’t assume every difference we see represents a real change.
What is a hypothesis?
A hypothesis is an idea or claim that we want to examine using data. In statistical testing, we usually begin with a starting assumption and ask whether the evidence is strong enough for us to question it.
What is the null hypothesis?
The null hypothesis is usually our starting assumption that there is no underlying difference or effect.
For example, if customer satisfaction rises from 78 to 81, we might initially assume that satisfaction hasn’t fundamentally changed and that the difference could simply reflect normal variation. We then examine whether the evidence gives us enough reason to question that assumption.
What is a p-value?
A p-value helps us judge how unusual our observed result would be if the null hypothesis were true.
A small p-value suggests that the result would be relatively unusual under that assumption, giving us evidence to question the null hypothesis. Importantly, a p-value does not tell us the probability that our hypothesis is correct, nor does it tell us whether the result matters to the business.
What does statistically significant mean?
A result is described as statistically significant when the evidence meets a chosen threshold for questioning the null hypothesis. You’ll often see 0.05 used as that threshold, although it is a convention rather than a magical dividing line between “true” and “false.” Statistical significance tells us something about the strength of the evidence. It does not tell us whether the difference is large or important.
What is practical significance?
Practical significance asks a different question: is the difference large enough to matter in the real world?
Saving six seconds on a task may be statistically significant, but whether it matters depends on how often the task is performed, what those six seconds are worth and what it cost to achieve the improvement.
What is sample size?
Sample size is simply the number of observations included in an analysis — for example, the number of customers surveyed or employees studied. Larger samples generally give us more precise estimates and make smaller differences easier to detect. This is also why a statistically significant result from a very large dataset may still represent a very small practical difference.































Comments