Your Data Shows a Pattern. But What Does It Really Mean?
- 1 day ago
- 8 min read
How to separate what the data shows from what we think it means

Imagine you're looking at your organisation's training data and you notice something interesting: employees who attend more training also tend to receive higher performance ratings. That's encouraging. Perhaps even encouraging enough to put on a management slide:
Employees who attend more training perform better.
But is that actually what the data told us?
Perhaps the training improved their performance, but there are other possibilities. Better performers may be more likely to be selected for training, ambitious employees may actively seek out development opportunities, or supportive managers may provide their teams with both more training and better conditions to perform.
The relationship may be real. What it means is a different question, and I think that distinction is one of the most important, and sometimes overlooked, parts of working with data.
What did the data show, and what do we think explains it?
When we analyse data, two very different questions can easily become mixed together:
What does the data show? What do I think explains it?
The first is about evidence. The second is about interpretation.
To see why that distinction matters, imagine a café owner notices that on hotter days, ice-cream sales tend to increase. That relationship makes intuitive sense. Temperature rises, people get hot, and more of them buy ice cream.
Now suppose the owner also discovers that cold-drink sales and ice-cream sales move very closely together. Does buying a cold drink make someone want an ice cream?
Probably not. A much more plausible explanation is sitting quietly behind both numbers: temperature. On hotter days, people buy more cold drinks and more ice cream, so the two sets of sales may show a very strong relationship even though neither is causing the other.

This is where correlation helps
Businesses look for relationships all the time, even if we don't use the word correlation. We might want to know whether customer satisfaction changes with waiting time, whether employee engagement moves with absenteeism, whether advertising expenditure relates to sales, or whether students with better attendance tend to achieve better results.
Correlation gives us a way to describe whether two numeric variables tend to move together, in which direction and how strongly.
A positive relationship means that as one tends to increase, the other tends to increase as well. A negative relationship means that as one increases, the other tends to decrease. Sometimes there may be very little consistent relationship between them at all.
What I find more important than memorising those labels, however, is the language we use afterwards.
We might say:
Higher training attendance is associated with higher performance ratings.
That is quite different from saying:
More training causes higher performance.
The first describes what we found. The second offers an explanation for why it happened, and the data may not yet give us enough evidence to make that leap.
Even a very strong relationship doesn't remove this problem. Go back to our café. On particularly hot days, cold-drink sales and ice-cream sales might move almost perfectly together. Making that relationship stronger still doesn't turn cold drinks into the cause of ice-cream sales. Temperature hasn't disappeared simply because the correlation became impressive.
So when I see a strong relationship, I'm interested in how strong it is, but I'm equally interested in another question:
What else could explain this?
That's the point where we move from calculating a relationship to thinking critically about it.
Where does regression fit into this?
Let's return to our training example.
Suppose we've established that employees who attend more training tend to receive higher performance ratings. That's useful, but we already know there could be other explanations. Job role might matter. Tenure might matter. Previous performance and manager support could matter too.
If we look only at training and performance, we cannot see what those other factors might be doing. This is where regression becomes useful.
Rather than looking only at whether two variables move together, regression allows us to examine an outcome while considering several factors at the same time. Instead of asking only whether training attendance and performance are related, we can ask a richer question:
How is training attendance related to performance when we also consider factors such as role, tenure and previous performance?

Regression allows us to investigate the relationship more carefully, but there's an important caution here: it doesn't magically turn a relationship into causation.
It can help us account for factors we know about and have measured. It cannot automatically account for something we never measured, something we don't know about, or a poor assumption about how the business actually works. In other words, regression gives us a better way to investigate the question. It doesn't give us permission to stop thinking.
So when would I use correlation, and when would I use regression?
I wouldn't begin by asking which technique is better. I'd begin with what you're actually trying to learn.
If you simply want to know whether two numeric variables tend to move together, correlation may be exactly what you need. Perhaps you want to explore whether training attendance and performance ratings are related, or whether waiting time tends to move with customer satisfaction.
But suppose the conversation changes. Management now wants to know whether the relationship between training and performance is still present when we consider tenure, job role and previous performance. The business question has become richer, and regression becomes much more useful because it allows those factors to be examined together.
Neither technique is automatically more sophisticated or more correct. They are answering different questions. Sometimes correlation is a useful place to start because it helps us discover relationships worth exploring. Regression can then help us investigate some of those relationships more deeply. But neither, by itself, settles the question of why the relationship exists.
Can I ever say that X caused Y?
Sometimes, yes, but usually we need stronger evidence than simply observing that X and Y move together.
Think again about our employees. If we want to know whether training actually improves performance, we may want to examine performance before and after training, compare employees who received the training with an appropriate comparison group, or in some situations design an experiment that helps us rule out alternative explanations.
We would also want to understand how employees were selected for training in the first place, because that can matter enormously.
Imagine managers tend to nominate their most promising employees for development programmes. Six months later, those employees outperform their colleagues. It is certainly possible that the programme helped, but those employees were already considered promising enough to be selected.
The challenge is therefore not simply to find a statistical technique that produces a more convincing number. We need to separate the possible effect of the programme from differences that may have existed before it.
Notice what has happened here. We haven't improved the analysis simply by choosing a more complicated technique. We've improved the question.
The calculation is becoming the easy part
There was a time when running statistical analyses required considerably more technical effort. Today, Excel, analytics platforms, statistical software and increasingly AI can perform many of these calculations very quickly.
AI can calculate a correlation. It can explain regression output and probably give you a reasonable description of what the coefficients mean.
That's useful, but I think it makes human judgement more important rather than less.
The difficult question increasingly isn't whether we can produce the calculation. It's whether we understand what that calculation actually allows us to say.
A correlation of 0.8 may look impressive, but I'd still want to know what was measured, how it was measured, who was included, over what period and what else might explain the relationship. A statistically significant regression result may look convincing, but I'd still want to know whether the relationship is large enough to matter in the real world.
The software can produce the number. Someone still has to make sense of it.
When you find a relationship, don't rush past it
One of the easiest mistakes in analytics is to move too quickly from finding a pattern to explaining it, and from explaining it to recommending what the organisation should do. I prefer to put a deliberate pause between those steps.

When you find an interesting relationship, first ask what the data actually shows. Then consider what might explain it. Could another factor be influencing both variables? Could the direction work the other way around? Does the relationship make sense given what you already know about the business?
Only then should we decide what additional analysis or evidence we need.
Sometimes that will mean regression. Sometimes it will mean collecting more data, comparing different groups or looking at what happened before and after an intervention. And sometimes the most useful conclusion will simply be:
“We've found a relationship worth investigating further.”
That may sound less dramatic than claiming you've discovered the cause, but careful conclusions tend to age better than exciting ones.
Perhaps the real skill is knowing how far the evidence lets you go
When people learn statistics, much of the attention naturally goes to techniques: correlation, regression, t-tests, ANOVA. Knowing how and when to use them matters, but there's another skill sitting underneath all of them: knowing where the evidence ends and our interpretation begins.
That's why I keep coming back to the same two questions:
What does the data show? What do I think explains it?
Sometimes the evidence supporting both will be strong. Sometimes we'll have a convincing relationship but several plausible explanations. And sometimes the most intellectually responsible thing we can say is, “We don't know yet.”
That isn't a weakness in the analysis. It's knowing what the analysis can and cannot tell us.
So the next time your data shows a relationship, don't rush to explain it. Understand the pattern, look for alternative explanations, decide what additional evidence you need, and then make the strongest conclusion the evidence can genuinely support.
Because finding a relationship is often not the end of an analysis.
It's where the interesting questions begin.
New to Correlation and Regression? A Quick Reference
If some of the statistical terms in this article are unfamiliar, here's the plain-English version.
What is correlation?
Correlation describes how closely two numeric variables tend to move together. A positive correlation means they generally move in the same direction, while a negative correlation means one tends to rise as the other falls. A correlation close to zero suggests there is little consistent linear relationship between them. The important word is relationship. Correlation by itself does not tell us why that relationship exists.
What is a correlation coefficient?
A correlation coefficient is a number ranging from -1 to +1 that describes the direction and strength of a relationship. Values closer to +1 indicate a stronger positive relationship, while values closer to -1 indicate a stronger negative relationship. Values around zero suggest little linear relationship.
It is useful as a summary, but the number should always be interpreted alongside the data and business context.
What is regression?
Regression examines how an outcome is related to one or more other variables.
For example, instead of looking only at whether training and performance move together, we might examine how performance relates to training while also considering tenure, role and previous performance.
Regression can help us investigate a relationship more deeply, but regression alone does not prove causation.
What is a confounding variable?
A confounding variable is another factor that may influence the relationship we're observing.
In our café example, cold-drink sales and ice-cream sales rise together, but temperature may be influencing both. Confounding variables are one reason we need to be careful when moving from “these things are related” to “this caused that.”
What does causation mean?
Causation means that a change in one factor actually contributes to a change in another.
Establishing causation generally requires stronger evidence than simply finding a correlation or fitting a regression model. The exact evidence needed depends on the question and situation, but the central challenge is ruling out credible alternative explanations.































Comments