Do you like to garden?

Mia does! She tracks the growth of her plants by recording their height (in centimetres) at different weeks.
She organised her data in the following table:
| Week | Height (cm) |
|---|---|
| 1 | 9.7 |
| 2 | 21.1 |
| 3 | 26.6 |
| 4 | 38.8 |
| 5 | 54.1 |
| 6 | 61.2 |
She then plotted a graph with the weeks on the x-axis (the horizontal 'left to right' one) and height on the y-axis (the vertical 'up and down' one):

This is an example of a scatter graph.
A scatter graph is a graphical representation of a set of bivariate data.

That sounds a little complicated!
What is bivariate data?
Bivariate data is data where we measured two variables (i.e. some quantities that vary).
That makes sense as we can see the root of the word 'bivariate' is vary (i.e. it's connected to our variables!).
Then there is just the prefix bi- which, if we think about bilingual meaning two languages or biweekly meaning twice a week, means two!
So bivariate means 'of two variables'!

For example, Mia measured two variables: the heights of the plants and which week it was!
Let's have a look at another example:
The relationship between the cost of rent of a double room and distance from London of the commuter is being investigated.
The following data was collected:
| Distance from London (km) | Rent Cost (GBP) |
|---|---|
| 10 | 1500 |
| 20 | 1300 |
| 30 | 1100 |
| 40 | 900 |
| 50 | 700 |
| 60 | 2000 |
| 70 | 600 |
A scatter graph was then constructed:

We can see that overall the points are going down - which makes sense - the further away from London we are, the less rent will cost!
We call this negative correlation.
Negative correlation is a relationship between two variables where as one variable goes up, the other goes down.
Mia's heights of plants and weeks on the other hand displayed a positive correlation:

As the weeks went up, so did the heights of the plants!
.jpg)
Sometimes we have no correlation.
That is when there is no apparent relationship between the two variables, i.e. one variable doesn't affect the other!
The scatter graph would then look something like this:

We can see that the points aren't going up as in positive correlation or down as with negative correlation - they're 'random'!

Let's come back to our rent vs distance from London graph one more time:

We can see there is one 'random point' at (60/2,000) that doesn't really fit our negative correlation.
We call this an outlier because it lies outside the trend we've spotted (points going down).
For example, here we have an outlier at (60/2,000) so we have rent of £2,000 at 60 km away from London.
Maybe there is a particularly beautiful or popular town 60 km from London and that's why we have an outlier there!
We omit any outliers when we draw the line of best fit:

A line of best fit (or a regression line or a trend line) is a straight line that best represents the relationship between a set of data points on a scatter graph.
We can then use it to predict what values we could have for the parts of the graph where we have no data.
For example, let's say we were looking to live at most 35 km from London but we wanted to get the cheapest rent possible.
Then we can use the scatter graph with a line of best fit to see what rent would be 35 km away from London based on the trend:

So we can see that the cheapest rent we can expect is just above £1,000!
Ready to have a go at some questions?




