Statistics glossary
Short, exact definitions. Each term links to a tool where you can see it in action.
- Bayes’ theorem
A rule for reversing conditional probabilities: P(A | B) = P(B | A) · P(A) / P(B). It updates a prior belief in light of new evidence.
- Bell curve
Informal name for the symmetric, bell-shaped graph of the normal distribution.
- Binomial distribution
The distribution of the number of successes in n independent trials that each succeed with the same probability p.
- Box plot
A chart of the five number summary: a box from Q1 to Q3 with a line at the median, whiskers to the most extreme non-outliers, and dots for outliers.
- Central limit theorem
For large enough samples, the distribution of the sample mean is approximately normal with mean μ and standard deviation σ/√n, whatever the shape of the population.
- Combination
A selection of items where order does not matter. The number of ways to choose r from n is nCr = n! / (r!(n − r)!).
- Conditional probability
The probability of A given that B has happened: P(A | B) = P(A and B) / P(B).
- Confidence interval
A range computed from sample data by a method that captures the true parameter a stated percentage of the time, such as 95%.
- Correlation coefficient (r)
A number from −1 to 1 measuring the strength and direction of a linear relationship between two variables.
- Cumulative frequency
The running total of frequencies up to and including a class or value.
- Degrees of freedom
The number of values free to vary once a statistic is fixed. For a one-sample t procedure it is n − 1.
- Empirical rule
For normal data, about 68%, 95% and 99.7% of values lie within 1, 2 and 3 standard deviations of the mean.
- Expected value
The long-run average of a random variable, E(X) = Σ x · P(x): each outcome weighted by its probability.
- Exponential distribution
A continuous distribution for waiting times between events that occur at a constant average rate λ.
- Five number summary
Minimum, first quartile, median, third quartile and maximum of a data set.
- Frequency distribution
A table listing classes or values alongside how many observations fall into each.
- Geometric distribution
The distribution of the number of trials needed to get the first success, when each trial succeeds with probability p.
- Histogram
A chart of a numeric variable split into adjacent intervals (bins), with bar heights showing how many values fall in each.
- Independent events
Events where knowing one happened does not change the probability of the other, so P(A and B) = P(A) · P(B).
- Interquartile range (IQR)
Q3 − Q1, the spread of the middle 50% of the data. It is resistant to outliers.
- Law of large numbers
As the number of independent trials grows, the sample average gets closer to the expected value.
- Margin of error
Half the width of a confidence interval: critical value × standard error.
- Mean
The sum of the values divided by how many there are; the balance point of the data.
- Median
The middle value of sorted data, or the average of the two middle values when the count is even.
- Mode
The value that occurs most often. A data set can have one mode, several, or none.
- Mutually exclusive events
Events that cannot both happen, so P(A and B) = 0 and P(A or B) = P(A) + P(B).
- Normal distribution
A continuous, symmetric, bell-shaped distribution fully described by its mean μ and standard deviation σ.
- Null hypothesis
The default claim in a significance test, usually "no effect" or "no difference", that the data are tested against.
- Outlier
A value unusually far from the rest. A common rule flags values more than 1.5 × IQR below Q1 or above Q3.
- P-value
The probability, assuming the null hypothesis is true, of a result at least as extreme as the one observed.
- Parameter
A number describing a whole population, such as μ or σ. It is usually unknown and estimated by a statistic.
- Percentile
The value below which a given percentage of the data falls. The 90th percentile has 90% of values at or below it.
- Permutation
An arrangement of items where order matters. The number of ways to arrange r of n items is nPr = n! / (n − r)!.
- Poisson distribution
The distribution of the number of events in a fixed interval when events occur independently at a constant average rate λ.
- Population
The entire group of individuals or items that a question is about.
- Probability
A number from 0 (impossible) to 1 (certain) describing how likely an event is.
- Quartiles
The three values Q1, Q2 (the median) and Q3 that split sorted data into four equal parts.
- Random variable
A numerical outcome of a random process, such as the number of heads in ten flips.
- Range
The largest value minus the smallest value.
- Regression line
The least-squares line ŷ = a + bx that minimises the sum of squared vertical distances from the points.
- Relative frequency
A frequency divided by the total count: the proportion of observations in a class.
- Sample
The subset of a population that is actually observed or measured.
- Sampling distribution
The distribution of a statistic, such as the sample mean, across all possible samples of the same size.
- Skewness
A measure of asymmetry. Right (positive) skew has a long tail to the right; left (negative) skew has a long tail to the left.
- Standard deviation
The square root of the variance; roughly the typical distance of a value from the mean, in the original units.
- Standard error
The standard deviation of a statistic's sampling distribution. For a mean it is σ/√n, estimated by s/√n.
- Standard normal distribution
The normal distribution with mean 0 and standard deviation 1, tabulated in the z table.
- Statistic
A number calculated from a sample, such as x̄ or s, often used to estimate a population parameter.
- Statistical significance
A result is called significant when its p-value falls below a chosen level α, such as 0.05.
- Stem and leaf plot
A display that splits each number into a stem and a final-digit leaf, keeping every original value visible.
- t distribution
A bell-shaped distribution with heavier tails than the normal, used when σ is estimated from a sample.
- Uniform distribution
A distribution where every value in an interval is equally likely; its density is a flat rectangle.
- Variance
The average squared deviation from the mean, dividing by n − 1 for a sample or N for a population.
- Z-score
The number of standard deviations a value lies from the mean: z = (x − μ) / σ.
Questions students ask
What is the difference between a population and a sample?
The population is the entire group you want to know about; a sample is the subset you actually measure. Statistics computed from a sample (x̄, s) estimate parameters of the population (μ, σ).
What is the difference between a statistic and a parameter?
A parameter describes a population and is usually unknown, like the true mean μ. A statistic is calculated from a sample, like x̄, and is used to estimate the parameter.
What is the difference between descriptive and inferential statistics?
Descriptive statistics summarise the data you have. Inferential statistics use a sample to draw conclusions, with a stated uncertainty, about a larger population.
What is a random variable?
A rule that assigns a number to each outcome of a random process, such as the number of heads in 10 flips. It is discrete if it takes countable values and continuous if it can take any value in an interval.