Math and natural sciences
How to learn statistics from scratch
Analysts, researchers, doctors, marketers and anyone who reads the news with numbers in it need statistics. The easiest way to learn it from zero is in three layers — describe the data, understand probability, draw conclusions from a sample — and to calculate on real data at every layer. Here is the order of topics, how to practise and the traps that catch even experienced people.
In this article
What you need before statistics#
For an applied course, school algebra is enough: fractions, powers, roots and reading formulas with sums. The summation sign Σ is scary only for the first couple of days — it just means "add up everything on the list". For mathematical statistics at university level you will also need integrals and some linear algebra, but you do not have to start there: it is better to build intuition on simple examples first. If fractions and percentages feel shaky, go through percentage word problems before anything else.
Layer 1: describing data#
Types of data (numerical and categorical), the mean, median and mode, range, variance and standard deviation, quartiles. Then charts: the histogram, the box plot, the scatter plot. The main lesson of this layer: one number does not describe data. The mean salary and the median salary can differ several times over, and understanding why is already half of statistical literacy.
Layer 2: probability#
The classical definition of probability, the addition and multiplication rules, conditional probability and Bayes' theorem, random variables, expected value. Then distributions: binomial, normal, Poisson. Do not memorise density formulas: it matters more to understand what process produces each distribution. The number of heads in ten coin tosses is binomial, adult height is close to normal, the number of calls to a help desk in an hour is Poisson.
Layer 3: drawing conclusions from a sample#
Sample and population, the central limit theorem, confidence intervals, hypothesis testing and the p-value, the t-test, the chi-squared test, then correlation and linear regression. This layer is the most important and the most treacherous. Before you calculate anything, put it into words: what question am I asking, what is the null hypothesis, what will the result mean.
Practice: calculate from day one#
Statistics without data turns abstract quickly. Pick a tool — a spreadsheet (Excel, LibreOffice Calc or Google Sheets), the R language, or Python with the pandas library — and for every topic work through a small example by hand, then repeat it in the program on a real dataset. Open datasets are published by national statistics offices, city open-data portals and sites such as Kaggle. It helps to pick one subject you care about — say, the weather in your city — and come back to it at every stage.
Common beginner mistakes#
- Confusing correlation with causation: ice cream sales and drownings both rise in summer, but one does not cause the other.
- Reading the p-value as the probability that the hypothesis is true. It is not: it is the probability of getting data this extreme or more extreme if the null hypothesis is true.
- Testing dozens of hypotheses and keeping the "significant" ones — with that many tries, chance matches are guaranteed.
- Looking only at the mean and never plotting the distribution.
How to tell you are making progress#
A good test is to take a popular-science article or a news story with numbers and pick it apart: what the sample was, what exactly was compared, whether there is a confidence interval, whether a correlation is being passed off as a cause. If you can ask these questions and answer half of them, you have the basics.
Step-by-step plan
- Weeks 1–3 — descriptive statisticsMean, median, variance, quartiles, a histogram and a box plot on data of your own.
- Weeks 4–7 — probabilityConditional probability, Bayes' theorem, random variables and the main distributions.
- Weeks 8–12 — inferenceThe central limit theorem, confidence intervals, hypothesis testing, the t-test.
- Weeks 13–16 — relationships and regressionCorrelation, linear regression, and a close reading of a real article with numbers.
Start learning this in your own space
The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.
Check yourself
1.Find the mean of 2, 4, 6, 8, 10.
2.Find the median of 3, 7, 1, 9, 5.
3.A coin is tossed twice. What is the probability of getting heads both times? (as a fraction or a decimal)
Sources
-
Seeing TheoryA free visual introduction to probability and statistics from Brown Universityfree
-
OpenStax Introductory StatisticsA free statistics textbook with exercises and answersfree
-
Khan Academy — Statistics and probabilityA free course with practice problemsfree
Was this helpful?