Programming and IT
How to become a data analyst from scratch
A data analyst answers business questions with numbers — why sales dropped, which version of a page works better, where customers drop off. The tools of the job are simple and well documented in the open: spreadsheets, SQL, Python and a bit of statistics. Below is the order to learn them in and how to build a portfolio from your own studies.
In this article
What an analyst's work is made of#
The cycle is almost always the same: clarify the question, find the data, clean it, compute the metrics, check the conclusions and present the result so that the person who asked can make a decision. You need technical skills at every step, but what matters most is framing the question well and not drawing a wrong conclusion from correct numbers.
Step 1: spreadsheets (2–3 weeks)#
Excel or Google Sheets is the first tool, and it is still used every day. Learn
sorting and filters, the formulas SUMIF, COUNTIF, VLOOKUP or XLOOKUP,
pivot tables and simple charts. Exercise: export six months of your own spending
and answer three questions — what takes the most money, how spending changes
month to month, and which expenses were one-offs.
Step 2: SQL (4–6 weeks)#
Company data lives in databases, and SQL is an analyst's main language. The
order: SELECT and WHERE, sorting, the aggregates COUNT, SUM, AVG,
GROUP BY and HAVING, JOINs, subqueries, then window functions. A typical
analyst query:
SELECT date_trunc('month', created_at) AS month,
count(DISTINCT user_id) AS buyers,
sum(amount) AS revenue
FROM orders
GROUP BY month
ORDER BY month;
A detailed plan for the language is in how to learn SQL from scratch.
Step 3: statistics without fear (4 weeks)#
The minimum: the mean, the median and why they drift far apart when there are outliers; spread and standard deviation; distributions; correlation and why it does not prove causation; samples and confidence intervals; hypothesis testing and A/B tests. Check every term on real data: compute the mean and median salary in an open job postings dataset and explain the difference.
Step 4: Python for analysis (6–8 weeks)#
After the basics of the language, move on to the libraries: pandas for tables,
matplotlib or seaborn for charts, Jupyter notebooks as your workspace. The
key pandas operations are loading a CSV, filtering, grouping, merging tables,
handling missing values and dates:
=
=
If Python is new to you, start with how to learn Python from scratch.
Portfolio: three studies#
Employers look not at certificates but at how you reason. Do two or three studies on open data — city statistics, housing prices, taxi trip data, datasets from competition platforms like Kaggle. Each one needs a question, a description of the data, cleaning, calculations, charts, conclusions and limitations. Keep the notebooks in a Git repository with a half-page summary in the README.
How to check your conclusions#
- Recompute the key number a second way — in a spreadsheet and in SQL; the numbers must match.
- Check how many rows you had before and after cleaning, and explain the difference.
- Ask yourself which other explanation would produce the same picture.
- Give the study to a friend: do they understand the conclusion without your explanations?
Step-by-step plan
- Weeks 1–3 — spreadsheetsFormulas, filters, pivot tables and charts on your own data.
- Weeks 4–9 — SQLFrom SELECT to JOIN and window functions; queries against a practice database.
- Weeks 10–13 — statisticsMean and median, spread, correlation, hypothesis testing, A/B tests.
- Weeks 14–21 — Python and pandasLoading, cleaning and grouping data, charts in Jupyter.
- Weeks 22–26 — portfolioTwo or three studies of open data with conclusions and limitations in a repository.
Start learning this in your own space
The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.
Check yourself
1.Find the arithmetic mean of 2, 4, 6, 8.
2.Find the median of the set 3, 1, 9, 7, 5.
3.Salaries in a team are 50, 55, 60, 60 and 500 thousand. Which measure better describes a typical salary?
Sources
-
pandas documentationThe official user guide and API referencefree
-
PostgreSQL documentationAn SQL tutorial and a reference for functionsfree
-
Khan Academy — Statistics and probabilityA free course in statistics and probabilityfree
-
Kaggle DatasetsOpen datasets for practice studiesfree
Was this helpful?