Programming and IT
How to learn machine learning from scratch
Machine learning means building models that find patterns in data and make predictions from them. Starting with neural networks is tempting, but it is safer to go from the bottom up — from Python and math to classic algorithms, honest evaluation and only then deep learning. Without a technical background, expect six months to a year of regular study.
In this article
What you need before your first model#
Python at a confident level: functions, collections, working with files, plus the NumPy (arrays and vectorised operations) and pandas (tables) libraries. Without them, all your time goes into fighting the code instead of understanding the models. If you are not there yet, start with how to learn Python from scratch.
Math — enough to understand, not enough to prove:
- linear algebra — vectors, matrices, matrix multiplication;
- calculus — derivatives, the gradient, the idea of optimisation;
- probability and statistics — distributions, mean and variance, conditional probability.
You do not have to cover all of it in advance: it is convenient to study a math topic at the moment it first shows up in an algorithm.
Your working environment#
Jupyter notebooks locally or in a free cloud notebook service such as Google
Colab. Libraries: numpy, pandas, matplotlib, scikit-learn. A separate
virtual environment for your studies saves you from version conflicts.
Classic machine learning (2–3 months)#
Problem types: regression (predict a number — a price, demand), classification (pick a class — spam or not), clustering (split objects into groups without ready answers). Algorithms in order: linear and logistic regression, k-nearest neighbours, decision trees, random forest, gradient boosting.
The main idea to grasp right away is evaluating a model on data it has not seen during training:
, =
, , , =
=
Next come cross-validation, hyperparameter tuning, overfitting and regularisation, and feature preparation: missing values, categorical features, scaling.
Metrics: how not to fool yourself#
Accuracy is misleading on imbalanced data: if 95% of emails are not spam, a model that says "nothing is spam" scores 0.95 and is useless. Learn precision, recall, F1 and ROC AUC for classification, and MAE and RMSE for regression. Always compare your model with the simplest baseline — for example, predicting the most frequent class or the mean value.
Neural networks (the next 2–3 months)#
After the classics — fully connected networks, backpropagation, activation functions, optimisers. Then convolutional networks for images and the transformer architecture for text. Framework: PyTorch or TensorFlow/Keras. Start by training a small network to recognise handwritten digits and getting it not to overfit.
Projects and self-checks#
Your portfolio is two or three projects with the full cycle: problem statement,
exploratory data analysis, a baseline, several models, a comparison by metrics,
conclusions. Ideas: predicting house prices from public listings, classifying
reviews as positive or negative, predicting customer churn on an open dataset.
Keep the code in Git and fix random_state so results can be reproduced.
How to check yourself:
- take part in practice competitions on platforms like Kaggle — scoring on hidden data shows your model's real quality;
- reproduce a result from a textbook or a paper from scratch without peeking at the code;
- implement linear regression and gradient descent in NumPy yourself — after that, library models stop being a black box;
- explain why the model got specific examples wrong.
Step-by-step plan
- Months 1–2 — Python and mathNumPy, pandas, linear algebra basics, derivatives, probability.
- Months 3–4 — classic algorithmsRegression, classification, trees, ensembles in scikit-learn; train/test splits.
- Month 5 — model qualityMetrics, cross-validation, overfitting, feature work, a baseline.
- Months 6–8 — neural networksPyTorch or Keras, fully connected and convolutional networks, transformer basics.
- PortfolioTwo or three full-cycle projects and practice competitions.
Start learning this in your own space
The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.
Check yourself
1.You need to predict the price of an apartment in dollars. What type of problem is this?
2.A model classified 90 objects out of 100 correctly. What is its accuracy (as a fraction from 0 to 1)?
3.A model flagged 10 emails as spam, and 8 of them really are spam. What is its precision (as a fraction from 0 to 1)?
Sources
-
scikit-learnDocumentation and user guide for classic algorithmsfree
-
Kaggle LearnFree short courses on pandas and machine learningfree
-
PyTorchA neural network framework with official tutorialsfree
-
Mathematics for Machine LearningA free book covering the math behind MLfree
Was this helpful?