Programming and IT

How to learn machine learning from scratch

Machine learning means building models that find patterns in data and make predictions from them. Starting with neural networks is tempting, but it is safer to go from the bottom up — from Python and math to classic algorithms, honest evaluation and only then deep learning. Without a technical background, expect six months to a year of regular study.

Updated
In this article

What you need before your first model#

Python at a confident level: functions, collections, working with files, plus the NumPy (arrays and vectorised operations) and pandas (tables) libraries. Without them, all your time goes into fighting the code instead of understanding the models. If you are not there yet, start with how to learn Python from scratch.

Math — enough to understand, not enough to prove:

  • linear algebra — vectors, matrices, matrix multiplication;
  • calculus — derivatives, the gradient, the idea of optimisation;
  • probability and statistics — distributions, mean and variance, conditional probability.

You do not have to cover all of it in advance: it is convenient to study a math topic at the moment it first shows up in an algorithm.

Your working environment#

Jupyter notebooks locally or in a free cloud notebook service such as Google Colab. Libraries: numpy, pandas, matplotlib, scikit-learn. A separate virtual environment for your studies saves you from version conflicts.

Classic machine learning (2–3 months)#

Problem types: regression (predict a number — a price, demand), classification (pick a class — spam or not), clustering (split objects into groups without ready answers). Algorithms in order: linear and logistic regression, k-nearest neighbours, decision trees, random forest, gradient boosting.

The main idea to grasp right away is evaluating a model on data it has not seen during training:

from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score

X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=42)

model = RandomForestClassifier(random_state=42).fit(X_train, y_train)
print(accuracy_score(y_test, model.predict(X_test)))

Next come cross-validation, hyperparameter tuning, overfitting and regularisation, and feature preparation: missing values, categorical features, scaling.

Metrics: how not to fool yourself#

Accuracy is misleading on imbalanced data: if 95% of emails are not spam, a model that says "nothing is spam" scores 0.95 and is useless. Learn precision, recall, F1 and ROC AUC for classification, and MAE and RMSE for regression. Always compare your model with the simplest baseline — for example, predicting the most frequent class or the mean value.

Neural networks (the next 2–3 months)#

After the classics — fully connected networks, backpropagation, activation functions, optimisers. Then convolutional networks for images and the transformer architecture for text. Framework: PyTorch or TensorFlow/Keras. Start by training a small network to recognise handwritten digits and getting it not to overfit.

Projects and self-checks#

Your portfolio is two or three projects with the full cycle: problem statement, exploratory data analysis, a baseline, several models, a comparison by metrics, conclusions. Ideas: predicting house prices from public listings, classifying reviews as positive or negative, predicting customer churn on an open dataset. Keep the code in Git and fix random_state so results can be reproduced.

How to check yourself:

  • take part in practice competitions on platforms like Kaggle — scoring on hidden data shows your model's real quality;
  • reproduce a result from a textbook or a paper from scratch without peeking at the code;
  • implement linear regression and gradient descent in NumPy yourself — after that, library models stop being a black box;
  • explain why the model got specific examples wrong.

Step-by-step plan

  1. Months 1–2 — Python and mathNumPy, pandas, linear algebra basics, derivatives, probability.
  2. Months 3–4 — classic algorithmsRegression, classification, trees, ensembles in scikit-learn; train/test splits.
  3. Month 5 — model qualityMetrics, cross-validation, overfitting, feature work, a baseline.
  4. Months 6–8 — neural networksPyTorch or Keras, fully connected and convolutional networks, transformer basics.
  5. PortfolioTwo or three full-cycle projects and practice competitions.

Start learning this in your own space

The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.

Start the plan

Check yourself

1.You need to predict the price of an apartment in dollars. What type of problem is this?

2.A model classified 90 objects out of 100 correctly. What is its accuracy (as a fraction from 0 to 1)?

3.A model flagged 10 emails as spam, and 8 of them really are spam. What is its precision (as a fraction from 0 to 1)?

Sources

Was this helpful?

More articles

Programming and IT How to learn Python from scratch Python is a good first programming language: code reads almost like text, and the standard library covers most everyday tasks. This plan takes you from installing the interpreter to your own scripts covered by tests in about four months, at roughly an hour a day. Programming and IT How to learn SQL from scratch SQL is the query language of relational databases. Developers, analysts, testers and managers who want to pull numbers themselves all need it. Basic queries take a few weeks to learn; working confidently with complex reports takes two or three months of practice. Below is the order of topics and ways to train on a real database. Programming and IT How to learn Linux from scratch Linux runs most servers, containers and countless devices, so developers, testers, analysts and system administrators all need the command line. The easiest way to learn it is not by reading lists of commands but by working in the terminal every day and solving small practical tasks. Below is a sequence of topics for two to three months. Programming and IT How to learn JavaScript from scratch JavaScript runs in every browser and, through Node.js, on the server too. It is easy to start — you can run code right in the browser console — but it has plenty of surprising corners, from type coercion to asynchronous code. The plan below takes three to four months and assumes you already know basic HTML and CSS or are learning them alongside. Programming and IT How to learn Java from scratch Java is a strictly typed language behind banking systems, the servers of large services and Android apps. The strictness slows you down at first, but the compiler catches many mistakes before the program ever runs. This plan takes about six months at an hour a day and leads from your first program to a small backend application. Programming and IT C++ from scratch C++ is a compiled language used wherever speed and direct access to memory matter — game engines, browsers, databases, firmware. It is harder to get into than Python, but you can build your first working program on the very first evening.

More solutions