Artificial intelligence

How neural networks work

A neural network is simpler than its name suggests. Inside there are no thoughts or images — there are tables of numbers multiplied by the input values, and a procedure that keeps nudging those numbers until the answers get more accurate.

Updated
In this article

One neuron: multiply, add, bend#

An artificial neuron takes several numbers, multiplies each by its own coefficient — a weight — adds up the results, adds a bias and passes the sum through a simple function.

output = f(w1·x1 + w2·x2 + w3·x3 + b)

The function f is there so the network does not collapse into one big linear formula. The most common choice is ReLU: negative values become zero, positive ones pass through unchanged. A worked example: weights 0.5, −1.2, 0.3, inputs 2, 1, 4, bias 0.1. The sum is 0.5·2 − 1.2·1 + 0.3·4 + 0.1 = 1.0 − 1.2 + 1.2 + 0.1 = 1.1. ReLU keeps 1.1.

There is no biology behind this — the resemblance to a living neuron survives only in the name.

Layers: each one sees the features of the previous one#

Neurons are grouped into layers, and the output of one layer becomes the input of the next. The first layer works with raw numbers — pixel brightness, sensor readings. The second combines them into simple patterns: edges, changes in brightness. The third builds combinations of combinations: corners, textures, then shapes. In image networks this ladder is easy to see if you visualise what each layer responds to.

The important part is that nobody programmed "edge" or "corner". These features appeared on their own because they turned out to be useful for the final answer. Deep learning simply means many layers: the more layers, the more composite the features a network can assemble.

Training: error, derivative, a small step#

At first the weights are random and the network outputs nonsense. Then a cycle repeats.

  1. Run examples through the network and get predictions.
  2. Compute the loss function — a number showing how far the answers are from the correct ones.
  3. Work out which way to move each weight so that the loss goes down. This is backpropagation: derivatives are computed from the end of the network back to the start.
  4. Move every weight a small step in that direction. The size of the step is called the learning rate.

There are millions of such steps. A useful picture is walking down a slope in fog: you cannot see the bottom of the valley, but you feel the slope under your feet and step downhill. Steps too small and the descent takes forever; steps too large and you jump right over the dip.

Once training is finished, the weights are frozen. From then on the network only computes: the input is multiplied by the finished numbers and out comes an answer. That is why a running model does not change because of your questions — the difference between training and inference is covered in what is artificial intelligence.

Data matters more than architecture#

A network adapts to exactly what it is shown. If the cats in the examples were photographed indoors and the dogs outdoors, the network may learn to tell the backgrounds apart rather than the animals, and get the first indoor dog wrong. Shortcuts like this are found by the network on its own and stay invisible until you hit the right example.

Overfitting is the second typical problem: the network has memorised the training examples with all their quirks and does worse on new ones. The fix is simple in principle and tedious in practice: part of the data is set aside and never shown during training, so you can honestly measure quality on something unfamiliar.

Different tasks, different architectures#

Convolutional networks are good for images: the same small filter slides across the picture, so a feature is recognised in any corner of the frame. Recurrent networks handled sequences by processing elements one at a time. Transformers, which language models are built on, use attention instead of a queue: when processing each element, the network weighs all the others and decides for itself what matters most right now. That made parallel computation possible and let models take long texts into account; details are in what is an LLM.

Limits and risks#

There is nowhere to pull the reason for an answer from. Inside there are millions of numbers, not rules. There are methods that highlight which inputs had the strongest influence, but a network does not give a full explanation of "why", and that hurts in areas where a decision must be justified: medicine, lending, hiring.

Biases in the training data move into the model wholesale. No data about a group of people, a dialect or a rare defect means no quality there either. The network will not warn you; it will just be confidently wrong.

Generative networks make things up. Text models produce wrong facts and links; image models produce extra fingers, unreadable lettering and fake artist signatures. The mistake looks as neat as a correct result.

A network does not tell you it has met something unfamiliar. Give it something that was not in training and it will still answer with high confidence, because the question "have I seen anything like this?" is not part of the computation.

And in general: any answer from such a system about medicine, law or money is material for a conversation with a doctor, lawyer or financial adviser, not a replacement for one. Where a mistake costs health, money or rights, the decision is made by a person who is accountable for it.

Step-by-step plan

  1. Compute a neuron by handTake three weights, three inputs and a bias and get the output with ReLU — the mechanism stops being just words.
  2. Play with a network in a sandboxIn a browser sandbox, add layers and neurons and watch the decision boundary change.
  3. Understand the loss functionSee on one example why training comes down to making one number smaller.
  4. Spot overfittingCompare quality on the training set and on a held-out set — the gap is overfitting.
  5. Build your first networkRepeat a ready-made tutorial example in one of the frameworks and change one parameter in it.

Start learning this in your own space

The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.

Start the plan

Check yourself

1.What gets adjusted while a neural network is trained?

2.A network works very well on the training examples and badly on new ones. What is this called?

3.Why is a non-linear activation function placed between layers?

Sources

Was this helpful?

More articles

More solutions