Artificial intelligence

What is an LLM

An LLM (large language model) is a program trained to continue text. Everything it does comes down to one operation: look at what has been written so far, predict the next small piece, and repeat.

Updated
In this article

One operation repeated a thousand times#

A chat looks like a conversation, but something else is going on inside. The system glues the developer's instructions, the conversation history and your latest question into one long text and asks the model to continue it. The model produces one piece, the piece is appended to the text, and the whole thing repeats — until a special end-of-text marker comes out.

That explains an everyday observation: the model starts answering before it has "thought of" the ending. It has no plan for the whole answer. If the first sentence takes a wrong turn, the model will faithfully keep going down that turn — one reason a rephrased question often gets a better answer than "try again".

Tokens: neither words nor letters#

Text is cut into tokens — chunks of roughly three to four characters for English, and often shorter for languages that are less represented in the model's vocabulary. A common word can be a single token; a rare or compound word splits into several; punctuation and spaces take up room too.

Tokens matter for three reasons. They measure the length of a request and an answer. They are what paid models charge for. And tokenization is why models struggle with letter-level tasks: counting how many r's are in a word is hard when the model does not see the word as a string of letters.

Where something that looks like knowledge comes from#

The model was trained on a huge collection of text: websites, books, documentation, forums, code. During training it adjusted its internal numbers to get better at guessing a hidden continuation. No database of facts is built along the way — what you get is statistics of which words and ideas tend to go together.

That is why a model's knowledge is blurry. It reproduces common knowledge reliably, rare things approximately, and things that barely appeared in its training text it fills in by pattern. The mechanics of adjusting those numbers are explained separately in how neural networks work.

After the main training, the model is further tuned on example dialogues and human ratings — so that it answers questions instead of continuing them with more questions, as a plain text-continuer would. This tuning changes the manner, but it does not add a reference book to the model.

The context window is a desk#

The context window is the limit of how much text the model can take into account at once: the system instructions, the whole conversation, attached files and the answer itself. It is measured in tokens.

The comparison is simple: it is not memory, it is a desk. While a sheet is on the desk, the model sees all of it. When the desk overflows, old sheets get dropped — the interface usually silently cuts the start of the conversation or replaces it with a shorter summary. That is the familiar feeling that in a long chat the model has "forgotten" what you agreed at the start: the sheet was taken off the desk.

Adding outside data — RAG, retrieval-augmented generation — works through the same desk. A search component finds relevant fragments of documents and puts them into the request, and the model answers from them. What changes is what is on the desk, not the model.

Why the same question gets different answers#

At each step the model produces not one continuation but a probability distribution over its entire vocabulary, and one token is then picked from it. Temperature is the setting for that choice. Near zero, the most likely token is almost always taken, and answers become uniform and predictable; higher up, the model more often takes less likely options, and the text becomes livelier and less predictable.

The practical upshot: low temperature for extracting data, fixed formats and code; higher temperature for brainstorming headlines. It also explains why "it answered me differently yesterday" is not a malfunction.

Limits: what to always check#

A fabrication looks exactly like the truth. Links, page numbers, paper titles, quotes, product codes, standard numbers — the model generates all of these the same way as the rest of the text, that is, by plausibility. Open the links, find the quotes in the source, check the numbers.

A confident tone means nothing. The model has no information about the edges of its own ignorance; a hedge like "possibly" in an answer is just another phrase chosen by probability.

What you send to someone else's service is no longer entirely yours. Conversations, attached documents and pieces of work code are stored by the service owner under the owner's rules.

Medicine, law and money are not its territory. The wording can be smooth and sound like a consultation, but nobody carries responsibility for the consequences. Where a mistake costs health, money or rights, the decision belongs to a doctor, lawyer or financial adviser with a licence and access to your actual situation.

Finally, the properties of specific products — ChatGPT, Claude, Gemini and others — change with every update: window size, available tools, behaviour. Comparing them by "intelligence" is pointless; test them on your own task.

Step-by-step plan

  1. Look at tokenizationOpen any online tokenizer and see how your text is cut into pieces and how many there are.
  2. Hit the context limitPush a long chat to the point where the model loses early agreements — the window stops being abstract.
  3. Play with temperatureAsk one question at low and high temperature and compare how much the answers vary.
  4. Check the links in an answerAsk for five sources on a narrow topic and open each one: some will not exist.
  5. Write your own ruleNote which tasks you trust an answer for straight away and which require checking.

Start learning this in your own space

The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.

Start the plan

Check yourself

1.What does a model's context window mean?

2.Temperature has been lowered almost to zero. What changes?

3.Why can a model give you a link to a paper that does not exist?

Sources

Was this helpful?

More articles

Artificial intelligence AI agents An agent is not a special, smarter model but a harness around an ordinary one: the model gets a list of tools and is allowed to call them in a loop until the task is finished — or until patience runs out. Artificial intelligence How neural networks work A neural network is simpler than its name suggests. Inside there are no thoughts or images — there are tables of numbers multiplied by the input values, and a procedure that keeps nudging those numbers until the answers get more accurate. Artificial intelligence Prompt engineering Prompt engineering is choosing the input text so that a model more often produces the result you need. It is not a set of magic spells but ordinary engineering work — a hypothesis, a test on a set of examples, rolling back changes that did not help. Artificial intelligence What is artificial intelligence Artificial intelligence is not one technology but the name of a whole field. In everyday speech the word covers programs that do tasks which used to need a person — recognising speech, translating, writing text, picking an answer. Artificial intelligence How to write a prompt A good prompt differs from a bad one not in length or politeness but in having everything the work needs — the task, the material, the shape of the answer and what a usable result looks like. Artificial intelligence How to use AI Getting started is simpler than it looks: you need one service, one real task from your own work and the habit of checking the result. Everything else is detail you pick up along the way.

More solutions