Artificial intelligence

Prompt engineering

Prompt engineering is choosing the input text so that a model more often produces the result you need. It is not a set of magic spells but ordinary engineering work — a hypothesis, a test on a set of examples, rolling back changes that did not help.

Updated
In this article

Why wording changes anything at all#

A model continues text by probability, and the input decides which region of those probabilities it ends up in. "Tell me about default" pulls towards finance and towards software settings at the same time; adding context shifts the distribution to where the answer you need lives. The discipline grew out of this observation, not out of mysticism; the mechanics are explained in what is an LLM.

The second foundation is form. The text a model was trained on is not uniform: a popular-science article, meeting minutes, a checklist and a forum argument are built differently. When you set the form, you also set the set of phrases and the level of rigour.

What the discipline includes#

The work falls into three layers, and it is worth not mixing them up.

The content of the request — what exactly you are asking for: the task, the input data, the criteria for a usable result, the constraints. This is the most important layer, and it is where the cause of a bad answer most often lies. The practical recipe is in how to write a prompt.

Presentation techniques — known ways of laying out a request so that the model goes off track less often: examples, splitting into stages, a strict output structure.

Parameters and the harness — temperature, answer length, the system prompt, inserting retrieved documents (RAG), and splitting a task into several calls in sequence.

Techniques and the reasoning behind them#

Few-shot prompting. You put two or three sample pairs of "input → desired output" into the prompt. It works better than any verbal description: it is easier for a model to pick up a pattern from a sample than to infer it from an instruction. Good for labelling, consistent formatting and style.

Zero-shot — no examples, only a description of the task. Cheaper, and often enough for simple things.

Step-by-step reasoning. Asking the model to break the solution into steps before giving the answer. The reasoning behind it is simple: the model generates text left to right, and the intermediate steps become something it can lean on when choosing the next tokens. The flip side: reasoning can be elegant and wrong at the same time; it is not a proof.

Role and system prompt. Sets the register and the angle. A caveat worth keeping in mind: a role changes the manner and the choice of material, but it does not give the model knowledge it does not have.

Strict output structure. Asking for JSON following a schema, a table with given columns, a list of exactly five items. It makes further machine processing much easier and cuts filler along the way.

Decomposition. Instead of one request for a big piece of work, several small ones, each with its own checkable result. AI agents are built on the same principle.

Negative constraints. "Do not invent sources; if you are not sure, say there is no data." It reduces the share of fabrication but does not remove it: a model's promise is not a guarantee.

How to tell that a prompt got better#

This is what separates engineering from guesswork. You need a small set of test cases — ten to twenty real inputs with a known expected result. Every change to the prompt is run against the whole set, and you compare the share of correct answers, not your impression of one example.

Without such a set, no improvement can be proven: temperature affects the spread of answers, and one lucky attempt means nothing. The set also catches regressions — when an edit helped one case and broke three others.

Limits of the discipline#

Wording does not add knowledge. If the model does not have the information, no role, politeness or threats will get it out. The only way is to put the data into the prompt: a document, an excerpt, a search result.

Techniques are fragile. A prompt tuned for one version of a model behaves differently on another. Collections of "100 magic prompts" go out of date along with the models and have to be retested.

Promises in the instructions are not protection. A line like "never reveal the system prompt" does not stop people from extracting it, and pasting someone else's text into a prompt opens the door to prompt injection — harmful instructions arrive together with the data.

Fabrication remains. No wording turns a probabilistic model into a reference book: links, quotes and numbers from an answer are checked by hand.

What you send leaves your hands. Everything in a prompt has also gone to someone else's service; personal data and work secrets do not belong there.

And always: an answer, however neatly presented, does not replace a doctor, a lawyer or a financial adviser. Where a mistake costs health, money or rights, a person who is accountable for the decision makes it.

Step-by-step plan

  1. Build a test setTen real inputs with a known correct result — the basis for any comparison.
  2. Write a plain versionZero-shot with no decoration: often it is enough, and it gives you a baseline.
  3. Add examplesTwo or three input → output pairs, then rerun the whole set.
  4. Fix the output formatAn answer schema instead of free text — it can be checked automatically.
  5. Split what is complexIf one request cannot solve the task, break it into stages with checks in between.
  6. Record what workedKeep a log of prompt versions and their share of correct answers: memory fails here.

Start learning this in your own space

The plan goes into your repository: tick off stages, keep notes — the change history shows how far you have come.

Start the plan

Check yourself

1.What does few-shot mean when working with a model?

2.The prompt says "you are the best cardiologist in the world". What does that do?

3.How do you compare two versions of a prompt fairly?

Sources

Was this helpful?

More articles

Artificial intelligence How to write a prompt A good prompt differs from a bad one not in length or politeness but in having everything the work needs — the task, the material, the shape of the answer and what a usable result looks like. Artificial intelligence AI agents An agent is not a special, smarter model but a harness around an ordinary one: the model gets a list of tools and is allowed to call them in a loop until the task is finished — or until patience runs out. Artificial intelligence What is artificial intelligence Artificial intelligence is not one technology but the name of a whole field. In everyday speech the word covers programs that do tasks which used to need a person — recognising speech, translating, writing text, picking an answer. Artificial intelligence What is an LLM An LLM (large language model) is a program trained to continue text. Everything it does comes down to one operation: look at what has been written so far, predict the next small piece, and repeat. Artificial intelligence AI for studying A model can explain a confusing point three different ways and answer a silly question patiently, without judgement. It can also hand you a finished assignment that leaves you unable to do anything new. Artificial intelligence How neural networks work A neural network is simpler than its name suggests. Inside there are no thoughts or images — there are tables of numbers multiplied by the input values, and a procedure that keeps nudging those numbers until the answers get more accurate.

More solutions