Call · 15 min
GlossaryMattia Esposito26 September 20266-minute read

LLM. The engine inside AI assistants, explained without jargon.

An LLM (large language model) is a program trained on huge amounts of text to predict the next piece of text, which makes it able to write, summarise, translate and answer in natural language. It's the engine inside assistants such as ChatGPT, Claude and Gemini.

In brief

An LLM predicts the most likely text, one piece at a time. That mechanism produces both the fluency of its writing and the mistakes it states with confidence.

The AI Act classes it among general-purpose AI models: trained with a large amount of data and capable of performing “a wide range of distinct tasks”.

In Italy, 16.4% of businesses with at least 10 employees use artificial intelligence, and of those, 59.1% use generative AI, according to Istat for 2025.

This entry is part of the AI and automation glossary, where every term has a short definition. Here the definition goes further: how an LLM works, the four words you need to understand it, what it can do in a small business, where it goes wrong and what European law says about it.

What an LLM is

An LLM is an artificial intelligence model trained on text, with billions of parameters, that produces new text from what it receives. “Large” refers to size: the amount of text used to train it and the number of parameters, the internal values adjusted during training.

The architecture almost every LLM is based on is the transformer, proposed in June 2017 in a paper by Ashish Vaswani and colleagues titled “Attention Is All You Need”. The key mechanism is attention: the model decides, word by word, which other parts of the text to look at.

How it works, in four words

Four words are enough to understand what goes on inside an LLM, and each has a practical consequence for a business using it: what it costs, what it knows, what it keeps in mind during a task, and how repeatable its answers are.

WordWhat it meansWhy it matters to you
Tokenthe unit of text

The piece of text the model reads and writes: a word, part of a word, a symbol.

Usage is paid for by the token, and the length of a document is measured in tokens.

Trainingwhen it learns

The phase in which the model learns from text by adjusting its parameters. It happens beforehand, once, and it's expensive.

The model doesn't know what happened afterwards, or anything about your documents.

Context windowwhat it keeps in mind

How much text the model takes into account at once: the question, attached documents, the conversation.

Whatever isn't in the window doesn't exist as far as the model is concerned.

Temperaturehow much it varies

The setting that controls how predictable or varied the answer is.

For business work it's kept low: less imagination, more repeatable answers.

What it can do in an SME

According to Istat, among Italian businesses using artificial intelligence in 2025, 70.8% use it to extract information from text documents and 59.1% use generative AI, whether written, spoken or visual. Those are the two jobs where an LLM delivers most: reading and writing.

In practice, an LLM reads an email in any language and understands what it's asking, extracts the data from an order or invoice, drafts a reply, a product sheet or a quote, summarises a long document, and translates while keeping the tone. Deciding whether the draft is right stays with a person.

Where it goes wrong

An LLM predicts probable text, and probable text can be false. When it doesn't know the answer, it tends to produce a plausible one anyway: that's a hallucination. In an article of 5 September 2025, OpenAI traces it to training and evaluation procedures that reward guessing over admitting uncertainty.

The second limit is the date. The model knows the world only up to its training, and doesn't know your documents. Both limits are kept in check by giving the model the right sources at the moment of the question, using the technique called RAG, and by putting a person wherever a mistake is costly: the human in the loop.

LLM, chatbot and agent

The LLM is the engine; a chatbot and an agent are two ways of using it. A chatbot puts the model in a conversation and produces answers. An agent gives it tools to act with, such as reading an inbox or writing to business software. The difference between the two is explained on the page AI agent or chatbot.

What the AI Act says

The AI Act, in Article 3, point 63, defines a “general-purpose AI model”: a model trained “with a large amount of data using self-supervision at scale”, displaying significant generality and capable of performing a wide range of distinct tasks. Large LLMs fall under it, and the main obligations fall on the companies that develop them.

For a business using them, two rules already in force matter most. Staff training, under Article 4, from 2 February 2025, explained on the page about the training obligation. And transparency towards people talking to a system, under Article 50, from 2 August 2026.

How Itria uses it

Itria uses LLMs inside systems with rules written in advance, never on their own. In Inbox AI the model reads enquiries and prepares a draft: on a test bench on 4 September 2026, with 100 enquiries across 4 channels and 5 languages, the draft was ready in a median of 7.290 seconds, in the right language in 97 cases out of 100.

The same models prepare the multilingual drafts and extract data in document entry. In every case, anything that commits the business goes through a person before sending, and the list of models used is available on request, as stated on the AI transparency page.

Related terms

AI hallucinations

A false statement made with confidence by a model. The best-known limit of LLMs, and the easiest to keep in check.

RAG

The technique that gives the model the right documents at the moment of the question, so it answers from them.

Prompt

The written instruction given to the model. The quality of the answer depends a lot on how it's written.

Generative artificial intelligence

The family of systems that produce new text, images or sound. LLMs are the part that writes.

Questions and answers

What is an LLM in simple terms?

An LLM, a large language model, is a program trained on huge amounts of text to predict the next piece of text.

Everything else comes from that ability: writing, summarising, translating, extracting data from a document, answering a question. It's the engine inside assistants such as ChatGPT, Claude and Gemini.

How does an LLM work?

It reads the text it receives broken into tokens, meaning words or parts of words, and works out which token is most likely to come next, one at a time, until the answer is complete. It learned those probabilities during training on text.

Almost all LLMs are based on the transformer, the architecture proposed in 2017 that uses a mechanism called attention to link the parts of a text together.

What's the difference between an LLM and ChatGPT?

The LLM is the model, the engine. ChatGPT is a product that puts a model inside a conversation, with an interface, rules and extra tools.

The same model can be used in other ways: inside a system that reads a business's emails, fills in business software or prepares drafts for approval.

Can an LLM get things wrong?

Yes. It predicts the most likely text, and when it doesn't know the answer it tends to produce a plausible one anyway: that's a hallucination. In an article of 5 September 2025, OpenAI traces this to training and evaluation procedures that reward guessing over admitting uncertainty.

You keep it in check by giving the model the right sources at the moment of the question and putting a person wherever a mistake is costly.

What does the AI Act say about LLMs?

Large LLMs fall within the definition of a general-purpose AI model in Article 3, point 63, of Regulation (EU) 2024/1689, and the main obligations for these models fall on the companies that develop them.

For a business using them, what matters most is staff training, Article 4, from 2 February 2025, and transparency towards people talking to a system, Article 50, from 2 August 2026.

Notes on sources

  1. The transformer was proposed in Vaswani et al., Attention Is All You Need, first version dated 12 June 2017.
  2. The AI Act quotations come from Article 3, point 63, of Regulation (EU) 2024/1689, quoted from the official English text. The dates for Articles 4 and 50 are set out, with sources, on our AI Act pages.
  3. The shares on the use of artificial intelligence come from Istat, Imprese e ICT, 2025 (in Italian), businesses with at least 10 employees. The 70.8% and 59.1% are calculated only on businesses that use artificial intelligence.
  4. The explanation of hallucinations comes from OpenAI, Why language models hallucinate, 5 September 2025, read on 26 September 2026. It's the view of the people who build the models, published alongside a research paper.
  5. The timings and the right-language figure are an Itria measurement on a test bench on 4 September 2026, with 100 test enquiries. It's a test system, not work carried out for a client, and the test can be rerun for anyone who asks.
·The next step

A model can read and write. What matters is where you put it in your work.

The first step with Itria is a fifteen-minute video call: we look at which texts your business reads and writes every day, and which could arrive already drafted. Drop us a line about what's slowing you down. We'll make the first move: we'll look at what a customer sees when they search for you, and tell you what we found. Even if we never end up working together.