LLMs concept guide
Large Language Models
Large neural networks trained on text and other modalities to predict, generate, reason over, and transform information.
Linked research
12
Published MachinoAI explainers currently connected to this concept.
Questions and answers
Search Large Language Models questions
What is a Large Language Model (LLM)?
A Large Language Model LLM is an AI model trained on very large amounts of text so it can understand and generate human-like language. Most modern LLMs learn by processing text as tokens —small units such as words, parts of words, or characters—and predicting which token is likely to come next. For example, given: an LLM may assign the highest probability to: By learning this prediction task across enormous datasets, the model develops statistical representations of language, facts, patterns, writing styles, and relationships between concepts. What can an LLM do? Depending on how it is trained and used, an LLM can perform tasks such as: - Answering questions - Writing and rewriting text - Summarizing documents - Translating languages - Generating code - Extracting or classifying information - Holding multi-turn conversations Why is it called "large"? The word large usually refers to a combination of scale: - A large number of model parameters - Large training datasets - Significant computing resources used during training Parameters are the learned numerical values that allow the model to represent patterns in the data. Does an LLM actually understand language? An LLM does not understand language in the same way a human does. It generates outputs by using patterns learned during training and the context provided in the prompt. Because of this, an LLM can sometimes produce confident but incorrect information, commonly called a hallucination . In short: an LLM is a large-scale language model that learns patterns from massive text datasets and uses those patterns to predict and generate language for many different tasks.
How does a Large Language Model work?
A Large Language Model LLM works by converting text into tokens, processing those tokens through a neural network—usually a Transformer—and predicting which token should come next. At a high level, the process looks like this: 1. Text is tokenized The input is split into tokens. A token may represent a word, part of a word, punctuation, or another text unit. 2. Tokens are converted into numerical representations Each token is mapped to a vector called an embedding so the model can process it mathematically. 3. The Transformer processes the context Transformer layers examine relationships between tokens using mechanisms such as self-attention . This helps the model determine which earlier parts of the input are most relevant to the current prediction. 4. The model produces scores for possible next tokens After processing the context, the model outputs a score for every token in its vocabulary. These scores are converted into probabilities. 5. A next token is selected The model chooses one token based on those probabilities and the decoding strategy being used. 6. The process repeats The newly generated token becomes part of the context, and the model predicts another token. Repeating this step produces sentences, paragraphs, code, or other text. For example: The model may select Tokyo , append it to the input, and then predict the following token. How does the LLM learn to do this? During training, the model sees enormous amounts of text and repeatedly tries to predict missing or next tokens, depending on the training objective. Its internal parameters are adjusted when its predictions are wrong. Over many training examples, those parameters learn statistical patterns involving grammar, facts, writing styles, code structures, and relationships between concepts. The model does not search a database for each word it produces. During normal generation, it uses the patterns encoded in its learned parameters together with the text currently in its context. In short: an LLM works by turning text into tokens, processing their relationships through a Transformer, predicting a probability distribution over possible next tokens, selecting one, and repeating the process to generate text.
Why are Large Language Models called “large”?
Large Language Models are called “large” because they are built and trained at a very large scale. The word large mainly refers to three things: - Many parameters: LLMs can contain millions, billions, or even more learned parameters. These parameters are the numerical values adjusted during training that help the model represent patterns in language. - Large training datasets: They are trained on enormous collections of text and, in some cases, other data such as code. - Large amounts of computation: Training an LLM typically requires substantial computing power to process the data and update its parameters. Does “large” mean only parameter count? Parameter count is one of the most common ways to describe an LLM's size, but it is not the only factor. Two models with similar parameter counts can perform differently because of differences in: - Training data quality - Model architecture - Training methods - Amount of training computation - Fine-tuning or post-training techniques There is also no universal parameter threshold at which a language model officially becomes a Large Language Model . For a detailed explanation of model parameters, see “What are parameters in an LLM?” In short: Large Language Models are called “large” because they operate at large scale—especially in their number of parameters, training data, and computational requirements.
What can Large Language Models do?
Large Language Models can perform many tasks that involve understanding, transforming, or generating language. Common capabilities include: - Answering questions based on information learned during training or provided in the prompt - Generating text such as explanations, articles, emails, stories, and reports - Summarizing long documents into shorter versions - Rewriting and editing text for clarity, tone, grammar, or style - Translating between languages - Classifying text , such as identifying sentiment, topic, or intent - Extracting information such as names, dates, entities, or structured fields - Generating and explaining code - Following instructions in instruction-tuned or chat-oriented models - Holding multi-turn conversations while using the available conversation context Why can one model perform so many tasks? An LLM is trained on a wide variety of language patterns rather than being built for only one fixed task. Because many tasks can be expressed as text instructions, the same model can often adapt its behavior based on the prompt. For example: The underlying model is the same, but the instruction changes the task. What are the limitations? LLMs are powerful, but they are not guaranteed to be correct. They can: - Generate inaccurate or fabricated information - Misinterpret ambiguous instructions - Struggle with information outside their context or training - Produce different answers depending on the prompt and decoding settings Their capabilities also depend heavily on the model's training, size, architecture, context window, and post-training methods. In short: Large Language Models can handle a broad range of language-based tasks, including answering, writing, summarizing, translating, extracting, classifying, coding, and conversation, all through the same general language-modeling system.
What are the main components of a Large Language Model?
A Large Language Model LLM is made of several core components that work together to convert text into predictions and generate new text. The main components are: 1. Tokenizer The tokenizer converts raw text into tokens —the units the model processes. These may be whole words, parts of words, punctuation, or other text pieces. 2. Token embeddings Each token is converted into a numerical vector called an embedding . This gives the model a mathematical representation it can process. 3. Positional information Because token order matters, the model also needs information about where each token appears in the sequence. Different LLM architectures encode position in different ways. 4. Transformer layers Most modern LLMs are built from many stacked Transformer layers. Each layer typically contains: - Self-attention , which allows tokens to use information from other relevant tokens in the context - Feed-forward networks , which transform the representation of each token - Residual connections and normalization , which help information flow through deep networks and make training more stable 5. Output layer / language-model head After the Transformer layers process the input, the model converts the final hidden representation into scores for tokens in its vocabulary. These scores are turned into probabilities for possible next tokens. 6. Learned parameters The embeddings, attention layers, feed-forward networks, and output layers contain numerical parameters learned during training. These parameters encode the patterns the model has learned from data. A simplified flow is: For deeper explanations, see “What is a token in an LLM?” , “What are parameters in an LLM?” , and “How does an LLM predict the next token?” In short: the main components of an LLM are the tokenizer, embeddings, positional information, stacked Transformer layers, an output layer, and the learned parameters that connect everything together.
How does an LLM understand text?
An LLM “understands” text by converting it into numerical representations and using patterns learned during training to determine how the parts of the text relate to one another. It does not understand language in the same way a human does. Instead, it builds an internal mathematical representation of the input and uses that representation to predict useful outputs. 1. Text is converted into tokens The input is first split into tokens . For example: might be represented as several token IDs. For a detailed explanation, see “What is a token in an LLM?” 2. Tokens become embeddings Each token is converted into a vector called an embedding . An embedding is a list of numbers that represents information about the token in a form the neural network can process. 3. The model uses context A token can mean different things depending on the surrounding text. For example: The word bank has a different meaning in each sentence. Transformer-based LLMs use self-attention to examine relationships between tokens and build context-dependent representations. 4. Representations are refined through many layers As the input passes through multiple Transformer layers, the model continuously updates its internal representation of each token. Earlier layers may capture relatively simple patterns, while deeper layers can represent more complex relationships involving syntax, meaning, references, and context. 5. The representation is used to make predictions The model uses the final internal representation to predict an output, such as the most likely next token. For example: the context makes France a highly likely continuation. Does this mean an LLM understands like a human? Not necessarily. An LLM learns statistical and structural patterns from its training data. These representations can support impressive language behavior, but the model does not experience meaning, intention, or the world in the same way people do. In short: an LLM processes text by turning tokens into numerical vectors, using Transformer layers and self-attention to build context-aware representations, and then using those representations to predict appropriate outputs.
How does an LLM generate text?
An LLM generates text one token at a time . It starts with the text already available in its context, processes that context, predicts a probability distribution over possible next tokens, selects one token, appends it to the sequence, and repeats the process. 1. The prompt becomes context Suppose the prompt is: The model converts this text into tokens and processes them through its neural network. 2. The model predicts possible next tokens The model produces a score for every token in its vocabulary. Those scores are converted into probabilities. For example: The model is not directly choosing an entire sentence. It is deciding what token should come next. For a dedicated explanation, see “How does an LLM predict the next token?” 3. A token is selected The next token can be selected in different ways. A system may: - Choose the highest-probability token - Sample from several likely tokens - Adjust randomness using settings such as temperature - Restrict choices using decoding methods such as top-k or top-p sampling Because token selection can involve sampling, the same prompt can sometimes produce different outputs. 4. The selected token becomes part of the context If the model selects: the sequence now becomes: The model processes the updated context and predicts the next token again. 5. The cycle continues This process repeats: Generation stops when the model produces a stopping token, reaches a configured length limit, or the application stops generation. In short: an LLM generates text autoregressively by repeatedly predicting and selecting the next token based on everything already present in its context.
How does an LLM predict the next token?
An LLM predicts the next token by turning the current text into internal numerical representations, processing those representations through Transformer layers, and then assigning a probability to every token in its vocabulary. Suppose the input is: The model might produce probabilities such as: The model then selects one token according to the decoding strategy being used. 1. The input is tokenized The text is first converted into token IDs. For a dedicated explanation, see “What is tokenization in LLMs?” 2. Tokens are converted into vectors Each token ID is mapped to an embedding. Positional information is also added so the model knows where each token appears in the sequence. 3. Transformer layers process the context The embeddings pass through many Transformer layers. Using self-attention , the model determines which earlier tokens are relevant to the current prediction and builds a context-aware representation. For example, in: the representation of later tokens depends on relationships with earlier ones. 4. The final hidden state is converted into logits For the position where the next token must be predicted, the model produces a vector of scores called logits . There is typically one logit for each token in the model's vocabulary. For example: A higher logit means the model considers that token more suitable in the current context. 5. Softmax converts logits into probabilities The logits are transformed into probabilities using the softmax function. Conceptually: The probabilities across the vocabulary sum to 1. 6. A token is selected The application can choose the highest-probability token or sample from the probability distribution using decoding settings such as temperature, top-k, or top-p. The selected token is appended to the context, and the model repeats the same process to predict the following token. In short: an LLM predicts the next token by processing the current context through Transformer layers, producing a logit for every vocabulary token, converting those logits into probabilities, and selecting the next token from that distribution.
What is next-token prediction in LLMs?
Next-token prediction is the task of predicting which token is most likely to come next given the tokens that came before it. For example, given: an LLM may assign high probability to: and lower probabilities to other possible tokens. How does it work? For the current context, the model produces a probability distribution over its entire vocabulary. Conceptually: The model then selects a token according to the decoding strategy being used. After selecting Paris , the sequence becomes: The model then predicts the next token again. This repeated process allows an LLM to generate complete sentences, paragraphs, code, and conversations. Why is next-token prediction important? For many autoregressive LLMs, next-token prediction is also the core training objective. During training, the model repeatedly receives a sequence of tokens and tries to predict the correct next token. For example: If the model gives a low probability to the correct target token, the training process adjusts its parameters to reduce that error. By performing this task across very large datasets, the model learns patterns involving grammar, facts, syntax, style, code, and relationships between concepts. Next-token prediction vs. text generation They are closely related but not identical. Next-token prediction is the individual prediction step. Text generation is the repeated application of that step: For the detailed mechanics behind the probability calculation, see “How does an LLM predict the next token?” In short: next-token prediction is the process of estimating which token should follow the current context. Repeating this prediction step is what allows autoregressive LLMs to generate text.
What is autoregressive language modeling?
Autoregressive language modeling is a way of modeling text where the model predicts each token using only the tokens that came before it. In other words, the model generates a sequence from left to right: At each step, the model estimates the probability of the next token given the previous tokens. Mathematically, a sequence can be represented as: Here, each token depends on the tokens that appear before it. Example Suppose the sequence begins with: The model predicts probabilities for the next token, such as: If help is selected, the context becomes: The model then predicts the next token again. Why is it called autoregressive? The term autoregressive comes from the idea that each new prediction depends on earlier values in the same sequence. For language models, that means previously observed or generated tokens become the input used to predict the next token. How is it used in LLMs? Many generative LLMs use autoregressive modeling during both training and generation. During training, the model learns to predict the next token from previous tokens. During generation, it repeatedly: This is closely related to next-token prediction . See “What is next-token prediction in LLMs?” for the prediction objective itself. Important limitation Because generation happens token by token, later outputs depend on earlier generated tokens. An incorrect or poor token choice can influence the rest of the generated sequence. In short: autoregressive language modeling predicts a text sequence one token at a time, with each prediction conditioned on the tokens that came before it.
Page 1 of 4
Research papers
Papers that connect to Large Language Models
Evaluation and Benchmarking of LLM Agents: A Survey
Evaluation and Benchmarking of LLM Agents: A Survey: Reliability, safety and capability evaluation.
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering: Coding agents and agent-computer interfaces.
OpenHands: An Open Platform for AI Software Developers as Generalist Agents
OpenHands: An Open Platform for AI Software Developers as Generalist Agents: Production-oriented coding agents.
Cognitive Architectures for Language Agents
Cognitive Architectures for Language Agents: Memory, planning and action architecture.
The Rise and Potential of Large Language Model Based Agents: A Survey
The Rise and Potential of Large Language Model Based Agents: A Survey: Complete agent landscape.
WebArena: A Realistic Web Environment for Building Autonomous Agents
WebArena: A Realistic Web Environment for Building Autonomous Agents: Web and computer-use agents.
Voyager: An Open-Ended Embodied Agent with Large Language Models
Voyager: An Open-Ended Embodied Agent with Large Language Models: Long-term learning and skill library.
MemGPT: Towards LLMs as Operating Systems
MemGPT: Towards LLMs as Operating Systems: Long-term contextual memory.
AgentBench: Evaluating LLMs as Agents
AgentBench: Evaluating LLMs as Agents: General agent evaluation.
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face: LLM as orchestrator.
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation: Multi-agent collaboration.
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society: Role-playing multi-agent systems.