LLMs concept guide
Large Language Models
Large neural networks trained on text and other modalities to predict, generate, reason over, and transform information.
Linked research
12
Published MachinoAI explainers currently connected to this concept.
Questions and answers
Search Large Language Models questions
What is a causal language model?
A causal language model is a language model that predicts each token using only the tokens that come before it in the sequence. It is called causal because information is allowed to flow in one direction: from earlier tokens to later tokens. For example, given: the model can use all of those earlier tokens to predict the next token, such as: How does a causal language model enforce this? In Transformer-based causal language models, self-attention uses a causal mask . The mask prevents a token from looking at future tokens during training. For a sequence such as: the model can use: But when predicting love , it is not allowed to see machine learning . Why is this useful? This training setup matches how text is generated at inference time. During generation, future tokens do not exist yet, so the model must predict the next token using only the current context. That makes causal language models well suited for: - Text generation - Conversation - Code generation - Story completion - Autocomplete - Instruction-following systems Is a causal language model autoregressive? Usually, yes. A causal language model generates text autoregressively by repeatedly predicting the next token from previous tokens. The key distinction is: - Causal describes the attention constraint: the model cannot use future tokens. - Autoregressive describes the sequence modeling process: each prediction depends on earlier tokens. For the broader sequence-generation idea, see “What is autoregressive language modeling?” Causal vs. masked language modeling A causal language model predicts tokens from left-to-right context. A masked language model can use tokens on both sides of a masked position. That distinction is covered separately in “What is a masked language model?” In short: a causal language model predicts each token using only earlier tokens, typically through a causal attention mask that blocks access to future positions.
What is a masked language model?
A masked language model is a language model trained to predict tokens that have been hidden, or masked , inside a sequence. Unlike a causal language model, which predicts the next token using only earlier tokens, a masked language model can usually use context from both the left and the right of the masked position. For example: The model uses the surrounding words to predict: How does masked language modeling work? During training, some tokens in a sentence are hidden or replaced. For example: The model must predict the missing token: Because the model can examine words before and after the masked position, it learns representations that capture context from both directions. Why is this useful? Masked language modeling is especially useful for tasks where understanding the full input matters more than generating long text token by token. Examples include: - Text classification - Sentiment analysis - Named-entity recognition - Search and retrieval - Extractive question answering - Learning contextual text representations Models such as BERT are well-known examples of masked language models. How is it different from causal language modeling? Consider: A masked language model can use both: and: to predict the missing token. A causal language model cannot normally look at future tokens when predicting the current token. The full comparison is covered separately in “What is the difference between causal and masked language models?” In short: a masked language model learns by predicting hidden tokens using surrounding context, often from both directions, which makes it particularly effective for language-understanding tasks.
What is the difference between causal and masked language models?
The main difference between causal and masked language models is the context they are allowed to use when predicting tokens. A causal language model predicts the next token using only the tokens that came before it. A masked language model predicts hidden tokens using surrounding context, often from both the left and the right. Simple example Consider this sentence: A causal language model might learn: When predicting sleeping , it cannot look ahead at on the sofa . A masked language model might instead see: It can use both: and: to predict sleeping . Key differences | Aspect | Causal language model | Masked language model | |---|---|---| | Prediction objective | Predict the next token | Predict hidden tokens | | Context used | Previous tokens only | Usually both left and right context | | Attention pattern | Causal / one-directional | Bidirectional | | Typical strength | Text generation | Text understanding | | Common example | GPT-style models | BERT-style models | Why causal models are good at generation During text generation, future tokens do not exist yet. A causal language model is trained under that same restriction, so it naturally generates text one token at a time: This makes causal models especially suitable for chat, writing, code generation, and other generative tasks. Why masked models are good at understanding Masked language models can examine the full surrounding sentence when predicting a missing token. That bidirectional context is useful for tasks such as: - Text classification - Named-entity recognition - Sentiment analysis - Search and retrieval - Extractive question answering Are they completely separate ideas? They use different training objectives, but both learn statistical relationships between tokens and can be built with Transformer architectures. For the individual concepts, see “What is a causal language model?” and “What is a masked language model?” In short: causal language models predict future tokens from past context and are naturally suited to generation, while masked language models predict hidden tokens using surrounding context and are especially useful for understanding-oriented tasks.
What is a base LLM?
A base LLM is a language model that has completed its main pretraining stage but has not yet been specifically tuned to behave like an instruction-following assistant or chat model . Its core objective is usually to predict the next token from the text that comes before it. For example, if you give a base LLM: it may continue with: But if you give it an instruction such as: a base model may not reliably follow the instruction in the way a chat or instruction-tuned model would. It may instead continue the text in a style similar to patterns it saw during pretraining. How is a base LLM created? A base LLM is typically trained on a very large corpus of text using a language-modeling objective such as next-token prediction. During this stage, the model learns patterns involving: - Grammar and syntax - Word and concept relationships - Facts and associations present in the training data - Writing styles - Code patterns - General statistical structure of language This training gives the model broad language capability, but it does not automatically make the model good at following user instructions. What happens after the base model? A base model can be further trained using post-training techniques such as: - Instruction tuning , where it learns from instruction-response examples - Preference or alignment training , where responses are optimized toward desired behavior - Domain-specific fine-tuning , where the model is adapted to a specialized task or field After this additional training, the resulting model may become an instruction-tuned LLM or a chat LLM . Base LLM vs. chat LLM A base LLM is mainly optimized to model and continue text. A chat LLM is additionally optimized to respond helpfully to conversational instructions. The detailed differences are covered separately in “What is the difference between a base model and a chat model?” In short: a base LLM is the pretrained language model before instruction-following or chat-specific post-training is applied. It provides the general language capability that later tuning builds upon.
What is an instruction-tuned LLM?
An instruction-tuned LLM is a pretrained language model that has been further trained to follow human instructions more reliably. A base LLM mainly learns to predict the next token. Instruction tuning adds another training stage where the model sees examples of instructions paired with suitable responses. For example: By training on many instruction-response examples, the model learns that a user prompt is not just text to continue—it is often a task to complete. How instruction tuning works A common approach is supervised fine-tuning SFT . The model is trained on datasets containing examples such as: During training, the model's parameters are adjusted so that its responses become more consistent with the target answers. What does instruction tuning improve? Instruction tuning can make a model better at: - Following explicit requests - Answering questions directly - Producing requested formats - Summarizing or rewriting text - Performing classification or extraction from instructions - Handling multi-step tasks more consistently For example, a base model given: might simply continue the surrounding text pattern. An instruction-tuned model is more likely to respond directly with: Is instruction tuning the same as alignment? Not exactly. Instruction tuning teaches the model how to follow tasks and response patterns. Additional post-training methods may then optimize qualities such as helpfulness, safety, or preference alignment using human or model-generated feedback. Instruction-tuned LLM vs. base LLM A base LLM has general language capability from pretraining. An instruction-tuned LLM starts from that base model and is further trained to interpret prompts as instructions and respond accordingly. For the foundation model stage, see “What is a base LLM?” In short: an instruction-tuned LLM is a pretrained language model that has undergone additional training on instruction-response examples so it can follow user requests more effectively.
What is a chat LLM?
A chat LLM is a language model that has been adapted to interact with users in a conversational format. It usually starts from a pretrained base model and then undergoes additional post-training so it can follow instructions, respond to questions, maintain conversational context, and behave more consistently as an assistant. How is a chat LLM different from a base LLM? A base LLM is mainly trained to predict the next token in text. A chat LLM is additionally trained to interpret structured conversations such as: This conversational structure helps the model learn how to respond appropriately to different roles and turns. What training is used? A chat LLM may use several post-training stages, including: - Instruction tuning , where the model learns from instruction-response examples - Conversation fine-tuning , where it learns from multi-turn dialogues - Preference or alignment training , where responses are optimized toward desired behavior - Safety tuning , where the model is trained to avoid or handle certain unsafe requests appropriately The exact post-training pipeline differs between model families. How does a chat LLM handle a conversation? The model receives the available conversation history as part of its context. For example: The second question depends on the earlier conversation. The model uses that previous context to understand that he refers to Alexander Fleming. A chat LLM does not permanently remember every conversation by default. It can only use information included in its current context unless the application provides an external memory system. Is every instruction-tuned model a chat model? Not necessarily. An instruction-tuned model may be optimized to follow individual tasks, while a chat model is usually further adapted for conversational interaction, role-based message formats, and multi-turn dialogue. In short: a chat LLM is a pretrained language model that has been further tuned to follow instructions and participate effectively in multi-turn conversations.
What is the difference between a base model and a chat model?
A base model and a chat model usually share the same underlying language-model architecture, but they are trained for different stages and behaviors. A base model is primarily trained to predict the next token from large amounts of text. A chat model starts from a base model and is further post-trained to follow instructions, respond conversationally, and handle multi-turn dialogue more effectively. Key differences | Aspect | Base model | Chat model | |---|---|---| | Main training stage | Pretraining | Pretraining + post-training | | Primary behavior | Continue or model text | Respond to user instructions | | Conversation handling | Not specifically optimized | Optimized for dialogue | | Instruction following | May be inconsistent | Usually much stronger | | Role-based messages | Usually not a core training format | Commonly trained with system/user/assistant roles | | Typical use | Research, fine-tuning, model development | Assistants, chatbots, interactive applications | Example Suppose the prompt is: A base model may continue the text according to patterns it learned during pretraining, but it may not reliably respect the instruction or requested format. A chat model is more likely to interpret the prompt as a task and return a direct two-sentence explanation. Why does the behavior change? The difference mainly comes from post-training . A chat model may undergo: - Instruction tuning - Conversation fine-tuning - Preference or alignment training - Safety tuning These stages teach the model how to respond as an assistant rather than simply continue text. Does a chat model have different knowledge? Usually, the chat model inherits most of its broad knowledge from the base model. Post-training mainly changes how that knowledge is used and how the model behaves. However, additional fine-tuning data can also add or reinforce certain capabilities and patterns. For the individual concepts, see “What is a base LLM?” and “What is a chat LLM?” In short: a base model is the pretrained language model before assistant-specific tuning, while a chat model is that model after additional training designed to improve instruction following and conversational behavior.
What is a foundation language model?
A foundation language model is a large, broadly trained language model that serves as a general-purpose starting point for many downstream applications. Instead of being trained for only one task, it is pretrained on large and diverse datasets so it can learn general language patterns, knowledge, and representations that can later be adapted for different uses. Why is it called a foundation model? It is called a foundation model because other systems can be built on top of it. A foundation language model can be adapted through methods such as: - Instruction tuning - Fine-tuning - Domain-specific training - Retrieval-augmented generation RAG - Tool use - Prompting For example, the same underlying foundation model might be adapted into: - A customer-support assistant - A coding assistant - A legal-domain model - A summarization system - A conversational chatbot Is every foundation language model a chat model? No. A foundation language model is the broad underlying model. A chat model is usually a foundation or base model that has undergone additional post-training so it can follow instructions and interact effectively in conversations. Foundation model vs. base LLM The terms can overlap, but they emphasize slightly different ideas. A base LLM usually refers to the pretrained model before instruction or chat tuning. A foundation model emphasizes that the pretrained model is general-purpose and can support many downstream tasks or specialized models. A base LLM can therefore also be a foundation language model. Why are foundation language models useful? Training a large model from scratch is expensive. A foundation model allows organizations and developers to start with broad pretrained capabilities and adapt them instead of rebuilding language understanding from the beginning. In short: a foundation language model is a broadly pretrained, general-purpose language model that provides the underlying capabilities for many specialized models and applications.
How does an LLM represent language internally?
An LLM represents language internally as numerical vectors that are continuously transformed as text passes through the model. The model does not store words and sentences as plain text internally. Instead, it converts tokens into numbers and builds increasingly context-aware representations of them. 1. Tokens become embeddings Each token is first mapped to an embedding : a vector of numbers. For example, a token such as: is represented by a vector such as: The exact numbers are learned during training. Tokens used in similar contexts often develop related representations. 2. Context changes the representation The initial embedding is only the starting point. As the token passes through Transformer layers, its representation is updated based on the surrounding context. For example: The token bank starts from the same token representation, but the model can produce different contextual representations because the surrounding words are different. 3. Transformer layers refine hidden states At each layer, self-attention and feed-forward networks transform the vectors. These intermediate vectors are often called hidden states . They can encode information related to: - Token identity - Position - Syntax - Semantic relationships - References between words - Broader context This information is distributed across many dimensions rather than stored in one single location. 4. Knowledge is distributed across parameters and activations An LLM's learned language patterns are represented partly in its parameters and partly in the temporary hidden states created while processing a specific input. The parameters contain patterns learned during training, while the hidden states represent the current context. 5. Final representations are used for prediction For a generative LLM, the final hidden representation is passed to the output layer, which produces scores for possible next tokens. The model then uses those scores to continue the sequence. For related concepts, see “What are parameters in an LLM?” and “What is a token in an LLM?” In short: an LLM represents language as high-dimensional numerical vectors. These vectors are repeatedly transformed by Transformer layers so they capture context, relationships, and patterns needed for prediction.
What are parameters in an LLM?
Parameters in an LLM are the learned numerical values inside the model that determine how it transforms input tokens and makes predictions. They include the weights used throughout components such as: - Token embeddings - Self-attention layers - Feed-forward networks - Normalization layers - Output projection layers How are parameters learned? At the start of training, most parameters are initialized with numerical values. The model processes training examples, makes predictions, measures how wrong those predictions are using a loss function, and then updates its parameters through optimization. Over many training steps, the parameters gradually encode useful patterns about language. What do parameters represent? A single parameter usually does not correspond to one specific fact or concept. Instead, information is typically distributed across many parameters working together. For example, the model's ability to recognize grammar, relationships between concepts, or common factual associations emerges from patterns spread across large parts of the network. Simple example A neural operation can be simplified as: Here, weight and bias are parameters. An LLM contains enormous numbers of these learned values arranged across many layers. Parameters vs. activations Parameters are the model's learned values and usually remain fixed during normal inference. Activations or hidden states are temporary values created while the model processes a particular prompt. So: Why do parameters matter? The number and quality of learned parameters affect how much structure the model can represent, but parameter count alone does not determine model quality. Training data, architecture, optimization, and post-training also matter. For the size-related concept, see “What does parameter count mean in an LLM?” In short: parameters are the learned numerical values that control how an LLM processes information. Training adjusts these values so the model becomes better at predicting and generating language.
Page 2 of 4
Research papers
Papers that connect to Large Language Models
Evaluation and Benchmarking of LLM Agents: A Survey
Evaluation and Benchmarking of LLM Agents: A Survey: Reliability, safety and capability evaluation.
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering: Coding agents and agent-computer interfaces.
OpenHands: An Open Platform for AI Software Developers as Generalist Agents
OpenHands: An Open Platform for AI Software Developers as Generalist Agents: Production-oriented coding agents.
Cognitive Architectures for Language Agents
Cognitive Architectures for Language Agents: Memory, planning and action architecture.
The Rise and Potential of Large Language Model Based Agents: A Survey
The Rise and Potential of Large Language Model Based Agents: A Survey: Complete agent landscape.
WebArena: A Realistic Web Environment for Building Autonomous Agents
WebArena: A Realistic Web Environment for Building Autonomous Agents: Web and computer-use agents.
Voyager: An Open-Ended Embodied Agent with Large Language Models
Voyager: An Open-Ended Embodied Agent with Large Language Models: Long-term learning and skill library.
MemGPT: Towards LLMs as Operating Systems
MemGPT: Towards LLMs as Operating Systems: Long-term contextual memory.
AgentBench: Evaluating LLMs as Agents
AgentBench: Evaluating LLMs as Agents: General agent evaluation.
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face: LLM as orchestrator.
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation: Multi-agent collaboration.
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society: Role-playing multi-agent systems.