LLMs concept guide
Large Language Models
Large neural networks trained on text and other modalities to predict, generate, reason over, and transform information.
Linked research
12
Published MachinoAI explainers currently connected to this concept.
Questions and answers
Search Large Language Models questions
What does parameter count mean in an LLM?
Parameter count in an LLM is the total number of learned numerical values inside the model. These parameters include the weights used in components such as: - Token embeddings - Self-attention layers - Feed-forward networks - Normalization layers - Output projection layers For example, a model described as having: contains roughly: learned numerical values. What does a higher parameter count mean? A larger parameter count generally gives the model more capacity to represent complex patterns. With more parameters, a model may be able to learn richer relationships involving language, knowledge, reasoning patterns, and code. However, more parameters do not automatically guarantee a better model. Performance also depends on factors such as: - Training data quality and quantity - Model architecture - Training objective - Optimization - Amount of training compute - Post-training methods A smaller, well-trained model can sometimes outperform a larger but poorly trained one. Where do all these parameters come from? Transformer-based LLMs contain many layers. Each layer includes large matrices of weights used by self-attention and feed-forward networks. If a model has many layers and each layer contains millions or billions of weights, the total parameter count can become very large. Does parameter count equal model file size? Not exactly, but they are closely related. Model storage depends on both: - Number of parameters - Number of bits used to store each parameter For example, storing parameters in lower-precision formats such as 8-bit or 4-bit representations can significantly reduce memory requirements. Are all parameters always used for every token? In a standard dense model, most parameters participate in each forward pass. Some architectures, such as Mixture-of-Experts MoE models, contain many total parameters but activate only a subset of them for each token. So a model can have a very large total parameter count without using every parameter for every input. For the underlying concept, see “What are parameters in an LLM?” In short: parameter count tells you how many learned numerical values an LLM contains. It is an important measure of model scale, but it does not by itself determine model quality or capability.
Why do LLMs have billions of parameters?
LLMs have billions of parameters because language is highly complex, and large numbers of learned parameters give the model enough capacity to represent many different patterns at the same time. These patterns can include: - Grammar and sentence structure - Meanings and relationships between words - Facts and associations found in training data - Writing styles - Code patterns - Long-range relationships between tokens - Context-dependent meanings Why are so many parameters needed? An LLM does not store one rule per parameter. Instead, useful behavior emerges from large networks of parameters working together. For example, understanding the word: requires different representations depending on whether the context is: or: Across millions of different linguistic situations, the model needs enough capacity to represent many overlapping relationships. Where do the billions of parameters come from? Transformer-based LLMs contain many layers, and each layer contains large matrices of learned weights. Major parameter-heavy components include: - Token embedding matrices - Attention projection matrices - Feed-forward network matrices - Output projection weights A simplified Transformer block may contain millions of parameters by itself. When dozens or hundreds of such layers are stacked together, the total can reach billions. Why does increasing model size help? With sufficient training data and compute, increasing model capacity can allow the model to learn more complex patterns and perform a wider range of tasks. However, simply adding parameters is not enough. A large model still needs: - High-quality training data - Effective architecture - Sufficient training compute - Good optimization - Appropriate post-training So billions of parameters provide capacity , not guaranteed intelligence or accuracy. Do all LLMs need billions of parameters? No. Smaller language models can perform very well for specific tasks, especially when they are efficiently trained or specialized for a particular domain. For the meaning of model size itself, see “What does parameter count mean in an LLM?” In short: LLMs often have billions of parameters because modeling the enormous variety and complexity of language requires substantial representational capacity. Those parameters work together to encode the patterns the model learns during training.
Does a larger parameter count always make an LLM better?
No. A larger parameter count does not automatically make an LLM better . More parameters give a model greater capacity to represent complex patterns, but model quality depends on several other factors. What else affects performance? Important factors include: - Training data quality and diversity - Amount of training data - Model architecture - Training compute - Optimization methods - Post-training and instruction tuning - Context handling - Evaluation task A smaller model trained more effectively can outperform a larger model on certain tasks. Why can larger models still help? When enough high-quality data and compute are available, increasing model size often improves performance because the model has more capacity to learn complex relationships. This is one reason modern LLM development often considers model size together with: rather than looking at parameter count alone. Bigger models also have costs Larger models generally require more: - Memory - Compute - Energy - Inference time - Deployment cost So the best model is not always the largest one. For a focused task, a smaller specialized model may be faster, cheaper, and accurate enough. What about Mixture-of-Experts models? Parameter count can also be misleading for Mixture-of-Experts MoE architectures. An MoE model may contain a very large number of total parameters but activate only a subset for each token. So total parameter count does not always indicate how much computation is used during inference. For the meaning of model size itself, see “What does parameter count mean in an LLM?” In short: more parameters can increase an LLM's capacity, but model quality depends on the combination of architecture, data, training compute, optimization, and post-training. Parameter count alone is not a reliable measure of which model is better.
Where is knowledge stored in an LLM?
An LLM does not store knowledge in a database-like list of facts. Most of what it learns during training is encoded across its learned parameters . Those parameters are the billions of numerical weights inside components such as: - Token embeddings - Self-attention layers - Feed-forward networks - Output layers Knowledge is distributed A fact or concept is usually not stored in one single parameter. Instead, information is represented across many parameters and model components working together. For example, the model's ability to complete: with: comes from patterns learned during training and distributed throughout the network. There is usually no single location that can be pointed to as: Parameters store learned patterns During training, the model repeatedly predicts tokens and adjusts its parameters when its predictions are wrong. Over time, these updates encode patterns involving: - Language structure - Word relationships - Factual associations - Common reasoning patterns - Writing styles - Code structures These learned patterns are sometimes called the model's parametric knowledge . What about the current prompt? Not everything the model uses comes from its parameters. When an LLM processes a prompt, it creates temporary internal representations called activations or hidden states . So it is useful to distinguish: For example, if you tell the model: the model can use that information during the current interaction even if it was never part of its original training. Can an LLM retrieve external knowledge? Yes, an application can also connect an LLM to external sources such as databases, search systems, tools, or retrieval-augmented generation RAG . That information is not stored inside the model's parameters . It is supplied to the model as additional context when needed. Is knowledge perfectly stored in the parameters? No. The model learns statistical patterns rather than maintaining a perfectly structured knowledge base. As a result, knowledge can be incomplete, approximate, outdated, or incorrectly recalled. For the underlying mechanism, see “What are parameters in an LLM?” In short: most knowledge learned during LLM training is distributed across the model's parameters, while prompt-specific and externally retrieved information is provided through the model's current context.
How do LLMs learn language patterns?
LLMs learn language patterns by training on very large amounts of text and repeatedly adjusting their internal parameters to make better predictions. For many generative LLMs, the core training task is next-token prediction . For example, the model may see: and try to predict: If its prediction is poor, the training process measures the error and updates the model's parameters so similar predictions become more accurate in the future. What kinds of patterns are learned? By repeating this process across enormous datasets, the model learns statistical patterns involving: - Grammar and sentence structure - Relationships between words and concepts - Common phrases and writing styles - Factual associations present in the training data - Code syntax and programming patterns - How context changes the meaning of a token - Long-range relationships across a sequence These patterns are not usually stored as explicit hand-written rules. How are the parameters updated? The model makes a prediction and compares it with the correct target. A loss function measures the prediction error. Training then uses backpropagation and an optimizer to adjust the model's parameters in a direction that reduces that loss. Conceptually: This process can happen across billions or trillions of training tokens. Does the model memorize everything? Not exactly. LLMs learn broad statistical patterns that allow them to generalize to text they have not seen before. However, models can sometimes memorize specific training examples, especially when data is repeated or distinctive. Their useful behavior comes from a combination of learned patterns, generalization, and—in some cases—memorization. Why does scale matter? With enough data, compute, and model capacity, the model can learn increasingly complex relationships. That is why modern LLM training usually depends on the combination of: For the prediction objective itself, see “What is next-token prediction in LLMs?” In short: LLMs learn language patterns by repeatedly predicting tokens, measuring their errors, and updating billions of parameters until those parameters capture useful statistical relationships in language.
What is a token in an LLM?
A token is a unit of text that an LLM processes. Before an LLM can work with text, a tokenizer converts the text into a sequence of tokens. Depending on the tokenizer, a token may represent: - A whole word - Part of a word - Punctuation - A space or other character pattern - A special control symbol For example, a tokenizer might split: into smaller pieces such as: The exact split depends on the tokenizer being used. Why does an LLM use tokens? Neural networks operate on numbers rather than raw text. Each token is therefore mapped to a numeric token ID . Conceptually: The model then converts each token ID into an embedding vector and processes the sequence through its Transformer layers. Is a token the same as a word? No. Some common words may correspond to one token, while uncommon or complex words may be split into multiple tokens. For example: could be represented as one token in one tokenizer but multiple tokens in another. This is why token count and word count are not the same thing . Why do tokens matter? Tokens affect several practical aspects of an LLM: - How much text fits inside the context window - How much input and output is processed - The computational cost of inference - API usage when pricing is based on token counts - How efficiently different languages or text patterns are represented What are special tokens? Some tokenizers also include special tokens used for structural purposes, such as marking the beginning or end of text, separating messages, or representing other control information. These tokens are part of the model's vocabulary even though they may not correspond to ordinary visible words. For the process that converts text into tokens, see “What is tokenization in LLMs?” In short: a token is the basic text unit an LLM processes. Text is split into tokens, each token is converted into a numeric ID, and those IDs are transformed into embeddings that the model can work with.
What is tokenization in LLMs?
Tokenization in LLMs is the process of converting raw text into smaller units called tokens that the model can process. A tokenizer may split text into: - Whole words - Parts of words - Punctuation - Spaces or character patterns - Special control tokens For example, the text: might be split into tokens such as: The exact split depends on the tokenizer. Why is tokenization needed? LLMs do not process raw text directly. They operate on numbers. After tokenization, each token is mapped to a numeric token ID . Conceptually: The model then converts those token IDs into embedding vectors and processes them through its Transformer layers. How are token boundaries chosen? Modern tokenizers usually use subword-based methods that try to balance two goals: - Keep common text pieces as single tokens - Break rare or unfamiliar words into smaller reusable pieces This allows the model to represent a very large vocabulary without needing a separate token for every possible word. For example, a rare word might be split into several known subword pieces rather than treated as an unknown word. Why does tokenization matter? Tokenization affects: - How much text fits in the context window - How many tokens an input or output consumes - Inference cost and latency - How efficiently different languages are represented - How words, numbers, code, and punctuation are segmented Two tokenizers can convert the same sentence into different numbers of tokens. Tokenization vs. tokens A token is one unit produced by the tokenizer. Tokenization is the process that converts text into those units. For the individual unit itself, see “What is a token in an LLM?” In short: tokenization converts raw text into a sequence of token IDs that an LLM can turn into embeddings and process mathematically.
How does an LLM tokenizer work?
An LLM tokenizer works by converting raw text into a sequence of token IDs that the language model can process. It does this using a fixed vocabulary and a set of tokenization rules learned or defined before the LLM is trained. 1. The tokenizer receives raw text Suppose the input is: The tokenizer first examines the characters and text patterns in the input. 2. It splits the text into known token pieces Instead of requiring every possible word to exist in the vocabulary, modern tokenizers usually work with reusable subword or byte-level pieces. A tokenizer might represent: as something like: while a common word might remain a single token. The exact result depends on the tokenizer and its vocabulary. 3. Each token is mapped to an ID Every token in the tokenizer vocabulary has a numeric identifier. Conceptually: The LLM receives these token IDs rather than the original text. 4. Special tokens may be added A tokenizer can also insert special tokens used by the model or application. These can mark things such as: - Beginning or end of a sequence - Message boundaries - System, user, or assistant roles - Padding or other control information Which special tokens exist depends on the model. 5. Token IDs are converted into embeddings Inside the LLM, each token ID is used to look up a learned embedding vector. The overall pipeline is: How is the tokenizer vocabulary created? Before model training, a tokenizer is typically trained or configured using a large text corpus. Algorithms such as Byte Pair Encoding BPE , WordPiece , or Unigram can identify useful recurring text pieces and build a vocabulary of manageable size. Different models may use different tokenization methods. How does text come back from tokens? Tokenization also works in reverse. After the model generates token IDs, the tokenizer's decoding process maps those IDs back into token pieces and combines them into readable text. For the broader concept, see “What is tokenization in LLMs?” In short: an LLM tokenizer breaks text into vocabulary pieces, maps those pieces to numeric token IDs, handles any required special tokens, and can later decode generated token IDs back into text.
Why do LLMs use tokens instead of words?
LLMs use tokens instead of whole words because tokens provide a more flexible and efficient way to represent language. If a model used only complete words, it would need an enormous vocabulary containing every possible word, spelling variation, number, name, technical term, and newly invented expression. Tokens solve this by allowing text to be represented using reusable pieces. 1. Tokens can represent parts of words A tokenizer can split an uncommon word into smaller pieces. For example: might be represented as pieces such as: The exact split depends on the tokenizer. This means the model does not need a separate vocabulary entry for every possible word. 2. Tokens can handle unfamiliar words Suppose the model encounters a rare company name, scientific term, or newly created word. Even if the entire word is not in the vocabulary, the tokenizer can often represent it using smaller known pieces. That makes token-based vocabularies much more flexible than word-only vocabularies. 3. Tokens keep the vocabulary manageable A pure word-level vocabulary could grow to millions of entries. A tokenizer instead uses a fixed vocabulary containing common words, subwords, punctuation, characters, or byte-level pieces. This keeps the model's embedding and output layers more manageable. 4. Tokens work better across different kinds of text LLMs process much more than ordinary sentences. They may need to represent: - Different languages - Numbers - URLs - Source code - Emojis - Punctuation - Misspellings - Technical terms Tokenization provides a common mechanism for representing all of these. Why not use individual characters? Character-level tokenization would avoid unknown words, but sequences would become much longer. For example: contains eight characters but might require only one or a few subword tokens. Modern tokenizers therefore usually aim for a compromise: Does every LLM use the same tokens? No. Different models can use different tokenizer algorithms and vocabularies, so the same sentence may produce different tokens and token counts across models. For the tokenization process itself, see “What is tokenization in LLMs?” In short: LLMs use tokens instead of only whole words because tokens keep the vocabulary manageable while still allowing the model to represent rare words, multiple languages, code, numbers, and unfamiliar text efficiently.
What is a tokenizer vocabulary?
A tokenizer vocabulary is the fixed collection of tokens that a tokenizer knows how to represent. Each token in the vocabulary is assigned a unique numeric token ID . For example, a simplified vocabulary might contain entries such as: When text is tokenized, the tokenizer breaks it into pieces that exist in this vocabulary and converts those pieces into their corresponding IDs. What can a tokenizer vocabulary contain? Depending on the tokenizer, the vocabulary may include: - Whole words - Subword pieces - Individual characters - Byte-level units - Punctuation - Whitespace patterns - Numbers or common number fragments - Special control tokens Modern LLM tokenizers often use a mixture of common words and smaller reusable pieces. Why not include every possible word? Language contains an enormous number of possible words, names, spellings, technical terms, and newly created expressions. A vocabulary containing every possible word would be impractical. Instead, tokenizers use smaller reusable units. For example, if a rare word is not available as one token, it can often be represented using several known pieces. The exact split depends on the tokenizer. How is the vocabulary created? Before the LLM is trained, a tokenizer is usually built from a large text corpus. Algorithms such as Byte Pair Encoding BPE , WordPiece , or Unigram identify useful recurring text pieces and select a fixed vocabulary. The vocabulary size is therefore a design choice. Why does vocabulary size matter? A larger vocabulary can represent more text pieces directly, potentially reducing the number of tokens needed for some text. However, it also increases the size of components such as the token embedding and output layers. A smaller vocabulary uses fewer entries but may split text into longer token sequences. So tokenizer design involves a tradeoff between: Is the tokenizer vocabulary the same as the model's knowledge? No. The vocabulary only defines the text units the model can process. The model's learned language patterns and knowledge are encoded primarily in its trained parameters, not in the tokenizer vocabulary. For the full tokenization process, see “How does an LLM tokenizer work?” In short: a tokenizer vocabulary is the fixed set of text pieces and special tokens that a tokenizer can map to numeric IDs for an LLM to process.
Page 3 of 4
Research papers
Papers that connect to Large Language Models
Evaluation and Benchmarking of LLM Agents: A Survey
Evaluation and Benchmarking of LLM Agents: A Survey: Reliability, safety and capability evaluation.
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering: Coding agents and agent-computer interfaces.
OpenHands: An Open Platform for AI Software Developers as Generalist Agents
OpenHands: An Open Platform for AI Software Developers as Generalist Agents: Production-oriented coding agents.
Cognitive Architectures for Language Agents
Cognitive Architectures for Language Agents: Memory, planning and action architecture.
The Rise and Potential of Large Language Model Based Agents: A Survey
The Rise and Potential of Large Language Model Based Agents: A Survey: Complete agent landscape.
WebArena: A Realistic Web Environment for Building Autonomous Agents
WebArena: A Realistic Web Environment for Building Autonomous Agents: Web and computer-use agents.
Voyager: An Open-Ended Embodied Agent with Large Language Models
Voyager: An Open-Ended Embodied Agent with Large Language Models: Long-term learning and skill library.
MemGPT: Towards LLMs as Operating Systems
MemGPT: Towards LLMs as Operating Systems: Long-term contextual memory.
AgentBench: Evaluating LLMs as Agents
AgentBench: Evaluating LLMs as Agents: General agent evaluation.
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face: LLM as orchestrator.
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation: Multi-agent collaboration.
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society: Role-playing multi-agent systems.