LLMs concept guide
Large Language Models
Large neural networks trained on text and other modalities to predict, generate, reason over, and transform information.
Linked research
12
Published MachinoAI explainers currently connected to this concept.
Questions and answers
Search Large Language Models questions
What is vocabulary size in an LLM?
Vocabulary size in an LLM is the total number of distinct tokens that the model's tokenizer can represent. Each token in the vocabulary has its own numeric token ID . For example, if a tokenizer has a vocabulary size of: then its token IDs may range roughly from: depending on how the tokenizer is designed. What is included in the vocabulary? A tokenizer vocabulary can contain: - Whole words - Subword pieces - Characters or byte-level pieces - Punctuation - Whitespace patterns - Numbers or number fragments - Special tokens For example, a tokenizer might contain separate entries for: Why does vocabulary size matter? Vocabulary size affects how text is represented. A larger vocabulary can represent more common words or text patterns as single tokens, which may shorten token sequences. A smaller vocabulary uses fewer token entries but may split text into more pieces, producing longer sequences. This creates a tradeoff: Does a larger vocabulary always make an LLM better? No. A larger vocabulary can improve tokenization efficiency for some text, but it also increases the number of parameters needed in parts of the model such as the embedding and output layers. Tokenizer quality, training data, language coverage, and tokenization strategy are more important than vocabulary size alone. Is vocabulary size the same as parameter count? No. Vocabulary size tells us how many different tokens the tokenizer knows. Parameter count tells us how many learned numerical values exist inside the model. They are related because vocabulary size influences the size of embedding and output matrices, but they measure different things. For the underlying concept, see “What is a tokenizer vocabulary?” In short: vocabulary size is the number of distinct tokens available to an LLM tokenizer. It influences how efficiently text is represented and how large some model components need to be.
What is subword tokenization?
Subword tokenization is a tokenization method that represents text using pieces that can be smaller than complete words but larger than individual characters. The goal is to keep common words or word pieces compact while still being able to represent rare or previously unseen words. For example, a tokenizer might split: into pieces such as: The exact split depends on the tokenizer and its vocabulary. Why use subwords? A word-level tokenizer would need a very large vocabulary to represent every possible word. A character-level tokenizer avoids that problem, but it produces much longer sequences. Subword tokenization provides a compromise: How does it handle uncommon words? Suppose a tokenizer does not contain the complete word: It may still be able to represent it using known pieces such as: This allows the model to process new names, technical terms, word variations, and other uncommon text without requiring every possible word to exist as a separate vocabulary entry. How is the subword vocabulary created? A tokenizer is usually built from a large text corpus. The tokenizer analyzes recurring text patterns and chooses useful pieces for its vocabulary. Different tokenization algorithms use different rules for doing this. Common approaches include: - Byte Pair Encoding BPE - WordPiece - Unigram tokenization These methods are covered separately in their dedicated articles. Does every word get split? No. A frequent word may already exist as a single token, while a less common word may require several subword tokens. For example: The exact result varies between tokenizers. Why does this matter for LLMs? Subword tokenization helps keep the vocabulary manageable while supporting a large variety of text. It also affects: - Token counts - Context-window usage - Inference cost - Multilingual efficiency - How words and code fragments are represented For the broader process, see “What is tokenization in LLMs?” In short: subword tokenization breaks text into reusable pieces that are often smaller than words, allowing an LLM to represent both common and rare text without requiring an enormous word-level vocabulary.
Why is subword tokenization used in LLMs?
Subword tokenization is used in LLMs because it provides a practical balance between word-level and character-level tokenization. Instead of requiring every possible word to exist in the vocabulary, the tokenizer can represent unfamiliar or rare words using smaller reusable pieces. For example: might be represented as: The exact split depends on the tokenizer. 1. It keeps the vocabulary manageable A word-level tokenizer would need a huge vocabulary to cover: - Common words - Rare words - Names - Technical terms - Spelling variations - Newly created words Subword tokenization avoids this by reusing smaller pieces across many words. 2. It handles rare and unseen words If a complete word is not in the vocabulary, the tokenizer can often split it into known subwords. For example: This greatly reduces the problem of unknown words. 3. It avoids very long character sequences Character-level tokenization can represent almost any text, but it often creates long sequences. For example: contains eight characters, while a subword tokenizer might represent it with one or only a few tokens. Shorter token sequences reduce the amount of sequence processing the model must perform. 4. It captures useful word structure Subword pieces can also capture recurring linguistic patterns such as: These reusable pieces can help the model learn relationships between related words. 5. It supports many kinds of text Subword tokenization works well for: - Multiple languages - Code - Numbers - Compound words - Technical terminology - Names and uncommon expressions The exact efficiency varies by tokenizer and language. Why not just use whole words? Whole-word tokenization is simple, but it struggles with vocabulary growth and unknown words. Subword tokenization instead allows: For the basic concept, see “What is subword tokenization?” In short: LLMs use subword tokenization because it keeps vocabulary size practical, handles rare or unseen text, and avoids the very long sequences produced by purely character-level tokenization.
Page 4 of 4
Research papers
Papers that connect to Large Language Models
Evaluation and Benchmarking of LLM Agents: A Survey
Evaluation and Benchmarking of LLM Agents: A Survey: Reliability, safety and capability evaluation.
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering: Coding agents and agent-computer interfaces.
OpenHands: An Open Platform for AI Software Developers as Generalist Agents
OpenHands: An Open Platform for AI Software Developers as Generalist Agents: Production-oriented coding agents.
Cognitive Architectures for Language Agents
Cognitive Architectures for Language Agents: Memory, planning and action architecture.
The Rise and Potential of Large Language Model Based Agents: A Survey
The Rise and Potential of Large Language Model Based Agents: A Survey: Complete agent landscape.
WebArena: A Realistic Web Environment for Building Autonomous Agents
WebArena: A Realistic Web Environment for Building Autonomous Agents: Web and computer-use agents.
Voyager: An Open-Ended Embodied Agent with Large Language Models
Voyager: An Open-Ended Embodied Agent with Large Language Models: Long-term learning and skill library.
MemGPT: Towards LLMs as Operating Systems
MemGPT: Towards LLMs as Operating Systems: Long-term contextual memory.
AgentBench: Evaluating LLMs as Agents
AgentBench: Evaluating LLMs as Agents: General agent evaluation.
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face: LLM as orchestrator.
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation: Multi-agent collaboration.
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society: Role-playing multi-agent systems.