Concepts

ML concept guide

Machine Learning

Systems that learn patterns from data and use those patterns to make predictions, decisions, or recommendations.

Linked research

3

Published MachinoAI explainers currently connected to this concept.

Questions and answers

Search Machine Learning questions

Machine LearningTransformer Architecture

A Transformer is a neural-network architecture designed to process sequences by using attention rather than relying on recurrence. Core idea For each token, the model creates three learned representations: query Q , key K , and value V . Self-attention compares queries with keys to determine how strongly tokens should attend to one another, then combines the corresponding values. A common form is: This lets the model connect information across a sequence directly. For example, in the sentence “The animal didn't cross the street because it was tired,” attention can help the representation of “it” incorporate information from “animal.” Main building blocks A Transformer typically combines: - Token embeddings to represent input tokens as vectors. - Positional information so token order is represented. - Multi-head self-attention so different attention heads can learn different relationships. - Feed-forward networks that transform each token representation. - Residual connections and normalization that help train deep networks. Encoder and decoder variants The original Transformer contains both an encoder and a decoder . Modern models often use only part of that design: - Encoder-only models are commonly used for understanding or representation tasks. - Decoder-only models generate tokens autoregressively and are widely used for LLMs. - Encoder-decoder models are useful when transforming one sequence into another, such as translation. Why Transformers matter Unlike recurrent neural networks, Transformers can process many sequence positions in parallel during training. Self-attention also provides a direct mechanism for modeling relationships between distant tokens. The tradeoff is that standard self-attention requires computing interactions between pairs of sequence positions, giving it roughly quadratic attention cost with sequence length . This becomes important when designing systems for long contexts. In short, a Transformer repeatedly uses attention and feed-forward computation to build context-aware token representations. This architecture became the foundation for many modern language models and has also been adapted to vision, audio, and multimodal systems.

Machine LearningEmbeddings and Vector Representations

Embeddings are numerical vector representations of data—such as words, sentences, images, users, or products—designed so that items with similar meaning or characteristics are located closer together in a vector space. For example, a text embedding model might represent: The real vectors usually contain hundreds or thousands of dimensions. The important property is not an individual number; it is the relative position of vectors . How embeddings are created An embedding model learns a function: During training, the model learns representations that capture useful patterns in its training objective. In language models, token embeddings are learned parameters that map token IDs to vectors. Separate embedding models can also convert complete sentences or documents into vectors intended for semantic comparison. How similarity is measured A common metric is cosine similarity , which compares the direction of two vectors: A higher cosine similarity generally indicates that two vectors are more similar according to the embedding model. A simple Python example: In a production system, an embedding model generates these vectors rather than developers manually defining them. Why embeddings are useful Embeddings make unstructured information searchable and comparable mathematically. Common applications include: - Semantic search: retrieve documents based on meaning rather than exact keyword matches. - RAG: embed documents and queries, retrieve relevant chunks, then provide them to an LLM. - Recommendation systems: represent users and items so similar preferences can be matched. - Clustering: group semantically related documents or products. - Classification: use embeddings as features for downstream models. - Duplicate detection: identify text or other content with very similar representations. Embeddings in RAG A typical retrieval workflow is: The query and stored documents should normally be embedded using compatible representations from the same embedding model or model family intended for that retrieval task. Embeddings vs. tokens Tokens and embeddings are related but different. A token is a discrete unit produced by tokenization. An embedding is a continuous numerical vector representing a token or larger piece of information. For example: In short, embeddings turn information into vectors that machine-learning systems can compare and process mathematically while preserving useful learned relationships.

Page 1 of 1

Research papers

Papers that connect to Machine Learning