The Math Behind LLMS

Large language models (LLMs) like ChatGPT have revolutionized how we interact with artificial intelligence, enabling machines to understand and generate human-like text. At the heart of this capability lies a sophisticated pipeline of processes that transform raw input into meaningful output. The journey begins with tokenization, where text is broken down into smaller units called tokens. These tokens are then converted into embeddings, high-dimensional vectors that capture semantic relationships between words, preserving context and meaning. Embeddings are further refined through transformer architectures, particularly attention blocks, which dynamically adjust the meaning of tokens based on their surrounding context. This interplay of tokenization, embeddings, and transformers forms the backbone of how LLMs process and generate language. Below, we detail some of the mechanisms, history, and math behind AI models, with much attention on attention blocks specifically, explaining both their purpose and how they function. This puts into context what makes LLMs so powerful, but also very different from real human thinking. The main mathematical tool utilized is linear algebra, specifically dot products and matrix multiplication.
The first step for anyone using a large language model is providing text input. The first step in breaking down and processing this input for the model however, is known as tokenization. Essentially, the input is broken in pieces, but not necessarily into the individual words. For example, words like “the” may be a token, while a word like “multilingual” might be broken into two: “multi” and “lingual”. Rather than seeing your input as words, the model sees it as tokens. The way models break words into tokens, is through a process called byte pair encoding (BPE). BPE was invented in 1994 by Philip Gage as a way to compress data. Essentially, the algorithm takes a string of letters, and assigns a token to the most common pair, and repeats this until the string cannot be shortened more by the algorithm. By applying this to training data for neural networks, you no longer need to pre-define a dictionary of words beforehand, and can instead use shorter tokens that are created by the algorithm.
Now while tokens are useful in shortening the length of a model’s dictionary, they do not provide context to the associations between words. For this, large language models use embeddings. Embeddings are large vectors that capture meaning by putting similar tokens positionally close in high-dimensional vector space. Vectors are lists of numbers that identify points in a corresponding vector space. By positioning tokens of similar meaning near each other in a vector space, semantics are preserved while also retaining the data efficiency of tokens. An example of a vector space that can be visualized is in three dimensions, where a vector [1,2,3] would identify a point with coordinates (1,2,3). However, all large language models use much higher dimensional vector spaces. For example, ChatGPT’s embeddings are around 12,228 dimensions.
The major breakthrough in the application of embeddings to LLMs was done by Tomas Mikolov, a computer scientist at Google. By training a neural network and embedding model to predict words from neighboring words and looking at the embedding vectors produced, they found structure that the model had “learned”. Essentially, the vectors the model produced held semantic value that corresponded to their real meaning. For example, the vector for “France” subtracted by the vector for “Paris” then added to the vector for “Italy” yielded the vector for “Rome”. It is important to note however, that this does not represent actual “thinking” in the human sense, as these relationships evolved purely based on the frequency of which words appear close to each other, which allows the model to make statistical predictions on what words will appear next.
With embeddings as described, words are assigned a general meaning. However, in natural language, the meaning of words is informed by the ones around it. For example, “quill” has very different meanings depending on if you’re talking about porcupines or the writing tool. To interpret this meaning on a case by case basis, LLMs use transformers, and more specifically attention blocks within transformers. The block looks at each embedding vector, and modifies the vector based on the surrounding embedding vectors in the string of text. So, in the sentence “the bank of the river was muddy”, the word “bank” would be modified by the words “river” and “muddy”, to indicate that this is a riverbank, not some other bank.
To understand how attention blocks work, the concept of parameters is essential. LLM’s have billions of values called parameters that influence how data is processed and how predictions are made. Within attention blocks, the embeddings of tokens are multiplied by key, query, and value matrices to give key, query, and value vectors for each token. Each of these matrices is made up of parameters that are determined by the model’s training, and the resulting vectors are much smaller in dimension than the embeddings. The dot product of the key and query vectors for each token is then computed, and the values of these products for each query is converted to a probability distribution using a function called softmax, which will be explained later. This allows the model to see which tokens are relevant to other tokens. Now, to update the values of the embedding vectors, the value vectors are added to the embedding vector of the relevant token. To better grasp how this works, consider the example sentence “The big green car was expensive.” Assuming the attention block in this example is adjusting nouns based on the adjectives that describe them. Then, the dot product between the key vectors for “big” and “green” and the query vector for “car” would be large, and be assigned a high probability after applying softmax. Then, the corresponding probabilities would be multiplied by the value vector for “big” and “green”, and added to the embedding vector for “car”. This yields an adjusted embedding vector that encodes a more detailed meaning for the word “car” in this sentence. It is important to note however, that for individual attention blocks, the key, query, and value matrices are unique, since for every kind of contextual updating, the parameters of the matrices would be different. Additionally, attention blocks usually are not assigned as specific as a purpose as seeing what adjectives affect nouns, and it is important to remember that tokens are not always complete words.
After passing through the attention block, each token’s embedding vector individually is passed through a feed forward network (FFN), to refine its values. FFN’s use associations learned in training data and applies it to the vector to give a more accurate representation of what the meaning is. Combined together, an attention block and a feed forward network create one transformer layer. Large language models then stack layers, essentially repeating the process over and over. Finally, after the last layer, the last embedding vector is multiplied by what is called an unembedding matrix. Each row of this matrix corresponds to an embedding vector in the model’s dictionary, so for ChatGPT this matrix would have 12,228 columns. The resulting vector after this multiplication is then normalized into a probability distribution using the aforementioned softmax. For context, a probability distribution contains values between 0 and 1, and all the values together must add up to 1. The way softmax works to achieve this, is the mathematical constant e is raised to the power of each entry of the vector, and then is divided by the sum of all these exponentiated values. Then, a constant called the temperature divides the value of the exponent, which allows adjustability to the probability. A higher temperature gives more weight (a higher probability) to the lower values , and a lower temperature gives more weight to the higher values. The model finally chooses the next word using this probability distribution, with methods varying by LLM.
To conclude how all of this works in a practical context where the LLM functions as a chatbot, the model is fed something like “A conversation between a helpful AI chatbot and a user is below:”, followed by the user’s input. The model then predicts what comes next and produces a text result.
The architecture of large language models is a marvel of modern computer science, combining tokenization, embeddings, and transformer layers to process and generate text. While this technology has serious applications, there are also concerns about the cost of this technology. First, the training data sets used for LLMs are made up of pentabytes of internet data, not all of which is copyright free. For example Anthropic, the creator of Claude, settled for $1.5 billion after a lawsuit was filed against them for using copyrighted materials to train their AI models. Secondly, the large matrix multiplication described throughout is highly computationally taxing, and the context size, which is the maximum amount of tokens an LLM can process in a single input, is still quite limited and not scalable. This also explains why some models, especially older ones, would forget previous things told to them and make mistakes. All in all however, the implications of this technology on our world socially make it crucial to understand how it works. Understanding the foundations of how LLMs work is crucial for the continued development of AI, and in understanding how we can make models more computationally efficient in an effort to mitigate environmental and legal effects.
Citations
[IBM Topic] V. Winland, What is self-attention?, IBM Think Topics (2024), available at [https://www.ibm.com/think/topics/self-attention](https://www.ibm.com/think/topics/self-attention).
[YouTube — 3Blue1Brown (Chapter 5)] G. Sanderson, Transformers, the tech behind LLMs, Deep Learning, Chapter 5, 3Blue1Brown (2024), video resource, available at [https://www.youtube.com/watch?v=wjZofJX0v4M](https://www.youtube.com/watch?v=wjZofJX0v4M).
[YouTube — 3Blue1Brown (Chapter 6)] G. Sanderson, Attention in transformers, step-by-step, Deep Learning, Chapter 6, 3Blue1Brown (2024), video resource, available at [https://www.youtube.com/watch?v=eMlx5fFNoYc](https://www.youtube.com/watch?v=eMlx5fFNoYc).
[YouTube — Syntax] CJ, LLMs Explained: Tokens, Embeddings, Transformers and More, Syntax (2024), video resource, available at [https://www.youtube.com/watch?v=YmLp8qe87A0](https://www.youtube.com/watch?v=YmLp8qe87A0).
[arXiv Paper] T. Mikolov, K. Chen, G. Corrado, and J. Dean, Efficient estimation of word representations in vector space, preprint, arXiv:1301.3781 [cs.CL] (2013).

