66 cards · AI & LLM Engineering · answer each one, then read the explanation. Your score tallies below.
0 / 66 answered · 0 correct
Link copied — send it to a friend!
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 001/066medium
In the Transformer architecture introduced by Vaswani et al. (2017) in "Attention Is All You Need," the dot products between queries and keys are divided by the square root of the key dimension before the softmax is applied. What is the primary reason for this scaling step?
AIt converts the attention weights into a probability distribution that sums to one across the sequence
BIt reduces the number of parameters needed in the query, key, and value projection matrices
CIt counteracts dot products growing large in magnitude for larger key dimensions, which would push the softmax into regions with extremely small gradients
DIt allows the same attention weights to be reused across all heads in multi-head attention
Correct answer: .
Vaswani et al. (2017) scale the raw dot products by the square root of the key dimension because, for larger key dimensions, the dot products between query and key vectors tend to grow large in magnitude; feeding these large values into the softmax pushes it into saturated regions where the gradient with respect to its inputs becomes extremely small, hurting learning. Dividing by the square root of the dimension keeps the values in a range where the softmax remains well-behaved. The option about producing a probability distribution describes what the softmax function itself does, not what the scaling step contributes before the softmax is applied. The option about reducing parameter counts is wrong because scaling is a fixed arithmetic operation on activations, not a change to the learned projection matrices. The option about reusing weights across heads is wrong because each attention head computes its own independent set of attention weights from its own query and key projections; scaling does not link heads together in any way.
Source: Vaswani et al., "Attention Is All You Need" (2017), arXiv:1706.03762, Section 3.2.1 (Scaled Dot-Product Attention)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 002/066easy
Vaswani et al. (2017) use multi-head attention instead of a single attention function operating on the full-dimensional queries, keys, and values. What was the main motivation for splitting attention into multiple heads?
AIt lets the model jointly attend to information from different representation subspaces at different positions, since a single attention head would average this into one weighted combination
BIt doubles the effective context window length by processing two halves of the input sequence in parallel
CIt removes the need for positional encoding, because each head learns to represent a different position offset
DIt ensures that every attention head produces an identical set of attention weights, providing an ensembling effect through redundancy
Correct answer: .
The paper motivates multi-head attention by noting that a single attention function averages over all the information relevant to a query into one weighted combination, which can wash out distinct relationships that exist at different representation subspaces and different positions. By projecting the queries, keys, and values into several smaller subspaces and running attention independently in each, then concatenating the results, the model can capture several different kinds of relationships simultaneously rather than blending them into one. The option about context window length is wrong because multi-head attention changes how a fixed-length sequence is processed, not how many tokens can be included in the sequence. The option about eliminating positional encoding is wrong because positional information is added separately to the input embeddings regardless of how many heads are used. The option describing identical weights across heads is wrong because the heads are given separate learned projection matrices specifically so they can learn to attend differently, not so they replicate each other.
Source: Vaswani et al., "Attention Is All You Need" (2017), arXiv:1706.03762, Section 3.2.2 (Multi-Head Attention)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 003/066easy
The original Transformer architecture adds sinusoidal positional encodings to the input token embeddings before the first attention layer. Why is this necessary given how self-attention itself processes a sequence?
ABecause the feed-forward sublayers in each block are recurrent and require explicit position indices to update their hidden state between positions
BBecause self-attention computes a weighted sum over all positions regardless of order, so without added positional information the model would treat the input as an unordered set of tokens
CBecause the token embedding layer cannot represent more than a fixed vocabulary size unless positional offsets are added to each embedding
DBecause positional encodings replace the residual connections that would otherwise be required around each sublayer
Correct answer: .
Self-attention computes, for each position, a weighted sum over the value vectors at every position in the sequence, and that computation is inherently permutation-invariant: reshuffling the input tokens and their positions would reshuffle the output in exactly the same way, with no built-in sense of order. Because the architecture contains no recurrence or convolution to encode sequence order implicitly, Vaswani et al. add sinusoidal positional encodings to the embeddings so the model can distinguish 'token A followed by token B' from 'token B followed by token A.' The option about recurrent feed-forward sublayers is wrong because the position-wise feed-forward network is a simple non-recurrent transformation applied independently to each position. The option about vocabulary size is wrong because vocabulary capacity is determined by the embedding table's size, unrelated to position information. The option about replacing residual connections is wrong because residual connections around each sublayer serve a separate purpose (easing gradient flow) and remain present alongside positional encoding, not instead of it.
Source: Vaswani et al., "Attention Is All You Need" (2017), arXiv:1706.03762, Section 3.5 (Positional Encoding)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 004/066medium
Sennrich et al. (2016), "Neural Machine Translation of Rare Words with Subword Units," introduced byte-pair encoding (BPE) as a tokenization method for handling rare and out-of-vocabulary words. How does the BPE algorithm build its subword vocabulary?
AIt assigns a fixed-length numeric code to each whole word in a predetermined dictionary, discarding any word not already present in that dictionary
BIt randomly samples character sequences of varying length from the training corpus until a target vocabulary size is reached
CIt splits text purely along whitespace and punctuation boundaries, performing no further segmentation within a word
DIt starts from individual characters and repeatedly merges the most frequent adjacent pair of symbols into a new symbol, learning a fixed number of merge operations to build a subword vocabulary
Correct answer: .
Sennrich et al. describe BPE as starting with each word represented as a sequence of characters plus an end-of-word symbol, with the initial symbol vocabulary consisting of all individual characters. The algorithm then repeatedly counts all adjacent symbol pairs across the training corpus, merges the single most frequent pair into a new symbol, and adds that symbol to the vocabulary; this repeats for a chosen number of merge operations, gradually building larger subword units out of frequent character sequences and even whole common words. This lets rare or unseen words be represented as sequences of known subwords instead of a single out-of-vocabulary token. The option describing whole-word dictionary codes is wrong because that is exactly the closed-vocabulary approach BPE was designed to avoid. The option about random sampling is wrong because merges are chosen deterministically by frequency, not randomly. The option describing pure whitespace/punctuation splitting is wrong because BPE explicitly performs further segmentation within words based on learned merges.
Source: Sennrich, Haddow & Birch, "Neural Machine Translation of Rare Words with Subword Units" (2016), arXiv:1508.07909, Section 3.2
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 005/066medium
Decoder-only autoregressive language models such as GPT (Radford et al., 2018) use masked self-attention during training. What does this masking do, and why is it necessary?
AIt prevents the model from attending to tokens more than a fixed window size away, purely to reduce the compute cost of attention
BIt forces every attention head to attend only to the current token itself, ignoring the representations at all other positions
CIt sets the attention scores for future positions to negative infinity before the softmax, so a token's representation depends only on itself and earlier tokens, matching the left-to-right generation process
DIt removes the value projection for positions beyond the current one, while still letting their key vectors influence the attention weights
Correct answer: .
Causal (masked) self-attention works by setting the attention scores between a given position and every later position to negative infinity before the softmax is computed, so after the softmax those later positions receive essentially zero weight. This means each position's output representation is built only from itself and the positions before it. This is necessary because these models are trained to predict the next token from only the preceding tokens, matching how the model must generate text left-to-right at inference time -- without this masking, the model could trivially 'see' the answer it is supposed to predict during training, which would not reflect real generation conditions. The option describing a fixed nearby window is wrong because causal masking blocks only future positions, not distant past ones. The option about attending only to the current token is wrong because earlier positions remain fully visible. The option about removing only the value projection while keeping keys visible is wrong because masking blocks the attention score itself, not one specific projection.
Source: Radford et al., "Improving Language Understanding by Generative Pre-Training" (2018); Vaswani et al., "Attention Is All You Need" (2017), arXiv:1706.03762, Section 3.2.3 (masked multi-head attention)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 006/066hard
Hoffmann et al. (2022), "Training Compute-Optimal Large Language Models" (the Chinchilla paper), trained over 400 models to study how model size and training data should be allocated for a fixed compute budget. What was their central finding?
AFor a fixed training compute budget, model parameters and training tokens should be scaled at roughly the same rate, implying substantially larger training datasets relative to model size than many earlier large models had used
BModel performance becomes essentially independent of the number of training tokens once the parameter count exceeds a few billion
CCompute-optimal training requires increasing model size much faster than training data, since parameter count is the dominant driver of capability
DThe optimal ratio of training tokens to parameters decreases as total training compute increases, so very large models need proportionally less data
Correct answer: .
Hoffmann et al. found that for compute-optimal training, model size and the number of training tokens should be scaled at roughly equal rates -- for every doubling of model size, training tokens should also roughly double, yielding a compute-optimal ratio of around 20 training tokens per parameter. Applying this, their 70-billion-parameter Chinchilla model, trained on about 1.4 trillion tokens, outperformed several much larger contemporary models that had been trained on comparatively fewer tokens relative to their size, showing that many earlier large models were undertrained relative to their parameter count. The option claiming performance becomes independent of token count past a few billion parameters is wrong because the paper's whole point is that additional training tokens keep mattering as models scale. The option claiming model size should scale much faster than data is wrong because it describes the pre-Chinchilla assumption the paper revised, not its finding. The option claiming the optimal token-to-parameter ratio decreases with scale is wrong because Hoffmann et al. found this ratio to be roughly constant across the compute budgets they studied.
Source: Hoffmann et al., "Training Compute-Optimal Large Language Models" (2022), arXiv:2203.15556
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 007/066medium
Ouyang et al. (2022), "Training Language Models to Follow Instructions with Human Feedback" (the InstructGPT paper), describes a three-stage pipeline for aligning a pretrained language model with human preferences. What are the three stages, in order?
APretrain on a labeled classification dataset, then deploy directly without further tuning, since pretraining alone is assumed to capture human preferences
BSupervised fine-tuning on human-written demonstrations, followed by training a reward model on human rankings of sampled model outputs, followed by reinforcement learning (PPO) that optimizes the policy against that reward model
CTrain a reward model on raw internet text first, then use it to filter the pretraining corpus before any supervised fine-tuning occurs
DFine-tune the model using only reinforcement learning against a fixed rule-based reward function, without using any human-labeled data at any stage
Correct answer: .
Ouyang et al. describe reinforcement learning from human feedback (RLHF) as a three-stage process applied on top of an already-pretrained language model. First, the model is supervised fine-tuned on a dataset of human-written demonstrations of desired outputs. Second, human labelers rank multiple model outputs for the same prompts by quality, and this comparison data is used to train a separate reward model that predicts which output humans would prefer. Third, the fine-tuned model is further optimized using Proximal Policy Optimization (PPO), treating the reward model's score as the reward signal to maximize. The option describing pretraining alone followed by direct deployment is wrong because it skips both the demonstration-based fine-tuning and the preference-based reward modeling entirely. The option starting with a reward model trained on raw internet text is wrong because the reward model is trained on human comparison rankings of model outputs, not used to pre-filter the pretraining corpus. The option describing pure rule-based reinforcement learning with no human data is wrong because both the demonstration data and the comparison rankings are human-generated and essential to the pipeline.
Source: Ouyang et al., "Training Language Models to Follow Instructions with Human Feedback" (2022), arXiv:2203.02155, Section 3
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 008/066easy
Each sublayer in the Transformer encoder and decoder (Vaswani et al., 2017) is wrapped with a residual connection before layer normalization is applied. What role do these residual connections play in training deep stacks of transformer blocks?
AThey reduce the total number of attention heads needed by allowing heads to share weights across different layers
BThey eliminate the need for a softmax normalization step within the self-attention mechanism
CThey replace the position-wise feed-forward sublayer with a simple identity function during inference, to speed up generation
DThey add each sublayer's input directly to its output before normalization, giving gradients a direct path backward through the addition and easing optimization of many stacked layers
Correct answer: .
In the Transformer, each sublayer's output is computed as the input added to the sublayer's own transformation of that input, with layer normalization applied to that sum; this addition is the residual (skip) connection. Because the input is added directly to the output, gradients computed at a later layer can flow backward through that addition largely unimpeded, rather than being forced entirely through the sublayer's transformation, which mitigates vanishing gradients and makes it practical to stack many such blocks and train them successfully. The option about sharing attention head weights across layers is wrong because residual connections operate on sublayer inputs and outputs, not on how heads are parameterized. The option about eliminating softmax is wrong because softmax remains a required step inside self-attention regardless of residual connections. The option about replacing the feed-forward sublayer with an identity function at inference is wrong because the feed-forward sublayer is still computed in full; the residual connection only adds its output to its input rather than bypassing the computation.
Source: Vaswani et al., "Attention Is All You Need" (2017), arXiv:1706.03762, Section 3.1
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 009/066easy
Ba, Kiros and Hinton (2016), "Layer Normalization," introduced a normalization technique widely used in transformer blocks. How does layer normalization compute its statistics, and how does this differ from batch normalization?
AIt normalizes each feature by computing statistics across every example in the mini-batch, exactly as batch normalization does, but applies the result at a different point in the network
BIt normalizes activations using running statistics collected only during a separate calibration pass performed after training has finished
CIt computes the mean and variance across the feature dimension for each individual training example, making it independent of batch size, whereas batch normalization computes statistics across the batch for each feature
DIt normalizes only the query and key projections used in self-attention, leaving the value projection and feed-forward outputs unnormalized
Correct answer: .
Ba et al. define layer normalization as computing the mean and variance across the feature dimension for a single training example, then using those statistics to normalize that example's activations; because the computation only ever looks within one example, it does not depend on how many examples are in a mini-batch. This differs from batch normalization, which instead computes the mean and variance for each feature across all examples in the current mini-batch. This batch-size independence is why layer normalization suits transformers and recurrent networks, where batch statistics can be unstable or awkward to define. The option describing batch-wide statistics is wrong because that describes batch normalization, not layer normalization. The option describing a post-training calibration pass is wrong because layer normalization's statistics are computed on the fly for every forward pass, during both training and inference. The option restricting normalization to only the query and key projections is wrong because layer normalization is applied to entire sublayer outputs, not to specific attention projections.
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 010/066easy
In addition to the self-attention sublayer, each Transformer block (Vaswani et al., 2017) contains a position-wise feed-forward network. How is this feed-forward network applied across the sequence?
AIt is applied identically and independently to each position in the sequence, typically expanding to a larger inner dimension through a nonlinearity before projecting back down
BIt combines information across all positions simultaneously in the same way self-attention does, making it functionally redundant with the attention sublayer
CIt only operates on the final position of the sequence, producing a single pooled representation for the entire input
DIt shares its weights with the token embedding layer so that no additional parameters are introduced by including it
Correct answer: .
Vaswani et al. describe the position-wise feed-forward network as consisting of two linear transformations with a nonlinearity between them, applied separately and identically to each position of the sequence -- the same learned weights are reused at every position, but each position's vector is transformed independently of the others. It typically expands the representation to a larger inner dimension before projecting it back down to the model's original dimension. This is distinct from self-attention, which is the sublayer responsible for mixing information across positions; the feed-forward sublayer instead adds per-position nonlinear transformation capacity. The option describing cross-position combination is wrong because that describes what self-attention does, not the feed-forward sublayer, which processes each position on its own. The option about operating only on the final position is wrong because it is applied at every position, not just the last one. The option about sharing weights with the embedding layer is wrong because the feed-forward network has its own separate learned parameters.
Source: Vaswani et al., "Attention Is All You Need" (2017), arXiv:1706.03762, Section 3.3
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 011/066hard
Shazeer et al. (2017), "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer," introduced a technique now used in some large language models to scale model capacity without a proportional increase in per-token compute. How does a sparsely-gated mixture-of-experts (MoE) layer achieve this?
AA gating network activates every expert sub-network for every input and averages their outputs, providing an ensembling effect at the cost of proportionally higher compute
BA trainable gating network routes each input to a sparse subset of expert feed-forward sub-networks, letting total parameter count scale far beyond what would be affordable if every expert were computed for every input
CEach expert is a full independent copy of the entire transformer stack, and the gating network selects which single copy runs the whole forward pass for a given input
DThe gating network is fixed at random initialization and never trained, relying purely on random routing to balance load across experts
Correct answer: .
Shazeer et al. introduce a sparsely-gated mixture-of-experts layer consisting of up to thousands of feed-forward expert sub-networks, paired with a trainable gating network that, for each input, selects and combines only a small sparse subset of those experts rather than running all of them. Because only the selected experts do any computation for a given input, the total number of parameters in the layer can grow enormously -- the paper reports capacity increases of over 1000x -- while the compute cost per input stays close to that of a much smaller dense network. The option describing activation of every expert with averaged outputs is wrong because that describes a dense ensemble, which is exactly the proportional-compute-cost approach this sparse gating avoids. The option describing experts as full copies of the entire transformer stack is wrong because the experts in this layer are feed-forward sub-networks, not complete duplicated models. The option describing an untrained, randomly fixed gate is wrong because the gating network's parameters are learned jointly with the rest of the model.
Source: Shazeer et al., "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer" (2017), arXiv:1701.06538
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 012/066easy
Extending a transformer-based language model's context window to accept much longer input sequences is architecturally expensive without further modifications to the standard attention mechanism. What property of self-attention, as defined by Vaswani et al. (2017), causes this?
ACompute and memory for self-attention scale linearly with sequence length, so extending the context window costs proportionally the same amount regardless of length
BThe context window is limited only by the size of the token vocabulary, not by any property of the attention computation itself
CSelf-attention's cost depends solely on the number of stacked layers in the model and is unaffected by how many tokens are in the input sequence
DSelf-attention computes a pairwise compatibility score between every two positions in the sequence, so both compute and memory scale quadratically with sequence length, making much longer context windows substantially more expensive
Correct answer: .
Standard self-attention, as defined by Vaswani et al., computes an attention score between every pair of positions in the input sequence by taking the dot product of each query with every key. For a sequence of length n, this means roughly n-squared score computations and an n-by-n matrix of attention weights must be produced and stored, so both the compute required and the memory needed for the attention matrix grow quadratically as sequence length increases. This is why simply increasing context length without architectural changes (such as sparse or approximate attention variants) becomes disproportionately expensive as sequences get longer. The option claiming linear scaling is wrong because it describes the cost profile that alternative, non-standard attention mechanisms aim to achieve, not the standard mechanism from the original paper. The option tying context length purely to vocabulary size is wrong because vocabulary size determines how many distinct tokens can be represented, not how many positions attention can efficiently process. The option claiming cost depends only on the number of layers is wrong because within each layer, cost still grows with sequence length regardless of depth.
Source: Vaswani et al., "Attention Is All You Need" (2017), arXiv:1706.03762, Section 4 (comparison of computational complexity per layer)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 013/066medium
In Section 3.4 of "Attention Is All You Need" (Vaswani et al., 2017), the authors describe sharing weights between two embedding layers and the pre-softmax linear transformation, then multiplying the embedding weights by the square root of the model dimension. What is the primary purpose of this weight-sharing scheme?
AIt lets the encoder and decoder share the same set of self-attention weights, so a single set of query/key/value matrices is reused in every layer of the network
BIt reduces the depth of the network by merging the final feed-forward sublayer directly into the output embedding matrix
CIt ties the same learned matrix to both converting tokens into vectors and converting the decoder's final vectors back into vocabulary logits, reducing the total parameter count and linking the input and output token representations
DIt guarantees that every possible output token receives an identical probability at initialization, which is later corrected during fine-tuning
Correct answer: .
Vaswani et al. tie a single learned weight matrix to both roles: converting input and output tokens into d_model-dimensional vectors and, transposed, converting the decoder's final vectors into logits over the vocabulary before the softmax; because one matrix serves both purposes, the model needs fewer total parameters than if the embedding and output projection were learned separately, and the input and output token representations remain linked to each other, an idea the paper credits to prior work on tying input and output embeddings. The embedding weights are then multiplied by the square root of the model dimension to keep their scale comparable to the positional encodings added to them. The option describing shared encoder/decoder self-attention weights is wrong because the tied weight matrix here is the token embedding/output-projection matrix, not the query, key, or value projections used inside attention, which remain separate learned parameters in each layer. The option describing merging the feed-forward sublayer into the embedding matrix is wrong because the position-wise feed-forward network is a separate sublayer with its own weights that this tying scheme does not touch. The option claiming every output token receives an identical initial probability is wrong because the tied matrix is initialized with the same random values used for embeddings, not with values that force a uniform output distribution.
Source: Vaswani et al., "Attention Is All You Need" (2017), arXiv:1706.03762, Section 3.4 (Embeddings and Softmax)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 014/066easy
The Transformer decoder (Vaswani et al., 2017) contains an "encoder-decoder attention" sublayer in addition to its own masked self-attention sublayer. Where do the queries, keys, and values for this encoder-decoder attention sublayer come from?
AThe queries come from the decoder's own previous sublayer, while the keys and values come from the output of the encoder stack, letting every decoder position attend over the entire input sequence
BThe queries, keys, and values all come from the encoder's final layer, and the decoder only reads the resulting attention output without contributing any of its own vectors
CThe queries and keys come from the encoder output while the values come from the decoder's own previous sublayer, reversing the usual roles of queries and keys
DThe queries, keys, and values all come from the decoder's own previous sublayer, identical to the masked self-attention sublayer that precedes it
Correct answer: .
In the encoder-decoder attention sublayer, Vaswani et al. take the queries from the output of the decoder's own preceding sublayer, while the keys and values come from the output of the encoder stack; this lets every position in the decoder attend over all positions in the input sequence, mirroring the query-key-value structure used in earlier sequence-to-sequence attention mechanisms. The option in which queries, keys, and values all come from the encoder is wrong because the decoder must still contribute its own queries so that each decoder position can ask a different question of the encoder output. The option that swaps queries/keys and values between encoder and decoder is wrong because Vaswani et al. specifically place the keys together with the values from the same source (the encoder) rather than splitting queries and keys across two different sources. The option describing queries, keys, and values all coming from the decoder is wrong because that describes the masked self-attention sublayer earlier in the decoder block, not the separate encoder-decoder attention sublayer that follows it.
Source: Vaswani et al., "Attention Is All You Need" (2017), arXiv:1706.03762, Section 3.2.3 (Applications of Attention in our Model)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 015/066hard
Su et al. (2021), "RoFormer: Enhanced Transformer with Rotary Position Embedding," introduce rotary position embedding (RoPE) as an alternative to adding sinusoidal positional vectors to token embeddings. How does RoPE encode positional information?
AIt concatenates a learned position index as an extra dimension appended to every query and key vector before the dot product is computed
BIt rotates each query and key vector by an angle that depends on the token's position and a per-dimension frequency, so that the dot product between a rotated query and key naturally depends on their relative distance rather than their absolute positions alone
CIt replaces the dot-product attention score with a lookup table indexed by the absolute position of the query, discarding the key vector's position entirely
DIt multiplies the attention output by a fixed sinusoidal mask after the softmax has already been applied, without altering the query or key vectors themselves
Correct answer: .
Su et al. encode position by treating pairs of dimensions within each query and key vector as coordinates and rotating them by an angle that is a function of the token's absolute position and a frequency that varies by dimension pair, before the query and key are used to compute attention scores. Because rotating both a query and a key vector by an amount tied to their own positions leaves the angle between them dependent only on the difference between those positions, the resulting dot product ends up encoding relative distance even though each vector was rotated using only its own absolute position, with the paper additionally noting that this relative dependency decays as tokens grow further apart. The option describing an appended learned position index is wrong because RoPE does not add any extra dimension to the vectors; it transforms the existing dimensions through rotation. The option describing a lookup table indexed only by the query's absolute position is wrong because RoPE's dot product depends on both vectors' positions through the same rotation mechanism, not a table keyed to one side only. The option describing a post-softmax sinusoidal mask is wrong because the rotation is applied directly to the query and key vectors before the dot product and softmax are computed, not to the attention output afterward.
Source: Su et al., "RoFormer: Enhanced Transformer with Rotary Position Embedding" (2021), arXiv:2104.09864
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 016/066medium
Press, Smith and Lewis (2021), "Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation," introduce ALiBi as a way to let a model trained on short sequences generalize to much longer ones at inference. How does ALiBi represent position, and how does this differ from adding sinusoidal or learned positional embeddings to the input?
AIt adds a learned positional embedding to each token exactly as the original Transformer does, but recomputes those embeddings after every training step so they extrapolate to unseen lengths
BIt removes positional information entirely, relying only on the order in which tokens are fed into the model during training to implicitly teach it sequence order
CIt replaces the query and key projections with position-specific weight matrices, so each position in the sequence uses a completely separate set of learned parameters
DIt adds no positional embeddings to the word embeddings at all; instead, it subtracts a static, non-learned penalty from each attention score that grows in proportion to the distance between the query and key positions, with the penalty's steepness set differently per attention head
Correct answer: .
Press, Smith and Lewis describe ALiBi as adding no positional embeddings to the word embeddings whatsoever; instead, after computing the raw query-key dot product, a static bias is subtracted that is proportional to the distance between the two positions, with each attention head using its own fixed slope for that penalty, so nearby tokens are penalized less than distant ones and the mechanism naturally extends to sequence lengths never seen during training. The option describing recomputed learned positional embeddings is wrong because ALiBi's bias is a fixed, non-learned function of distance, not a learned embedding table at all. The option describing reliance purely on feed order with no explicit signal is wrong because ALiBi does inject an explicit, though non-learned, positional signal directly into the attention score computation. The option describing position-specific query/key weight matrices is wrong because ALiBi keeps the same query and key projections at every position and instead modifies only the attention scores after the dot product, adding no new per-position parameters.
Source: Press, Smith & Lewis, "Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation" (2021), arXiv:2108.12409
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 017/066easy
GPT-2 and GPT-3's position-wise feed-forward sublayers use the Gaussian Error Linear Unit (GELU), defined by Hendrycks and Gimpel (2016), instead of the plain ReLU activation used in some earlier architectures. What distinguishes GELU from ReLU?
AGELU weights an input by the value of the standard Gaussian cumulative distribution function evaluated at that input, producing a smooth curve that can pass through small negative values rather than clamping every negative input to exactly zero
BGELU is identical to ReLU for all positive inputs and outputs a small fixed negative constant for every negative input, similar to Leaky ReLU
CGELU replaces the two linear transformations in the feed-forward sublayer with a single linear transformation, removing the nonlinearity between them entirely
DGELU is an activation function that operates only on the attention scores after the softmax, rather than within the position-wise feed-forward sublayer
Correct answer: .
Hendrycks and Gimpel define GELU as weighting an input by the value of the standard Gaussian cumulative distribution function evaluated at that same input, which produces a smooth, differentiable curve that, unlike ReLU, does not hard-clamp every negative input to exactly zero -- small negative inputs pass through scaled by a small positive weight rather than being zeroed out entirely. GPT-2 and GPT-3 use this activation between the two linear transformations of their position-wise feed-forward sublayers. The option describing a small fixed negative constant for every negative input is wrong because that describes Leaky ReLU's behavior, a different activation with a fixed negative slope, not GELU's Gaussian-weighted curve. The option describing removal of the nonlinearity is wrong because GELU is itself the nonlinearity inserted between the two linear transformations, not a replacement that eliminates nonlinearity. The option placing this activation after the softmax on attention scores is wrong because GELU is used inside the feed-forward sublayer's two linear transformations, entirely separate from the self-attention computation and its softmax.
Source: Hendrycks & Gimpel, "Gaussian Error Linear Units (GELUs)" (2016), arXiv:1606.08415; used in Radford et al., "Language Models are Unsupervised Multitask Learners" (GPT-2, 2019) and Brown et al., "Language Models are Few-Shot Learners" (GPT-3, 2020)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 018/066easy
Decoder-only language models such as GPT are pretrained on a simple self-supervised objective before any instruction tuning or reinforcement learning stage is applied. What is that pretraining objective?
APredicting whether two randomly paired sentences from the corpus originally appeared next to each other in the source text, using a binary classification loss
BReconstructing a small number of randomly masked-out tokens scattered throughout an otherwise-visible input sequence, using a masked-token classification loss
CPredicting the next token in a sequence given only the tokens that came before it, using a cross-entropy loss between the model's predicted probability distribution and the actual next token, repeated across every position in the training corpus
DPredicting a scalar reward score for an entire generated sequence, using a regression loss trained on human preference comparisons
Correct answer: .
GPT-style decoder-only models are pretrained with a next-token prediction objective: at every position in a training sequence, the model produces a probability distribution over the vocabulary conditioned only on the preceding tokens (enforced by causal masking), and a cross-entropy loss compares that predicted distribution against the actual next token in the corpus, averaged across all positions and all training sequences; this simple, self-supervised setup requires no labeled data beyond raw text. The option describing next-sentence-pair classification is wrong because that describes a separate auxiliary objective used by some encoder models, not the autoregressive next-token objective decoder-only GPT-style models are pretrained on. The option describing reconstruction of scattered masked tokens is wrong because that describes the masked-language-modeling objective used by bidirectional encoder models, which is incompatible with the strictly left-to-right causal masking decoder-only models use. The option describing a scalar reward regression trained on preference comparisons is wrong because that describes the separate reward-model training stage used later in an RLHF pipeline, not the initial self-supervised pretraining objective.
Source: Radford et al., "Improving Language Understanding by Generative Pre-Training" (2018), Section 3.1; Radford et al., "Language Models are Unsupervised Multitask Learners" (2019)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 019/066medium
When an autoregressive Transformer decoder generates text one token at a time, naively recomputing self-attention from scratch at every step would repeat a large amount of work. What does key-value (KV) caching do to avoid this, and what does it not need to recompute?
AIt stores the entire attention-score matrix from the very first generation step and reuses those exact scores unchanged for every later token, regardless of what new token is generated
BIt skips computing new query vectors for later tokens, reusing the query vector from the first generated token for every subsequent generation step
CIt discards the key and value vectors after each step and instead caches only the final output logits, replaying them directly for the next step
DIt stores the key and value vectors computed for every previously generated token so that, at each new step, only the new token's query, key, and value need to be computed, with that new query then attended over the cached keys and values, rather than recomputing keys and values for the whole sequence so far
Correct answer: .
Because each previously generated token's key and value vectors do not change as generation proceeds -- they depend only on that token's own representation, which is fixed once it has been produced -- an autoregressive decoder can cache those key and value vectors after they are first computed. At each new decoding step, only the newest token needs a fresh query, key, and value computed; that new query then attends over the cached keys and values from every earlier position plus its own, avoiding the need to recompute keys and values for the entire sequence from scratch at every step. The option describing reused unchanged attention scores from the first step is wrong because the attention scores themselves depend on the current query, which is different at every new step, so the scores cannot simply be replayed. The option describing reuse of the first token's query vector is wrong because every new token still needs its own freshly computed query to attend correctly; only the keys and values of earlier tokens are cached, not any query. The option describing caching only final output logits is wrong because caching logits alone would provide no way to compute attention scores for new tokens against earlier positions at all.
Source: Shazeer, "Fast Transformer Decoding: One Write-Head is All You Need" (2019), arXiv:1911.02150, Section 1 (Introduction)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 020/066medium
Xiong et al. (2020), "On Layer Normalization in the Transformer Architecture," compare two ways of placing layer normalization relative to each sublayer's residual connection: Post-LN, used in the original Transformer, and Pre-LN. What did they find about training stability, and what practical consequence follows for Post-LN training?
APre-LN and Post-LN produce identical gradient magnitudes at initialization, so the choice between them has no measurable effect on how training proceeds
BPost-LN, which applies layer normalization after the residual addition, produces expected gradients near the output layer that grow large at initialization, which the authors show makes a learning-rate warm-up stage necessary; placing layer normalization inside the residual block instead (Pre-LN) keeps gradients well-behaved at initialization without requiring warm-up
CPre-LN requires a longer warm-up stage than Post-LN because normalizing before each sublayer slows down how quickly gradient magnitudes stabilize during the first training steps
DThe placement of layer normalization only affects inference-time computation cost and has no bearing on gradient behavior or the need for a warm-up stage during training
Correct answer: .
Using mean field theory, Xiong et al. show that in the Post-LN Transformer -- where layer normalization is applied after the residual addition, as in the original architecture -- the expected gradients near the output layer are large at initialization, and they demonstrate that this is why a learning-rate warm-up stage has empirically been necessary to train Post-LN Transformers successfully. When layer normalization is instead placed inside each residual block, before the sublayer's transformation (Pre-LN), the authors show gradients are well-behaved at initialization, and Pre-LN Transformers can be trained without the warm-up stage while reaching comparable results in less training time. The option claiming identical gradient magnitudes for both placements is wrong because the paper's central contribution is precisely that the two placements behave very differently at initialization. The option claiming Pre-LN needs a longer warm-up is wrong because the paper's finding is the reverse: Pre-LN is the placement that removes the need for warm-up, not the one that requires more of it. The option claiming the placement only affects inference cost is wrong because the paper's analysis and experiments concern training-time gradient behavior and optimization stability, not inference-time computation.
Source: Xiong et al., "On Layer Normalization in the Transformer Architecture" (2020), arXiv:2002.04745
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 021/066easy
Brown et al. (2020), "Language Models are Few-Shot Learners," describe evaluating GPT-3 in a few-shot setting on many tasks. What does "few-shot" mean in this evaluation, and what happens to the model's weights during it?
AA small number of labeled examples for the task are used to fine-tune GPT-3's weights with a few additional gradient-descent steps before the model is evaluated on new examples
BA small number of task examples are included as text directly within the prompt given to the model at inference time, and GPT-3 is evaluated on new examples using this prompt alone, with no gradient updates or fine-tuning of its weights performed for the task
CGPT-3's weights are duplicated into several smaller copies, each fine-tuned on a few examples from a different task, and the copy that performs best on a validation set is kept
DA small, separate classifier head is attached to GPT-3 and trained from scratch on a few labeled examples, while the rest of the pretrained model is frozen
Correct answer: .
Brown et al. describe few-shot evaluation as conditioning GPT-3 purely through text: a handful of example input-output pairs for the task are written directly into the prompt, followed by a new input, and the model is asked to produce the corresponding output -- all of this happens through ordinary text generation at inference time, with no gradient updates, fine-tuning, or any change whatsoever to the model's weights performed for the task. The option describing a few additional gradient-descent steps is wrong because that describes traditional few-shot fine-tuning, which the paper explicitly contrasts with its in-context approach applied to GPT-3. The option describing duplicating and separately fine-tuning several model copies is wrong because only a single frozen copy of GPT-3 is used, conditioned solely through the prompt, not multiple task-specific fine-tuned copies. The option describing an attached, separately trained classifier head is wrong because no new parameters of any kind are added to or trained on top of GPT-3 for the task; the entire model remains exactly as pretrained.
Source: Brown et al., "Language Models are Few-Shot Learners" (2020), arXiv:2005.14165
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 022/066hard
Shazeer (2019), "Fast Transformer Decoding: One Write-Head is All You Need," introduces multi-query attention as a modification to standard multi-head attention aimed at speeding up autoregressive decoding. What does multi-query attention change relative to standard multi-head attention, and why does this help?
AIt reduces the number of query heads down to a single shared query head, while each head keeps its own separate key and value projections, cutting the number of query computations performed at each step
BIt removes the key and value projections entirely, having every attention head operate directly on the raw token embeddings instead of projected keys and values
CIt increases the number of key and value heads beyond the number of query heads, giving each query head access to several redundant copies of the same key and value vectors
DIt keeps multiple separate query heads but has all of them share a single set of key and value projections, shrinking the size of the cached keys and values that must be stored and read back at every decoding step, which reduces the memory-bandwidth cost that dominates incremental decoding
Correct answer: .
Shazeer's multi-query attention keeps the multiple, separately learned query heads of standard multi-head attention, but has all of those heads share one single set of key and value projections instead of each head computing and caching its own; because the memory bandwidth needed to read the cached keys and values back from memory at every incremental decoding step is the dominant cost of autoregressive generation, shrinking the cached keys and values down to a single shared copy per layer substantially reduces that bandwidth cost, at a modest expense in model quality. The option describing a single shared query head with separate keys and values per head is wrong because it inverts the paper's actual change: multiple query heads are kept, and it is the keys and values that are shared, not the queries. The option describing removal of key and value projections entirely is wrong because keys and values are still computed and cached, just from one shared projection rather than one per head. The option describing more key/value heads than query heads is wrong because multi-query attention reduces the key/value heads down to exactly one, rather than increasing them beyond the query head count.
Source: Shazeer, "Fast Transformer Decoding: One Write-Head is All You Need" (2019), arXiv:1911.02150
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 023/066medium
Dao et al. (2022), "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness," speed up the standard self-attention computation on GPUs without changing the attention mechanism's mathematical result. What kind of optimization does FlashAttention make, and what does it NOT do?
AIt reduces the number of reads and writes between the GPU's high-bandwidth memory and its much faster on-chip SRAM by tiling the computation and avoiding materializing the full attention-score matrix, while still computing exactly the same attention output as the standard formula, not an approximation of it
BIt approximates the full quadratic attention computation with a sparse or low-rank attention pattern, trading some accuracy in the attention output for a reduction in the number of floating-point operations performed
CIt reduces the number of floating-point operations attention requires by lowering the numerical precision of the query and key vectors, at the cost of a less numerically exact attention output
DIt restructures the attention computation to run entirely within CPU memory instead of GPU memory, trading GPU compute for cheaper CPU compute
Correct answer: .
Dao et al. observe that standard attention implementations are limited less by the number of floating-point operations they perform and more by how much data must be moved between the GPU's large but comparatively slow high-bandwidth memory (HBM) and its much smaller but much faster on-chip SRAM; FlashAttention restructures the computation using tiling, computing attention in blocks and using recomputation during the backward pass, so that it reads from and writes to HBM far less often, all while still computing exactly the same attention output the standard formula would produce -- it is explicitly an exact algorithm, not an approximation. The option describing a sparse or low-rank approximation is wrong because FlashAttention preserves the exact standard attention computation rather than substituting a different, approximate attention pattern. The option describing lower numerical precision as the source of the speedup is wrong because FlashAttention's gains come from reducing memory movement through tiling, not from reducing precision, and it does not sacrifice numerical exactness of the result. The option describing running entirely in CPU memory is wrong because FlashAttention's optimization is specifically about the memory hierarchy within the GPU itself (HBM versus on-chip SRAM), not about moving computation to the CPU.
Source: Dao et al., "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness" (2022), arXiv:2205.14135
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 024/066easy
Radford et al. (2019), "Language Models are Unsupervised Multitask Learners" (the GPT-2 paper), tokenize text using a byte-level variant of byte-pair encoding rather than applying BPE directly to Unicode characters. What problem does operating on bytes instead of Unicode characters solve?
AIt allows the tokenizer to skip the merge-based vocabulary-building process entirely, since every byte value can be mapped directly to a token without any merges
BIt increases the base vocabulary size well beyond 100,000 symbols before any merges are added, giving the model more starting granularity to work with
CIt keeps the base vocabulary, before any merges, small and fixed at 256 symbols (one for every possible byte value), avoiding the far larger base vocabulary that applying BPE directly to Unicode characters would require, and letting the model assign a probability to any input without ever needing an unknown-token symbol
DIt removes the need for any subword merging at all, since GPT-2 processes each byte independently as its own token throughout generation
Correct answer: .
Radford et al. note that applying BPE directly at the Unicode character level would start from a base vocabulary of over 130,000 symbols before any merges are learned, far larger than the 32,000-to-64,000-token vocabularies typically used; operating on bytes instead means the base vocabulary before any merges is fixed at exactly 256 symbols, one for every possible byte value, and the byte-pair merge process is then run on top of that small, fixed base, eventually growing GPT-2's vocabulary to 50,257 tokens. Because every possible input can be represented as some sequence of bytes, the model can assign a probability to any string at all, removing the need for an unknown-token fallback. The option describing skipping the merge process entirely is wrong because GPT-2 still runs the same iterative BPE merge algorithm on top of the byte-level base vocabulary; only the starting point changes. The option claiming the base vocabulary grows beyond 100,000 symbols is wrong because the entire point of operating on bytes is to keep that base vocabulary small, fixed at 256, not to expand it. The option claiming each byte remains its own separate token throughout generation is wrong because the merge operations combine frequently co-occurring byte sequences into larger subword tokens, exactly as in ordinary BPE.
Source: Radford et al., "Language Models are Unsupervised Multitask Learners" (2019), Section 2.2 (Input Representation)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 025/066medium
Ainslie et al. (2023), "GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints," introduce grouped-query attention as an intermediate design between standard multi-head attention and multi-query attention. How does GQA position itself between these two extremes?
AIt keeps a single shared key and value head for the entire attention layer, exactly as multi-query attention does, but restores full quality by giving every query head its own independent query projection matrix
BIt divides the query heads into a fixed number of groups, with all query heads in a group sharing one key and value head; setting the group count equal to the number of query heads recovers standard multi-head attention, and setting it to one recovers multi-query attention
CIt keeps every query head paired with its own key and value head as in standard multi-head attention, but reduces the total number of query heads to cut memory bandwidth
DIt replaces the key and value projections with a single shared feed-forward network applied after attention, removing the need for separate key and value weight matrices altogether
Correct answer: .
Grouped-query attention divides a layer's query heads into a fixed number of groups and has every query head within a group share one key and value head, rather than each query head owning its own key and value projection as in standard multi-head attention, or all query heads sharing a single key and value head as in multi-query attention. This group count is the tunable knob that lets GQA interpolate between the two extremes: setting the number of groups equal to the number of query heads recovers ordinary multi-head attention exactly, while setting it to one recovers multi-query attention exactly, and any intermediate group count trades off some of multi-head attention's quality for much of multi-query attention's memory-bandwidth savings during autoregressive decoding. The option describing a single shared key and value head with independently-projected query heads is simply describing multi-query attention with an invented extra detail, not the grouped design GQA actually introduces. The option describing a reduction in the number of query heads themselves is backwards: GQA leaves the query heads and their projections untouched and instead reduces the number of distinct key and value heads. The option describing a shared feed-forward network replacing the key and value projections describes a mechanism the paper does not use at all; GQA still computes ordinary key and value projections, just fewer of them, shared across groups of query heads.
Source: Ainslie et al., "GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints" (2023), arXiv:2305.13245
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 026/066hard
Kaplan et al. (2020), "Scaling Laws for Neural Language Models," and Hoffmann et al. (2022), "Training Compute-Optimal Large Language Models" (the Chinchilla paper), both fit power-law relationships between loss, model size, and training data, but reached different practical recommendations for how to spend a fixed compute budget. What is the key difference between their conclusions?
AKaplan et al. concluded model size and data should scale in roughly equal proportion as compute grows, while Hoffmann et al. concluded model size should grow much faster than data
BBoth papers reached identical conclusions about compute-optimal allocation; Hoffmann et al. only confirmed Kaplan et al.'s original recipe with a larger set of models
CKaplan et al. found that data size mattered more than model size for a fixed compute budget, while Hoffmann et al. found the opposite, that model size should be prioritized above all else
DKaplan et al.'s fitted power laws recommended training very large models on a comparatively modest amount of data, stopping well short of convergence, whereas Hoffmann et al.'s larger and more careful set of training runs found that model size and training data should instead be scaled in roughly equal proportion, implying many contemporary large models were oversized and undertrained relative to their compute budgets
Correct answer: .
Kaplan et al.'s original power-law fits, based on a smaller and less systematically-swept set of training runs, recommended spending a fixed compute budget mostly on making the model larger, training on a comparatively modest amount of data and stopping well short of the point where the model's loss on that data would have converged. Hoffmann et al. later trained a far larger and more carefully controlled sweep of models and found that Kaplan et al.'s fitted curves had systematically under-weighted the value of training data: compute-optimal training instead scales model size and training-data size in roughly equal proportion as the compute budget grows. Because most large models trained under the earlier recipe followed Kaplan et al.'s guidance, Hoffmann et al.'s finding implied that many of those models, though larger, were undertrained relative to what a compute-optimal allocation of the same budget would have used, and a smaller model trained on more data could match or beat one of them. The option claiming Kaplan et al. favored equal scaling and Hoffmann et al. favored model-size-heavy scaling has the two papers' conclusions swapped. The option claiming the two papers reached identical conclusions ignores the entire reason the Chinchilla paper is considered a correction to the earlier scaling recipe. The option claiming Kaplan et al. prioritized data size over model size states the opposite of what that paper's fitted allocation curve recommended.
Source: Kaplan et al., "Scaling Laws for Neural Language Models" (2020), arXiv:2001.08361; Hoffmann et al., "Training Compute-Optimal Large Language Models" (2022), arXiv:2203.15556
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 027/066easy
Kudo and Richardson (2018), introducing the SentencePiece toolkit, implement a Unigram language model tokenization algorithm as an alternative to byte-pair encoding (BPE) for building a subword vocabulary. How does the Unigram algorithm build that vocabulary, in contrast to BPE?
AIt starts from a large seed vocabulary of candidate subwords and iteratively removes the subwords that contribute least to the likelihood of the training corpus under a unigram language model, continuing until the target vocabulary size is reached, rather than building the vocabulary up from individual characters by repeatedly merging the most frequent adjacent pair
BIt builds the vocabulary bottom-up by repeatedly merging the two most frequent adjacent symbols into a new symbol, identical to BPE, but differs only in how the final tokenizer chooses a single segmentation at inference time
CIt requires the input text to already be split into whitespace-separated words before subword modelling begins, unlike BPE which can operate on raw, undelimited text
DIt assigns every character its own fixed, permanent token and never combines characters into larger subword units, guaranteeing a constant vocabulary size regardless of corpus
Correct answer: .
SentencePiece's Unigram algorithm works in the opposite direction from byte-pair encoding: it begins with a large seed vocabulary of candidate subword pieces, then repeatedly measures how much the overall likelihood of the training corpus under a unigram language model over subwords would drop if each piece were removed, and prunes the least useful pieces until the vocabulary shrinks to the target size. BPE instead starts from individual characters and builds upward, repeatedly merging whichever adjacent pair of symbols is most frequent in the training data into a single new symbol, growing the vocabulary from the bottom up rather than shrinking it from the top down. The option describing bottom-up merging of the most frequent adjacent symbols is actually describing BPE's own mechanism, not what distinguishes Unigram from it. The option claiming SentencePiece requires text to already be split into whitespace-separated words gets the comparison backwards: a major point of the SentencePiece paper is that it can train directly on raw, undelimited sentences, unlike tools that assume pre-tokenized word boundaries. The option claiming every character keeps its own permanent token and is never combined describes neither algorithm; both BPE and Unigram build multi-character subword units.
Source: Kudo & Richardson, "SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing" (2018), arXiv:1808.06226
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 028/066medium
Rafailov et al. (2023), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model," propose DPO as an alternative to the reward-model-plus-PPO pipeline used in InstructGPT-style RLHF. What does DPO change about this pipeline?
AIt replaces the pairwise human-preference comparisons used to train the reward model with a fully automated LLM-as-judge scoring system, while keeping the separate reward model and the PPO reinforcement-learning step unchanged
BIt trains the reward model exactly as InstructGPT does, but replaces the PPO optimizer with a simpler supervised fine-tuning step on the reward model's top-scoring completions
CIt uses a closed-form solution to the KL-regularized reward-maximization objective to reparameterize the reward directly in terms of the policy's own output probabilities relative to a reference policy, letting the model be optimized directly on preference pairs with a simple classification-style loss, without ever training a separate reward model or running an online reinforcement-learning sampling loop
DIt removes the need for any human or model preference data at all, instead deriving alignment purely from the pretraining objective
Correct answer: .
DPO's key move is algebraic: it takes the closed-form solution to the KL-regularized reward-maximization problem that RLHF is ultimately trying to solve, and uses it to express the reward function directly in terms of the ratio between the policy being trained and a fixed reference policy. Substituting this reparameterized reward into the standard preference-modelling loss produces a simple classification-style objective that can be optimized directly on pairs of preferred and dispreferred completions, with no separate reward model ever trained and no online reinforcement-learning sampling loop such as PPO ever run. The option describing an automated LLM-as-judge replacing human comparisons while keeping the reward model and PPO step intact misidentifies what DPO removes; DPO's contribution is eliminating the reward model and the RL step themselves, not just changing how preference labels are collected. The option describing supervised fine-tuning on the reward model's top-scoring completions describes a best-of-n-style distillation approach, not DPO, and still keeps a separate reward model, which DPO does not train at all. The option claiming DPO removes the need for preference data entirely contradicts the method's basic setup: DPO is trained directly on preference pairs; it just removes the intermediate reward-model and RL stages, not the preference data itself.
Source: Rafailov et al., "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023), arXiv:2305.18290
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 029/066hard
Wei et al. (2022), "Emergent Abilities of Large Language Models," reported that model performance on certain tasks stays near chance level until a threshold scale, then rises sharply. Schaeffer et al. (2023), "Are Emergent Abilities of Large Language Models a Mirage?," challenge this framing. What is Schaeffer et al.'s core argument?
AThey show that the sharp jumps disappear entirely once larger models are trained, proving that Wei et al.'s reported tasks were measurement errors caused by insufficient training data
BThey argue that the apparent sharp, unpredictable jumps are largely an artifact of using nonlinear or discontinuous metrics, such as strict exact-match accuracy on multi-step tasks; when the same underlying model outputs are instead scored with a smoother, partial-credit metric, performance improves gradually and predictably with scale rather than jumping
CThey argue that emergent abilities are real and predictable in advance from a model's parameter count alone, and propose a formula that lets any lab compute the exact scale at which a new ability will appear before training a model
DThey argue that emergent abilities only ever appear in decoder-only architectures and never in encoder-decoder models, regardless of how performance is measured
Correct answer: .
Schaeffer et al.'s central claim is a measurement critique, not a claim that emergence never occurs at all: they argue that many of the sharp, unpredictable jumps Wei et al. reported arise because the tasks in question were scored with metrics such as strict exact-match accuracy, which are nonlinear or outright discontinuous functions of a model's per-token error rate on multi-step problems, so a smoothly improving underlying model can still show what looks like a sudden jump on that metric. When Schaeffer et al. re-score the same models' outputs using smoother, partial-credit metrics, the previously sharp jumps become gradual and predictable trends instead, which supports interpreting emergence as substantially a property of the chosen metric rather than a discontinuity in the model itself. The option claiming the jumps disappear once larger models are trained mischaracterizes the argument as being about model scale rather than about metric choice. The option claiming Schaeffer et al. propose a formula that predicts, in advance and from parameter count alone, exactly when any new ability will appear overstates their contribution well beyond what the paper claims; their point is that the timing looked sharp mainly due to measurement choice, not that it becomes precisely predictable from parameter count. The option restricting the argument to decoder-only architectures invents an architectural claim that does not appear in either paper.
Source: Wei et al., "Emergent Abilities of Large Language Models" (2022), arXiv:2206.07682; Schaeffer et al., "Are Emergent Abilities of Large Language Models a Mirage?" (2023), arXiv:2304.15004
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 030/066easy
Fedus, Zoph and Shazeer (2021), "Switch Transformer: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity," simplify the sparsely-gated mixture-of-experts design that Shazeer et al. (2017) had used with a top-k gating function. What change does the Switch Transformer make to expert routing?
AIt removes expert routing altogether and instead sends every token through every expert, then averages all of the experts' outputs
BIt increases the number of experts activated per token from Shazeer et al.'s original two experts to a much larger fixed number, such as sixteen, to improve output quality
CIt replaces the learned gating network with a fixed, hand-designed rule that assigns tokens to experts based on the token's position in the sequence rather than its content
DIt routes each token to exactly one expert (top-1 routing) instead of combining outputs from several experts per token, reducing routing computation and communication cost while relying on a load-balancing loss to keep experts evenly utilized
Correct answer: .
The Switch Transformer's main simplification of the earlier sparsely-gated mixture-of-experts design is to route each token to exactly one expert instead of combining the weighted outputs of several experts per token, cutting the routing computation and the cross-device communication needed to gather multiple experts' outputs; a differentiable load-balancing auxiliary loss is added to keep tokens spread roughly evenly across experts despite the router now making a single hard choice per token. The option describing every token passing through every expert and then averaging all outputs describes a dense model with no sparsity at all, the opposite of what makes a mixture-of-experts layer computationally cheap relative to its total parameter count. The option describing an increase to a larger fixed number of experts activated per token, such as sixteen, moves in the opposite direction from what the paper actually does, which is to reduce the number of activated experts per token to one. The option describing a fixed, position-based routing rule invents a mechanism the paper does not use; the router remains a learned, content-based gating network, it is simply restricted to choosing a single expert per token rather than a weighted combination of several.
Source: Fedus, Zoph & Shazeer, "Switch Transformer: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity" (2021), arXiv:2101.03961
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 031/066easy
Beltagy, Peters and Cohan (2020), "Longformer: The Long-Document Transformer," replace full self-attention with a combination of attention patterns to process much longer sequences efficiently. What is the core mechanism behind Longformer's efficiency gain?
AMost tokens use a fixed-size sliding-window attention, attending only to nearby tokens within that window, which reduces the memory and compute cost from quadratic in sequence length to linear; a small number of tokens are additionally given full global attention to preserve task-relevant long-range information
BEvery token still attends to every other token exactly as in standard self-attention, but the attention scores are computed in a lower floating-point precision to save memory, with no change to which tokens attend to which
CTokens are grouped into fixed-size non-overlapping blocks, and attention is computed only between tokens in the same block, with no mechanism for any token to see information outside its own block
DThe sequence is first compressed into a much shorter fixed-length summary using a separate pooling network, and standard full attention is then applied only to that shortened summary
Correct answer: .
Longformer's efficiency gain comes from replacing full attention, where every token attends to every other token at quadratic cost in sequence length, with a fixed-size sliding-window attention pattern for most tokens, so each token only attends to a bounded number of nearby neighbors and the cost grows linearly with sequence length instead; a small set of task-specific tokens, such as a classification token, are additionally given full global attention to every other token, so the model still has a channel for long-range information to flow through despite most positions only seeing a local window. The option describing lower floating-point precision with an unchanged attention pattern describes a numerical-precision optimization unrelated to Longformer's actual contribution, which is about which token pairs are attended to at all, not the numeric format used. The option describing fixed non-overlapping blocks with absolutely no mechanism for information to cross block boundaries omits the global-attention tokens that Longformer specifically adds to avoid exactly that limitation. The option describing compressing the sequence with a separate pooling network before applying full attention describes a different long-context approach entirely, not the sliding-window-plus-global-attention design Longformer introduces.
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 032/066medium
Chen et al. (2023), "Extending Context Window of Large Language Models via Positional Interpolation," extend the usable context length of RoPE-based pretrained models such as LLaMA with only minimal fine-tuning. What does their Position Interpolation method actually do to the model's position indices?
AIt extrapolates position indices beyond the range seen during pretraining, feeding the model position values larger than any it encountered in training, and relies on RoPE's periodicity to generalize correctly to these unseen positions
BIt discards the rotary position embedding scheme entirely and replaces it with the original sinusoidal position encoding from Vaswani et al. (2017), which the authors show generalizes to longer sequences without any fine-tuning
CIt linearly rescales the new, longer sequence's position indices down so that the largest index still falls within the range the model saw during pretraining, then fine-tunes briefly on this rescaled range, avoiding the out-of-distribution position values that naive extrapolation would produce
DIt increases the model's embedding dimension so that each position can be represented with more precision, allowing longer sequences to be distinguished without changing how position indices are computed
Correct answer: .
Position Interpolation's core idea is to linearly rescale the position indices of a longer input sequence down so that even the largest index still lands within the range of position values the model actually saw during pretraining, and then briefly fine-tune the model on sequences using these rescaled positions; because the model never has to process a position value larger than anything in its training distribution, this avoids the extremely large, out-of-distribution attention scores that naive extrapolation beyond the trained range produces, which the paper shows can badly break the attention mechanism. The option describing feeding the model position values larger than any seen in training, relying on RoPE's periodicity to cope, describes the naive extrapolation approach the paper explicitly identifies as producing catastrophic attention scores, essentially the failure mode Position Interpolation is designed to avoid, not the method itself. The option describing discarding RoPE for the original sinusoidal encoding invents a change the paper does not make; the method is specifically designed to extend RoPE-based models, not replace their position embedding scheme. The option describing an increase in embedding dimension to add positional precision is not how the paper extends context length at all; it rescales existing position indices rather than changing model dimensions.
Source: Chen et al., "Extending Context Window of Large Language Models via Positional Interpolation" (2023), arXiv:2306.15595
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 033/066easy
Leviathan, Kalman and Matias (2023), "Fast Inference from Transformers via Speculative Decoding," speed up autoregressive generation from a large target model without changing its output distribution. How does their method achieve this?
AIt replaces the large target model with a smaller model for the entire generation, only occasionally calling the large model to spot-check a sample of the output afterward, which changes the output distribution slightly but saves the most compute
BA smaller, cheaper draft model proposes several candidate next tokens; the large target model then verifies these candidates in a single parallel forward pass and accepts the longest prefix consistent with its own probability distribution via a rejection-sampling scheme, so the final output is distributed exactly as if the large model had generated every token itself, just with fewer serial large-model calls
CIt caches the large model's key and value tensors from previous generation steps so that they never need to be recomputed at later steps, removing redundant self-attention computation
DIt reorders the large model's attention computation to be more IO-aware, fusing operations to avoid writing large intermediate attention matrices to slow GPU memory, without changing which tokens are generated
Correct answer: .
Speculative decoding uses a small, cheap draft model to propose several candidate next tokens at once, then runs the large target model once, in parallel, over that whole candidate sequence to obtain the target model's own probabilities for each position; a rejection-sampling procedure then accepts the longest prefix of draft tokens that is consistent with the target model's distribution and, when a draft token is rejected, resamples correctly from the target model's distribution at that position, so the tokens produced are distributed exactly as if the target model alone had generated every one of them, just computed with fewer sequential large-model forward passes. The option describing the large model only occasionally spot-checking a small model's output describes an approximate scheme that would change the output distribution, which contradicts the paper's explicit goal of matching the large model's output exactly. The option describing caching previously-computed key and value tensors so they are never recomputed describes key-value caching, a different and complementary inference optimization, not speculative decoding. The option describing IO-aware fusion of attention operations to avoid writing large intermediate matrices to slow memory describes FlashAttention, another distinct optimization that changes how attention is computed, not how tokens are proposed and verified.
Source: Leviathan, Kalman & Matias, "Fast Inference from Transformers via Speculative Decoding" (2023), arXiv:2211.17192
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 034/066easy
Hinton, Vinyals and Dean (2015), "Distilling the Knowledge in a Neural Network," train a small "student" network using the outputs of a larger, already-trained "teacher" network. What specifically does the student learn to match, and why does raising the softmax temperature matter?
AThe student is trained only on the teacher's single most likely predicted class for each training example, exactly as it would be trained on the original ground-truth hard labels, and temperature is only used to speed up the teacher's inference
BThe student learns to reproduce the teacher's internal hidden-layer activations exactly, layer by layer, and temperature controls how many of the teacher's layers the student is required to match
CThe student is trained to minimize the difference between its own raw, un-normalized output scores and the teacher's raw output scores, with temperature having no effect on training since it is only applied at test time
DThe student is trained to match the teacher's full "soft target" probability distribution over all classes, not just the single correct label; raising the softmax temperature when computing both the teacher's and the student's distributions during training softens the probabilities, revealing the relative probabilities the teacher assigns to incorrect classes ("dark knowledge") that a hard label alone would hide, and the student reverts to a temperature of 1 at deployment
Correct answer: .
Knowledge distillation trains the small student network to reproduce the large teacher network's full soft-target probability distribution across all classes, not merely the single class the teacher ranks highest; computing both the teacher's and the student's output distributions at a raised softmax temperature during this training softens the probabilities enough to expose the relative likelihoods the teacher assigns to the classes it considers plausible but incorrect, information the paper calls dark knowledge that a single hard label could never convey, and after training the student is deployed using the normal temperature of one. The option describing training only on the teacher's single most-likely class, identical to training on ground-truth hard labels, misses the entire soft-target mechanism the paper introduces; that approach would give the student no more information than ordinary supervised training already provides. The option describing matching hidden-layer activations layer by layer, with temperature controlling how many layers must match, describes a different, feature-based distillation strategy, not the output-distribution matching Hinton et al. actually propose. The option describing matching raw, un-normalized output scores while claiming temperature has no effect during training gets the mechanism backwards: temperature is applied precisely during training, to both networks' distributions, and removing it removes the extra information distillation is designed to transfer.
Source: Hinton, Vinyals & Dean, "Distilling the Knowledge in a Neural Network" (2015), arXiv:1503.02531
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 035/066medium
Zhang and Sennrich (2019), "Root Mean Square Layer Normalization," propose RMSNorm as a simplification of LayerNorm later adopted by architectures such as LLaMA and Mistral. What does RMSNorm remove relative to standard LayerNorm, and what property does it rely on instead?
ARMSNorm removes the mean-centering (re-centering) step and does not subtract the mean of the summed inputs; it instead rescales activations only by their root-mean-square value, keeping re-scaling invariance while dropping re-centering invariance
BRMSNorm removes the learnable scale (gain) parameter entirely, normalizing activations using only their mean and variance exactly as LayerNorm does but without any parameters left to learn
CRMSNorm removes normalization from all but the final transformer block, applying LayerNorm's full mean-and-variance computation only once at the network's output instead of at every block
DRMSNorm removes normalization within a layer and instead normalizes each token's embedding relative to every other token's embedding in the same batch, a form of batch normalization applied along the sequence axis
Correct answer: .
RMSNorm keeps only the re-scaling half of LayerNorm's computation: it divides each activation by the root-mean-square of the summed inputs to that layer, without first subtracting the mean, so it never computes or removes the mean statistic that ordinary LayerNorm needs for re-centering. Zhang and Sennrich hypothesized that the re-centering invariance LayerNorm provides is not essential to its benefits, and that dropping it yields comparable model quality while cutting the normalization computation, which is why RMSNorm appears in later architectures such as LLaMA, T5, and Mistral. The option claiming RMSNorm drops the learnable gain parameter is wrong because RMSNorm still keeps a learned per-feature scale; what disappears is the centering step and its associated bias parameter, not the multiplicative gain. The option describing normalization applied only at the final block misdescribes both LayerNorm and RMSNorm, which are applied at every normalization point throughout the network, not just once at the end. The option describing normalization across tokens in a batch confuses RMSNorm with batch normalization; RMSNorm, like LayerNorm, normalizes across the feature dimension of a single token's own activations, independent of other tokens or batch members.
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 036/066easy
Mikolov et al. (2013) introduced word2vec, which learns a single fixed vector for each vocabulary word from a large text corpus. Why do transformer-based language models rely on contextual embeddings instead of word2vec-style static embeddings alone?
AStatic embeddings like word2vec cannot be trained on large text corpora at all, so no fixed-vector representation of word meaning can be learned without first having a transformer architecture available
BA static embedding assigns a word exactly one vector regardless of context, so a polysemous word keeps the same representation in every sentence; self-attention instead produces a representation for each token that is computed from the other tokens surrounding it, letting the same word take on different representations in different contexts
CStatic embeddings are only capable of representing nouns, while contextual embeddings are required to represent verbs, adjectives, and every other part of speech
DContextual embeddings are simply static embeddings computed with a larger vocabulary size, so the difference is purely how many words the vocabulary contains rather than how each vector is computed
Correct answer: .
Word2vec assigns each vocabulary word exactly one learned vector, fixed once training finishes, so a polysemous word such as one meaning both a river bank and a financial bank gets a single blended representation used everywhere it appears, regardless of the sentence around it. Self-attention breaks that limitation because every token's output representation at each layer is computed as a function of the other tokens it attends to, so the same word produces a different vector depending on its surrounding context, letting the model distinguish the different senses a static embedding table could never separate. The option claiming static embeddings cannot be trained on large corpora at all is factually backwards, since word2vec was specifically designed to train efficiently on very large corpora using only shallow, non-transformer neural networks. The option restricting static embeddings to nouns is invented; word2vec assigns a vector to every vocabulary token regardless of part of speech, static or contextual. The option claiming the only difference is vocabulary size ignores the actual mechanism entirely: the distinguishing feature is that contextual embeddings are computed per-occurrence from surrounding tokens, not that more words are covered.
Source: Mikolov et al., "Efficient Estimation of Word Representations in Vector Space" (2013), arXiv:1301.3781; Vaswani et al., "Attention Is All You Need" (2017)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 037/066easy
Holtzman et al. (2019), "The Curious Case of Neural Text Degeneration," observed that greedy decoding and beam search from a language model often produce bland, repetitive text even though the model assigns that text high likelihood. What does their proposed nucleus (top-p) sampling method do differently at each decoding step?
AIt always samples the single most probable token at each step, identical to greedy decoding, and then re-ranks the resulting full sequence afterward using a separate language model
BIt restricts sampling to a fixed number of the k highest-probability tokens at every step, where k is a constant chosen once before generation begins and never changes
CIt sorts tokens by probability and samples from the smallest set of tokens whose cumulative probability reaches a chosen threshold p, so the size of the candidate set shrinks when the model is confident and grows when the model is uncertain, avoiding both the unreliable low-probability tail and an arbitrary fixed-size cutoff
DIt disables sampling entirely during generation and instead deterministically outputs the sequence with the single highest total log-probability across the whole output, computed exactly via dynamic programming
Correct answer: .
Nucleus sampling first sorts the vocabulary by the probability the model assigns at that decoding step, then adds tokens from the top down until the cumulative probability first reaches the chosen threshold p, and samples the next token only from that dynamically sized set; because the set expands when the model spreads probability across many plausible continuations and shrinks when the model is confident in a narrow set of options, nucleus sampling avoids drawing from the long, unreliable low-probability tail that Holtzman et al. tie to greedy and beam search's tendency toward degenerate, repetitive text, without imposing an arbitrary fixed cutoff on how many tokens are eligible. The option describing sampling the single most probable token and re-ranking afterward with another model matches neither greedy decoding's own mechanism nor nucleus sampling, which never collapses to one deterministic choice by design. The option describing a fixed number k of top tokens chosen once before generation describes top-k sampling, the alternative the paper explicitly contrasts nucleus sampling against, since a constant k is either too restrictive when the model is uncertain or too permissive when it is confident. The option describing an exact, deterministic highest-log-probability search over the whole sequence describes neither sampling nor how beam search itself actually works, since beam search is an approximate search, not an exact one.
Source: Holtzman, Buys, Du, Forbes & Choi, "The Curious Case of Neural Text Degeneration" (2019), arXiv:1904.09751
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 038/066hard
Loshchilov and Hutter (2017), "Decoupled Weight Decay Regularization," introduce AdamW to fix how the original Adam optimizer interacts with weight decay. What specifically does AdamW change relative to plain Adam trained with an L2 penalty added to the loss?
AAdamW removes weight decay from the optimizer entirely, relying only on early stopping to prevent overfitting during large-scale pretraining
BAdamW replaces Adam's per-parameter adaptive learning rates with a single global learning rate shared by every parameter, which is what allows weight decay to work correctly
CAdamW applies weight decay only to the bias terms of the model and excludes every weight matrix from decay, reversing which parameters L2 regularization normally targets
DAdamW applies weight decay as a separate, direct shrinkage of the parameters at each update step, outside of the gradient-based adaptive moment estimation; with plain Adam, adding an L2 penalty to the loss instead folds the decay term into the gradient, where Adam's per-parameter adaptive scaling then distorts its effective strength, weakening decay on frequently-updated parameters and strengthening it on rarely-updated ones
Correct answer: .
Loshchilov and Hutter show that for adaptive optimizers such as Adam, adding an L2 penalty to the loss is not equivalent to true weight decay the way it is for plain SGD: the L2 term gets folded into the gradient before Adam's per-parameter adaptive scaling is applied, so parameters with large accumulated squared gradients (frequently or strongly updated ones) end up with their effective decay weakened, while parameters with small gradient history get decayed more strongly than intended, an uneven and unintended effect. AdamW fixes this by decoupling weight decay from the gradient-based update entirely, applying it as a separate multiplicative shrinkage directly to the parameters at each step, so every parameter receives the same, uniform decay regardless of its gradient history. The option claiming AdamW removes weight decay entirely is backwards, since decoupled decay is AdamW's defining feature, not its absence. The option claiming AdamW switches to a single global learning rate misdescribes the fix; AdamW keeps Adam's per-parameter adaptive learning rates for the gradient-based update and only decouples the separate decay term. The option describing decay applied only to bias terms and excluding weight matrices inverts common practice, where weight decay is typically applied to weight matrices and often excluded from biases and normalization parameters, not the reverse.
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 039/066easy
Large language model pretraining runs such as GPT-3's (Brown et al., 2020) commonly use a learning rate schedule that starts near zero, rises linearly for an initial warmup period, and then decays smoothly (often along a cosine curve) for the rest of training. Why is that initial warmup period used, rather than starting immediately at the peak learning rate?
AEarly in training the model's weights are far from any well-conditioned region and gradients (and an adaptive optimizer's moment estimates) are still unreliable; taking large steps immediately at the full learning rate risks unstable updates or divergence, so warmup ramps the step size up gradually while the optimizer's statistics stabilize
BWarmup exists purely to save compute cost, since a smaller learning rate at the very start of training requires fewer floating-point operations per step than a larger one would
CWarmup is required only when using plain stochastic gradient descent without momentum, and serves no purpose whatsoever once the optimizer already has per-parameter adaptive learning rates such as Adam or AdamW
DWarmup gradually increases the model's context window length from a short initial length up to the full training sequence length, rather than gradually increasing the learning rate itself
Correct answer: .
At the very start of training, the model's weights are close to their random initialization, gradients can be large and noisy, and an adaptive optimizer's running estimates of the gradient's mean and variance have seen very few updates and are still poorly calibrated; taking full-sized steps immediately can push the model into a badly conditioned region or cause outright divergence. A short warmup period, during which the learning rate rises gradually from near zero to its peak, lets the optimizer's statistics accumulate and the model settle into a more stable region before the largest steps are taken, after which a smooth decay (often cosine-shaped) gradually reduces the step size again as training approaches convergence. The option claiming warmup saves compute is wrong because the number of floating-point operations per step in a transformer's forward and backward pass does not depend on the learning rate at all; the learning rate only scales how large a step is taken, not how much computation the step costs. The option claiming warmup is unnecessary for adaptive optimizers is wrong, since Adam and AdamW's own moment estimates are exactly what warmup protects while they are still unreliable early on. The option describing a gradually growing context window describes a completely different technique (sequence-length curriculum), not the learning rate schedule the question asks about.
Source: Brown et al., "Language Models are Few-Shot Learners" (2020), arXiv:2005.14165, Appendix B (training details)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 040/066medium
Micikevicius et al. (2017), "Mixed Precision Training," describe training large neural networks using FP16 (half precision) for most computations while still matching FP32 (single precision) training accuracy. What two techniques does their method rely on to prevent this switch to FP16 from degrading accuracy?
AStoring the model's activations in FP16 while performing the optimizer's weight update directly on FP16 weights, discarding any higher-precision copy of the weights once training begins
BMaintaining a master copy of the weights in FP32 that accumulates each step's small update (since FP16 cannot represent many small updates precisely enough), and multiplying the loss by a scaling factor before backpropagation so that small gradient values do not underflow to zero in FP16's limited numeric range, then unscaling the gradients before the optimizer step
CRounding every FP32 value in the network down to the nearest representable FP16 value using stochastic rounding alone, with no other change needed to preserve accuracy
DRunning two independent copies of the entire model, one in FP16 and one in FP32, and averaging their two sets of final predictions together at inference time
Correct answer: .
Micikevicius et al. keep an FP32 master copy of every weight specifically because FP16's limited precision cannot faithfully accumulate the very small per-step updates an optimizer applies over many thousands of steps; each step's update is computed and applied to this FP32 copy, which is then rounded down to FP16 for the forward and backward computation. Separately, they multiply the loss by a scaling factor before backpropagation so that gradient values, which can be very small, are shifted into FP16's representable range instead of underflowing to zero, and then divide the gradients by the same factor before the optimizer step uses them, a technique called loss scaling. The option describing FP16-only weight storage with no higher-precision copy is exactly the failure mode the FP32 master weights are designed to avoid, since updates would be lost to FP16 rounding. The option describing stochastic rounding of every value with no other change ignores the loss-scaling half of the method entirely, which is necessary because rounding alone does not address gradient underflow. The option describing two independently trained model copies averaged at inference describes model ensembling, an unrelated technique that does not address numeric precision during training at all.
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 041/066hard
Chen et al. (2016), "Training Deep Nets with Sublinear Memory Cost," introduce a technique now commonly called gradient checkpointing (or activation recomputation) for training very deep networks under tight memory budgets. What trade-off does this technique make, and how?
AIt reduces memory by storing activations in a lower numeric precision than the rest of the network, without changing how many activations are stored or when they are computed
BIt reduces memory by discarding the gradients of early layers entirely after they are used once, so those layers stop being updated for the rest of training
CIt reduces the memory needed to store activations for backpropagation by saving only a subset of activations ("checkpoints") during the forward pass and recomputing the discarded intermediate activations on demand during the backward pass, trading roughly one extra forward pass's worth of compute for memory that can scale close to the square root of the number of layers instead of growing linearly with depth
DIt reduces memory by splitting the model across multiple GPUs so that each device only ever holds a fraction of the total activations, without performing any extra recomputation on any device
Correct answer: .
Ordinary backpropagation through a deep network keeps every layer's forward-pass activations in memory until they are needed by the backward pass, so peak memory grows linearly with depth. Gradient checkpointing instead keeps only a sparse set of checkpointed activations during the forward pass and discards the rest; when the backward pass reaches a discarded activation, it recomputes it by re-running the forward computation from the nearest earlier checkpoint. Chen et al. show this lets memory scale roughly with the square root of the number of layers rather than linearly, at the cost of roughly one additional forward pass's worth of compute across the whole network. The option describing lower-precision activation storage describes mixed precision training, a different technique that does not change which activations are recomputed or when. The option describing discarding early layers' gradients after one use describes freezing layers, which would stop them from being trained further, not the memory-for-compute trade-off gradient checkpointing actually makes, since checkpointing still trains every layer normally. The option describing splitting the model across GPUs with no recomputation describes model parallelism, a distinct memory-reduction strategy that gradient checkpointing does not require and can be combined with rather than substitute for.
Source: Chen, Xu, Zhang & Guestrin, "Training Deep Nets with Sublinear Memory Cost" (2016), arXiv:1604.06174
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 042/066medium
Mistral 7B (Jiang et al., 2023) replaces full self-attention with sliding window attention (SWA), where each token attends to at most W tokens from the previous layer. How does stacking multiple layers of this fixed local window still let the model capture dependencies far beyond W tokens away, and what does this enable for its key-value cache?
AEvery layer's window is centered on a different, randomly chosen segment of the input, so across many layers the union of all windows eventually covers the entire sequence within a single layer's own computation
BIt does not actually extend the effective receptive field beyond W tokens at all; Mistral instead relies entirely on a separate global-attention mechanism layered on top of SWA to reach tokens farther away, and its key-value cache still grows without bound
CSliding window attention removes the need for a key-value cache entirely, since each token only ever looks at a small, fixed number of neighboring tokens and can discard all cache entries after every single step
DBecause a token at a given layer can attend to tokens within W positions of it at the previous layer, and each of those tokens could itself attend up to W positions further back at the layer before that, the effective receptive field compounds with depth, reaching roughly W multiplied by the number of layers; this fixed window also lets Mistral use a rolling buffer cache of fixed size, overwriting the oldest entries as the window slides forward instead of letting the key-value cache grow without bound
Correct answer: .
A token at layer k can attend to tokens within W positions of it at layer k-1, but each of those tokens was itself computed by attending up to W positions further back at layer k-2, and so on down the stack; this recursive access means information from a token up to roughly W times the number of layers away can reach the current position by the final layer, even though no single layer ever attends beyond its own fixed window W. Because each layer only ever needs the most recent W tokens' worth of key-value pairs rather than the full history, Mistral can use a rolling buffer cache of fixed size W, overwriting the oldest cached entries as the window slides forward, instead of a cache that grows with total sequence length. The option describing randomly centered windows covering the whole sequence within one layer misdescribes sliding window attention, which is a fixed local window around each token's own position, not a randomized or global one. The option claiming SWA does not extend the receptive field and relies on a separate global-attention mechanism describes a different design (closer to Longformer's local-plus-global hybrid), not Mistral's approach, which achieves its extended range purely through layer stacking. The option claiming no key-value cache is needed at all is wrong, since a rolling cache is still required and is exactly what the fixed window size enables to be bounded rather than eliminated.
Source: Jiang et al., "Mistral 7B" (2023), arXiv:2310.06825
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 043/066easy
Transformer-based language models are commonly grouped into three architecture families based on which parts of the original Vaswani et al. (2017) encoder-decoder design they keep: encoder-only, decoder-only, and encoder-decoder. Which option correctly matches each family to a representative model and its typical use?
AEncoder-only models such as BERT use bidirectional self-attention over the whole input and are typically used for tasks that classify or extract information from existing text; decoder-only models such as GPT use causal (masked) self-attention and are typically used for open-ended text generation; encoder-decoder models such as the original Transformer and T5 use an encoder to process the input and a separate decoder, conditioned on the encoder's output, to generate the output, and are typically used for sequence-to-sequence tasks such as translation or summarization
BEncoder-only models such as GPT generate text autoregressively one token at a time; decoder-only models such as BERT are used only for classification tasks and cannot generate any text at all; encoder-decoder models are a deprecated design no longer used in any modern language model
CThe three families differ only in how many layers they contain, with encoder-only models always having the fewest layers, decoder-only models a moderate number, and encoder-decoder models always having the most layers regardless of the task
DEncoder-only, decoder-only, and encoder-decoder all refer to the same self-attention computation; the labels only describe which software framework (such as TensorFlow, PyTorch, or JAX) was used to implement the model
Correct answer: .
Encoder-only models, of which BERT is the canonical example, apply bidirectional self-attention so every position can attend to every other position in the input, which suits tasks that classify or extract information from text the model has fully in front of it, such as sentiment classification or span extraction, but does not by itself define a natural way to generate new open-ended text. Decoder-only models, of which GPT is the canonical example, use causal self-attention so each position can only attend to earlier positions, which is exactly the constraint needed to generate text one token at a time from left to right. Encoder-decoder models, including the original Transformer and later models such as T5, keep both halves: an encoder that processes the full input with bidirectional attention, and a separate decoder that attends causally over its own generated output while also attending to the encoder's output, a structure well suited to sequence-to-sequence tasks like translation or summarization where an input sequence is transformed into a different output sequence. The option swapping GPT and BERT's roles and calling encoder-decoder models deprecated inverts the correct assignment and misstates encoder-decoder models' continued use. The option claiming the families differ only by layer count and the option claiming the labels only refer to software frameworks both ignore the actual architectural and attention-pattern differences the three families are defined by.
Source: Vaswani et al., "Attention Is All You Need" (2017); Devlin et al., "BERT" (2018); Radford et al., "Improving Language Understanding by Generative Pre-Training" (GPT, 2018); Raffel et al., "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer" (T5, 2020)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 044/066easy
BERT (Devlin et al., 2018) is pretrained using a masked language modeling (MLM) objective, while GPT-style decoder-only models are pretrained using causal (next-token) language modeling. How does BERT's MLM objective differ from GPT's pretraining objective, and what does BERT's approach require during training that GPT's does not?
ABoth objectives are identical: at every training step, both BERT and GPT predict only the very next token after everything seen so far, using leftward context alone in both cases
BBERT randomly replaces roughly 15% of input tokens (mostly with a special mask token, with the remainder left unchanged or swapped for a random token) and trains the model to predict the original identity of those replaced tokens using bidirectional context from both directions; GPT instead predicts each token from only the tokens before it, using causal self-attention that never looks ahead, so GPT's objective needs no masking of the input tokens themselves, only the causal attention mask restricting what each position can see
CBERT predicts every single token in the input on every training step, with none of them masked, while GPT is trained to predict only a randomly chosen 15% of tokens and ignores the rest
DBERT and GPT use the identical causal self-attention pattern during pretraining; the only real difference between them is that BERT is trained on longer documents than GPT
Correct answer: .
BERT's masked language modeling objective randomly selects about 15% of input tokens, of which most are replaced with a special mask token, a smaller portion are replaced with a random other token, and the remainder are left unchanged, and the model is trained to predict the original token at each of those selected positions using bidirectional self-attention over the entire surrounding sequence in both directions. GPT's causal language modeling objective instead predicts each token from only the tokens that came before it, enforced by a causal attention mask that prevents any position from attending to later positions, so nothing in the input itself needs to be corrupted or masked the way BERT's does; the restriction is purely on what each position is allowed to attend to. The option claiming both objectives are identical, using only leftward context, ignores BERT's defining bidirectional mechanism entirely. The option claiming BERT predicts every unmasked token while GPT predicts only a random 15% inverts which model actually uses the 15% masking rate, which belongs to BERT, not GPT. The option claiming both models use identical causal self-attention and differ only in training document length ignores the fundamental difference in attention direction (bidirectional versus causal) that defines the two objectives.
Source: Devlin, Chang, Lee & Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding" (2018), arXiv:1810.04805
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 045/066medium
In a transformer-based language model, what is the primary purpose of the softmax function applied to the attention scores (the dot products of queries and keys)?
AIt normalises the attention scores into a probability distribution that sums to 1, ensuring each token attends to all other tokens with weights that reflect their relative relevance rather than their raw dot-product magnitudes
BIt eliminates all negative attention scores so that only positively correlated tokens influence each other, preventing destructive interference in the hidden representations
CIt replaces the attention mechanism with a simple averaging operation, giving every token in the sequence an equal weight regardless of the query-key similarity
DIt is used exclusively during training to compute the cross-entropy loss and has no role during inference — at inference time the raw dot products are used directly without normalisation
Correct answer: .
The softmax function converts the raw attention scores (logits from query-key dot products) into a valid probability distribution where all weights are non-negative and sum to 1. This means each token's representation is a weighted combination of all value vectors, with the weights reflecting learned relevance. Without softmax, the raw scores could be arbitrarily large or negative, making the weighted sum unstable and uninterpretable. The option about eliminating negatives is partially true (softmax outputs are non-negative) but mischaracterises the purpose — softmax normalises, it does not simply clip negatives. The option about equal weights describes mean pooling, not attention. The option restricting softmax to training is wrong — softmax is applied at every forward pass, including inference.
Source: Vaswani et al. (2017), 'Attention Is All You Need', Section 3.2.1
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 046/066medium
A team is deciding between using a 7-billion-parameter open-weight model fine-tuned on their domain data versus a 400-billion-parameter general-purpose API model with in-context learning (no fine-tuning). Which factor most strongly favours the smaller fine-tuned model?
AThe task requires consistently following a narrow, domain-specific output format with specialised terminology, and the team has a high-quality labelled dataset of several thousand examples — a fine-tuned smaller model can learn this distribution directly, often outperforming a much larger model that must infer the pattern from a handful of in-context examples
BThe team wants maximum flexibility to handle a wide variety of unpredictable user queries across many domains without retraining
CThe team has no labelled data and cannot invest in creating any, making supervised fine-tuning impossible
DThe task requires broad world knowledge and multi-step reasoning across topics the team cannot anticipate in advance
Correct answer: .
Fine-tuning a smaller model on a high-quality domain-specific dataset encodes the desired behavior directly into the model's weights, producing reliable outputs for narrow tasks at lower inference cost. Research consistently shows that a well-fine-tuned small model can match or exceed a much larger general model on the specific task it was trained for. The option about maximum flexibility favours the larger general-purpose model, since fine-tuning specialises a model and can reduce its breadth. The option about no labelled data makes fine-tuning infeasible, again favouring in-context learning with the larger model. The option about broad multi-step reasoning also favours the larger model, which has more parameters to store world knowledge and more capacity for complex reasoning chains.
Source: Anthropic, 'When to Fine-tune vs. Prompt', docs.anthropic.com; OpenAI Fine-tuning Guide
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 047/066medium
Many modern open-weight LLMs (e.g. LLaMA, PaLM, Mistral) replace the original Transformer's ReLU-activated position-wise feed-forward network with a SwiGLU-based feed-forward network, following Shazeer's "GLU Variants Improve Transformer." What structurally distinguishes a SwiGLU feed-forward sublayer from a standard single-activation feed-forward sublayer?
AIt replaces the two-layer feed-forward network with a single linear projection, removing the need for a nonlinearity at all, since the residual connection provides all necessary nonlinearity
BIt adds a third feed-forward layer stacked after the original two, deepening the sublayer without changing how any individual layer computes its output
CIt computes two separate linear projections of the input, applies the Swish (SiLU) activation to one of them, and multiplies the two projections element-wise before the final output projection, gating how much of each unit passes through
DIt replaces the feed-forward sublayer's learned weight matrices with a fixed, non-trainable gating function computed directly from the token's position in the sequence
Correct answer: .
Shazeer's SwiGLU sublayer computes two separate linear projections of the sublayer's input, applies the Swish (SiLU) activation function to one of these projections, and takes the element-wise product of that activated projection with the other (un-activated) projection, before a final output projection maps the result back down; this element-wise product means the un-activated projection's units are "gated" by how large the corresponding Swish-activated unit is, letting the network learn to pass through more or less information per unit rather than applying one fixed nonlinearity uniformly. The option describing removal of any nonlinearity is wrong because a nonlinearity is still applied, just to only one of the two projections before gating, and residual connections don't substitute for this role. The option describing an added third layer is wrong because SwiGLU changes what the existing feed-forward sublayer computes internally rather than adding a wholly new stacked layer. The option describing a fixed, position-derived gating function is wrong because the gating values come from a fully learned linear projection of the token's own hidden representation, not from a non-trainable function of position, which would instead describe something closer to a positional encoding scheme.
Source: Shazeer, "GLU Variants Improve Transformer" (arXiv:2002.05202, 2020); adoption confirmed in LLaMA (Touvron et al., 2023) and PaLM (Chowdhery et al., 2022) architecture sections.
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 048/066easy
Before RoPE and ALiBi, Shaw, Uszkoreit and Vaswani (2018) proposed an alternative to the original Transformer's approach of adding sinusoidal positional vectors to token embeddings before the first layer. How does their relative position representation approach encode positional information instead?
AIt removes positional information entirely, relying only on the order in which tokens are fed through recurrent connections between layers
BIt concatenates a one-hot vector of the token's absolute position to the token embedding before the first attention layer, doubling the embedding dimension
CIt multiplies each attention head's query vector by a rotation matrix whose angle depends on the query's absolute position in the sequence
DIt learns a separate embedding vector for each pairwise distance between a query and key position (clipped at a maximum), and adds this distance-indexed vector directly into the attention score and value computations at every layer, rather than adding any position vector to the input embeddings
Correct answer: .
Shaw et al. (2018) modify self-attention itself: for each pair of positions, they look up a learned embedding indexed by the clipped relative distance between the query and key positions (distances beyond a maximum window are clamped to the same embedding), and add this distance-dependent term into both the attention score computation and the weighted sum over value vectors at every layer, rather than injecting position information once at the input via a fixed or learned absolute positional vector. This lets attention directly represent how far apart two tokens are, at every layer, rather than requiring the network to reconstruct relative distance indirectly from absolute positions added at the bottom. The option describing reliance on recurrent connections is wrong because the Transformer, with or without this relative-position modification, contains no recurrence; sequence order must come from some explicit positional signal, and relative position representations are exactly one such signal. The option describing a concatenated one-hot absolute-position vector is wrong because it describes a naive absolute-position scheme quite different from Shaw et al.'s pairwise, relative-distance embeddings, and one-hot position encoding was never their method. The option describing multiplying a query vector by a rotation matrix whose angle depends on absolute position describes RoPE (Su et al., 2021), a distinct, later technique that encodes position through rotation rather than through Shaw et al.'s learned additive distance embeddings.
Source: Shaw, Uszkoreit & Vaswani, "Self-Attention with Relative Position Representations" (NAACL 2018), arXiv:1803.02155
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 049/066easy
Lee et al. (2022), "Deduplicating Training Data Makes Language Models Better," found that large web-scraped pretraining corpora such as C4 contain a substantial number of near-duplicate documents and repeated substrings (in one case, a single sentence repeated tens of thousands of times). What did they find deduplicating this training data actually does to the resulting language models?
AIt has no measurable effect on either memorization or downstream accuracy, since duplicate documents are rare enough in web-scraped corpora to be statistically negligible
BIt substantially reduces how often trained models emit memorized text copied verbatim from the training data, while letting models reach the same or better accuracy in fewer training steps
CIt increases verbatim memorization of the surviving unique documents, because removing duplicates concentrates more training passes on each remaining example
DIt only affects evaluation, by removing train-test overlap, without changing anything about the trained model's own generation behavior
Correct answer: .
Lee et al. found that large web-scraped corpora such as C4 contain many near-duplicate documents and long repeated substrings, and that models trained on this duplicated data emit memorized text copied verbatim from training data at a measurably higher rate. After deduplicating the training data with their near-duplicate and exact-substring detection tools, they found models emit memorized text roughly ten times less often, and reach the same or better validation accuracy using fewer training steps overall -- deduplication is not just a memorization fix but also improves training efficiency. The option claiming no measurable effect is wrong because the paper documents concrete, large duplication rates (for example, over a percent of C4 being near-duplicate) and a corresponding measurable reduction in memorization after deduplication. The option claiming deduplication increases memorization is wrong and inverts the paper's actual finding; concentrating training on fewer, non-duplicated examples did not increase verbatim memorization of what remained. The option restricting the effect to evaluation alone is wrong because, while the paper separately notes that deduplication reduces train-test overlap and thus makes evaluation more accurate, it also reports the training-side effects on memorization rate and training efficiency described above, not just an evaluation-side correction.
Source: Lee et al., "Deduplicating Training Data Makes Language Models Better" (ACL 2022), arXiv:2107.06499
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 050/066easy
Petrov et al. (2023), "Language Model Tokenizers Introduce Unfairness Between Languages," measured how many tokens the same piece of text requires once translated into different languages, using tokenizers from widely deployed LLMs. What did they find, and why does it matter?
ATokenizing the same content in different languages can require up to roughly 15 times as many tokens depending on the language, even for tokenizers explicitly trained to support multiple languages, which raises the cost, processing latency, and effective context length available to speakers of the more heavily-tokenized languages
BTokenizer length varies only trivially (well under 10%) across languages once a tokenizer's vocabulary includes any coverage of that language's script, so the disparity is not practically significant
CThe disparity is fully explained by differing character-set sizes alone, and disappears entirely once text is measured in bytes rather than tokens
DThe disparity only affects languages whose tokenizer support was added after initial release, and multilingual tokenizers trained from scratch on balanced multilingual data eliminate it completely
Correct answer: .
Petrov et al. measured tokenized length of parallel text (the same content translated across many languages) using tokenizers from widely deployed LLMs, and found disparities of up to roughly 15 times as many tokens for the same content depending on the language, a disparity that persists even for tokenizers explicitly trained with multilingual support in mind. Because commercial LLM usage is commonly billed and rate-limited per token, and because a model's fixed context window holds a fixed number of tokens rather than a fixed amount of text, this means speakers of more heavily-tokenized languages pay more, wait longer, and can fit less content into the same context window for equivalent content, which the paper frames as a fairness concern. The option describing only a trivial, sub-10%-variation is wrong and understates the paper's actual, much larger measured disparities. The option attributing the disparity entirely to character-set size and claiming it vanishes when measured in bytes is wrong; the paper's point is specifically about token counts under real deployed tokenizers, and byte-level measurement is a different, tokenizer-independent yardstick that doesn't resolve the token-based cost and context disparities users actually face. The option claiming multilingual training from scratch eliminates the disparity is wrong because the paper found the unfairness persists even in tokenizers already built for multilingual support, not only in ones retrofitted after the fact.
Source: Petrov, La Malfa, Torr & Bibi, "Language Model Tokenizers Introduce Unfairness Between Languages" (NeurIPS 2023), arXiv:2305.15425
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 051/066medium
In "Attention Is All You Need," Vaswani et al. (2017) compare self-attention, recurrent, and convolutional layers by three properties: computational complexity per layer, the minimum number of sequential operations required, and the maximum path length between any two input and output positions. What do they report for self-attention on the last two of these three properties, and why did this motivate using self-attention?
ASelf-attention requires O(n) sequential operations, the same as a recurrent layer, but achieves a shorter maximum path length of O(log n) by processing the sequence in a tree-structured order
BSelf-attention and convolutional layers are identical on both properties, since both process the full sequence within a single layer's receptive field once enough layers are stacked
CSelf-attention requires a constant (O(1)) number of sequentially executed operations, and its maximum path length between any two positions is also O(1) regardless of distance, letting it learn long-range dependencies more easily than a layer type whose maximum path length grows with sequence length
DSelf-attention requires more sequential operations to complete than a recurrent layer does, but its shorter maximum path length only becomes a practical advantage once sequence length n exceeds the representation dimensionality d
Correct answer: .
Vaswani et al.'s Table 1 reports that a self-attention layer connects all positions with a constant, O(1), number of sequentially executed operations (all pairwise interactions are computed in parallel within one layer), and that its maximum path length between any two input and output positions is likewise O(1), the same short distance regardless of how far apart the two positions are in the sequence; a recurrent layer, in contrast, needs O(n) sequential operations and has a maximum path length of O(n), since information must pass step-by-step through every intervening position. A shorter, position-independent path length matters because it makes it easier for gradients and information to flow between distant positions during learning, which the paper cites as one motivation for self-attention over recurrence for capturing long-range dependencies. The option describing O(n) sequential operations for self-attention is wrong and instead describes the recurrent layer's own property; self-attention's sequential-operation count in the table is constant, not linear in sequence length. The option treating self-attention and convolutional layers as identical is wrong because the table gives convolutional layers a different, kernel-size-dependent maximum path length (logarithmic in sequence length for stacked dilated convolutions, not constant) and a different per-layer complexity expression. The option claiming self-attention requires more sequential operations than a recurrent layer is wrong and reverses the table's actual comparison; the paper's caveat about sequence length exceeding representation dimensionality applies to the separate complexity-per-layer comparison (self-attention's per-layer complexity is favorable when n is smaller than d), not to the sequential-operations or path-length comparisons this question asks about.
Source: Vaswani et al., "Attention Is All You Need" (2017), arXiv:1706.03762, Table 1 and Section 4
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 052/066hard
Xiao et al. (2023), "Efficient Streaming Language Models with Attention Sinks," studied why simply evicting the oldest tokens from a fixed-size KV cache (a sliding window) causes a language model's quality to collapse once the window has evicted the first few tokens of a long stream, even though those first tokens carry little relevant content for later text. What did they find causes this collapse, and how does StreamingLLM avoid it?
AThe collapse happens because positional encodings become invalid once any token is evicted, so StreamingLLM avoids it by recomputing sinusoidal positional encodings for the entire remaining cache after every eviction
BThe collapse happens because a sliding window shrinks the effective vocabulary the model can attend to, so StreamingLLM avoids it by periodically re-inserting a random sample of previously evicted tokens back into the cache
CThe collapse happens purely from a mismatch between training and inference sequence lengths, so StreamingLLM avoids it by fine-tuning the model on sequences as long as any stream it will process at inference time
DThe first few tokens act as "attention sinks" that absorb a disproportionate share of attention across layers and heads regardless of their semantic content, so evicting them removes a stabilizing target for attention scores; StreamingLLM keeps these first few tokens permanently in the cache alongside a sliding window of the most recent tokens, which restores stable performance without any fine-tuning
Correct answer: .
Xiao et al. found that in autoregressive language models, the first few tokens of a sequence consistently absorb a disproportionately large share of attention across many layers and attention heads, regardless of whether those tokens carry any relevant semantic content for later text -- a phenomenon they term "attention sinks," which likely exists because the softmax over attention scores must sum to one, and the model learns to dump unneeded attention mass onto these easily-identified early positions rather than distributing it meaningfully. Once a naive sliding-window cache evicts these first tokens to make room for new ones, this stabilizing target disappears and the model's perplexity and output quality collapse. StreamingLLM fixes this cheaply, without any fine-tuning, by always retaining a handful of these initial "sink" tokens in the KV cache in addition to the usual sliding window of recent tokens, which keeps attention well-behaved indefinitely. The option about positional encodings is wrong because the paper's fix does not recompute positional encodings for the whole cache; it instead keeps positions consistent within the retained window and sink tokens. The option about vocabulary shrinkage is wrong because a sliding window doesn't restrict which tokens the model's vocabulary contains, only which positions remain visible to attention, and StreamingLLM's fix doesn't involve randomly reinserting evicted tokens. The option attributing the collapse purely to a training/inference length mismatch and fixing it via fine-tuning on very long sequences is wrong because StreamingLLM's headline contribution is precisely that it avoids any such fine-tuning, extending models to effectively unbounded stream lengths using only this fixed-size cache change.
Source: Xiao, Tian, Chen, Han & Lewis, "Efficient Streaming Language Models with Attention Sinks" (ICLR 2024), arXiv:2309.17453
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 053/066medium
Power et al. (2022), "Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets," trained small transformer models on algorithmic tasks (such as modular arithmetic) well past the point where training accuracy had already reached ~100% while validation accuracy remained near chance level. What did they observe happens if training is continued far beyond this point, and what term did they give this phenomenon?
AValidation accuracy never improves beyond chance level no matter how much further training continues, since the model has already fully overfit to the training set and cannot recover; they term this phenomenon "memorization collapse"
BValidation accuracy can suddenly rise from near chance level to near-perfect generalization long after training accuracy has already saturated, a delayed transition they term "grokking"
CTraining accuracy itself starts to decline the longer training continues past saturation, as the optimizer is forced to trade training performance for validation performance; they term this phenomenon "double descent"
DBoth training and validation accuracy oscillate periodically forever once training accuracy saturates, never settling at a stable value in either direction; they term this phenomenon "catastrophic forgetting"
Correct answer: .
Power et al. observed that on small algorithmic datasets, after a network's training accuracy has already reached (or nearly reached) 100% while its validation accuracy remains stuck near chance level, continuing to train far past this point can eventually produce a sudden, delayed rise in validation accuracy up to near-perfect generalization, well after the point where the training loss had already appeared to converge; they named this delayed transition "grokking." They also found smaller training datasets required substantially more optimization steps before this transition occurred. The option claiming validation accuracy never improves is wrong and describes the opposite of the paper's central, titular finding. The option claiming training accuracy declines and calling this "double descent" is wrong on two counts: it misdescribes what happens to training accuracy (which stays high, not declining) and misapplies "double descent," which is a different, well-known phenomenon about test error non-monotonically depending on model or dataset size, not this training-duration-based transition. The option describing endless oscillation and calling it "catastrophic forgetting" is wrong; catastrophic forgetting refers to a trained model losing previously learned capability when trained on new data or tasks, an unrelated phenomenon to the single-task, delayed-generalization pattern this paper documents.
Source: Power, Burda, Edwards, Babuschkin & Misra, "Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets" (2022), arXiv:2201.02177
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 054/066easy
Berglund et al. (2023), "The Reversal Curse: LLMs Trained on 'A is B' Fail to Learn 'B is A,'" fine-tuned GPT-3 and Llama-1 on fictitious factual statements of the form "[Name] is the [role] of [thing]" (e.g., a fictitious composer of a fictitious piece) and then tested the models on the reversed question. What did they find, and under what condition does this failure not occur?
AThe models answered the reversed question correctly at the same rate as the original direction, showing fine-tuned factual associations generalize symmetrically regardless of which direction was stated during training
BThe models answered the reversed question correctly only when the statement described a real, pretraining-era fact rather than a fictitious one invented for the fine-tuning set
CThe models trained on "[Name] is the [role] of [thing]" statements failed to correctly answer the reversed question at a rate far better than chance, even though they could state the original direction correctly; this reversal failure did not occur when both directions of the relationship were presented together within the same context window at inference time, since the model could then read off the reverse relationship directly rather than needing to recall it
DThe models succeeded on the reversed direction only for statements about numerical or quantitative relationships, and failed exclusively on statements about named entities and their roles
Correct answer: .
Berglund et al. fine-tuned GPT-3 and Llama-1 on invented statements of the form "[Name] is the [role] of [thing]" and found that, despite the models correctly answering questions about the original direction they were trained on, they failed to answer the reversed direction (e.g., asking who holds that role for that thing) at a rate no better than chance, a failure they term the "Reversal Curse." Crucially, they also found this failure specifically applies to recalling the relationship purely from parameters learned during fine-tuning: if both directions of the same relationship are given together within the model's context at inference time, the model can read off the reverse relationship directly from what's in front of it rather than needing to recall it from training, so the reversal failure specifically concerns a gap in what fine-tuning generalizes into the model's weights, not a gap in the model's in-context reasoning ability generally. The option claiming symmetric generalization is wrong and states the opposite of the paper's central finding. The option about real versus fictitious facts is wrong; the paper deliberately used fictitious statements to rule out the possibility that models were merely recalling the reverse direction from separate exposure during pretraining, and this is a control, not the condition under which reversal succeeds. The option restricting the failure to non-numerical, named-entity statements is wrong; the paper's core claim is about the general direction-dependence of what fine-tuning teaches, not a distinction between quantitative and named-entity content.
Source: Berglund et al., "The Reversal Curse: LLMs Trained on 'A is B' Fail to Learn 'B is A'" (ICLR 2024), arXiv:2309.12288
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 055/066hard
Yang et al. (2017), "Breaking the Softmax Bottleneck: A High-Rank RNN Language Model," argue that standard softmax-based neural language models are fundamentally limited in how well they can approximate the true conditional next-word distribution, no matter how the model is trained. What is the "softmax bottleneck" they identify, and how does their proposed Mixture of Softmaxes (MoS) address it?
AThe bottleneck is that softmax computation is too slow at large vocabulary sizes to be practical, and MoS addresses it by approximating the softmax's normalizing constant instead of computing it exactly
BThe bottleneck is that softmax outputs cannot represent probabilities smaller than a fixed minimum value due to floating-point precision, and MoS addresses it by computing several softmax passes at different numeric precisions and averaging them
CThe bottleneck is that the model's final hidden state is tied to too many output word embeddings, causing the model to overfit; MoS addresses it by using a smaller hidden dimension so fewer parameters need to be learned per predicted word
DFormulating language modeling as a matrix factorization problem, they show that a standard softmax layer's context-to-word log-probability matrix has an output rank bounded by the model's hidden dimension, which is too low to express the true, effectively higher-rank distribution over next words given diverse contexts; MoS instead computes several separate softmax distributions from the same hidden state and combines them in a learned weighted mixture, achieving an effectively higher-rank output than a single softmax can reach
Correct answer: .
Yang et al. formalize the problem of predicting the next word given a context as a matrix factorization problem: the ideal matrix of true log-probabilities, over all contexts and all vocabulary words, has some true rank, but a standard softmax layer's context-to-word log-probability matrix is mathematically constrained to have rank no higher than the size of the model's hidden dimension, since it's produced from a single linear projection of a single hidden vector per context; when the true distribution's effective rank exceeds this bound, no amount of additional training data or optimization can let a single softmax represent it exactly, a limitation they call the "softmax bottleneck." Their Mixture of Softmaxes (MoS) works around this by computing several distinct softmax distributions from the same underlying hidden state (each using its own linear projection), then combining these distributions in a learned, context-dependent weighted average, which can represent a higher-rank matrix than any single softmax term alone, improving perplexity on standard benchmarks. The option about softmax computation speed is wrong; the paper's bottleneck is about the softmax's representational capacity (its output rank), not its computational cost, and MoS doesn't approximate any normalizing constant. The option about floating-point precision is wrong and misidentifies an unrelated numerical-precision issue as the paper's actual, rank-based argument. The option describing the bottleneck as overfitting from weight tying, fixed by shrinking the hidden dimension, is wrong and inverts the paper's proposed fix, which combines multiple softmaxes to increase expressiveness rather than reducing the hidden dimension to reduce parameters.
Source: Yang, Dai, Salakhutdinov & Cohen, "Breaking the Softmax Bottleneck: A High-Rank RNN Language Model" (ICLR 2018), arXiv:1711.03953
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 056/066easy
Chinchilla's scaling analysis (Hoffmann et al., 2022) identifies, for a fixed training compute budget, the model size and token count that minimizes training loss. Touvron et al.'s LLaMA paper (2023) deliberately trained smaller models on far more tokens than this Chinchilla-optimal ratio would recommend for their training compute budget alone. What is the actual justification for training a smaller model "beyond Chinchilla-optimal," per this line of reasoning?
AChinchilla's original analysis contained an arithmetic error later corrected by Touvron et al., who found the truly compute-optimal ratio of tokens to parameters was substantially higher than Hoffmann et al. reported
BA smaller model trained on more tokens always achieves strictly lower training loss than a larger model trained Chinchilla-optimally on the same compute budget, making the Chinchilla recommendation simply incorrect
CChinchilla's ratio only minimizes loss for a given training compute budget; it does not account for inference cost, and a smaller model is cheaper to run at inference time, so if a model will be served enough times, training it for longer than Chinchilla-optimal on more tokens to reach the best possible quality at a smaller, cheaper-to-serve size can be worth the extra training compute
DChinchilla's analysis assumed a fixed dataset size that no longer applies now that far larger web-scraped training corpora are available, so its ratio is simply obsolete rather than being deliberately overridden by any inference-cost argument
Correct answer: .
Chinchilla's scaling analysis (Hoffmann et al., 2022) answers a specific, narrower question: for a fixed training compute budget, what model size and token count minimizes training loss. It says nothing about inference cost, which for a widely-deployed model can dwarf training cost over the model's lifetime, since inference cost scales with the number of parameters served times the number of requests, not with how the model was trained. Touvron et al.'s LLaMA paper explicitly reasoned about this: given that a smaller model is cheaper and faster to run at inference time, it can be worth spending more training compute than Chinchilla-optimal to train that smaller model on many more tokens than Chinchilla's ratio would recommend, continuing to see loss improve well past the Chinchilla-optimal point for that smaller size, in exchange for a model that is both strong and cheap to serve at scale. The option claiming Chinchilla contained a corrected arithmetic error is wrong; Touvron et al. did not claim to have found an error in Hoffmann et al.'s analysis, only to be optimizing for a different objective (inference cost included) than pure training-compute-optimality. The option claiming a smaller, longer-trained model always strictly beats a larger Chinchilla-optimal model on training loss for the same compute is wrong; Chinchilla's whole point is that its ratio is loss-minimizing for a given training compute budget, so deliberately deviating from it to train a smaller model longer means using more total training compute, not matching it. The option attributing the change to now-larger available datasets making Chinchilla's analysis obsolete is wrong; the actual argument for training beyond Chinchilla-optimal is about weighing inference cost against training cost, not about dataset size having outgrown the original analysis's assumptions.
Source: Touvron et al., "LLaMA: Open and Efficient Foundation Language Models" (2023), arXiv:2302.13971, Section 1; Hoffmann et al., "Training Compute-Optimal Large Language Models" (2022), arXiv:2203.15556
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 057/066medium
Schuster and Nakajima (2012) introduced WordPiece, the subword tokenization algorithm Devlin et al. (2018) later adopted for BERT. WordPiece builds its vocabulary by iteratively merging the pair of adjacent symbols that most increases the likelihood of the training corpus under a language model, rather than Sennrich et al.'s byte-pair encoding (BPE), which merges whichever adjacent pair occurs most frequently at each step. What is the practical consequence of choosing merges by likelihood increase instead of raw frequency?
AWordPiece can still only ever merge the pair that appears more often than any other pair in the corpus, exactly like BPE, so the two algorithms always produce identical vocabularies when trained on the same text
BA pair can be merged before a more frequent pair if merging it yields a disproportionately larger boost to the corpus's likelihood under the language model, so WordPiece can favor statistically informative subwords over merely common ones that BPE's frequency-only criterion would pick first
CWordPiece replaces merging entirely with a fixed-size character n-gram model, never combining learned subwords into longer units the way BPE's iterative merges do
DWordPiece requires a human-labeled corpus of correct word segmentations to learn each merge, whereas BPE only needs unlabeled raw text
Correct answer: .
WordPiece scores each candidate merge by how much it would increase the likelihood of the training corpus under a language model, computed as the probability of the merged unit divided by the product of the probabilities of its two parts considered separately. This means a pair that is only moderately frequent can still be merged ahead of a more frequent pair if combining it captures a strong statistical dependency between its parts, letting WordPiece favor subwords that are informative rather than merely common. BPE, by contrast, always merges whichever adjacent pair has the highest raw frequency, with no notion of how much that merge actually helps predict the data. The option claiming the two algorithms always produce identical vocabularies is wrong because it collapses WordPiece's likelihood criterion into BPE's frequency criterion, erasing the exact distinction the question asks about. The option describing a fixed-size character n-gram model is wrong because WordPiece, like BPE, builds its vocabulary through iterative merging of symbols into longer subwords rather than using static n-grams. The option requiring human-labeled word segmentations is wrong because WordPiece's likelihood criterion is computed directly from unlabeled training text using a language model fit to that text, the same kind of raw corpus BPE itself trains on.
Source: Schuster & Nakajima, "Japanese and Korean Voice Search" (2012); Devlin et al., "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding" (2018), arXiv:1810.04805, Section 3.2
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 058/066medium
Child et al. (2019), "Generating Long Sequences with Sparse Transformers," train Transformers on sequences up to 16,384 tokens long by replacing full self-attention, whose compute and memory grow quadratically with sequence length, with a sparse factorization split across attention heads: a local pattern and a strided pattern. How does combining these two fixed sparse patterns across several layers let the model still capture a dependency between two arbitrary, far-apart positions, while each individual attention operation only considers a sparse subset of positions?
AEach head attends to literally every position in the sequence but at reduced numerical precision, which lowers memory use without changing which positions are attended to
BThe local pattern lets each position attend only to a fixed number of its nearest neighbors, and the strided pattern is discarded during training, used only to speed up inference afterward
CThe strided pattern alone is sufficient to connect every pair of positions directly within a single layer, so the local pattern is added only to reduce the number of learned parameters, not to extend which positions can be reached
DThe local pattern connects each position to its nearby neighbors while the strided pattern connects positions spaced at regular intervals; stacking several layers lets information reach any other position through a path of local and strided hops, so full connectivity emerges across depth even though each layer's attention only ever considers a sparse subset of positions
Correct answer: .
Sparse Transformer's factorized attention splits the full attention pattern across heads into a local pattern, where each position attends to a fixed-size window of nearby positions, and a strided pattern, where each position attends to positions spaced at regular fixed intervals further back in the sequence. Neither pattern alone directly connects every pair of positions within one layer, but because the model stacks several such layers, information from a given position can reach a far-away position by hopping through intermediate positions reached by local connections in some layers and strided connections in others, much like how stacking several layers of short-range dilated convolutions builds a long receptive field without any single layer looking at the whole input. This is what lets Child et al. reduce complexity from the quadratic cost of full attention to O(n√n) while still approximating dense attention's effective connectivity. The option describing full attention at reduced precision is wrong because it keeps every position attending to every other position, which is exactly the quadratic cost the sparse factorization is designed to avoid. The option claiming the strided pattern is discarded during training is wrong because both patterns are used throughout training, not added only at inference. The option claiming the strided pattern alone connects every pair of positions within a single layer is wrong because the strided pattern only reaches positions at fixed stride intervals, not arbitrary positions, which is precisely why the local pattern and depth are both needed for full reachability.
Source: Child, Gray, Radford & Sutskever, "Generating Long Sequences with Sparse Transformers" (2019), arXiv:1904.10509, Section 4
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 059/066hard
DeepSeek-V2 (DeepSeek-AI, 2024) introduces Multi-head Latent Attention (MLA) specifically to shrink the key-value cache needed for autoregressive inference, which standard multi-head attention grows linearly with both sequence length and the number of attention heads. Rather than having groups of heads share key/value heads the way multi-query or grouped-query attention do, what does MLA actually store in the KV cache for each token, and how does it still produce distinct per-head keys and values at attention time?
AMLA compresses each token's keys and values into a single low-dimensional latent vector via a learned down-projection, stores only that compact latent vector in the KV cache, and reconstructs the full per-head keys and values on demand at attention time using learned up-projection matrices, so cache size depends on the latent dimension rather than on the number of heads
BMLA stores the full per-head keys and values exactly as standard multi-head attention does, but compresses them afterward using 8-bit quantization before writing them to the cache, so no up-projection is needed at attention time
CEvery attention head in MLA shares one single key head and one single value head, identical to multi-query attention, and the only difference from multi-query attention is that MLA additionally applies rotary position embeddings directly to the shared keys
DMLA removes the key-value cache entirely by recomputing every previous token's keys and values from scratch at every new decoding step, trading all of the saved memory for additional compute
Correct answer: .
MLA's low-rank joint compression projects each token's keys and values down into a single compact latent vector using a learned down-projection matrix, and only this latent vector is written to the KV cache, rather than separate full-dimensional keys and values for every head. At attention time, learned up-projection matrices expand that one stored latent vector back out into the distinct per-head keys and values each head needs, so the heads still attend using different effective keys and values even though only one shared low-dimensional vector was ever cached per token. Because cache size scales with the latent dimension rather than with the number of heads times the per-head dimension, DeepSeek-V2 reports a far smaller KV cache than standard multi-head attention at the same model size, reported as roughly a 93% reduction relative to its predecessor. The option describing 8-bit quantization of full per-head keys and values is wrong because MLA's saving comes from storing a lower-dimensional latent representation, not from quantizing already-full-dimensional cached tensors, and it still requires up-projection to reconstruct per-head values, not none at all. The option describing a single shared key head and value head across all heads is wrong because that describes multi-query attention, which MLA is explicitly contrasted against; MLA reconstructs distinct per-head keys and values from the shared latent rather than having every head literally use the same key and value. The option describing full recomputation at every step is wrong because MLA still caches something, the compressed latent vector, rather than caching nothing and recomputing everything from scratch.
Source: DeepSeek-AI, "DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model" (2024), arXiv:2405.04434, Section 2.1 (Multi-Head Latent Attention)
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 060/066easy
Shoeybi et al. (2019), introducing Megatron-LM, describe tensor parallelism (also called intra-layer model parallelism) as an alternative to data parallelism for training a single large transformer across multiple GPUs. In ordinary data parallelism, every GPU holds a full copy of the model and processes a different batch of training data. What does tensor parallelism change instead?
AIt keeps a full copy of the model on every GPU, exactly as data parallelism does, and only changes how the training data is split into batches across GPUs
BIt assigns entirely different layers of the model to different GPUs, so each GPU runs a different consecutive block of the network's depth on the same data as it flows through in sequence
CIt splits individual weight matrices within a layer, such as those in the attention and feed-forward sublayers, across GPUs, so that computing a single layer's output requires those GPUs to collaborate and exchange partial results, rather than each GPU holding a full copy of every weight matrix
DIt trains several smaller independent models in parallel on different GPUs, then averages their final weights together only once training finishes
Correct answer: .
Tensor parallelism, as Shoeybi et al. implement it in Megatron-LM, splits the individual weight matrices inside a layer, such as the matrices in a feed-forward or attention sublayer, into column-wise or row-wise partitions distributed across GPUs, so that no single GPU holds a complete copy of every weight matrix; computing that layer's output requires the participating GPUs to exchange partial results, typically via an all-reduce operation, before the computation can proceed. This differs from data parallelism, where every GPU already holds a full copy of the whole model and parallelism instead comes from splitting the training batch across GPUs. The option describing data parallelism itself, with the model fully replicated and only the batch split, is wrong because it describes the baseline approach tensor parallelism is introduced as an alternative to, not what tensor parallelism changes. The option describing different layers assigned to different GPUs that process data sequentially describes pipeline (inter-layer) parallelism, a distinct distributed-training strategy that splits the model by depth rather than splitting individual weight matrices within a layer. The option describing training several independent models and averaging their final weights is wrong because tensor-parallel GPUs are jointly computing the output of one single model's layers in real time during every forward and backward pass, not training separate models that are only reconciled at the end.
Source: Shoeybi, Patwary, Puri, LeGresley, Casper & Catanzaro, "Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism" (2019), arXiv:1909.08053, Section 3
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 061/066easy
Olsson et al. (2022), "In-context Learning and Induction Heads," argue that a specific, identifiable attention circuit is the primary mechanism behind a transformer's ability to improve at a task from examples given only in its prompt, with no weight updates involved. What pattern does an induction head implement, and what did the authors observe happens during training as these circuits form?
AAn induction head retrieves the single training example from pretraining that is most similar to the current prompt, and in-context learning ability appears gradually and smoothly throughout training, with no identifiable turning point
BAn induction head looks back for a previous occurrence of the current token, finds what token came immediately after it that earlier time, and predicts that same token will come next again; the authors observed a sharp phase change early in training where induction heads form and in-context learning ability improves abruptly, across model sizes from small attention-only models up to 13-billion-parameter models
CAn induction head only operates on numerical and tabular data within the prompt, and plays no role in the few-shot text-completion behavior described by Brown et al. for GPT-3
DAn induction head is a component added only during a separate supervised fine-tuning stage after pretraining, and does not exist in a purely pretrained, non-instruction-tuned language model
Correct answer: .
Olsson et al. identify induction heads as attention heads that implement a simple copy-and-complete pattern: given the current token, the head searches back through the context for an earlier occurrence of that same token, finds whatever token immediately followed it that earlier time, and boosts the probability that the same token will follow again now. This single mechanism accounts for a large share of a transformer's ability to pick up patterns from examples placed in its prompt, with no gradient update needed. The authors further report a sharp phase change early in training, a narrow window during which induction heads form and in-context learning ability jumps abruptly, a pattern they observed consistently from small, attention-only toy models up through 13-billion-parameter language models. The option describing retrieval of the single most similar pretraining example is wrong because induction heads operate entirely within the current context window by matching tokens already present in the prompt, not by searching pretraining data, and the paper's central finding is an abrupt phase change, not smooth, gradual improvement. The option restricting induction heads to numerical and tabular data is wrong because the copy-and-complete pattern they implement operates on tokens generally, including the ordinary natural-language few-shot completion behavior GPT-3 was shown to exhibit. The option claiming induction heads only appear after a separate fine-tuning stage is wrong because Olsson et al. study them forming during ordinary pretraining itself, well before any instruction-tuning or fine-tuning stage would occur.
Source: Olsson, Elhage, Nanda, Joseph, DasSarma, Henighan et al. (Anthropic), "In-context Learning and Induction Heads" (2022), Transformer Circuits Thread
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 062/066medium
Gu and Dao (2023), introducing Mamba, propose a selective state space model (SSM) as a subquadratic alternative to the Transformer's attention mechanism. Prior structured SSMs, such as Gu et al.'s S4, use state-transition matrices that are fixed learned constants, the same for every input token regardless of its content. What change does Mamba make to these matrices, and what does this change restore?
AIt makes the SSM's transition and input-projection matrices functions of the current input token itself, rather than fixed constants shared across all inputs, restoring the model's ability to selectively propagate or discard information based on each token's content
BIt discards the state space formulation entirely in favor of full quadratic self-attention over the whole sequence, trading away linear-time scaling in sequence length to regain content-based reasoning
CIt keeps the transition matrices fixed and input-independent exactly as in S4, but adds a separate content-independent LSTM-style gate on top of the recurrent state update
DIt enlarges the fixed, input-independent transition matrices to a much higher dimension, relying on that extra fixed capacity alone to encode content-dependent behavior without making the matrices themselves depend on the input
Correct answer: .
Mamba's selection mechanism makes the state space model's transition and input-projection matrices depend on the current input token, rather than holding them fixed as constants shared across every token the way linear-time-invariant (LTI) SSMs such as S4 do. This restores content-based reasoning: the model can selectively propagate, compress, or discard information in its recurrent state depending on what each token actually is, something an LTI system mathematically cannot do since it must treat every input the same way. Making the matrices input-dependent breaks the convolutional, FFT-parallelizable view that LTI SSMs rely on for efficient training, so Mamba instead introduces a hardware-aware parallel scan algorithm to keep training efficient, while still achieving linear-time scaling in sequence length, unlike attention's quadratic cost. The option describing a switch to full quadratic self-attention is wrong because it would surrender exactly the subquadratic scaling Mamba is designed to preserve, and the paper's whole motivation is to match attention's modeling power without paying its quadratic cost. The option describing a fixed, content-independent LSTM-style gate added on top of unchanged constant matrices is wrong because it does not make the transition dynamics themselves depend on the input at all; Mamba's change is to the matrices' values, not an added fixed gating layer around still-constant matrices. The option describing enlarged but still fixed matrices is wrong because the paper attributes the modeling weakness specifically to the matrices being input-independent, not to insufficient size; scaling up fixed constants cannot make them respond to a given token's content.
Source: Gu, A. & Dao, T. (2023), "Mamba: Linear-Time Sequence Modeling with Selective State Spaces," arXiv:2312.00752
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 063/066easy
Huang et al. (2018), introducing GPipe, describe pipeline (inter-layer) parallelism as a way to train a model too large for any single accelerator's memory, by assigning different consecutive layers to different accelerators that then process data as it flows through in sequence. Naively running this layer-by-layer pipeline on one training batch at a time leaves most accelerators idle for most of each step, waiting their turn. What technique does GPipe introduce to reduce this idle time, and how does it work?
AGPipe gives every accelerator a full copy of every layer, so each one can process an entire training batch independently without waiting on any other accelerator
BGPipe splits each layer's individual weight matrices across accelerators, so that every accelerator participates in computing every layer's output simultaneously rather than waiting for upstream layers to finish
CGPipe splits each training batch into smaller micro-batches and pipelines them through the sequence of accelerators, so that once an accelerator finishes a micro-batch and passes its output downstream, it can immediately start on the next micro-batch instead of sitting idle
DGPipe discards intermediate activations after each micro-batch's forward pass and recomputes them during the backward pass instead of storing them, which is a separate technique GPipe also uses to reduce memory use but does not by itself address idle pipeline time
Correct answer: .
GPipe splits each training mini-batch into a number of smaller micro-batches and feeds them one after another into the layer-partitioned pipeline, so that as soon as an accelerator finishes processing one micro-batch's forward or backward computation and passes its output to the next (or previous) stage, it can immediately begin the next micro-batch rather than sitting idle waiting for the whole batch to clear the pipeline. GPipe schedules all micro-batches' forward passes first, then runs all their backward passes; the remaining idle time at the start and end of each step, the pipeline 'bubble,' shrinks as the number of micro-batches grows relative to the number of pipeline stages, though it is never eliminated entirely. The option describing a full model copy on every accelerator is wrong because it describes data parallelism, the opposite setup from the layer-partitioned scheme the question describes, and it would not even fit if no single accelerator can hold the whole model. The option describing splitting individual weight matrices within each layer across accelerators is wrong because it describes tensor (intra-layer) parallelism, the Megatron-LM style of splitting a layer's own matrices rather than pipelining whole layers across micro-batches; it solves a different problem (fitting one oversized layer) via a different mechanism. The option describing discarding and recomputing activations is wrong because that is GPipe's separate activation-recomputation (re-materialization) technique for reducing memory footprint, not its mechanism for keeping pipeline stages busy during a training step.
Source: Huang, Y., Cheng, Y., Bapna, A. et al. (2018), "GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism," arXiv:1811.06965
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 064/066easy
Sennrich et al.'s byte-pair encoding (BPE) merges adjacent symbol pairs deterministically according to a fixed learned merge list, so the same word is always split into the exact same sequence of subword tokens every time it appears. Provilkov, Emelianenko and Voita's BPE-dropout instead trains on multiple different segmentations of that same word, using the same BPE vocabulary. How does BPE-dropout actually produce these multiple segmentations?
AIt replaces BPE's merge list with a probabilistic Unigram language model, of the kind SentencePiece implements, which assigns a likelihood to each possible segmentation and samples among the highest-likelihood options
BIt trains a separate neural segmentation model to propose alternative splits of each word, which are then merged with BPE's own output before tokenization
CIt keeps BPE's merges fully deterministic but randomly shuffles the order in which otherwise-valid merges are applied, producing a different final sequence of subword tokens purely from that reordering
DAt each step of the standard BPE merge procedure, it randomly drops some otherwise-applicable merges with a fixed probability, forcing the algorithm to fall back to smaller subwords at those points, so the same word can end up split differently from one training pass to the next
Correct answer: .
BPE-dropout keeps the same learned BPE vocabulary and merge rules, but during training it randomly drops out, with a small fixed probability, some of the merges that would otherwise apply at each step of the standard BPE procedure; when a merge is dropped, the algorithm is forced to leave that position split into smaller subwords instead, so the same word can come out segmented differently from one training pass to the next, all while remaining fully compatible with ordinary deterministic BPE used unchanged at inference. Training on this variety of segmentations makes the model less dependent on any single fixed split of a word and more robust to segmentation variation, which Provilkov et al. report improves translation quality by up to 3 BLEU over standard BPE. The option describing a switch to a Unigram language model is wrong because that describes Kudo's distinct SentencePiece algorithm, which the paper explicitly contrasts as the prior way of obtaining multiple segmentations; BPE-dropout instead stays within the BPE merge procedure itself rather than switching algorithms. The option describing a separate neural segmentation model is wrong because no additional model is trained at all; the randomness is injected directly into the existing deterministic BPE merge steps. The option describing reordering merges is wrong because BPE-dropout does not reorder anything; it stochastically omits certain merges outright, which changes the final token boundaries, rather than merely processing an unchanged set of merges in a different sequence.
Source: Provilkov, I., Emelianenko, D. & Voita, E. (2020), "BPE-Dropout: Simple and Effective Subword Regularization," ACL 2020, arXiv:1910.13267
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 065/066hard
Katharopoulos et al. (2020), "Transformers are RNNs," rewrite self-attention by replacing the softmax similarity between queries and keys with a kernel feature map applied separately to each query and key, then relying on the associativity of matrix multiplication to reorder the computation. What complexity does this reformulation achieve for computing attention over a sequence of length N, and what additional capability does it unlock for autoregressive generation?
AIt achieves the same quadratic time complexity as standard softmax attention, since the kernel feature map is applied only after the full pairwise query-key similarity matrix has already been computed, but it reduces memory use by never fully materializing that matrix on-chip
BIt reduces attention's time complexity from quadratic to linear in sequence length, and because the result can be expressed as a running sum of key-value outer products, it lets autoregressive generation update that running sum incrementally one token at a time, like the hidden state of a recurrent network, instead of reprocessing the full sequence at every step
CIt reduces attention's time complexity to log-linear by applying the kernel feature map only to a sparse, fixed subset of key positions chosen ahead of time for each query, similar to the local-plus-strided patterns used by sparse attention
DIt leaves attention's time complexity unchanged at quadratic, but achieves its reported speedup entirely by fusing the softmax and matrix-multiplication steps into a single GPU kernel to reduce memory reads and writes
Correct answer: .
By replacing softmax(QK^T) with phi(Q)phi(K)^T for a kernel feature map phi, attention's output can be rewritten, via the associativity of matrix multiplication, as phi(Q) multiplied by a running sum of phi(K_i) outer-product V_i terms accumulated across positions, without ever forming the full N-by-N pairwise similarity matrix; this drops both the time and memory cost of computing attention over a sequence of length N from quadratic to linear. For autoregressive generation, that running sum plays the same role as an RNN's hidden state: at each new token the model updates the single accumulated sum with that position's contribution and reads off the next output in constant time per step, instead of recomputing attention over the whole growing sequence, which the authors report yields speedups of up to 4000x at long sequence lengths. The option describing an unchanged quadratic computation with the similarity matrix formed first is wrong because the entire point of the kernel reformulation is to avoid ever forming that matrix, not to compute it and then discard it from memory. The option describing a fixed sparse subset of key positions is wrong because that describes sparse or windowed attention approaches, such as the Sparse Transformer or Longformer, which change which positions get attended to; linear attention still lets every query attend to every prior key, just computed through the kernel trick rather than restricted. The option describing a fused GPU kernel that reduces memory reads and writes while keeping the same quadratic computation is wrong because that describes FlashAttention's exact, IO-aware approach, the opposite strategy from Katharopoulos et al.'s, who change the mathematical form of attention itself to eliminate quadratic cost rather than only optimizing its memory-access pattern.
Source: Katharopoulos, A., Vyas, A., Pappas, N. & Fleuret, F. (2020), "Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention," ICML 2020, arXiv:2006.16236
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 066/066medium
Grave et al. (2017) introduce adaptive softmax to speed up training and inference for neural language models with very large output vocabularies, where computing and normalizing a full softmax over every vocabulary word at each position is a major computational bottleneck, particularly on GPUs. Unlike the softmax bottleneck described elsewhere in this exam, which concerns the softmax's limited expressiveness regardless of how fast it runs, adaptive softmax is specifically a computational speed optimization. How does adaptive softmax actually reduce this computational cost?
AIt groups vocabulary words into clusters based on their frequency, computing the expensive full-dimensional softmax only over a small head cluster of the most frequent words plus one shorthand symbol per rarer cluster, and only expands a rarer cluster's full softmax over its member words on the comparatively few occasions a word from that cluster is actually the target
BIt replaces the softmax function with a sigmoid applied independently to each vocabulary word's score, turning the single multi-way classification into many independent binary classifications computed in parallel with no normalization step
CIt reduces the vocabulary itself by merging rare words into a single shared unknown-word token before training begins, so the softmax only ever needs to normalize over the smaller set of remaining frequent words
DIt keeps the full softmax over the entire vocabulary unchanged at every position, but moves that computation from the GPU to the CPU, where large matrix-vector products over a long vocabulary dimension are reportedly cheaper
Correct answer: .
Adaptive softmax exploits the highly unbalanced, roughly Zipfian frequency distribution of words in natural-language corpora: it partitions the vocabulary into a small head cluster containing the most frequent words plus a compact set of placeholder symbols, one per lower-frequency cluster, and one or more much larger tail clusters of progressively rarer words. Computing the head cluster's full softmax is cheap since it is small, and each much larger tail cluster's own full softmax over its member words only needs to be computed on the comparatively rare training examples whose actual target word falls in that cluster; this lowers the average computational cost per position far below that of one full softmax over the entire vocabulary every single time, and the clustering is deliberately structured around matrix-matrix operations efficient on GPUs. The option describing a sigmoid applied independently to each word is wrong because that describes a different strategy, closer to candidate sampling or noise-contrastive estimation, which discards the shared normalization structure entirely rather than adaptive softmax's actual approach of keeping normalized mini-softmaxes within a cluster hierarchy. The option describing merging rare words into a shared unknown-word token is wrong because adaptive softmax does not discard or collapse any vocabulary word; every word remains individually distinguishable, just scored through a cheaper cluster hierarchy rather than removed from the vocabulary. The option describing moving computation to the CPU is wrong because the method's explicit motivation, named in its own title, is to be efficient specifically on GPUs through its clustering structure, not to avoid the GPU altogether.
Source: Grave, E., Joulin, A., Cissé, M., Grangier, D. & Jégou, H. (2017), "Efficient softmax approximation for GPUs," ICML 2017, arXiv:1609.04309