passdrill

How LLMs Work: Transformers & Training

66 cards · AI & LLM Engineering · answer each one, then read the explanation. Your score tallies below.

0 / 66 answered · 0 correct

AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 001/066 medium

In the Transformer architecture introduced by Vaswani et al. (2017) in "Attention Is All You Need," the dot products between queries and keys are divided by the square root of the key dimension before the softmax is applied. What is the primary reason for this scaling step?

  1. It converts the attention weights into a probability distribution that sums to one across the sequence
  2. It reduces the number of parameters needed in the query, key, and value projection matrices
  3. It counteracts dot products growing large in magnitude for larger key dimensions, which would push the softmax into regions with extremely small gradients
  4. It allows the same attention weights to be reused across all heads in multi-head attention
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 002/066 easy

Vaswani et al. (2017) use multi-head attention instead of a single attention function operating on the full-dimensional queries, keys, and values. What was the main motivation for splitting attention into multiple heads?

  1. It lets the model jointly attend to information from different representation subspaces at different positions, since a single attention head would average this into one weighted combination
  2. It doubles the effective context window length by processing two halves of the input sequence in parallel
  3. It removes the need for positional encoding, because each head learns to represent a different position offset
  4. It ensures that every attention head produces an identical set of attention weights, providing an ensembling effect through redundancy
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 003/066 easy

The original Transformer architecture adds sinusoidal positional encodings to the input token embeddings before the first attention layer. Why is this necessary given how self-attention itself processes a sequence?

  1. Because the feed-forward sublayers in each block are recurrent and require explicit position indices to update their hidden state between positions
  2. Because self-attention computes a weighted sum over all positions regardless of order, so without added positional information the model would treat the input as an unordered set of tokens
  3. Because the token embedding layer cannot represent more than a fixed vocabulary size unless positional offsets are added to each embedding
  4. Because positional encodings replace the residual connections that would otherwise be required around each sublayer
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 004/066 medium

Sennrich et al. (2016), "Neural Machine Translation of Rare Words with Subword Units," introduced byte-pair encoding (BPE) as a tokenization method for handling rare and out-of-vocabulary words. How does the BPE algorithm build its subword vocabulary?

  1. It assigns a fixed-length numeric code to each whole word in a predetermined dictionary, discarding any word not already present in that dictionary
  2. It randomly samples character sequences of varying length from the training corpus until a target vocabulary size is reached
  3. It splits text purely along whitespace and punctuation boundaries, performing no further segmentation within a word
  4. It starts from individual characters and repeatedly merges the most frequent adjacent pair of symbols into a new symbol, learning a fixed number of merge operations to build a subword vocabulary
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 005/066 medium

Decoder-only autoregressive language models such as GPT (Radford et al., 2018) use masked self-attention during training. What does this masking do, and why is it necessary?

  1. It prevents the model from attending to tokens more than a fixed window size away, purely to reduce the compute cost of attention
  2. It forces every attention head to attend only to the current token itself, ignoring the representations at all other positions
  3. It sets the attention scores for future positions to negative infinity before the softmax, so a token's representation depends only on itself and earlier tokens, matching the left-to-right generation process
  4. It removes the value projection for positions beyond the current one, while still letting their key vectors influence the attention weights
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 006/066 hard

Hoffmann et al. (2022), "Training Compute-Optimal Large Language Models" (the Chinchilla paper), trained over 400 models to study how model size and training data should be allocated for a fixed compute budget. What was their central finding?

  1. For a fixed training compute budget, model parameters and training tokens should be scaled at roughly the same rate, implying substantially larger training datasets relative to model size than many earlier large models had used
  2. Model performance becomes essentially independent of the number of training tokens once the parameter count exceeds a few billion
  3. Compute-optimal training requires increasing model size much faster than training data, since parameter count is the dominant driver of capability
  4. The optimal ratio of training tokens to parameters decreases as total training compute increases, so very large models need proportionally less data
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 007/066 medium

Ouyang et al. (2022), "Training Language Models to Follow Instructions with Human Feedback" (the InstructGPT paper), describes a three-stage pipeline for aligning a pretrained language model with human preferences. What are the three stages, in order?

  1. Pretrain on a labeled classification dataset, then deploy directly without further tuning, since pretraining alone is assumed to capture human preferences
  2. Supervised fine-tuning on human-written demonstrations, followed by training a reward model on human rankings of sampled model outputs, followed by reinforcement learning (PPO) that optimizes the policy against that reward model
  3. Train a reward model on raw internet text first, then use it to filter the pretraining corpus before any supervised fine-tuning occurs
  4. Fine-tune the model using only reinforcement learning against a fixed rule-based reward function, without using any human-labeled data at any stage
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 008/066 easy

Each sublayer in the Transformer encoder and decoder (Vaswani et al., 2017) is wrapped with a residual connection before layer normalization is applied. What role do these residual connections play in training deep stacks of transformer blocks?

  1. They reduce the total number of attention heads needed by allowing heads to share weights across different layers
  2. They eliminate the need for a softmax normalization step within the self-attention mechanism
  3. They replace the position-wise feed-forward sublayer with a simple identity function during inference, to speed up generation
  4. They add each sublayer's input directly to its output before normalization, giving gradients a direct path backward through the addition and easing optimization of many stacked layers
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 009/066 easy

Ba, Kiros and Hinton (2016), "Layer Normalization," introduced a normalization technique widely used in transformer blocks. How does layer normalization compute its statistics, and how does this differ from batch normalization?

  1. It normalizes each feature by computing statistics across every example in the mini-batch, exactly as batch normalization does, but applies the result at a different point in the network
  2. It normalizes activations using running statistics collected only during a separate calibration pass performed after training has finished
  3. It computes the mean and variance across the feature dimension for each individual training example, making it independent of batch size, whereas batch normalization computes statistics across the batch for each feature
  4. It normalizes only the query and key projections used in self-attention, leaving the value projection and feed-forward outputs unnormalized
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 010/066 easy

In addition to the self-attention sublayer, each Transformer block (Vaswani et al., 2017) contains a position-wise feed-forward network. How is this feed-forward network applied across the sequence?

  1. It is applied identically and independently to each position in the sequence, typically expanding to a larger inner dimension through a nonlinearity before projecting back down
  2. It combines information across all positions simultaneously in the same way self-attention does, making it functionally redundant with the attention sublayer
  3. It only operates on the final position of the sequence, producing a single pooled representation for the entire input
  4. It shares its weights with the token embedding layer so that no additional parameters are introduced by including it
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 011/066 hard

Shazeer et al. (2017), "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer," introduced a technique now used in some large language models to scale model capacity without a proportional increase in per-token compute. How does a sparsely-gated mixture-of-experts (MoE) layer achieve this?

  1. A gating network activates every expert sub-network for every input and averages their outputs, providing an ensembling effect at the cost of proportionally higher compute
  2. A trainable gating network routes each input to a sparse subset of expert feed-forward sub-networks, letting total parameter count scale far beyond what would be affordable if every expert were computed for every input
  3. Each expert is a full independent copy of the entire transformer stack, and the gating network selects which single copy runs the whole forward pass for a given input
  4. The gating network is fixed at random initialization and never trained, relying purely on random routing to balance load across experts
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 012/066 easy

Extending a transformer-based language model's context window to accept much longer input sequences is architecturally expensive without further modifications to the standard attention mechanism. What property of self-attention, as defined by Vaswani et al. (2017), causes this?

  1. Compute and memory for self-attention scale linearly with sequence length, so extending the context window costs proportionally the same amount regardless of length
  2. The context window is limited only by the size of the token vocabulary, not by any property of the attention computation itself
  3. Self-attention's cost depends solely on the number of stacked layers in the model and is unaffected by how many tokens are in the input sequence
  4. Self-attention computes a pairwise compatibility score between every two positions in the sequence, so both compute and memory scale quadratically with sequence length, making much longer context windows substantially more expensive
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 013/066 medium

In Section 3.4 of "Attention Is All You Need" (Vaswani et al., 2017), the authors describe sharing weights between two embedding layers and the pre-softmax linear transformation, then multiplying the embedding weights by the square root of the model dimension. What is the primary purpose of this weight-sharing scheme?

  1. It lets the encoder and decoder share the same set of self-attention weights, so a single set of query/key/value matrices is reused in every layer of the network
  2. It reduces the depth of the network by merging the final feed-forward sublayer directly into the output embedding matrix
  3. It ties the same learned matrix to both converting tokens into vectors and converting the decoder's final vectors back into vocabulary logits, reducing the total parameter count and linking the input and output token representations
  4. It guarantees that every possible output token receives an identical probability at initialization, which is later corrected during fine-tuning
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 014/066 easy

The Transformer decoder (Vaswani et al., 2017) contains an "encoder-decoder attention" sublayer in addition to its own masked self-attention sublayer. Where do the queries, keys, and values for this encoder-decoder attention sublayer come from?

  1. The queries come from the decoder's own previous sublayer, while the keys and values come from the output of the encoder stack, letting every decoder position attend over the entire input sequence
  2. The queries, keys, and values all come from the encoder's final layer, and the decoder only reads the resulting attention output without contributing any of its own vectors
  3. The queries and keys come from the encoder output while the values come from the decoder's own previous sublayer, reversing the usual roles of queries and keys
  4. The queries, keys, and values all come from the decoder's own previous sublayer, identical to the masked self-attention sublayer that precedes it
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 015/066 hard

Su et al. (2021), "RoFormer: Enhanced Transformer with Rotary Position Embedding," introduce rotary position embedding (RoPE) as an alternative to adding sinusoidal positional vectors to token embeddings. How does RoPE encode positional information?

  1. It concatenates a learned position index as an extra dimension appended to every query and key vector before the dot product is computed
  2. It rotates each query and key vector by an angle that depends on the token's position and a per-dimension frequency, so that the dot product between a rotated query and key naturally depends on their relative distance rather than their absolute positions alone
  3. It replaces the dot-product attention score with a lookup table indexed by the absolute position of the query, discarding the key vector's position entirely
  4. It multiplies the attention output by a fixed sinusoidal mask after the softmax has already been applied, without altering the query or key vectors themselves
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 016/066 medium

Press, Smith and Lewis (2021), "Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation," introduce ALiBi as a way to let a model trained on short sequences generalize to much longer ones at inference. How does ALiBi represent position, and how does this differ from adding sinusoidal or learned positional embeddings to the input?

  1. It adds a learned positional embedding to each token exactly as the original Transformer does, but recomputes those embeddings after every training step so they extrapolate to unseen lengths
  2. It removes positional information entirely, relying only on the order in which tokens are fed into the model during training to implicitly teach it sequence order
  3. It replaces the query and key projections with position-specific weight matrices, so each position in the sequence uses a completely separate set of learned parameters
  4. It adds no positional embeddings to the word embeddings at all; instead, it subtracts a static, non-learned penalty from each attention score that grows in proportion to the distance between the query and key positions, with the penalty's steepness set differently per attention head
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 017/066 easy

GPT-2 and GPT-3's position-wise feed-forward sublayers use the Gaussian Error Linear Unit (GELU), defined by Hendrycks and Gimpel (2016), instead of the plain ReLU activation used in some earlier architectures. What distinguishes GELU from ReLU?

  1. GELU weights an input by the value of the standard Gaussian cumulative distribution function evaluated at that input, producing a smooth curve that can pass through small negative values rather than clamping every negative input to exactly zero
  2. GELU is identical to ReLU for all positive inputs and outputs a small fixed negative constant for every negative input, similar to Leaky ReLU
  3. GELU replaces the two linear transformations in the feed-forward sublayer with a single linear transformation, removing the nonlinearity between them entirely
  4. GELU is an activation function that operates only on the attention scores after the softmax, rather than within the position-wise feed-forward sublayer
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 018/066 easy

Decoder-only language models such as GPT are pretrained on a simple self-supervised objective before any instruction tuning or reinforcement learning stage is applied. What is that pretraining objective?

  1. Predicting whether two randomly paired sentences from the corpus originally appeared next to each other in the source text, using a binary classification loss
  2. Reconstructing a small number of randomly masked-out tokens scattered throughout an otherwise-visible input sequence, using a masked-token classification loss
  3. Predicting the next token in a sequence given only the tokens that came before it, using a cross-entropy loss between the model's predicted probability distribution and the actual next token, repeated across every position in the training corpus
  4. Predicting a scalar reward score for an entire generated sequence, using a regression loss trained on human preference comparisons
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 019/066 medium

When an autoregressive Transformer decoder generates text one token at a time, naively recomputing self-attention from scratch at every step would repeat a large amount of work. What does key-value (KV) caching do to avoid this, and what does it not need to recompute?

  1. It stores the entire attention-score matrix from the very first generation step and reuses those exact scores unchanged for every later token, regardless of what new token is generated
  2. It skips computing new query vectors for later tokens, reusing the query vector from the first generated token for every subsequent generation step
  3. It discards the key and value vectors after each step and instead caches only the final output logits, replaying them directly for the next step
  4. It stores the key and value vectors computed for every previously generated token so that, at each new step, only the new token's query, key, and value need to be computed, with that new query then attended over the cached keys and values, rather than recomputing keys and values for the whole sequence so far
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 020/066 medium

Xiong et al. (2020), "On Layer Normalization in the Transformer Architecture," compare two ways of placing layer normalization relative to each sublayer's residual connection: Post-LN, used in the original Transformer, and Pre-LN. What did they find about training stability, and what practical consequence follows for Post-LN training?

  1. Pre-LN and Post-LN produce identical gradient magnitudes at initialization, so the choice between them has no measurable effect on how training proceeds
  2. Post-LN, which applies layer normalization after the residual addition, produces expected gradients near the output layer that grow large at initialization, which the authors show makes a learning-rate warm-up stage necessary; placing layer normalization inside the residual block instead (Pre-LN) keeps gradients well-behaved at initialization without requiring warm-up
  3. Pre-LN requires a longer warm-up stage than Post-LN because normalizing before each sublayer slows down how quickly gradient magnitudes stabilize during the first training steps
  4. The placement of layer normalization only affects inference-time computation cost and has no bearing on gradient behavior or the need for a warm-up stage during training
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 021/066 easy

Brown et al. (2020), "Language Models are Few-Shot Learners," describe evaluating GPT-3 in a few-shot setting on many tasks. What does "few-shot" mean in this evaluation, and what happens to the model's weights during it?

  1. A small number of labeled examples for the task are used to fine-tune GPT-3's weights with a few additional gradient-descent steps before the model is evaluated on new examples
  2. A small number of task examples are included as text directly within the prompt given to the model at inference time, and GPT-3 is evaluated on new examples using this prompt alone, with no gradient updates or fine-tuning of its weights performed for the task
  3. GPT-3's weights are duplicated into several smaller copies, each fine-tuned on a few examples from a different task, and the copy that performs best on a validation set is kept
  4. A small, separate classifier head is attached to GPT-3 and trained from scratch on a few labeled examples, while the rest of the pretrained model is frozen
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 022/066 hard

Shazeer (2019), "Fast Transformer Decoding: One Write-Head is All You Need," introduces multi-query attention as a modification to standard multi-head attention aimed at speeding up autoregressive decoding. What does multi-query attention change relative to standard multi-head attention, and why does this help?

  1. It reduces the number of query heads down to a single shared query head, while each head keeps its own separate key and value projections, cutting the number of query computations performed at each step
  2. It removes the key and value projections entirely, having every attention head operate directly on the raw token embeddings instead of projected keys and values
  3. It increases the number of key and value heads beyond the number of query heads, giving each query head access to several redundant copies of the same key and value vectors
  4. It keeps multiple separate query heads but has all of them share a single set of key and value projections, shrinking the size of the cached keys and values that must be stored and read back at every decoding step, which reduces the memory-bandwidth cost that dominates incremental decoding
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 023/066 medium

Dao et al. (2022), "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness," speed up the standard self-attention computation on GPUs without changing the attention mechanism's mathematical result. What kind of optimization does FlashAttention make, and what does it NOT do?

  1. It reduces the number of reads and writes between the GPU's high-bandwidth memory and its much faster on-chip SRAM by tiling the computation and avoiding materializing the full attention-score matrix, while still computing exactly the same attention output as the standard formula, not an approximation of it
  2. It approximates the full quadratic attention computation with a sparse or low-rank attention pattern, trading some accuracy in the attention output for a reduction in the number of floating-point operations performed
  3. It reduces the number of floating-point operations attention requires by lowering the numerical precision of the query and key vectors, at the cost of a less numerically exact attention output
  4. It restructures the attention computation to run entirely within CPU memory instead of GPU memory, trading GPU compute for cheaper CPU compute
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 024/066 easy

Radford et al. (2019), "Language Models are Unsupervised Multitask Learners" (the GPT-2 paper), tokenize text using a byte-level variant of byte-pair encoding rather than applying BPE directly to Unicode characters. What problem does operating on bytes instead of Unicode characters solve?

  1. It allows the tokenizer to skip the merge-based vocabulary-building process entirely, since every byte value can be mapped directly to a token without any merges
  2. It increases the base vocabulary size well beyond 100,000 symbols before any merges are added, giving the model more starting granularity to work with
  3. It keeps the base vocabulary, before any merges, small and fixed at 256 symbols (one for every possible byte value), avoiding the far larger base vocabulary that applying BPE directly to Unicode characters would require, and letting the model assign a probability to any input without ever needing an unknown-token symbol
  4. It removes the need for any subword merging at all, since GPT-2 processes each byte independently as its own token throughout generation
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 025/066 medium

Ainslie et al. (2023), "GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints," introduce grouped-query attention as an intermediate design between standard multi-head attention and multi-query attention. How does GQA position itself between these two extremes?

  1. It keeps a single shared key and value head for the entire attention layer, exactly as multi-query attention does, but restores full quality by giving every query head its own independent query projection matrix
  2. It divides the query heads into a fixed number of groups, with all query heads in a group sharing one key and value head; setting the group count equal to the number of query heads recovers standard multi-head attention, and setting it to one recovers multi-query attention
  3. It keeps every query head paired with its own key and value head as in standard multi-head attention, but reduces the total number of query heads to cut memory bandwidth
  4. It replaces the key and value projections with a single shared feed-forward network applied after attention, removing the need for separate key and value weight matrices altogether
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 026/066 hard

Kaplan et al. (2020), "Scaling Laws for Neural Language Models," and Hoffmann et al. (2022), "Training Compute-Optimal Large Language Models" (the Chinchilla paper), both fit power-law relationships between loss, model size, and training data, but reached different practical recommendations for how to spend a fixed compute budget. What is the key difference between their conclusions?

  1. Kaplan et al. concluded model size and data should scale in roughly equal proportion as compute grows, while Hoffmann et al. concluded model size should grow much faster than data
  2. Both papers reached identical conclusions about compute-optimal allocation; Hoffmann et al. only confirmed Kaplan et al.'s original recipe with a larger set of models
  3. Kaplan et al. found that data size mattered more than model size for a fixed compute budget, while Hoffmann et al. found the opposite, that model size should be prioritized above all else
  4. Kaplan et al.'s fitted power laws recommended training very large models on a comparatively modest amount of data, stopping well short of convergence, whereas Hoffmann et al.'s larger and more careful set of training runs found that model size and training data should instead be scaled in roughly equal proportion, implying many contemporary large models were oversized and undertrained relative to their compute budgets
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 027/066 easy

Kudo and Richardson (2018), introducing the SentencePiece toolkit, implement a Unigram language model tokenization algorithm as an alternative to byte-pair encoding (BPE) for building a subword vocabulary. How does the Unigram algorithm build that vocabulary, in contrast to BPE?

  1. It starts from a large seed vocabulary of candidate subwords and iteratively removes the subwords that contribute least to the likelihood of the training corpus under a unigram language model, continuing until the target vocabulary size is reached, rather than building the vocabulary up from individual characters by repeatedly merging the most frequent adjacent pair
  2. It builds the vocabulary bottom-up by repeatedly merging the two most frequent adjacent symbols into a new symbol, identical to BPE, but differs only in how the final tokenizer chooses a single segmentation at inference time
  3. It requires the input text to already be split into whitespace-separated words before subword modelling begins, unlike BPE which can operate on raw, undelimited text
  4. It assigns every character its own fixed, permanent token and never combines characters into larger subword units, guaranteeing a constant vocabulary size regardless of corpus
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 028/066 medium

Rafailov et al. (2023), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model," propose DPO as an alternative to the reward-model-plus-PPO pipeline used in InstructGPT-style RLHF. What does DPO change about this pipeline?

  1. It replaces the pairwise human-preference comparisons used to train the reward model with a fully automated LLM-as-judge scoring system, while keeping the separate reward model and the PPO reinforcement-learning step unchanged
  2. It trains the reward model exactly as InstructGPT does, but replaces the PPO optimizer with a simpler supervised fine-tuning step on the reward model's top-scoring completions
  3. It uses a closed-form solution to the KL-regularized reward-maximization objective to reparameterize the reward directly in terms of the policy's own output probabilities relative to a reference policy, letting the model be optimized directly on preference pairs with a simple classification-style loss, without ever training a separate reward model or running an online reinforcement-learning sampling loop
  4. It removes the need for any human or model preference data at all, instead deriving alignment purely from the pretraining objective
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 029/066 hard

Wei et al. (2022), "Emergent Abilities of Large Language Models," reported that model performance on certain tasks stays near chance level until a threshold scale, then rises sharply. Schaeffer et al. (2023), "Are Emergent Abilities of Large Language Models a Mirage?," challenge this framing. What is Schaeffer et al.'s core argument?

  1. They show that the sharp jumps disappear entirely once larger models are trained, proving that Wei et al.'s reported tasks were measurement errors caused by insufficient training data
  2. They argue that the apparent sharp, unpredictable jumps are largely an artifact of using nonlinear or discontinuous metrics, such as strict exact-match accuracy on multi-step tasks; when the same underlying model outputs are instead scored with a smoother, partial-credit metric, performance improves gradually and predictably with scale rather than jumping
  3. They argue that emergent abilities are real and predictable in advance from a model's parameter count alone, and propose a formula that lets any lab compute the exact scale at which a new ability will appear before training a model
  4. They argue that emergent abilities only ever appear in decoder-only architectures and never in encoder-decoder models, regardless of how performance is measured
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 030/066 easy

Fedus, Zoph and Shazeer (2021), "Switch Transformer: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity," simplify the sparsely-gated mixture-of-experts design that Shazeer et al. (2017) had used with a top-k gating function. What change does the Switch Transformer make to expert routing?

  1. It removes expert routing altogether and instead sends every token through every expert, then averages all of the experts' outputs
  2. It increases the number of experts activated per token from Shazeer et al.'s original two experts to a much larger fixed number, such as sixteen, to improve output quality
  3. It replaces the learned gating network with a fixed, hand-designed rule that assigns tokens to experts based on the token's position in the sequence rather than its content
  4. It routes each token to exactly one expert (top-1 routing) instead of combining outputs from several experts per token, reducing routing computation and communication cost while relying on a load-balancing loss to keep experts evenly utilized
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 031/066 easy

Beltagy, Peters and Cohan (2020), "Longformer: The Long-Document Transformer," replace full self-attention with a combination of attention patterns to process much longer sequences efficiently. What is the core mechanism behind Longformer's efficiency gain?

  1. Most tokens use a fixed-size sliding-window attention, attending only to nearby tokens within that window, which reduces the memory and compute cost from quadratic in sequence length to linear; a small number of tokens are additionally given full global attention to preserve task-relevant long-range information
  2. Every token still attends to every other token exactly as in standard self-attention, but the attention scores are computed in a lower floating-point precision to save memory, with no change to which tokens attend to which
  3. Tokens are grouped into fixed-size non-overlapping blocks, and attention is computed only between tokens in the same block, with no mechanism for any token to see information outside its own block
  4. The sequence is first compressed into a much shorter fixed-length summary using a separate pooling network, and standard full attention is then applied only to that shortened summary
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 032/066 medium

Chen et al. (2023), "Extending Context Window of Large Language Models via Positional Interpolation," extend the usable context length of RoPE-based pretrained models such as LLaMA with only minimal fine-tuning. What does their Position Interpolation method actually do to the model's position indices?

  1. It extrapolates position indices beyond the range seen during pretraining, feeding the model position values larger than any it encountered in training, and relies on RoPE's periodicity to generalize correctly to these unseen positions
  2. It discards the rotary position embedding scheme entirely and replaces it with the original sinusoidal position encoding from Vaswani et al. (2017), which the authors show generalizes to longer sequences without any fine-tuning
  3. It linearly rescales the new, longer sequence's position indices down so that the largest index still falls within the range the model saw during pretraining, then fine-tunes briefly on this rescaled range, avoiding the out-of-distribution position values that naive extrapolation would produce
  4. It increases the model's embedding dimension so that each position can be represented with more precision, allowing longer sequences to be distinguished without changing how position indices are computed
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 033/066 easy

Leviathan, Kalman and Matias (2023), "Fast Inference from Transformers via Speculative Decoding," speed up autoregressive generation from a large target model without changing its output distribution. How does their method achieve this?

  1. It replaces the large target model with a smaller model for the entire generation, only occasionally calling the large model to spot-check a sample of the output afterward, which changes the output distribution slightly but saves the most compute
  2. A smaller, cheaper draft model proposes several candidate next tokens; the large target model then verifies these candidates in a single parallel forward pass and accepts the longest prefix consistent with its own probability distribution via a rejection-sampling scheme, so the final output is distributed exactly as if the large model had generated every token itself, just with fewer serial large-model calls
  3. It caches the large model's key and value tensors from previous generation steps so that they never need to be recomputed at later steps, removing redundant self-attention computation
  4. It reorders the large model's attention computation to be more IO-aware, fusing operations to avoid writing large intermediate attention matrices to slow GPU memory, without changing which tokens are generated
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 034/066 easy

Hinton, Vinyals and Dean (2015), "Distilling the Knowledge in a Neural Network," train a small "student" network using the outputs of a larger, already-trained "teacher" network. What specifically does the student learn to match, and why does raising the softmax temperature matter?

  1. The student is trained only on the teacher's single most likely predicted class for each training example, exactly as it would be trained on the original ground-truth hard labels, and temperature is only used to speed up the teacher's inference
  2. The student learns to reproduce the teacher's internal hidden-layer activations exactly, layer by layer, and temperature controls how many of the teacher's layers the student is required to match
  3. The student is trained to minimize the difference between its own raw, un-normalized output scores and the teacher's raw output scores, with temperature having no effect on training since it is only applied at test time
  4. The student is trained to match the teacher's full "soft target" probability distribution over all classes, not just the single correct label; raising the softmax temperature when computing both the teacher's and the student's distributions during training softens the probabilities, revealing the relative probabilities the teacher assigns to incorrect classes ("dark knowledge") that a hard label alone would hide, and the student reverts to a temperature of 1 at deployment
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 035/066 medium

Zhang and Sennrich (2019), "Root Mean Square Layer Normalization," propose RMSNorm as a simplification of LayerNorm later adopted by architectures such as LLaMA and Mistral. What does RMSNorm remove relative to standard LayerNorm, and what property does it rely on instead?

  1. RMSNorm removes the mean-centering (re-centering) step and does not subtract the mean of the summed inputs; it instead rescales activations only by their root-mean-square value, keeping re-scaling invariance while dropping re-centering invariance
  2. RMSNorm removes the learnable scale (gain) parameter entirely, normalizing activations using only their mean and variance exactly as LayerNorm does but without any parameters left to learn
  3. RMSNorm removes normalization from all but the final transformer block, applying LayerNorm's full mean-and-variance computation only once at the network's output instead of at every block
  4. RMSNorm removes normalization within a layer and instead normalizes each token's embedding relative to every other token's embedding in the same batch, a form of batch normalization applied along the sequence axis
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 036/066 easy

Mikolov et al. (2013) introduced word2vec, which learns a single fixed vector for each vocabulary word from a large text corpus. Why do transformer-based language models rely on contextual embeddings instead of word2vec-style static embeddings alone?

  1. Static embeddings like word2vec cannot be trained on large text corpora at all, so no fixed-vector representation of word meaning can be learned without first having a transformer architecture available
  2. A static embedding assigns a word exactly one vector regardless of context, so a polysemous word keeps the same representation in every sentence; self-attention instead produces a representation for each token that is computed from the other tokens surrounding it, letting the same word take on different representations in different contexts
  3. Static embeddings are only capable of representing nouns, while contextual embeddings are required to represent verbs, adjectives, and every other part of speech
  4. Contextual embeddings are simply static embeddings computed with a larger vocabulary size, so the difference is purely how many words the vocabulary contains rather than how each vector is computed
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 037/066 easy

Holtzman et al. (2019), "The Curious Case of Neural Text Degeneration," observed that greedy decoding and beam search from a language model often produce bland, repetitive text even though the model assigns that text high likelihood. What does their proposed nucleus (top-p) sampling method do differently at each decoding step?

  1. It always samples the single most probable token at each step, identical to greedy decoding, and then re-ranks the resulting full sequence afterward using a separate language model
  2. It restricts sampling to a fixed number of the k highest-probability tokens at every step, where k is a constant chosen once before generation begins and never changes
  3. It sorts tokens by probability and samples from the smallest set of tokens whose cumulative probability reaches a chosen threshold p, so the size of the candidate set shrinks when the model is confident and grows when the model is uncertain, avoiding both the unreliable low-probability tail and an arbitrary fixed-size cutoff
  4. It disables sampling entirely during generation and instead deterministically outputs the sequence with the single highest total log-probability across the whole output, computed exactly via dynamic programming
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 038/066 hard

Loshchilov and Hutter (2017), "Decoupled Weight Decay Regularization," introduce AdamW to fix how the original Adam optimizer interacts with weight decay. What specifically does AdamW change relative to plain Adam trained with an L2 penalty added to the loss?

  1. AdamW removes weight decay from the optimizer entirely, relying only on early stopping to prevent overfitting during large-scale pretraining
  2. AdamW replaces Adam's per-parameter adaptive learning rates with a single global learning rate shared by every parameter, which is what allows weight decay to work correctly
  3. AdamW applies weight decay only to the bias terms of the model and excludes every weight matrix from decay, reversing which parameters L2 regularization normally targets
  4. AdamW applies weight decay as a separate, direct shrinkage of the parameters at each update step, outside of the gradient-based adaptive moment estimation; with plain Adam, adding an L2 penalty to the loss instead folds the decay term into the gradient, where Adam's per-parameter adaptive scaling then distorts its effective strength, weakening decay on frequently-updated parameters and strengthening it on rarely-updated ones
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 039/066 easy

Large language model pretraining runs such as GPT-3's (Brown et al., 2020) commonly use a learning rate schedule that starts near zero, rises linearly for an initial warmup period, and then decays smoothly (often along a cosine curve) for the rest of training. Why is that initial warmup period used, rather than starting immediately at the peak learning rate?

  1. Early in training the model's weights are far from any well-conditioned region and gradients (and an adaptive optimizer's moment estimates) are still unreliable; taking large steps immediately at the full learning rate risks unstable updates or divergence, so warmup ramps the step size up gradually while the optimizer's statistics stabilize
  2. Warmup exists purely to save compute cost, since a smaller learning rate at the very start of training requires fewer floating-point operations per step than a larger one would
  3. Warmup is required only when using plain stochastic gradient descent without momentum, and serves no purpose whatsoever once the optimizer already has per-parameter adaptive learning rates such as Adam or AdamW
  4. Warmup gradually increases the model's context window length from a short initial length up to the full training sequence length, rather than gradually increasing the learning rate itself
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 040/066 medium

Micikevicius et al. (2017), "Mixed Precision Training," describe training large neural networks using FP16 (half precision) for most computations while still matching FP32 (single precision) training accuracy. What two techniques does their method rely on to prevent this switch to FP16 from degrading accuracy?

  1. Storing the model's activations in FP16 while performing the optimizer's weight update directly on FP16 weights, discarding any higher-precision copy of the weights once training begins
  2. Maintaining a master copy of the weights in FP32 that accumulates each step's small update (since FP16 cannot represent many small updates precisely enough), and multiplying the loss by a scaling factor before backpropagation so that small gradient values do not underflow to zero in FP16's limited numeric range, then unscaling the gradients before the optimizer step
  3. Rounding every FP32 value in the network down to the nearest representable FP16 value using stochastic rounding alone, with no other change needed to preserve accuracy
  4. Running two independent copies of the entire model, one in FP16 and one in FP32, and averaging their two sets of final predictions together at inference time
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 041/066 hard

Chen et al. (2016), "Training Deep Nets with Sublinear Memory Cost," introduce a technique now commonly called gradient checkpointing (or activation recomputation) for training very deep networks under tight memory budgets. What trade-off does this technique make, and how?

  1. It reduces memory by storing activations in a lower numeric precision than the rest of the network, without changing how many activations are stored or when they are computed
  2. It reduces memory by discarding the gradients of early layers entirely after they are used once, so those layers stop being updated for the rest of training
  3. It reduces the memory needed to store activations for backpropagation by saving only a subset of activations ("checkpoints") during the forward pass and recomputing the discarded intermediate activations on demand during the backward pass, trading roughly one extra forward pass's worth of compute for memory that can scale close to the square root of the number of layers instead of growing linearly with depth
  4. It reduces memory by splitting the model across multiple GPUs so that each device only ever holds a fraction of the total activations, without performing any extra recomputation on any device
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 042/066 medium

Mistral 7B (Jiang et al., 2023) replaces full self-attention with sliding window attention (SWA), where each token attends to at most W tokens from the previous layer. How does stacking multiple layers of this fixed local window still let the model capture dependencies far beyond W tokens away, and what does this enable for its key-value cache?

  1. Every layer's window is centered on a different, randomly chosen segment of the input, so across many layers the union of all windows eventually covers the entire sequence within a single layer's own computation
  2. It does not actually extend the effective receptive field beyond W tokens at all; Mistral instead relies entirely on a separate global-attention mechanism layered on top of SWA to reach tokens farther away, and its key-value cache still grows without bound
  3. Sliding window attention removes the need for a key-value cache entirely, since each token only ever looks at a small, fixed number of neighboring tokens and can discard all cache entries after every single step
  4. Because a token at a given layer can attend to tokens within W positions of it at the previous layer, and each of those tokens could itself attend up to W positions further back at the layer before that, the effective receptive field compounds with depth, reaching roughly W multiplied by the number of layers; this fixed window also lets Mistral use a rolling buffer cache of fixed size, overwriting the oldest entries as the window slides forward instead of letting the key-value cache grow without bound
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 043/066 easy

Transformer-based language models are commonly grouped into three architecture families based on which parts of the original Vaswani et al. (2017) encoder-decoder design they keep: encoder-only, decoder-only, and encoder-decoder. Which option correctly matches each family to a representative model and its typical use?

  1. Encoder-only models such as BERT use bidirectional self-attention over the whole input and are typically used for tasks that classify or extract information from existing text; decoder-only models such as GPT use causal (masked) self-attention and are typically used for open-ended text generation; encoder-decoder models such as the original Transformer and T5 use an encoder to process the input and a separate decoder, conditioned on the encoder's output, to generate the output, and are typically used for sequence-to-sequence tasks such as translation or summarization
  2. Encoder-only models such as GPT generate text autoregressively one token at a time; decoder-only models such as BERT are used only for classification tasks and cannot generate any text at all; encoder-decoder models are a deprecated design no longer used in any modern language model
  3. The three families differ only in how many layers they contain, with encoder-only models always having the fewest layers, decoder-only models a moderate number, and encoder-decoder models always having the most layers regardless of the task
  4. Encoder-only, decoder-only, and encoder-decoder all refer to the same self-attention computation; the labels only describe which software framework (such as TensorFlow, PyTorch, or JAX) was used to implement the model
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 044/066 easy

BERT (Devlin et al., 2018) is pretrained using a masked language modeling (MLM) objective, while GPT-style decoder-only models are pretrained using causal (next-token) language modeling. How does BERT's MLM objective differ from GPT's pretraining objective, and what does BERT's approach require during training that GPT's does not?

  1. Both objectives are identical: at every training step, both BERT and GPT predict only the very next token after everything seen so far, using leftward context alone in both cases
  2. BERT randomly replaces roughly 15% of input tokens (mostly with a special mask token, with the remainder left unchanged or swapped for a random token) and trains the model to predict the original identity of those replaced tokens using bidirectional context from both directions; GPT instead predicts each token from only the tokens before it, using causal self-attention that never looks ahead, so GPT's objective needs no masking of the input tokens themselves, only the causal attention mask restricting what each position can see
  3. BERT predicts every single token in the input on every training step, with none of them masked, while GPT is trained to predict only a randomly chosen 15% of tokens and ignores the rest
  4. BERT and GPT use the identical causal self-attention pattern during pretraining; the only real difference between them is that BERT is trained on longer documents than GPT
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 045/066 medium

In a transformer-based language model, what is the primary purpose of the softmax function applied to the attention scores (the dot products of queries and keys)?

  1. It normalises the attention scores into a probability distribution that sums to 1, ensuring each token attends to all other tokens with weights that reflect their relative relevance rather than their raw dot-product magnitudes
  2. It eliminates all negative attention scores so that only positively correlated tokens influence each other, preventing destructive interference in the hidden representations
  3. It replaces the attention mechanism with a simple averaging operation, giving every token in the sequence an equal weight regardless of the query-key similarity
  4. It is used exclusively during training to compute the cross-entropy loss and has no role during inference — at inference time the raw dot products are used directly without normalisation
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 046/066 medium

A team is deciding between using a 7-billion-parameter open-weight model fine-tuned on their domain data versus a 400-billion-parameter general-purpose API model with in-context learning (no fine-tuning). Which factor most strongly favours the smaller fine-tuned model?

  1. The task requires consistently following a narrow, domain-specific output format with specialised terminology, and the team has a high-quality labelled dataset of several thousand examples — a fine-tuned smaller model can learn this distribution directly, often outperforming a much larger model that must infer the pattern from a handful of in-context examples
  2. The team wants maximum flexibility to handle a wide variety of unpredictable user queries across many domains without retraining
  3. The team has no labelled data and cannot invest in creating any, making supervised fine-tuning impossible
  4. The task requires broad world knowledge and multi-step reasoning across topics the team cannot anticipate in advance
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 047/066 medium

Many modern open-weight LLMs (e.g. LLaMA, PaLM, Mistral) replace the original Transformer's ReLU-activated position-wise feed-forward network with a SwiGLU-based feed-forward network, following Shazeer's "GLU Variants Improve Transformer." What structurally distinguishes a SwiGLU feed-forward sublayer from a standard single-activation feed-forward sublayer?

  1. It replaces the two-layer feed-forward network with a single linear projection, removing the need for a nonlinearity at all, since the residual connection provides all necessary nonlinearity
  2. It adds a third feed-forward layer stacked after the original two, deepening the sublayer without changing how any individual layer computes its output
  3. It computes two separate linear projections of the input, applies the Swish (SiLU) activation to one of them, and multiplies the two projections element-wise before the final output projection, gating how much of each unit passes through
  4. It replaces the feed-forward sublayer's learned weight matrices with a fixed, non-trainable gating function computed directly from the token's position in the sequence
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 048/066 easy

Before RoPE and ALiBi, Shaw, Uszkoreit and Vaswani (2018) proposed an alternative to the original Transformer's approach of adding sinusoidal positional vectors to token embeddings before the first layer. How does their relative position representation approach encode positional information instead?

  1. It removes positional information entirely, relying only on the order in which tokens are fed through recurrent connections between layers
  2. It concatenates a one-hot vector of the token's absolute position to the token embedding before the first attention layer, doubling the embedding dimension
  3. It multiplies each attention head's query vector by a rotation matrix whose angle depends on the query's absolute position in the sequence
  4. It learns a separate embedding vector for each pairwise distance between a query and key position (clipped at a maximum), and adds this distance-indexed vector directly into the attention score and value computations at every layer, rather than adding any position vector to the input embeddings
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 049/066 easy

Lee et al. (2022), "Deduplicating Training Data Makes Language Models Better," found that large web-scraped pretraining corpora such as C4 contain a substantial number of near-duplicate documents and repeated substrings (in one case, a single sentence repeated tens of thousands of times). What did they find deduplicating this training data actually does to the resulting language models?

  1. It has no measurable effect on either memorization or downstream accuracy, since duplicate documents are rare enough in web-scraped corpora to be statistically negligible
  2. It substantially reduces how often trained models emit memorized text copied verbatim from the training data, while letting models reach the same or better accuracy in fewer training steps
  3. It increases verbatim memorization of the surviving unique documents, because removing duplicates concentrates more training passes on each remaining example
  4. It only affects evaluation, by removing train-test overlap, without changing anything about the trained model's own generation behavior
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 050/066 easy

Petrov et al. (2023), "Language Model Tokenizers Introduce Unfairness Between Languages," measured how many tokens the same piece of text requires once translated into different languages, using tokenizers from widely deployed LLMs. What did they find, and why does it matter?

  1. Tokenizing the same content in different languages can require up to roughly 15 times as many tokens depending on the language, even for tokenizers explicitly trained to support multiple languages, which raises the cost, processing latency, and effective context length available to speakers of the more heavily-tokenized languages
  2. Tokenizer length varies only trivially (well under 10%) across languages once a tokenizer's vocabulary includes any coverage of that language's script, so the disparity is not practically significant
  3. The disparity is fully explained by differing character-set sizes alone, and disappears entirely once text is measured in bytes rather than tokens
  4. The disparity only affects languages whose tokenizer support was added after initial release, and multilingual tokenizers trained from scratch on balanced multilingual data eliminate it completely
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 051/066 medium

In "Attention Is All You Need," Vaswani et al. (2017) compare self-attention, recurrent, and convolutional layers by three properties: computational complexity per layer, the minimum number of sequential operations required, and the maximum path length between any two input and output positions. What do they report for self-attention on the last two of these three properties, and why did this motivate using self-attention?

  1. Self-attention requires O(n) sequential operations, the same as a recurrent layer, but achieves a shorter maximum path length of O(log n) by processing the sequence in a tree-structured order
  2. Self-attention and convolutional layers are identical on both properties, since both process the full sequence within a single layer's receptive field once enough layers are stacked
  3. Self-attention requires a constant (O(1)) number of sequentially executed operations, and its maximum path length between any two positions is also O(1) regardless of distance, letting it learn long-range dependencies more easily than a layer type whose maximum path length grows with sequence length
  4. Self-attention requires more sequential operations to complete than a recurrent layer does, but its shorter maximum path length only becomes a practical advantage once sequence length n exceeds the representation dimensionality d
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 052/066 hard

Xiao et al. (2023), "Efficient Streaming Language Models with Attention Sinks," studied why simply evicting the oldest tokens from a fixed-size KV cache (a sliding window) causes a language model's quality to collapse once the window has evicted the first few tokens of a long stream, even though those first tokens carry little relevant content for later text. What did they find causes this collapse, and how does StreamingLLM avoid it?

  1. The collapse happens because positional encodings become invalid once any token is evicted, so StreamingLLM avoids it by recomputing sinusoidal positional encodings for the entire remaining cache after every eviction
  2. The collapse happens because a sliding window shrinks the effective vocabulary the model can attend to, so StreamingLLM avoids it by periodically re-inserting a random sample of previously evicted tokens back into the cache
  3. The collapse happens purely from a mismatch between training and inference sequence lengths, so StreamingLLM avoids it by fine-tuning the model on sequences as long as any stream it will process at inference time
  4. The first few tokens act as "attention sinks" that absorb a disproportionate share of attention across layers and heads regardless of their semantic content, so evicting them removes a stabilizing target for attention scores; StreamingLLM keeps these first few tokens permanently in the cache alongside a sliding window of the most recent tokens, which restores stable performance without any fine-tuning
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 053/066 medium

Power et al. (2022), "Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets," trained small transformer models on algorithmic tasks (such as modular arithmetic) well past the point where training accuracy had already reached ~100% while validation accuracy remained near chance level. What did they observe happens if training is continued far beyond this point, and what term did they give this phenomenon?

  1. Validation accuracy never improves beyond chance level no matter how much further training continues, since the model has already fully overfit to the training set and cannot recover; they term this phenomenon "memorization collapse"
  2. Validation accuracy can suddenly rise from near chance level to near-perfect generalization long after training accuracy has already saturated, a delayed transition they term "grokking"
  3. Training accuracy itself starts to decline the longer training continues past saturation, as the optimizer is forced to trade training performance for validation performance; they term this phenomenon "double descent"
  4. Both training and validation accuracy oscillate periodically forever once training accuracy saturates, never settling at a stable value in either direction; they term this phenomenon "catastrophic forgetting"
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 054/066 easy

Berglund et al. (2023), "The Reversal Curse: LLMs Trained on 'A is B' Fail to Learn 'B is A,'" fine-tuned GPT-3 and Llama-1 on fictitious factual statements of the form "[Name] is the [role] of [thing]" (e.g., a fictitious composer of a fictitious piece) and then tested the models on the reversed question. What did they find, and under what condition does this failure not occur?

  1. The models answered the reversed question correctly at the same rate as the original direction, showing fine-tuned factual associations generalize symmetrically regardless of which direction was stated during training
  2. The models answered the reversed question correctly only when the statement described a real, pretraining-era fact rather than a fictitious one invented for the fine-tuning set
  3. The models trained on "[Name] is the [role] of [thing]" statements failed to correctly answer the reversed question at a rate far better than chance, even though they could state the original direction correctly; this reversal failure did not occur when both directions of the relationship were presented together within the same context window at inference time, since the model could then read off the reverse relationship directly rather than needing to recall it
  4. The models succeeded on the reversed direction only for statements about numerical or quantitative relationships, and failed exclusively on statements about named entities and their roles
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 055/066 hard

Yang et al. (2017), "Breaking the Softmax Bottleneck: A High-Rank RNN Language Model," argue that standard softmax-based neural language models are fundamentally limited in how well they can approximate the true conditional next-word distribution, no matter how the model is trained. What is the "softmax bottleneck" they identify, and how does their proposed Mixture of Softmaxes (MoS) address it?

  1. The bottleneck is that softmax computation is too slow at large vocabulary sizes to be practical, and MoS addresses it by approximating the softmax's normalizing constant instead of computing it exactly
  2. The bottleneck is that softmax outputs cannot represent probabilities smaller than a fixed minimum value due to floating-point precision, and MoS addresses it by computing several softmax passes at different numeric precisions and averaging them
  3. The bottleneck is that the model's final hidden state is tied to too many output word embeddings, causing the model to overfit; MoS addresses it by using a smaller hidden dimension so fewer parameters need to be learned per predicted word
  4. Formulating language modeling as a matrix factorization problem, they show that a standard softmax layer's context-to-word log-probability matrix has an output rank bounded by the model's hidden dimension, which is too low to express the true, effectively higher-rank distribution over next words given diverse contexts; MoS instead computes several separate softmax distributions from the same hidden state and combines them in a learned weighted mixture, achieving an effectively higher-rank output than a single softmax can reach
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 056/066 easy

Chinchilla's scaling analysis (Hoffmann et al., 2022) identifies, for a fixed training compute budget, the model size and token count that minimizes training loss. Touvron et al.'s LLaMA paper (2023) deliberately trained smaller models on far more tokens than this Chinchilla-optimal ratio would recommend for their training compute budget alone. What is the actual justification for training a smaller model "beyond Chinchilla-optimal," per this line of reasoning?

  1. Chinchilla's original analysis contained an arithmetic error later corrected by Touvron et al., who found the truly compute-optimal ratio of tokens to parameters was substantially higher than Hoffmann et al. reported
  2. A smaller model trained on more tokens always achieves strictly lower training loss than a larger model trained Chinchilla-optimally on the same compute budget, making the Chinchilla recommendation simply incorrect
  3. Chinchilla's ratio only minimizes loss for a given training compute budget; it does not account for inference cost, and a smaller model is cheaper to run at inference time, so if a model will be served enough times, training it for longer than Chinchilla-optimal on more tokens to reach the best possible quality at a smaller, cheaper-to-serve size can be worth the extra training compute
  4. Chinchilla's analysis assumed a fixed dataset size that no longer applies now that far larger web-scraped training corpora are available, so its ratio is simply obsolete rather than being deliberately overridden by any inference-cost argument
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 057/066 medium

Schuster and Nakajima (2012) introduced WordPiece, the subword tokenization algorithm Devlin et al. (2018) later adopted for BERT. WordPiece builds its vocabulary by iteratively merging the pair of adjacent symbols that most increases the likelihood of the training corpus under a language model, rather than Sennrich et al.'s byte-pair encoding (BPE), which merges whichever adjacent pair occurs most frequently at each step. What is the practical consequence of choosing merges by likelihood increase instead of raw frequency?

  1. WordPiece can still only ever merge the pair that appears more often than any other pair in the corpus, exactly like BPE, so the two algorithms always produce identical vocabularies when trained on the same text
  2. A pair can be merged before a more frequent pair if merging it yields a disproportionately larger boost to the corpus's likelihood under the language model, so WordPiece can favor statistically informative subwords over merely common ones that BPE's frequency-only criterion would pick first
  3. WordPiece replaces merging entirely with a fixed-size character n-gram model, never combining learned subwords into longer units the way BPE's iterative merges do
  4. WordPiece requires a human-labeled corpus of correct word segmentations to learn each merge, whereas BPE only needs unlabeled raw text
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 058/066 medium

Child et al. (2019), "Generating Long Sequences with Sparse Transformers," train Transformers on sequences up to 16,384 tokens long by replacing full self-attention, whose compute and memory grow quadratically with sequence length, with a sparse factorization split across attention heads: a local pattern and a strided pattern. How does combining these two fixed sparse patterns across several layers let the model still capture a dependency between two arbitrary, far-apart positions, while each individual attention operation only considers a sparse subset of positions?

  1. Each head attends to literally every position in the sequence but at reduced numerical precision, which lowers memory use without changing which positions are attended to
  2. The local pattern lets each position attend only to a fixed number of its nearest neighbors, and the strided pattern is discarded during training, used only to speed up inference afterward
  3. The strided pattern alone is sufficient to connect every pair of positions directly within a single layer, so the local pattern is added only to reduce the number of learned parameters, not to extend which positions can be reached
  4. The local pattern connects each position to its nearby neighbors while the strided pattern connects positions spaced at regular intervals; stacking several layers lets information reach any other position through a path of local and strided hops, so full connectivity emerges across depth even though each layer's attention only ever considers a sparse subset of positions
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 059/066 hard

DeepSeek-V2 (DeepSeek-AI, 2024) introduces Multi-head Latent Attention (MLA) specifically to shrink the key-value cache needed for autoregressive inference, which standard multi-head attention grows linearly with both sequence length and the number of attention heads. Rather than having groups of heads share key/value heads the way multi-query or grouped-query attention do, what does MLA actually store in the KV cache for each token, and how does it still produce distinct per-head keys and values at attention time?

  1. MLA compresses each token's keys and values into a single low-dimensional latent vector via a learned down-projection, stores only that compact latent vector in the KV cache, and reconstructs the full per-head keys and values on demand at attention time using learned up-projection matrices, so cache size depends on the latent dimension rather than on the number of heads
  2. MLA stores the full per-head keys and values exactly as standard multi-head attention does, but compresses them afterward using 8-bit quantization before writing them to the cache, so no up-projection is needed at attention time
  3. Every attention head in MLA shares one single key head and one single value head, identical to multi-query attention, and the only difference from multi-query attention is that MLA additionally applies rotary position embeddings directly to the shared keys
  4. MLA removes the key-value cache entirely by recomputing every previous token's keys and values from scratch at every new decoding step, trading all of the saved memory for additional compute
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 060/066 easy

Shoeybi et al. (2019), introducing Megatron-LM, describe tensor parallelism (also called intra-layer model parallelism) as an alternative to data parallelism for training a single large transformer across multiple GPUs. In ordinary data parallelism, every GPU holds a full copy of the model and processes a different batch of training data. What does tensor parallelism change instead?

  1. It keeps a full copy of the model on every GPU, exactly as data parallelism does, and only changes how the training data is split into batches across GPUs
  2. It assigns entirely different layers of the model to different GPUs, so each GPU runs a different consecutive block of the network's depth on the same data as it flows through in sequence
  3. It splits individual weight matrices within a layer, such as those in the attention and feed-forward sublayers, across GPUs, so that computing a single layer's output requires those GPUs to collaborate and exchange partial results, rather than each GPU holding a full copy of every weight matrix
  4. It trains several smaller independent models in parallel on different GPUs, then averages their final weights together only once training finishes
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 061/066 easy

Olsson et al. (2022), "In-context Learning and Induction Heads," argue that a specific, identifiable attention circuit is the primary mechanism behind a transformer's ability to improve at a task from examples given only in its prompt, with no weight updates involved. What pattern does an induction head implement, and what did the authors observe happens during training as these circuits form?

  1. An induction head retrieves the single training example from pretraining that is most similar to the current prompt, and in-context learning ability appears gradually and smoothly throughout training, with no identifiable turning point
  2. An induction head looks back for a previous occurrence of the current token, finds what token came immediately after it that earlier time, and predicts that same token will come next again; the authors observed a sharp phase change early in training where induction heads form and in-context learning ability improves abruptly, across model sizes from small attention-only models up to 13-billion-parameter models
  3. An induction head only operates on numerical and tabular data within the prompt, and plays no role in the few-shot text-completion behavior described by Brown et al. for GPT-3
  4. An induction head is a component added only during a separate supervised fine-tuning stage after pretraining, and does not exist in a purely pretrained, non-instruction-tuned language model
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 062/066 medium

Gu and Dao (2023), introducing Mamba, propose a selective state space model (SSM) as a subquadratic alternative to the Transformer's attention mechanism. Prior structured SSMs, such as Gu et al.'s S4, use state-transition matrices that are fixed learned constants, the same for every input token regardless of its content. What change does Mamba make to these matrices, and what does this change restore?

  1. It makes the SSM's transition and input-projection matrices functions of the current input token itself, rather than fixed constants shared across all inputs, restoring the model's ability to selectively propagate or discard information based on each token's content
  2. It discards the state space formulation entirely in favor of full quadratic self-attention over the whole sequence, trading away linear-time scaling in sequence length to regain content-based reasoning
  3. It keeps the transition matrices fixed and input-independent exactly as in S4, but adds a separate content-independent LSTM-style gate on top of the recurrent state update
  4. It enlarges the fixed, input-independent transition matrices to a much higher dimension, relying on that extra fixed capacity alone to encode content-dependent behavior without making the matrices themselves depend on the input
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 063/066 easy

Huang et al. (2018), introducing GPipe, describe pipeline (inter-layer) parallelism as a way to train a model too large for any single accelerator's memory, by assigning different consecutive layers to different accelerators that then process data as it flows through in sequence. Naively running this layer-by-layer pipeline on one training batch at a time leaves most accelerators idle for most of each step, waiting their turn. What technique does GPipe introduce to reduce this idle time, and how does it work?

  1. GPipe gives every accelerator a full copy of every layer, so each one can process an entire training batch independently without waiting on any other accelerator
  2. GPipe splits each layer's individual weight matrices across accelerators, so that every accelerator participates in computing every layer's output simultaneously rather than waiting for upstream layers to finish
  3. GPipe splits each training batch into smaller micro-batches and pipelines them through the sequence of accelerators, so that once an accelerator finishes a micro-batch and passes its output downstream, it can immediately start on the next micro-batch instead of sitting idle
  4. GPipe discards intermediate activations after each micro-batch's forward pass and recomputes them during the backward pass instead of storing them, which is a separate technique GPipe also uses to reduce memory use but does not by itself address idle pipeline time
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 064/066 easy

Sennrich et al.'s byte-pair encoding (BPE) merges adjacent symbol pairs deterministically according to a fixed learned merge list, so the same word is always split into the exact same sequence of subword tokens every time it appears. Provilkov, Emelianenko and Voita's BPE-dropout instead trains on multiple different segmentations of that same word, using the same BPE vocabulary. How does BPE-dropout actually produce these multiple segmentations?

  1. It replaces BPE's merge list with a probabilistic Unigram language model, of the kind SentencePiece implements, which assigns a likelihood to each possible segmentation and samples among the highest-likelihood options
  2. It trains a separate neural segmentation model to propose alternative splits of each word, which are then merged with BPE's own output before tokenization
  3. It keeps BPE's merges fully deterministic but randomly shuffles the order in which otherwise-valid merges are applied, producing a different final sequence of subword tokens purely from that reordering
  4. At each step of the standard BPE merge procedure, it randomly drops some otherwise-applicable merges with a fixed probability, forcing the algorithm to fall back to smaller subwords at those points, so the same word can end up split differently from one training pass to the next
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 065/066 hard

Katharopoulos et al. (2020), "Transformers are RNNs," rewrite self-attention by replacing the softmax similarity between queries and keys with a kernel feature map applied separately to each query and key, then relying on the associativity of matrix multiplication to reorder the computation. What complexity does this reformulation achieve for computing attention over a sequence of length N, and what additional capability does it unlock for autoregressive generation?

  1. It achieves the same quadratic time complexity as standard softmax attention, since the kernel feature map is applied only after the full pairwise query-key similarity matrix has already been computed, but it reduces memory use by never fully materializing that matrix on-chip
  2. It reduces attention's time complexity from quadratic to linear in sequence length, and because the result can be expressed as a running sum of key-value outer products, it lets autoregressive generation update that running sum incrementally one token at a time, like the hidden state of a recurrent network, instead of reprocessing the full sequence at every step
  3. It reduces attention's time complexity to log-linear by applying the kernel feature map only to a sparse, fixed subset of key positions chosen ahead of time for each query, similar to the local-plus-strided patterns used by sparse attention
  4. It leaves attention's time complexity unchanged at quadratic, but achieves its reported speedup entirely by fusing the softmax and matrix-multiplication steps into a single GPU kernel to reduce memory reads and writes
AI & LLM Engineering · How LLMs Work: Transformers & Training · Card 066/066 medium

Grave et al. (2017) introduce adaptive softmax to speed up training and inference for neural language models with very large output vocabularies, where computing and normalizing a full softmax over every vocabulary word at each position is a major computational bottleneck, particularly on GPUs. Unlike the softmax bottleneck described elsewhere in this exam, which concerns the softmax's limited expressiveness regardless of how fast it runs, adaptive softmax is specifically a computational speed optimization. How does adaptive softmax actually reduce this computational cost?

  1. It groups vocabulary words into clusters based on their frequency, computing the expensive full-dimensional softmax only over a small head cluster of the most frequent words plus one shorthand symbol per rarer cluster, and only expands a rarer cluster's full softmax over its member words on the comparatively few occasions a word from that cluster is actually the target
  2. It replaces the softmax function with a sigmoid applied independently to each vocabulary word's score, turning the single multi-way classification into many independent binary classifications computed in parallel with no normalization step
  3. It reduces the vocabulary itself by merging rare words into a single shared unknown-word token before training begins, so the softmax only ever needs to normalize over the smaller set of remaining frequent words
  4. It keeps the full softmax over the entire vocabulary unchanged at every position, but moves that computation from the GPU to the CPU, where large matrix-vector products over a long vocabulary dimension are reportedly cheaper