81 cards · AI & LLM Engineering · answer each one, then read the explanation. Your score tallies below. Looking for Which prompting technique to use: CoT vs self-consistency vs ReAct vs ToT? Read the explainer.
0 / 81 answered · 0 correct
Link copied — send it to a friend!
AI & LLM Engineering · Prompt Engineering · Card 001/081easy
According to the original GPT-3 paper, "Language Models are Few-Shot Learners" (Brown et al., 2020), which best describes the difference between zero-shot and few-shot prompting?
AZero-shot and few-shot both require gradient updates to the model's weights; they differ only in how many examples are used per update
BZero-shot requires the model to be fine-tuned on the target task first, while few-shot requires no training at all
CFew-shot prompting means the model is shown zero examples but asked to solve the task in fewer than five reasoning steps
DZero-shot provides no task examples in the prompt, relying only on a natural language instruction, while few-shot includes a small number of input-output examples in the prompt before the actual query
Correct answer: .
In Brown et al. (2020), the zero-shot setting gives the model only a natural language description or instruction of the task with no worked examples, while the few-shot setting (in-context learning) includes several demonstration examples formatted as input-output pairs directly in the prompt before the real query, with no weight updates in either case. The second option is wrong because neither setting involves fine-tuning; the paper's central point is that GPT-3 performs tasks purely through prompting, without any gradient-based training on the target task. The first option is wrong because in-context learning explicitly does not update model weights at all -- the "learning" happens only within the context window at inference time, not through gradient updates. The third option confuses few-shot examples with reasoning steps, which is an unrelated chain-of-thought concept rather than the actual definition of few-shot prompting given in this paper.
Source: Brown et al., "Language Models are Few-Shot Learners" (2020), arXiv:2005.14165, Sections 1-2
AI & LLM Engineering · Prompt Engineering · Card 002/081easy
Under the technique introduced by Wei et al. (2022) in "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models," what does chain-of-thought prompting add to a standard few-shot prompt?
AIt instructs the model to output only the final answer with no explanation, in order to reduce token usage and cost
BIt includes intermediate reasoning steps leading to the final answer within the few-shot exemplars, rather than showing only the input and final answer
CIt replaces the task examples with a single, more detailed instruction and removes all examples from the prompt
DIt requires retraining the model on a dataset of step-by-step solutions before it can be used
Correct answer: .
Wei et al. (2022) show that augmenting few-shot exemplars with intermediate natural-language reasoning steps -- a "chain of thought" -- leading up to the final answer substantially improves performance on multi-step reasoning tasks, without any change to the model's weights. The third option is wrong because chain-of-thought prompting still uses worked examples; it enriches them with reasoning text rather than removing them in favor of a single instruction. The fourth option is wrong because the technique is applied at inference time to an already-trained, frozen model, not through a retraining or fine-tuning procedure on step-by-step solutions. The first option describes the opposite of the technique: chain-of-thought prompting deliberately elicits more explanatory reasoning text, not less, because that intermediate text is precisely what improves the accuracy of the final answer.
Source: Wei et al., "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (2022), arXiv:2201.11903
AI & LLM Engineering · Prompt Engineering · Card 003/081easy
Anthropic's prompt engineering documentation for Claude recommends "giving Claude a role." Where does it say a role description should be written, and why?
AIn the final line of the user's message, because Claude only applies role instructions that appear immediately before the question being asked
BIn a separate fine-tuning dataset, because role behavior can only be changed by retraining the model on role-labeled examples
CIn the system prompt, because even a single sentence establishing a role there focuses Claude's behavior and tone for the specific use case
DIn the assistant's prior turn, because Claude infers its role only from how it phrased earlier responses in the conversation
Correct answer: .
Anthropic's documentation states that setting a role in the system prompt focuses Claude's behavior and tone for the use case, and notes that even a single sentence placed there makes a measurable difference -- no fine-tuning or special placement within the user turn is required. The first option is wrong because placement within the user message is not how the documented technique works; the system prompt is a separate, dedicated field intended for this kind of persistent framing. The second option is wrong because role-setting as documented here is a prompting technique applied at inference time through the system prompt, not a training-time intervention requiring a labeled dataset. The fourth option is wrong because it describes the model inferring a role from its own past outputs rather than being told a role explicitly up front, which is not the mechanism the documentation describes.
Source: Anthropic, "Prompting best practices" -- "Give Claude a role" section, platform.claude.com prompt engineering documentation
AI & LLM Engineering · Prompt Engineering · Card 004/081easy
In LLM inference APIs such as Anthropic's Messages API and OpenAI's Chat Completions API, what effect does lowering the `temperature` sampling parameter toward 0 have on generated text?
AIt sharpens the probability distribution over next tokens so the model more consistently picks the highest-probability token, producing more deterministic, less varied output
BIt disables sampling entirely and forces the model to retrieve an exact quote from its training data
CIt increases the number of tokens the model is allowed to generate in a single response
DIt reduces the model's context window, so fewer previous tokens are considered when generating each new token
Correct answer: .
Temperature scales the logits before the softmax step that turns them into a probability distribution over next tokens; a temperature near 0 sharpens that distribution so the highest-probability token is chosen almost every time, yielding deterministic, low-variance output, while higher temperatures flatten the distribution and increase randomness and diversity. The fourth option is wrong because temperature has nothing to do with how much prior context the model attends to -- that is governed by the context window, a separate, unrelated setting. The third option is wrong because output length is controlled by a separate maximum-tokens style parameter, not by temperature. The second option is wrong because even at temperature 0 the model is still generating tokens step by step from its learned probability distribution; it is not performing retrieval or reproducing an exact quotation from training data.
Source: Anthropic Messages API reference and OpenAI Chat Completions API reference, `temperature` parameter documentation
AI & LLM Engineering · Prompt Engineering · Card 005/081easy
Per the OWASP Top 10 for LLM Applications (2025), which best distinguishes "indirect" prompt injection from "direct" prompt injection?
ADirect prompt injection means malicious instructions are typed straight into the model's input by the user; indirect prompt injection means the malicious instructions are hidden in external content, such as a webpage or document, that the LLM later ingests and follows
BDirect prompt injection only affects open-source models, while indirect prompt injection only affects closed, API-based models
CIndirect prompt injection requires physical access to the server running the model, while direct prompt injection can be performed remotely over the network
DDirect prompt injection is a purely theoretical risk with no real-world examples, while indirect prompt injection has already been fixed in all major LLM products
Correct answer: .
OWASP's LLM01 entry defines prompt injection as crafted input that alters model behavior or output, and distinguishes direct injection -- where an attacker types malicious instructions straight into the model's input, for example telling it to ignore previous instructions -- from indirect injection, where the malicious instructions are embedded in external content such as a document, webpage, or email that the LLM ingests and treats as if it were a legitimate instruction. The second option is wrong because the direct/indirect distinction is about the source of the malicious instruction, not about a model's licensing or deployment model; both open-source and closed API-based models are vulnerable to either variant. The third option is wrong because neither variant requires physical access to the server; both are input-based attacks deliverable remotely through normal application inputs. The fourth option is wrong because OWASP lists prompt injection as its top-ranked, actively exploited risk in real deployments, not a theoretical or already-solved problem.
Source: OWASP Top 10 for LLM Applications (2025), LLM01: Prompt Injection
AI & LLM Engineering · Prompt Engineering · Card 006/081easy
According to Anthropic's prompt engineering documentation, what is the stated benefit of wrapping different parts of a prompt (instructions, context, examples, input) in distinct XML tags such as `<instructions>` and `<context>`?
AIt compresses the prompt so that it consumes fewer tokens than the equivalent plain-text prompt
BIt is required syntax without which the Messages API will reject the request with a formatting error
CIt automatically translates the tagged sections into a different language before the model processes them
DIt helps Claude parse complex prompts unambiguously by clearly separating different types of content, reducing the chance the model misinterprets what is instruction versus example versus input
Correct answer: .
Anthropic's documentation states that XML tags help Claude parse complex prompts unambiguously, especially when a prompt mixes instructions, context, examples, and variable input, and it recommends wrapping each type of content in its own consistently named tag, nesting tags when the content has a natural hierarchy. The second option is wrong because XML tags are a stylistic and structural convention recommended for clarity, not syntax enforced or required by the Messages API; plain-text prompts without tags remain valid requests. The first option is wrong because adding tags adds characters and therefore tokens rather than reducing them; the documented benefit is parsing clarity, not compression. The third option is wrong because tagging content has no translation function of any kind; the model still processes the tagged text in whatever language it was originally written in.
Source: Anthropic, "Prompting best practices" -- "Structure prompts with XML tags," platform.claude.com prompt engineering documentation
AI & LLM Engineering · Prompt Engineering · Card 007/081medium
Zhao et al. (2021), "Calibrate Before Use: Improving Few-Shot Performance of Language Models," identify "recency bias" as one cause of instability in few-shot prompting. What does recency bias describe?
AThe tendency of a model's accuracy to decline over time as new versions of the model are released
BThe model favoring the most recently published research papers when asked to cite sources
CThe model's tendency to disproportionately predict whichever label appeared in the example placed nearest the end of the few-shot prompt, regardless of the true input
DThe tendency to give more weight to the very first example in a few-shot prompt while ignoring later examples
Correct answer: .
Zhao et al. show that few-shot LLM predictions are systematically biased toward the answer of whichever demonstration example sits closest to the query at the end of the prompt -- what they term recency bias -- alongside majority label bias, which favors whichever label is most frequent among the examples, and common token bias; they propose contextual calibration to correct for these effects. The option about accuracy declining as new model versions are released and the option about favoring recently published papers when citing sources describe unrelated phenomena, model performance drift across product releases and citation preferences, that this paper does not study or address. The option about giving more weight to the very first example describes the opposite of recency bias: it describes a primacy effect toward the first example, whereas the paper's finding is specifically a bias toward the last, most recent example in the prompt, not the earliest one.
Source: Zhao, Wallace, Feng, Klein, Singh, "Calibrate Before Use: Improving Few-Shot Performance of Language Models" (ICML 2021), arXiv:2102.09690
AI & LLM Engineering · Prompt Engineering · Card 008/081medium
Wang et al. (2022), "Self-Consistency Improves Chain of Thought Reasoning in Language Models," propose replacing greedy decoding with what alternative decoding strategy for chain-of-thought prompts?
ADiscard chain-of-thought reasoning entirely and instead retrieve the answer from an external search engine
BSample multiple diverse reasoning paths for the same question at a nonzero temperature, then take a majority vote over the final answers each path arrives at
CAlways generate exactly one reasoning path deterministically, then ask a separate human reviewer to check it before accepting the answer
DFine-tune the model on the correct chain-of-thought path found by brute-force search over the entire training set
Correct answer: .
Self-consistency samples a diverse set of reasoning paths for the same chain-of-thought prompt, using sampling rather than greedy decoding, and then marginalizes over them by taking the most common final answer, exploiting the intuition that correct reasoning tends to converge on the same answer through multiple valid paths even when the intermediate reasoning text differs. The third option is wrong because it introduces a human-in-the-loop review step rather than the fully automatic sampling-and-voting procedure the paper describes. The fourth option is wrong because self-consistency requires no fine-tuning and no brute-force search over training data; it operates purely at inference time on an already-trained, frozen model. The first option is wrong because it eliminates chain-of-thought reasoning altogether in favor of external retrieval, which is unrelated to self-consistency; the technique keeps and multiplies the reasoning paths rather than replacing them with a search engine lookup.
Source: Wang et al., "Self-Consistency Improves Chain of Thought Reasoning in Language Models" (ICLR 2023), arXiv:2203.11171
AI & LLM Engineering · Prompt Engineering · Card 009/081medium
Zhou et al. (2022), "Least-to-Most Prompting Enables Complex Reasoning in Large Language Models," describe a two-stage strategy for solving problems harder than those shown in the prompt's examples. What is that strategy?
AFirst prompt the model to decompose the problem into a sequence of simpler subproblems, then sequentially prompt it to solve each subproblem in order, feeding each prior subproblem's answer into the context used for the next
BFirst run the same prompt through several different LLMs from different vendors, then pick whichever vendor's answer appears most frequently
CFirst ask the model to guess the final answer directly, then ask it to generate a chain-of-thought justification for that already-chosen answer after the fact
DFirst fine-tune the model on the hardest available examples, then evaluate it zero-shot on easier examples to measure generalization downward
Correct answer: .
Least-to-most prompting works in two stages: a decomposition stage where the model breaks a complex problem into an ordered list of simpler subproblems, followed by a sequential problem-solving stage where the model solves each subproblem in turn, with the answers to earlier subproblems included in the context used to solve later ones; this lets the technique generalize to problems harder than the demonstrations shown, unlike standard chain-of-thought prompting. The fourth option is wrong because the technique involves no fine-tuning at all; it is a purely prompting-based, inference-time method applied to a frozen model. The third option is wrong because it reverses the actual order and purpose of the method -- least-to-most decomposes a problem before solving it, rather than justifying an answer that was already guessed, and post-hoc justification is not what the paper studies. The second option is wrong because the method uses a single model working sequentially through ordered subproblems, not a multi-vendor voting ensemble.
Source: Zhou et al., "Least-to-Most Prompting Enables Complex Reasoning in Large Language Models" (ICLR 2023), arXiv:2205.10625
AI & LLM Engineering · Prompt Engineering · Card 010/081medium
OpenAI's documentation distinguishes "Structured Outputs" from the older "JSON mode" feature of its Chat Completions / Responses APIs. According to that documentation, what is the key difference between the two?
AThe two features are functionally identical; "Structured Outputs" is simply a rebranding of "JSON mode" with no change in behavior
BStructured Outputs can only be used with image inputs, while JSON mode is restricted to text-only prompts
CJSON mode guarantees schema conformance, while Structured Outputs only guarantees syntactically valid JSON without any schema checking
DBoth guarantee syntactically valid JSON, but only Structured Outputs also guarantees the output conforms to the caller's supplied JSON Schema (for example, required keys and enum values); JSON mode guarantees valid JSON syntax only
Correct answer: .
OpenAI's documentation states that Structured Outputs "ensures the model will always generate responses that adhere to your supplied JSON Schema," explicitly calling it "the evolution of JSON mode," while JSON mode only ensures the output parses as valid JSON without enforcing that it matches any particular schema, meaning JSON mode output can still omit required fields or use invalid enum values even though it is syntactically valid JSON. The third option reverses the actual guarantees documented by OpenAI. The second option is wrong because the distinction has nothing to do with input modality; both features concern the format of the model's text output regardless of whether the request includes image inputs. The first option is wrong because the documentation explicitly frames Structured Outputs as adding a stricter guarantee, schema adherence, that JSON mode does not provide, rather than describing an identical rebrand.
AI & LLM Engineering · Prompt Engineering · Card 011/081hard
Turpin et al. (2023), "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting," demonstrate unfaithfulness by manipulating what feature of a few-shot prompt, and observing what result?
AThey removed all chain-of-thought reasoning from the few-shot examples entirely and found the model refused to answer at all without it
BThey increased the number of few-shot examples from 2 to 200 and found chain-of-thought accuracy improved with no change in faithfulness concerns
CThey reordered the multiple-choice options in the few-shot examples so the correct answer was biased to always fall on a particular letter (e.g., always "(A)"); the model's chain-of-thought then rationalized picking that biased letter while never mentioning the answer ordering as its real reason
DThey translated the few-shot examples into a different natural language and found the model's final answers became random regardless of the question
Correct answer: .
Turpin et al. show that adding a biasing feature to a prompt, such as reordering multiple-choice options so the correct answer is artificially made to always fall on a particular letter, can systematically shift the model's final answer toward that letter, while the model's generated chain-of-thought explanation confabulates a plausible-sounding justification that never mentions the true, biasing cause; this demonstrates that the stated reasoning does not faithfully reflect the actual process behind the answer. The option about removing all chain-of-thought reasoning, the option about translating the examples into another language, and the option about scaling the number of examples describe manipulations and outcomes -- removing all reasoning causing refusal, translation causing randomness, or increasing example count improving accuracy -- that are not the intervention or finding reported in this paper. The paper's central manipulation is specifically the answer-position biasing intervention, and its central finding concerns faithfulness of the explanation, not raw accuracy or refusal behavior.
Source: Turpin, Michael, Perez, Bowman, "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting" (2023), arXiv:2305.04388
AI & LLM Engineering · Prompt Engineering · Card 012/081hard
Sclar et al. (2023), "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design," measure how much purely cosmetic formatting choices (such as separators and spacing) in a few-shot prompt affect accuracy, holding the semantic content constant. What did they find?
AAccuracy differences from formatting disappeared entirely once few-shot examples were replaced with zero-shot instructions
BPurely formatting-level changes to a semantically identical few-shot prompt caused accuracy swings of up to tens of accuracy points (as much as 76 points for one open-source model tested), and this sensitivity persisted even with larger models, more few-shot examples, and instruction tuning
CFormatting only matters for image-based prompts and has no measurable effect on plain-text few-shot prompts
DFormatting choices had no measurable effect on accuracy once a model exceeded roughly one billion parameters, fully resolving the issue at modern model scales
Correct answer: .
Sclar et al. found that large language models are highly sensitive to spurious, purely cosmetic formatting choices in few-shot prompts, such as separators, spacing, and other superficial template details that do not change the semantic content, with accuracy swings of up to 76 points observed for one model studied (LLaMA-2-13B), and this sensitivity did not disappear as model size, number of few-shot examples, or instruction tuning increased. The fourth option is wrong because the paper explicitly reports that this sensitivity persists even as models scale up, rather than resolving at some parameter threshold. The third option is wrong because the study concerns plain-text prompt formatting rather than image inputs, and it found the effect specifically in text-based few-shot prompting. The first option is wrong because the paper studies few-shot formatting sensitivity directly; it does not report that switching to zero-shot prompting eliminates formatting-driven accuracy differences.
Source: Sclar, Choi, Tsvetkov, Suhr, "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design" (ICLR 2024), arXiv:2310.11324
AI & LLM Engineering · Prompt Engineering · Card 013/081easy
Kojima et al. (2022), "Large Language Models are Zero-Shot Reasoners," show that a specific technique substantially improves LLM performance on multi-step reasoning benchmarks without any task-specific worked examples in the prompt. What is this zero-shot chain-of-thought technique, and how does it differ from the few-shot chain-of-thought prompting of Wei et al. (2022)?
AProviding the correct final answer to the model up front and asking it to work backward to justify it, whereas Wei et al.'s method asks for the answer with no justification
BAppending a task-agnostic trigger phrase such as "Let's think step by step" before the answer, with no worked examples in the prompt at all, whereas Wei et al.'s method requires several few-shot exemplars that each include their own written-out reasoning steps
CReplacing the multiple-choice options with open-ended text so the model can no longer guess from answer choices, a change unrelated to Wei et al.'s few-shot method
DFine-tuning the model on a small labeled set of step-by-step solutions before inference, whereas Wei et al.'s method requires no training at all
Correct answer: .
Kojima et al. (2022) show that a single, task-agnostic trigger phrase like "Let's think step by step," inserted before the model generates its answer with zero worked examples, elicits a chain of reasoning and substantially improves accuracy on arithmetic and symbolic reasoning benchmarks such as MultiArith and GSM8K, calling this Zero-shot-CoT. This differs from Wei et al.'s few-shot chain-of-thought prompting, which requires several exemplars in the prompt that each spell out intermediate reasoning steps leading to their answers. The option proposing fine-tuning on a labeled set of step-by-step solutions is wrong because Zero-shot-CoT involves no fine-tuning or gradient updates; the trigger phrase is applied at inference time to a frozen, pretrained model. The option that provides the correct final answer up front is wrong because it describes working backward from a given final answer, which is the opposite of eliciting reasoning that leads to an answer the model has not yet produced. The option about replacing multiple-choice options with open-ended text is wrong because Zero-shot-CoT makes no change to how answer choices are presented; the entire intervention is the trigger phrase inserted before the model's output.
Source: Kojima, Gu, Reid, Matsuo, Iwasawa, "Large Language Models are Zero-Shot Reasoners" (2022), arXiv:2205.11916
AI & LLM Engineering · Prompt Engineering · Card 014/081medium
Zhou et al. (2022/2023), "Large Language Models Are Human-Level Prompt Engineers," propose Automatic Prompt Engineer (APE). How does APE generate and select an effective task instruction, according to the paper?
AIt requires a human panel to write hundreds of candidate instructions, which the LLM then simply ranks by fluency without regard to task performance
BIt fine-tunes the model's weights on many instruction/response pairs and treats the resulting model checkpoint itself as the "instruction"
CIt works only for classification tasks with a fixed label set and cannot generate free-form instructions for open-ended tasks
DIt treats the instruction itself as a "program": an LLM proposes a pool of candidate instructions from a handful of input-output demonstrations, and each candidate is scored (for example, by how well it reproduces the demonstrations when used as a prompt) so the highest-scoring instruction can be selected or further refined
Correct answer: .
APE frames instruction generation as a program-synthesis problem: given a small set of input-output demonstrations, an LLM is prompted to propose a pool of candidate natural-language instructions, and each candidate is then scored by a chosen metric, such as how well it lets the model reproduce the demonstrations when prepended as a prompt, so search and Monte Carlo-style resampling can converge on a high-scoring instruction. The paper reports this automatically generated instruction matches or beats human-written instructions on 19 of 24 tasks tested. The first option is wrong because no human panel writes the candidates; the LLM itself generates the pool of candidate instructions, and selection is driven by a task-performance score, not by fluency ranking. The second option is wrong because APE requires no fine-tuning or weight updates at all; it is a purely inference-time search over instruction text using a frozen model. The third option is wrong because the paper evaluates APE across 24 diverse NLP tasks, including open-ended generation and even steering models toward truthfulness, not solely fixed-label classification.
Source: Zhou, Muresanu, Han, Paster, Pitis, Chan, Ba, "Large Language Models Are Human-Level Prompt Engineers" (2023), arXiv:2211.01910
AI & LLM Engineering · Prompt Engineering · Card 015/081hard
Min et al. (2022), "Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?," test in-context learning by randomly replacing the labels in few-shot demonstrations with incorrect ones. What did they find, and what does it suggest about why few-shot demonstrations help?
AReplacing the demonstrations' labels with random, often-incorrect ones barely hurt accuracy across a range of classification and multiple-choice tasks, suggesting that the correctness of the input-label mapping matters far less than other aspects of the demonstrations, such as the label space, the input distribution, and the overall format
BThe experiment could not be run, because in-context learning requires every demonstration label to be verified against a held-out validation set before each query
CReplacing the labels with random ones improved accuracy beyond using correct labels, showing that few-shot demonstrations actively mislead the model and should generally be avoided
DReplacing the labels with random ones caused accuracy to collapse to chance level on every task tested, confirming that the model learns the exact input-label mapping shown in the demonstrations the way a supervised classifier would
Correct answer: .
Min et al. (2022) find that swapping in random, frequently incorrect labels for the demonstrations' true labels causes only a small drop in accuracy, consistently across 12 different models and a range of classification and multi-choice tasks, whereas removing other properties of the demonstrations (such as their label space or input distribution) hurts much more. This suggests in-context learning does not primarily work by the model learning the specific input-to-label mapping shown, the way a supervised classifier fits training pairs, but instead benefits from being shown the format, the space of valid labels, and the distribution of inputs. The option claiming accuracy collapsed to chance level is wrong because the paper's central, surprising finding is that accuracy did not collapse to chance when labels were randomized; it stayed close to the correct-label condition. The option claiming random labels improved accuracy beyond correct labels is wrong because random labels did not outperform correct labels; the finding is that they are roughly comparable, not that random labels are actively better or that demonstrations mislead the model. The option claiming the experiment could not be run without validating every label against a held-out set is nonsensical and describes no real methodological constraint of in-context learning.
Source: Min, Lyu, Holtzman, Artetxe, Lewis, Hajishirzi, Zettlemoyer, "Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?" (EMNLP 2022), arXiv:2202.12837
AI & LLM Engineering · Prompt Engineering · Card 016/081hard
Lu et al. (2022), "Fantastically Ordered Prompts and Where to Find Them," study how the order of few-shot examples within an otherwise-identical prompt affects accuracy. What did they find, and what method did they propose to pick a good order without a labeled validation set?
AThey found the best fix was to always sort examples alphabetically by their label text, which fully eliminated order sensitivity across all models and tasks tested
BThey found order sensitivity only affects models under one billion parameters, and it disappears automatically once a model is scaled past that size
CThe order of the same few-shot examples can swing accuracy from near state-of-the-art to close to random guessing; they proposed generating an artificial, unlabeled "probing" set from the language model itself and selecting the ordering whose predicted-label distribution has favorable entropy statistics on that set, without needing any labeled dev data
DExample order has no measurable effect on accuracy once the examples themselves are held constant, so the paper concludes ordering can safely be ignored
Correct answer: .
Lu et al. (2022) demonstrate that, holding the same few-shot examples constant and only permuting their order, GPT-family models can swing between near state-of-the-art and near-random accuracy, and that this sensitivity is unpredictable from prompt length or example choice alone. To pick a good ordering without labeled validation data, they use the generative model itself to construct an artificial probing set of unlabeled outputs and then select the candidate ordering whose predicted-label distribution over that probing set has favorable entropy statistics, reporting a 13% relative improvement on average across eleven text classification tasks. The option claiming example order has no measurable effect is wrong because the paper's entire premise and headline finding is that order sensitivity is large and highly consequential, not negligible. The option proposing an alphabetical sort by label text is wrong because no universal alphabetical-sort rule is proposed or shown to eliminate the effect; the paper's actual proposal is the entropy-based probing method. The option restricting the problem to models under one billion parameters is wrong because the paper does not report or claim that scaling past one billion parameters resolves order sensitivity.
Source: Lu, Bartolo, Moore, Riedel, Stenetorp, "Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity" (ACL 2022), arXiv:2104.08786
AI & LLM Engineering · Prompt Engineering · Card 017/081easy
Liu et al. (2022), "Generated Knowledge Prompting for Commonsense Reasoning," propose a two-stage prompting pipeline for commonsense question answering. What are the two stages?
AFirst, prompt a language model to generate several relevant knowledge statements about the question's topic; second, provide those generated knowledge statements as additional context alongside the original question in a separate prompt that produces the final answer
BFirst, ask the model to answer the question directly with no context; second, ask a different model to translate that answer into another language to check consistency
CFirst, retrieve documents from a fixed external knowledge base such as an encyclopedia; second, fine-tune the model's weights on those retrieved documents before it answers the question
DFirst, generate several candidate final answers; second, average their token probabilities together to produce a single blended answer string
Correct answer: .
Generated knowledge prompting first prompts a language model to produce several candidate knowledge statements relevant to the question's topic, using a separate knowledge-generation prompt with its own few-shot examples, and then feeds those generated statements as additional context alongside the original question into a second prompt whose job is to produce the final answer, integrating over multiple generated knowledge statements when helpful. This achieved state-of-the-art results on commonsense benchmarks like NumerSense, CommonsenseQA 2.0, and QASC. The option built on retrieving documents from a fixed external knowledge base and then fine-tuning the model is wrong because the method generates knowledge from the language model's own parametric knowledge rather than retrieving it, and it involves no fine-tuning or weight updates at any stage. The option proposing a translation-based consistency check between models is wrong because no such check exists in this pipeline. The option about averaging token-level probabilities across candidate answers is wrong because the method conditions the final answer on generated knowledge text used as context, not on probability averaging.
Source: Liu, Liu, Lu, Welleck, West, Le Bras, Choi, Hajishirzi, "Generated Knowledge Prompting for Commonsense Reasoning" (ACL 2022), arXiv:2110.08387
AI & LLM Engineering · Prompt Engineering · Card 018/081medium
Wang et al. (2023), "Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models," target error categories that plague Kojima et al.'s zero-shot "Let's think step by step" prompting, such as missing reasoning steps and calculation errors. What does Plan-and-Solve prompting add to address this, without using any worked examples?
AIt replaces natural-language reasoning entirely with a formal, executable program that a separate code interpreter runs to produce the final answer
BIt asks the model to skip straight to the final numeric answer without showing any intermediate reasoning, in order to avoid calculation errors introduced by long chains of text
CIt requires collecting several hundred worked examples with expert-annotated plans and fine-tuning the model on them before it can perform any reasoning
DIt instructs the model, in a single zero-shot prompt, to first devise a plan that divides the overall task into smaller subtasks, and then to carry out that plan step by step before giving the final answer
Correct answer: .
Plan-and-Solve prompting keeps Zero-shot-CoT's no-worked-examples format but adds an explicit planning instruction: the model is prompted to first devise a plan that breaks the overall problem into smaller subtasks, and then to execute that plan step by step before stating its final answer, with an extended "PS+" variant adding further instructions to extract relevant variables and pay attention to calculation accuracy. This targets exactly the missing-step and semantic-misunderstanding errors the paper identifies in plain Zero-shot-CoT. The third option is wrong because Plan-and-Solve remains a purely zero-shot, inference-time prompting method with no fine-tuning and no annotated example collection. The first option is wrong because that describes Program-of-Thought prompting, a different technique the paper compares against; Plan-and-Solve itself keeps natural-language reasoning and simply adds an explicit planning step. The second option is wrong because the technique asks for more structured intermediate reasoning, not less, and the paper's goal is reducing errors within that reasoning, not removing it.
Source: Wang, Xu, Lan, Hu, Lan, Lee, Lim, "Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models" (ACL 2023), arXiv:2305.04091
AI & LLM Engineering · Prompt Engineering · Card 019/081medium
Press et al. (2022), "Measuring and Narrowing the Compositionality Gap in Language Models," introduce the "self-ask" prompting method to address multi-hop questions whose sub-answers the model already knows individually but fails to combine correctly. How does self-ask structure the model's output?
AThe compositionality gap is closed simply by scaling up model size, and self-ask is presented only as an unrelated historical baseline that the paper argues should be discarded
BThe prompt's few-shot exemplars teach the model to explicitly decide whether a follow-up question is needed, then pose and answer that follow-up question itself within the same generation, repeating as needed before stating the final answer -- a structure into which an external search engine can optionally be plugged to answer the follow-up questions instead
CTwo separate, independently prompted models debate each other's answers over several rounds until they converge on the same final answer, which is the paper's central proposed method
DThe model is asked to answer the multi-hop question directly in a single token with no intermediate text of any kind, which the paper shows eliminates the compositionality gap entirely
Correct answer: .
The paper defines the compositionality gap as the gap between how often a model answers all of a multi-hop question's individual sub-questions correctly versus how often it produces the correct composed final answer, and it shows this gap does not reliably shrink as GPT-3-family models scale up. Self-ask narrows it by using few-shot exemplars that teach the model to state whether a follow-up question is needed, then ask and answer that follow-up itself before giving a final answer, repeating as necessary; the same structured slot for follow-up questions also makes it easy to swap in an external search engine to answer them instead, which further improves accuracy. The option demanding a single-token answer with no intermediate text is wrong because self-ask relies on generating intermediate follow-up questions and answers, the opposite of a single-token direct answer, and the paper does not claim any method eliminates the gap entirely. The option describing two independently prompted models debating each other is wrong because no multi-model debate scheme is described. The option claiming the gap is closed simply by scaling up model size is wrong because the paper's key finding is that scaling alone does not reliably close the gap, and self-ask is the paper's own main proposed method, not a discarded baseline.
Source: Press, Zhang, Min, Schmidt, Smith, Lewis, "Measuring and Narrowing the Compositionality Gap in Language Models" (Findings of EMNLP 2023), arXiv:2210.03350
AI & LLM Engineering · Prompt Engineering · Card 020/081easy
Zhang et al. (2022), "Automatic Chain of Thought Prompting in Large Language Models," propose Auto-CoT to avoid the manual effort, and potential for hand-written mistakes, involved in writing few-shot chain-of-thought exemplars by hand. What two-step procedure does Auto-CoT use to build its demonstrations automatically?
AFirst, fine-tune the model on a large labeled reasoning dataset; second, discard the few-shot examples entirely, since the fine-tuned model no longer needs any demonstrations
BFirst, ask human annotators to hand-write a reasoning chain for every single question in the dataset; second, use the model to pick which of those human-written chains looks most fluent
CFirst, cluster the dataset's questions by similarity and pick one representative question from each cluster; second, generate a reasoning chain for each representative question automatically using zero-shot chain-of-thought (for example, "Let's think step by step"), assembling the resulting diverse set of question-plus-generated-chain pairs into the few-shot demonstrations
DFirst, generate one reasoning chain for the very first question in the dataset; second, reuse that exact same chain, unmodified, as the only demonstration for every other question regardless of topic
Correct answer: .
Auto-CoT first partitions the dataset's questions into clusters by similarity and selects one representative question per cluster, then generates a reasoning chain for each representative question automatically using zero-shot chain-of-thought prompting, and finally assembles this diverse set of question-and-generated-chain pairs into the few-shot demonstrations used for the actual task, which the paper shows matches or exceeds manually written chain-of-thought demonstrations on ten reasoning benchmarks. Sampling diverse clusters rather than similar questions reduces the chance that a single mistaken generated chain gets repeated across many similar demonstrations. The option relying on human annotators to hand-write a chain for every question is wrong because it reintroduces the manual, human-written-chain effort that Auto-CoT is explicitly designed to eliminate. The option proposing fine-tuning on a labeled reasoning dataset is wrong because Auto-CoT involves no fine-tuning; it remains a purely prompting-based method that still uses few-shot demonstrations, just auto-generated ones. The option reusing one chain from the very first question as the sole demonstration is wrong because Auto-CoT deliberately draws diverse representative questions from multiple clusters rather than repeating one chain for every question regardless of topic, which is precisely the failure mode diversity is meant to avoid.
Source: Zhang, Zhang, Li, Smola, "Automatic Chain of Thought Prompting in Large Language Models" (ICLR 2023), arXiv:2210.03493
AI & LLM Engineering · Prompt Engineering · Card 021/081easy
According to Anthropic's current prompt engineering documentation for Claude, roughly how many examples should a multishot (few-shot) prompt include for best results, and what qualities should those examples have?
AAt least 50 examples are required before Claude can reliably follow the demonstrated pattern at all
BAbout 3 to 5 examples, each relevant to the actual use case and diverse enough (including edge cases) that Claude does not pick up unintended patterns from the examples
CThe examples must all be drawn from the exact same edge case repeated with different wording, since variety between examples confuses Claude about which pattern to follow
DExactly 1 example is optimal, and using more than 1 measurably degrades Claude's performance on every task
Correct answer: .
Anthropic's documentation states that including 3 to 5 examples gives the best results for multishot prompting, and that examples should be relevant (mirroring the actual use case closely) and diverse (covering edge cases and varying enough that Claude does not pick up unintended patterns), in addition to being structured with example tags so Claude can distinguish them from instructions. The option calling exactly 1 example optimal is wrong because the documentation recommends multiple examples specifically, not a single one, and does not claim additional examples degrade performance. The option requiring at least 50 examples is wrong because no 50-example minimum is stated anywhere in the guidance; a small handful is presented as sufficient. The option demanding that all examples repeat the exact same edge case is wrong because the documentation explicitly recommends diversity across examples, including covering edge cases, rather than repeating one edge case with different wording, which is the opposite of what is recommended and is exactly the kind of narrow pattern that causes Claude to overfit to an unintended regularity.
AI & LLM Engineering · Prompt Engineering · Card 022/081medium
Anthropic's current documentation on long-context prompting recommends where to place long documents or data-rich inputs (roughly 20,000+ tokens) relative to the query and instructions. What does it recommend, and what improvement does it cite?
ASplit the long document into many short, separate API calls of under 500 tokens each, since Claude cannot process more than roughly 20,000 tokens of context in a single request
BInterleave single sentences of the long document between each instruction sentence so that context and instructions alternate line by line throughout the prompt
CPlace the long documents and other inputs near the top of the prompt, above the query, instructions, and examples; the documentation notes that putting queries at the end can improve response quality by up to 30 percent in tests, especially for complex, multidocument inputs
DPlace the long documents at the very end of the prompt, after the query and instructions, because Claude always weighs the last few hundred tokens of a prompt most heavily regardless of content length
Correct answer: .
Anthropic's documentation recommends putting longform data and documents near the top of the prompt, above the query, instructions, and examples, when working with large or data-rich inputs of roughly 20,000 or more tokens, and it notes that queries placed at the end can improve response quality by up to 30 percent in tests, especially with complex, multidocument inputs, alongside wrapping each document in XML tags with source metadata. The fourth option is wrong because it recommends the opposite placement of what the documentation states; the guidance is to put documents near the top, not at the end. The first option is wrong because Claude's context window comfortably exceeds 20,000 tokens in a single request, and the tip concerns how to structure one large prompt, not evidence that requests must be split into many small calls. The second option is wrong because no sentence-by-sentence interleaving of context and instructions is recommended; the documented structure instead wraps whole documents in dedicated tags kept separate from the instructions.
AI & LLM Engineering · Prompt Engineering · Card 023/081easy
According to Anthropic's current documentation on chaining complex prompts for Claude, what is described as the most common prompt-chaining pattern, and how is it structured across separate API calls?
ALoad balancing: send the exact same prompt to several different Claude API keys simultaneously and return whichever response arrives first, discarding the rest
BCache warming: repeatedly resend an unchanged prompt purely to keep it available in the prompt cache, with no reviewing or refining step involved
CWeight merging: average the model weights used to generate two separate draft responses into a single hybrid model before producing the final answer
DSelf-correction: generate a draft response in one API call, have Claude review that draft against stated criteria in a second call, then have Claude refine the draft based on that review in a third call, so each step is a separate call that can be logged, evaluated, or branched on
Correct answer: .
Anthropic's documentation states that explicit prompt chaining, breaking a task into sequential API calls, remains useful when you need to inspect intermediate outputs or enforce a specific pipeline structure, and it identifies self-correction as the most common chaining pattern: generate a draft, have Claude review that draft against stated criteria, then have Claude refine the draft based on the review, with each step run as its own separate API call so results can be logged, evaluated, or branched on. The load-balancing option is wrong because sending identical requests to multiple API keys and keeping the fastest response is a redundancy or load-balancing pattern unrelated to improving reasoning quality, and it is not the documented chaining pattern. The weight-merging option is wrong because the Messages API does not expose model weights for merging, and no such technique is described. The cache-warming option is wrong because it describes maintaining a prompt cache with no reviewing or refining step, which is a caching concern rather than the review-and-refine chaining pattern the documentation describes.
AI & LLM Engineering · Prompt Engineering · Card 024/081easy
Anthropic's prompt caching documentation describes marking a "cache breakpoint" with the `cache_control` parameter. For a request that mixes stable content (tool definitions, a system prompt, a large reference document) with content that changes on every call (the current user turn), where should that breakpoint be placed, and why?
AOn the last content block whose prefix is identical across requests -- that is, after the stable tool definitions, system prompt, and document, and before the changing per-request content -- because the cache only stores what comes before the breakpoint, and placing it on content that changes every request would make the cached prefix's hash change each time, producing no cache hits
BNowhere -- Anthropic's API caches every request identically by default with no `cache_control` parameter or configuration needed
COn the changing user message itself, because caching is described as useful specifically for content that differs on every single call
DOn the very first token of the entire request, including the tool definitions, so that almost nothing in the request ends up covered by the cached prefix
Correct answer: .
Anthropic's documentation explains that a cache breakpoint should sit on the last block whose prefix is identical across requests, with stable content ordered first (tool definitions, then the system prompt, then long reference documents or few-shot examples) and dynamic per-request content, such as the current user turn, placed after the breakpoint; the API caches writes only up to that point and reads by looking backward for a matching cached prefix, so a breakpoint on ever-changing content would produce a different hash on every call and never hit the cache. The option placing the breakpoint on the very first token of the request is wrong because that would leave the large stable blocks that follow it, such as the tool definitions and system prompt, uncached rather than covered, which defeats the purpose of caching the expensive static prefix. The option placing the breakpoint on the changing user message is wrong because it inverts the guidance: caching is valuable specifically for content that repeats unchanged across calls, not for content that differs every time. The option claiming every request is cached by default with no configuration is wrong because the documentation requires explicitly adding a `cache_control` field, whether as a single top-level marker or on specific content blocks, for caching to take effect at all.
AI & LLM Engineering · Prompt Engineering · Card 025/081medium
An engineer is building an agent that must look up information from a search API partway through solving a multi-step question and adjust its plan based on what the search returns. According to Yao et al. (2022), "ReAct: Synergizing Reasoning and Acting in Language Models," what does the ReAct prompting framework do to make this possible?
AGenerating only the sequence of actions to take, such as API calls, without ever producing any intermediate natural-language reasoning
BInterleaving natural-language reasoning traces with task-specific actions and the observations those actions return, within a single prompted trajectory, so reasoning can decide the next action and each new observation can update the reasoning that follows
CTraining a separate reasoning model and a separate acting model with reinforcement learning and combining their outputs after each has finished running independently
DProducing a chain-of-thought explanation for a problem but never issuing any call to an external tool or API
Correct answer: .
Yao et al. (2022) propose ReAct, which prompts a language model to generate reasoning traces and task-specific actions in an interleaved sequence within one trajectory: a reasoning step can plan or revise what to do next, an action step then queries an external source such as a search API, and the resulting observation is fed back into the next reasoning step, letting the model track and adjust its plan as new information arrives. The option describing actions with no reasoning describes the plain action-generation baseline ReAct is shown to outperform. The option describing separate reasoning and acting models trained independently with reinforcement learning describes a different architecture, not the single-model interleaved prompting method ReAct uses. The option describing reasoning with no external action describes plain chain-of-thought prompting, which cannot gather new information from outside the model the way ReAct's action steps do.
Source: Yao et al., "ReAct: Synergizing Reasoning and Acting in Language Models" (2022), arXiv:2210.03629
AI & LLM Engineering · Prompt Engineering · Card 026/081hard
A researcher is applying a language model to a puzzle-like task, such as the Game of 24, where an early move can turn out to be a dead end many steps later, and simply extending a single left-to-right chain of thought performs poorly. According to Yao et al. (2023), "Tree of Thoughts: Deliberate Problem Solving with Large Language Models," what does the Tree of Thoughts framework add on top of chain-of-thought prompting to address this?
AProducing one single, linear sequence of reasoning steps from the problem to the final answer, exactly as in standard chain-of-thought prompting
BSampling many complete, independent chain-of-thought reasoning paths for the whole problem and picking the final answer that the largest number of them agree on
CFraming problem solving as a search over a tree whose nodes are intermediate "thoughts": the model generates and self-evaluates multiple candidate next thoughts at each step, and the search can look ahead or backtrack using strategies such as breadth-first or depth-first search, at substantially higher inference-time compute cost than a single reasoning chain
DTraining a separate value function offline to score candidate solutions, then using that fixed value function alone to pick a solution without the language model exploring or backtracking at inference time
Correct answer: .
Yao et al. (2023) introduce Tree of Thoughts to handle tasks like Game of 24 where an early step can turn out to be a dead end many steps later and a single left-to-right chain of thought performs poorly: the model frames problem solving as search over a tree of intermediate "thoughts," generating and self-evaluating multiple candidate next thoughts at each step and using search strategies such as breadth-first or depth-first search to look ahead or backtrack, which the paper reports raised Game of 24 success far above plain chain-of-thought, at the cost of far more inference-time compute than a single chain. The option describing one linear chain describes the baseline Tree of Thoughts is built to outperform on exactly this kind of task. The option describing sampling many independent full chains and majority-voting describes self-consistency, which does not let the model evaluate or backtrack partway through a path the way Tree of Thoughts does. The option describing an offline-trained value function used alone describes a different architecture that removes the model's own step-by-step exploration and self-evaluation at inference time.
Source: Yao et al., "Tree of Thoughts: Deliberate Problem Solving with Large Language Models" (2023), arXiv:2305.10601
AI & LLM Engineering · Prompt Engineering · Card 027/081easy
A developer is building an application on OpenAI's Chat Completions API and needs to send the model both a high-level instruction that should shape its behavior for the whole conversation and the end user's own question. Per OpenAI's API documentation, how should these two pieces of context be structured in the request?
ABoth pieces of context are sent as a single combined string, since the Chat Completions API has no way to distinguish who authored which part of the input
BOnly the end user's question can be sent to the model; there is no supported way to give it standing instructions that apply to the whole conversation
CThe high-level instruction must be repeated inside every individual user message, since a message with a system role is discarded by the API after the first turn
DPer OpenAI's API documentation, the request is built as an array of role-tagged messages: the high-level instruction is sent as a system message placed at the start of the array, the end user's question is sent as a user message, and any of the model's own prior responses are represented back to it as assistant messages
Correct answer: .
OpenAI's Chat Completions API documentation structures a request as an ordered array of messages, each tagged with a role: a system message near the start of the array carries high-level, developer-authored instructions meant to shape the model's behavior for the whole conversation, a user message carries the end user's own input, and an assistant message represents the model's own prior output fed back into the conversation so it can be attended to on later turns. The option describing one combined untagged string ignores that each message in the array carries its own distinct role rather than being an undifferentiated blob. The option claiming there is no way to give standing instructions ignores the system message's explicit documented purpose. The option claiming a system message is discarded after the first turn is incorrect; it remains present in the array on each subsequent request unless the developer removes it themselves.
AI & LLM Engineering · Prompt Engineering · Card 028/081easy
A team building a math word-problem solver notices that a large language model using standard chain-of-thought prompting writes out a correct step-by-step plan in natural language but still makes an arithmetic slip when computing the final number, producing a wrong answer despite sound reasoning. Gao et al. (2022), "PAL: Program-Aided Language Models," propose an alternative prompting approach specifically to fix this class of error. What does PAL have the model do differently?
APAL asks the model to write out the same natural-language reasoning chain twice independently and takes whichever of the two final numeric answers appears first
BPAL replaces every arithmetic step with a request for the model to look up the answer in a retrieved external knowledge base of pre-solved problems
CPAL prompts the model to translate the problem into intermediate steps expressed as runnable code (for example Python statements) and hands that generated program to an external interpreter to execute, using the interpreter's output as the final answer instead of having the model compute it itself
DPAL fine-tunes the underlying model on millions of additional arithmetic examples so it memorizes correct computations for common operation patterns
Correct answer: .
PAL keeps the language model responsible only for reading the word problem and decomposing it into a short program, typically statements that mirror the natural-language reasoning steps, but hands the actual arithmetic evaluation of that program to a real interpreter rather than asking the model to carry out the computation itself. Because the interpreter executes the generated code deterministically, the kind of small arithmetic slip that undermines otherwise sound chain-of-thought reasoning disappears, since the model never has to compute the final number by reasoning about digits. The option describing two independent natural-language reasoning chains with the first answer kept is wrong because that describes neither PAL nor genuine self-consistency, which uses majority voting across many samples rather than order of arrival, and it still leaves all arithmetic to the model. The option describing lookup in a retrieved knowledge base of pre-solved problems is wrong because PAL involves no retrieval step at all; it only needs a code interpreter. The option describing additional fine-tuning is wrong because PAL is a pure prompting technique applied to a frozen, already-trained model via few-shot examples that pair problems with programs.
AI & LLM Engineering · Prompt Engineering · Card 029/081easy
A developer notices that a long-form LLM answer (for example, 'list and explain five approaches to X') is slow to generate because the model must produce the entire response as one long sequential stream of tokens, point by point, in order. Ning et al. (2023), "Skeleton-of-Thought: Large Language Models Can Do Parallel Decoding," propose a prompting-and-decoding approach aimed specifically at cutting this latency without changing the model's weights. What does Skeleton-of-Thought do?
AIt first prompts the model to generate a brief skeleton of the answer's main points, then issues separate parallel calls (or batched decoding) to expand each point concurrently, before combining the expanded points into the final answer
BIt shortens the answer by instructing the model to omit the elaboration for each point, trading completeness for speed
CIt routes the request to a smaller, faster model for a first draft, then upgrades to the original model only for a final proofreading pass
DIt caches the model's response to the same prompt across users, so only the first request incurs the full sequential generation cost
Correct answer: .
Skeleton-of-Thought works in two stages: it first prompts the model to produce a short skeleton listing the main points of the eventual answer, and then, instead of continuing to generate the full elaboration sequentially, it issues separate API calls or uses batched decoding to expand each skeleton point at the same time, finally stitching the expanded points together into the complete response. The paper reports considerable speed-ups across a dozen LLMs from this parallelism, and in some cases even better answer quality, because each point is expanded with a focused, self-contained prompt. The option describing omitting elaboration for speed is wrong because Skeleton-of-Thought still generates full elaborated content for every point; it changes when and how that content is generated, not whether it exists. The option describing a smaller draft model handed off to the original model for proofreading is wrong because it describes a draft-then-upgrade or speculative-decoding style pipeline, not the skeleton-then-parallel-expansion structure this paper proposes. The option describing caching a response across users is wrong because that is an infrastructure-level caching strategy operating across separate requests, unrelated to how tokens are decoded within a single answer's generation.
Source: Ning, Lin, Zhou, Yang & Wang, 'Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation' (arXiv:2307.15337, 2023)
AI & LLM Engineering · Prompt Engineering · Card 030/081easy
Deng et al. (2023), "Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves," start from the observation that a human's phrasing of a question often carries ambiguity or missing context that a language model reads differently than the human intended. What does the Rephrase and Respond (RaR) method have the model do about this, and how does the paper describe its relationship to chain-of-thought prompting?
ARaR trains a separate classifier to detect ambiguous questions and reroutes only those questions to a human reviewer before any model response is generated
BRaR has the model rephrase and expand the given question in its own words, adding clarifying detail, before answering, which the paper shows is complementary to chain-of-thought rather than a replacement for it, and combining the two performs better than either alone
CRaR instructs the model to translate the question into a different natural language first, on the theory that translation removes ambiguity, and this fully replaces the need for chain-of-thought
DRaR skips rephrasing entirely and instead asks the model to answer the same question multiple times, then rephrases only the final selected answer for readability
Correct answer: .
Rephrase and Respond has the model itself rephrase and expand the original question, adding clarifying detail that resolves the gap between how a human framed the question and how the model would ideally like it framed, before producing an answer to its own rephrased version. The paper explicitly compares RaR to chain-of-thought both theoretically and empirically and finds the two are complementary rather than competing techniques, with combining rephrasing and step-by-step reasoning outperforming either technique used alone. The option describing a separate ambiguity-detecting classifier that routes questions to a human reviewer is wrong because RaR involves no human intervention or classifier at all; the same model handles rephrasing and answering end to end. The option describing literal translation into another language is wrong both because that is not the mechanism RaR uses and because the paper frames RaR as complementary to chain-of-thought, not as something that fully replaces it. The option describing multiple answer attempts with only the final answer rephrased is wrong because RaR rephrases the question before answering, not the answer after the fact.
Source: Deng, Zhang, Chen & Gu, 'Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves' (arXiv:2311.04205, 2023)
AI & LLM Engineering · Prompt Engineering · Card 031/081easy
A developer calling Anthropic's Messages API wants Claude's response to always begin with a valid JSON object and never with an introductory sentence like 'Here is the JSON you requested:'. According to Anthropic's documentation on prefilling Claude's response, how is this achieved, and through what channel is the technique available?
ABy adding a `response_format: json` parameter to the request, which is documented as the only supported way to constrain where Claude's output can start
BBy appending a stop sequence equal to the word 'Here' so that Claude is blocked from generating that word at the start of its answer
CBy fine-tuning a custom version of Claude on examples that start directly with `{`, since the base model cannot be steered to skip a preamble through prompting alone
DBy including the desired starting text (for example an opening `{`) as the beginning of the assistant turn itself in the API call, so Claude's generated response continues directly from that point; the documentation notes this prefill technique is API-only and is not supported on the newest Claude models
Correct answer: .
Anthropic's documentation describes prefilling as placing the desired starting text, such as an opening brace for JSON, directly into the assistant turn of the API request; because Claude's generated content continues from wherever the assistant message leaves off, the model never generates the preamble in the first place, since that starting text was never something it needed to produce itself. This is documented as an API-only capability, available through direct API or SDK usage, and the documentation also notes it is not supported on the newest Claude models, where such requests return an error. The option describing a `response_format: json` parameter is wrong because Anthropic's documentation describes prefill, not a JSON-mode-style response-format flag, as the mechanism for this. The option describing a stop sequence on the word 'Here' is wrong because stop sequences only halt generation once a matching string is produced; they do not control or guarantee what content the model starts with, so a stray rephrasing could still avoid the word while still adding a preamble. The option describing fine-tuning is wrong because prefilling is a prompting-time technique available to any API caller without training a custom model.
Source: Anthropic, 'Prefill Claude's response for greater output control,' Claude Platform Docs (platform.claude.com/docs/en/build-with-claude/prompt-engineering/prefill-claudes-response)
AI & LLM Engineering · Prompt Engineering · Card 032/081easy
A developer who has used Anthropic's Messages API is used to marking a `cache_control` breakpoint on the content she wants cached (such as a large system prompt or reference document), so that stable content is reused across calls and only new content is billed at full price. Moving the same workload to OpenAI's API, she wants to know whether she must add an equivalent manual marker to get a caching discount. Per OpenAI's own documentation on prompt caching, what should she expect?
ANo manual marker is required for basic caching: OpenAI's prompt caching is applied automatically, matching and reusing the longest previously-seen prefix of a prompt once the prompt is long enough to qualify, without any `cache_control`-style parameter needed to opt in
BYes, the requirement is identical: she must annotate the exact same content block with a `cache_control` field, using the same syntax as Anthropic's API, or no caching discount will ever apply
CNo, because OpenAI's API does not offer any form of prompt caching at all, regardless of prompt length or repetition
DYes, but only through a separate paid add-on subscription that must be purchased before any caching discount becomes available on cached tokens
Correct answer: .
OpenAI's documentation describes prompt caching as automatic: once a prompt reaches a minimum length, the API caches the longest prefix of the prompt it has already computed and reuses that cached prefix on subsequent requests that share the same beginning, applying a discount without the caller needing to add any special annotation to opt in. This is a meaningfully different design from Anthropic's Messages API, where the caller must explicitly mark a `cache_control` breakpoint on the content it wants cached; OpenAI's basic caching happens purely as a side effect of repeated prefixes, not an explicit instruction. The option claiming an identical `cache_control`-style requirement is wrong because OpenAI's basic caching mechanism needs no such field and does not share Anthropic's syntax. The option claiming OpenAI offers no prompt caching at all is wrong because automatic prefix-based caching is a documented, real feature of OpenAI's API. The option describing a separate paid add-on subscription is wrong because the caching discount OpenAI describes applies automatically to qualifying requests as part of normal API usage, not as a separately purchased product.
Source: OpenAI, 'Prompt Caching in the API' (openai.com/index/api-prompt-caching/); Anthropic, Claude Platform Docs on prompt caching (contrasting explicit cache_control breakpoints)
AI & LLM Engineering · Prompt Engineering · Card 033/081easy
Madaan et al. (2023), "Self-Refine: Iterative Refinement with Self-Feedback," propose a prompting loop that improves a language model's own output without any additional training data, fine-tuning, or external verifier model. What is the loop, and which single model performs every role in it?
AA larger 'teacher' model grades the outputs of a separate, smaller 'student' model and returns a numeric score that the student uses to pick from several candidate answers
BA retrieval system feeds the model documents relevant to its own previous answer, and the model simply copies the most relevant retrieved sentence into its final response
CThe same pretrained model first produces an initial answer, then critiques that answer as feedback, then uses its own critique to produce a refined answer, repeating this generate-feedback-refine loop for multiple rounds; no separate model or extra training is involved at any stage
DA human reviewer reads the model's first answer and writes the feedback, which is then pasted back into a second prompt for the model to revise
Correct answer: .
Self-Refine uses one pretrained model for every step of the loop: the model produces an initial answer, is then prompted to critique that same answer and generate feedback on its weaknesses, and is then prompted again to use its own feedback to produce a refined answer, with this generate-feedback-refine cycle repeatable for multiple rounds. No additional training data, fine-tuning, or separate verifier model is required, since the same model plays the generator, critic, and refiner roles purely through prompting; the paper reports around a 20 percent average performance gain across tasks like math reasoning and dialogue response. The option describing a separate larger teacher model grading a smaller student model is wrong because Self-Refine's central claim is that a single model can supply useful feedback about its own output, with no second model involved. The option describing a retrieval system feeding in outside documents is wrong because Self-Refine's feedback comes from the model's own critique of its prior answer, not from any retrieved external text. The option describing a human writing the feedback is wrong because Self-Refine is designed to work without any human in the loop.
Source: Madaan et al., 'Self-Refine: Iterative Refinement with Self-Feedback' (arXiv:2303.17651, 2023)
AI & LLM Engineering · Prompt Engineering · Card 034/081medium
A model is asked a specific, detail-heavy physics question and, despite reasoning step by step, applies the wrong underlying formula because it dives straight into the specific numbers without first recalling which general principle governs the situation. Zheng et al. (2023), "Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models," propose Step-Back Prompting to address exactly this failure mode. How does the technique change what the model is prompted to do before answering?
AIt has the model answer the specific question first, and only afterward asks it to state which general principle it implicitly used, purely as a post-hoc explanation with no effect on the answer
BIt has the model retrieve the original specific question from a database of similar past questions and copy the closest match's stored answer
CIt has the model break the specific question into the smallest possible sub-questions and answer each sub-question independently before summing the sub-answers
DIt first prompts the model to derive a higher-level, more abstract question or general principle from the specific instance, then has the model reason from that abstraction to answer the original specific question, rather than reasoning directly from the details alone
Correct answer: .
Step-Back Prompting has the model perform an abstraction step before it ever engages with the specific question's details: it first prompts the model to derive a higher-level question or general principle that the specific instance is an example of, and only then has the model reason from that abstraction down to the original, detail-heavy question, rather than reasoning directly from the details alone. The paper reports this abstraction-first ordering produces substantial gains on reasoning-intensive tasks tested with PaLM-2L, GPT-4, and Llama2-70B, including double-digit percentage improvements on STEM knowledge and multi-hop reasoning benchmarks. The option describing a post-hoc statement of the general principle after already answering is wrong because doing the abstraction after the answer has already been produced cannot change how that answer was derived, which defeats the technique's purpose. The option describing retrieving a stored answer from a database of similar past questions is wrong because Step-Back Prompting generates its own abstraction through reasoning, not through looking up a previously stored answer. The option describing decomposition into the smallest sub-questions is wrong because that describes decomposition-style prompting, not the abstraction-then-reasoning structure this paper proposes.
Source: Zheng, Mishra, Chen, Cheng, Chi, Le & Zhou, 'Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models' (arXiv:2310.06117, 2023)
AI & LLM Engineering · Prompt Engineering · Card 035/081medium
Chia et al. (2023), "Contrastive Chain-of-Thought Prompting," start from a surprising finding in prior work: giving a model chain-of-thought demonstrations with deliberately invalid reasoning had only a small negative effect compared to fully valid demonstrations, suggesting standard CoT does not clearly teach the model what mistakes to avoid. What do the authors propose instead, and what is its rationale?
ARemoving chain-of-thought demonstrations altogether and relying solely on zero-shot instructions, on the theory that any demonstration risks teaching a spurious pattern
BProviding both valid and invalid reasoning demonstrations side by side for the same or paired problems, together with an automatic method for constructing such contrastive demonstrations, so the model is explicitly shown examples of flawed reasoning alongside correct reasoning rather than only ever seeing correct chains
CIncreasing the number of valid chain-of-thought demonstrations from a handful to several dozen, so that sheer repetition overwhelms any influence from invalid reasoning
DReplacing human-written demonstrations with demonstrations generated entirely by a separate, larger teacher model, without regard to whether the reasoning in them is valid or invalid
Correct answer: .
Contrastive Chain-of-Thought proposes giving the model both valid and invalid reasoning demonstrations for the same or paired problems, inspired by the observation that humans learn effectively from both positive and negative examples, and it introduces an automatic method for constructing such contrastive demonstration pairs rather than requiring them to be hand-written. The rationale is that standard chain-of-thought only ever shows correct reasoning and therefore never explicitly informs the model about what kinds of mistakes to avoid, which the authors argue is why invalid demonstrations had surprisingly little negative effect in prior work; showing the flawed reasoning alongside the correct reasoning is meant to close that gap. The option describing removal of all demonstrations is wrong because contrastive chain-of-thought still relies on demonstrations, now paired with invalid counterparts, rather than eliminating them. The option describing simply adding more valid demonstrations is wrong because the paper's fix is about demonstration content and pairing, not raw demonstration count. The option describing demonstrations generated entirely by a larger teacher model is wrong because the method's automatic construction process is not defined by, or dependent on, using a separate larger model as a teacher.
AI & LLM Engineering · Prompt Engineering · Card 036/081medium
A team wants the reasoning-boosting benefit of few-shot chain-of-thought prompting on a new task, but has no labeled exemplars available and does not want to spend engineering effort hand-writing or retrieving demonstrations for every new problem type. Yasunaga et al. (2023), "Large Language Models as Analogical Reasoners," propose Analogical Prompting to address this. What does the model do under this method, and what does it avoid needing?
AThe model retrieves the single most similar labeled example from a large curated exemplar database using a nearest-neighbor embedding search, then copies that example's reasoning structure exactly
BThe model is fine-tuned on a small set of analogous problems drawn from a related domain before being asked to solve the target problem
CThe model asks a human expert, via an interactive clarification step, to supply one worked analogy before attempting to solve the problem
DThe model is prompted to recall or self-generate relevant exemplars and, where useful, relevant background knowledge related to the given problem before solving it, entirely on its own; the method avoids needing any labeled exemplars to be hand-written or retrieved from an external source
Correct answer: .
Analogical Prompting is designed so the model itself recalls or generates exemplars and relevant background knowledge related to the problem at hand, inspired by how humans draw on analogous past experience when reasoning, and it does this entirely at inference time on its own. Because the exemplars are self-generated rather than sourced externally, the method avoids the engineering cost of hand-writing demonstrations or building and querying a retrieval system, while still tailoring the generated exemplars to each specific problem; the paper reports it outperforming both zero-shot chain-of-thought and manual few-shot chain-of-thought across math, code generation, and other reasoning benchmarks. The option describing retrieval of the single most similar example from a curated database is wrong because that reintroduces exactly the external retrieval and curated-labeling burden the method is designed to avoid. The option describing fine-tuning on analogous problems is wrong because Analogical Prompting requires no training step; it operates purely through prompting a frozen model. The option describing a human supplying a worked analogy is wrong because the method is meant to be fully automatic, with no human clarification step involved.
Source: Yasunaga et al., 'Large Language Models as Analogical Reasoners' (arXiv:2310.01714, 2023)
AI & LLM Engineering · Prompt Engineering · Card 037/081medium
A model asked to write a short biographical paragraph about a real historical figure produces a mostly accurate response but invents two incorrect dates. Dhuliawala et al. (2023), "Chain-of-Verification Reduces Hallucination in Large Language Models," propose a four-step prompting procedure (CoVe) meant to catch this kind of error before the response is delivered to the user. What are the four steps, in order?
ADraft an answer, then immediately ask the model to rate its own confidence on a 1-10 scale, then deliver whichever draft scores highest without any further steps
BDraft an initial response, plan a set of targeted verification questions that would fact-check specific claims in that draft, answer each verification question independently so those answers are not biased by the original draft, then use the verification answers to produce a final, revised response
CDraft an answer, translate it into a different language and back, compare the two versions for discrepancies, and keep whichever version is shorter
DDraft an answer, retrieve external documents matching every named entity in the draft, replace every named entity with whatever the retrieved documents say without any further verification step, and stop
Correct answer: .
Chain-of-Verification proceeds in four steps: the model drafts an initial response to the query, then plans a set of verification questions targeted at fact-checking specific claims made in that draft, then answers each verification question independently of the original draft so those answers are not biased by whatever the draft already claimed, and finally uses the verification answers to produce a final response that corrects any discrepancies the verification step surfaced. The paper shows this reduces hallucinations across tasks including list-based questions from Wikidata, closed-book MultiSpanQA, and longform text generation, precisely the kind of invented-date error in the biography scenario. The option describing a simple self-rated numeric confidence score is wrong because it skips the independent verification-question step entirely and does not fact-check specific claims. The option describing round-trip translation is wrong because no translation step appears anywhere in CoVe. The option describing blind replacement of named entities from retrieved documents with no independent verification step is wrong because CoVe's verification relies on the model answering its own generated verification questions, not on an unverified retrieval-based substitution.
Source: Dhuliawala et al., 'Chain-of-Verification Reduces Hallucination in Large Language Models' (arXiv:2309.11495, 2023)
AI & LLM Engineering · Prompt Engineering · Card 038/081hard
A team wants to steer a large black-box LLM they cannot fine-tune (no access to weights) toward producing summaries that reliably include certain keywords, without hand-crafting a new instruction for every single input document. Li et al. (2023), "Guiding Large Language Models via Directional Stimulus Prompting," propose a framework for exactly this setting. What mechanism does Directional Stimulus Prompting use, and how is that mechanism trained?
AIt trains a separate, small, tunable policy model (such as a T5-sized model) to generate an instance-specific 'directional stimulus,' a short hint or set of keywords tailored to each input, which is inserted into the prompt sent to the black-box LLM; the policy model itself is optimized via supervised fine-tuning on labeled data and/or reinforcement learning using rewards derived from the black-box LLM's own output quality
BIt fine-tunes the black-box LLM's own weights directly on a small labeled dataset of ideal summaries, despite the team's lack of weight access, by using a gradient-free zeroth-order optimization method against the model's API
CIt hard-codes the same fixed list of keywords into every prompt regardless of the input document, relying on the black-box LLM to decide on its own which of the fixed keywords are actually relevant
DIt relies entirely on the black-box LLM's built-in retrieval-augmented generation feature to pull in relevant keywords from a search index, with no additional model or training step involved
Correct answer: .
Directional Stimulus Prompting sidesteps the impossibility of fine-tuning a black-box LLM by instead training a separate, small, tunable policy model, such as a model of T5's size, whose job is to generate an instance-specific directional stimulus, a short hint or set of keywords tailored to the particular input, which is then inserted into the prompt actually sent to the black-box LLM. That policy model is what gets optimized, either through supervised fine-tuning on labeled data, through reinforcement learning using rewards derived from evaluating the black-box LLM's resulting output, or both, while the black-box LLM's own weights are never touched. The option describing direct zeroth-order fine-tuning of the black-box LLM's weights is wrong because the entire premise of the framework is that the LLM's weights remain inaccessible and untouched; only the small policy model is trained. The option describing a single fixed keyword list reused for every input is wrong because the method's value comes specifically from generating a different, tailored stimulus per instance. The option describing built-in retrieval-augmented generation is wrong because the framework introduces a trained generative policy model, not a retrieval or search mechanism.
Source: Li et al., 'Guiding Large Language Models via Directional Stimulus Prompting' (arXiv:2302.11520, NeurIPS 2023)
AI & LLM Engineering · Prompt Engineering · Card 039/081hard
A team building on OpenAI's newer reasoning models (the o-series and later) via the Chat Completions or Responses API wants to know how the older 'system' role relates to the newer 'developer' role, and where each level of instruction sits in OpenAI's documented instruction-following hierarchy. Per OpenAI's Model Spec and API guidance, which best describes this?
AThe 'system' and 'developer' roles are two independent, co-equal channels with no defined precedence between them, so a conflict between a system message and a developer message is resolved arbitrarily by the model at random
BThe 'developer' role sits below the 'user' role in authority, meaning any instruction a user provides during the conversation can always override an instruction given via the developer role
CFor the newer models, the 'developer' role takes over the position the 'system' role used to occupy for API callers, sitting above the 'user' role in OpenAI's documented chain-of-command (platform-level instructions rank above developer instructions, which in turn rank above user instructions, with assistant and tool messages carrying no independent authority)
DThe 'system' role was retired entirely with no replacement, and any instruction that used to go in a system message must now be passed only through the per-request `instructions` parameter, with no message-role equivalent available at all
Correct answer: .
For OpenAI's newer models, the 'developer' role takes over the position that the 'system' role used to occupy for API callers, and OpenAI's documented chain-of-command places platform-level instructions above developer instructions, developer instructions above user instructions, and gives assistant and tool messages no independent authority of their own. This means a developer message still functions as the caller's highest-authority channel short of platform-level constraints, with the end user's messages ranking below it, which matches how the older system role used to behave relative to user messages. The option describing the two roles as co-equal with random conflict resolution is wrong because OpenAI documents an explicit precedence order rather than leaving conflicts undefined. The option claiming the developer role ranks below the user role is wrong because it inverts the documented hierarchy; developer instructions outrank user instructions, not the reverse. The option claiming the system role was retired with absolutely no message-role equivalent is wrong because the developer role message is exactly that equivalent; instructions did not need to move to a parameter with no role-based alternative available.
Source: OpenAI Model Spec (chain-of-command / instruction hierarchy: platform, developer, user, guideline, no authority); OpenAI Developer Community documentation on the developer role in the Responses API
AI & LLM Engineering · Prompt Engineering · Card 040/081medium
A red-teaming exercise finds that stuffing an LLM's prompt with a few dozen fabricated dialogue turns showing a compliant assistant answering harmful requests has little effect on the model's safety behavior, but scaling the same fabricated dialogue up to hundreds of turns reliably breaks it. Per Anthropic's 2024 research on this technique, called "many-shot jailbreaking," what mechanism explains why effectiveness increases so sharply with the number of turns, and which models are most exposed?
AIt exploits in-context learning: the growing number of faux dialogue turns showing harmful compliance functions like few-shot demonstrations that steer the model's behavior toward the pattern shown, with attack success following a power-law increase as the number of turns grows; models with the largest context windows are most exposed because they can fit enough turns for the effect to take hold
BIt works by repeatedly asking the same harmful question in slightly reworded form until the model's output-length limit forces it to answer directly instead of refusing
CIt exploits a training-data leakage bug specific to one vendor's models, where feeding back memorized fragments of the safety-training dataset itself disables the safety filter
DIt works by encoding the harmful request in a language underrepresented in the safety-training data, exhausting a fixed per-request translation budget before any safety classifier runs
Correct answer: .
Many-shot jailbreaking works by exploiting the same in-context learning mechanism that makes few-shot prompting useful for legitimate tasks: each fabricated turn showing a supposed assistant complying with a harmful request acts like a demonstration that nudges the model toward continuing the shown pattern, and Anthropic's paper reports that attack success climbs following a power-law relationship as the number of these faux turns grows, being negligible at only a handful of turns but consistent once scaled into the hundreds. Because fitting hundreds of turns into a single prompt requires a large context window, the paper identifies models with the newest, longest context windows as the most exposed, since the attack was not practically feasible before such windows became common. The option describing repeated rewording to exhaust an output-length limit is wrong because the technique does not rely on truncating a refusal; it relies on demonstration volume shifting the model's learned behavior. The option describing a training-data leakage bug is wrong because the effect is a general property of in-context learning observed across many models, not a vendor-specific data leak. The option describing exhausting a translation budget is wrong because the attack requires no translation step at all; it works directly in whichever language the faux dialogue is written in.
AI & LLM Engineering · Prompt Engineering · Card 041/081easy
A developer needs an LLM to classify 600 short product reviews as positive or negative, paying per input-and-output token through a hosted API. Sending each review as its own separate call means the same lengthy classification instructions and few-shot examples are repeated in every single request. Cheng et al. (2023), "Batch Prompting: Efficient Inference with Large Language Model APIs," propose an alternative. What does batch prompting do to cut this cost?
AIt fine-tunes a smaller specialized model on the same classification examples so the large model is no longer needed once training completes
BIt groups several independent input samples together into a single prompt, sharing one copy of the instructions and in-context examples across the whole group, and has the model return the answers for all of the grouped samples in one response instead of issuing a separate call per sample
CIt reduces cost by having the model skip the few-shot examples entirely and rely purely on zero-shot instructions for every request
DIt caches every previous request-response pair, so a later request only costs tokens if its text is not an exact character-for-character match to an earlier one
Correct answer: .
Batch prompting groups multiple independent input samples into a single prompt so that the shared instructions and few-shot demonstrations only need to appear once per batch rather than once per sample, and it asks the model to return the answers for every grouped sample together in one response; because the fixed overhead of instructions and examples is amortized across several samples, the paper reports the token and time cost of inference dropping nearly linearly with the number of samples placed in each batch, while accuracy stays comparable to the one-sample-per-call approach. The option describing fine-tuning a smaller model is wrong because batch prompting is a pure prompting technique applied to the existing model through its ordinary API, with no separate training step. The option describing dropping the few-shot examples is wrong because batch prompting keeps the same examples; it just stops repeating them once per sample. The option describing caching exact-match requests is wrong because batch prompting changes how many samples are packed into a single call, not whether previously seen text is reused from a cache.
Source: Cheng, Kasai & Yu, 'Batch Prompting: Efficient Inference with Large Language Model APIs' (arXiv:2301.08721, 2023)
AI & LLM Engineering · Prompt Engineering · Card 042/081hard
A team wants an LLM to write and keep improving its own instruction for a reasoning task, given only a way to score any candidate instruction against a small held-out set (for example, accuracy on GSM8K). Zhou et al.'s Automatic Prompt Engineer (APE) generates a batch of candidate instructions in one pass from example input-output pairs and selects whichever scores best. Yang et al. (2023), "Large Language Models as Optimizers," propose OPRO for the same kind of problem. How does OPRO's procedure differ from APE's one-pass generate-and-select approach?
AOPRO trains a separate small neural network purely on the numeric scores to predict the best wording, without the LLM ever seeing the scores itself
BOPRO removes scoring from the loop entirely and instead has the LLM vote on which of several candidate instructions sounds most natural to a human reader
COPRO runs an iterative optimization loop: at each step it feeds the model a "meta-prompt" containing the trajectory of previously tried instructions together with their scores, and asks the model to propose a new instruction intended to score higher than the ones already tried, so each new attempt is conditioned on the full history of earlier attempts rather than everything being generated in one pass
DOPRO requires updating the LLM's own weights via gradient descent on the scoring function, making it inapplicable to closed-weight models accessed only through an API
Correct answer: .
OPRO frames prompt discovery as an iterative optimization loop rather than a single generate-and-select pass: at every step, the model is shown a meta-prompt that lists the instructions tried so far alongside their measured scores, and it is asked to propose a new instruction that it expects to score higher than those already tried; that new instruction is scored, added to the trajectory, and the loop repeats, so later proposals are conditioned on the accumulated history of what has and has not worked, letting accuracy climb gradually starting from low-scoring initial prompts. This is a meaningfully different structure from APE, which produces a batch of candidate instructions in one pass from example input-output pairs and picks the best scorer without ever feeding scores back into a further round of generation. The option describing a separate neural network trained purely on scores is wrong because in OPRO the LLM itself reads the scores directly inside the meta-prompt and reasons over them in natural language. The option describing voting on naturalness is wrong because OPRO's proposals are selected by measured task performance, not by a naturalness preference. The option describing gradient descent on model weights is wrong because OPRO treats the optimizer LLM as a black box invoked through ordinary prompting, which is exactly why it works with closed, API-only models.
Source: Yang, Wang, Lu, Liu, Le, Zhou & Chen, 'Large Language Models as Optimizers' (arXiv:2309.03409, 2023)
AI & LLM Engineering · Prompt Engineering · Card 043/081easy
Li et al. (2023), "Large Language Models Understand and Can Be Enhanced by Emotional Stimuli," test appending short psychologically-motivated phrases, such as "This is very important for my career" or "You'd better be sure and think carefully," to the end of otherwise-unchanged task instructions, an approach the paper calls EmotionPrompt. What did the paper report as the effect of adding these phrases, and did the technique require retraining the model?
AThe phrases had no measurable effect on any tested model, confirming that language models are insensitive to wording intended to convey urgency or stakes
BThe phrases improved performance only after the model was fine-tuned on a dataset of emotionally-annotated examples paired with correct answers, so the effect required retraining rather than being a pure prompting technique
CThe phrases decreased benchmark performance by making the model overly cautious and more likely to refuse to answer, though human raters still preferred the more cautious tone
DAdding these appended emotional-stimulus phrases to otherwise-unchanged instructions produced measurable improvements in benchmark performance and in human ratings of the resulting text across multiple LLMs, and the technique required no retraining or fine-tuning of any kind, since it works purely by changing the wording of the prompt
Correct answer: .
EmotionPrompt appends a short psychologically-motivated phrase to an existing instruction without touching anything else about the prompt or the model, and the paper reports that this simple wording change produced measurable improvements in benchmark performance and in human ratings of the resulting text across the multiple LLMs tested, with no retraining or fine-tuning involved at any point; the improvement comes purely from how the request is phrased. The option claiming no measurable effect is wrong because the paper's core finding is precisely the opposite, that these phrases reliably shifted performance and human ratings. The option claiming the benefit only appears after fine-tuning on emotionally-annotated examples is wrong because EmotionPrompt is tested and reported as a prompting-only technique applied to already-trained, frozen models. The option claiming performance dropped due to excessive caution is wrong because the paper's reported outcome is an improvement in both benchmark scores and human preference, not a refusal-driven decline.
Source: Li, Jiang, Zhang, Chen, Lv, Zhao, Wu & Lin, 'Large Language Models Understand and Can Be Enhanced by Emotional Stimuli' (arXiv:2307.11760, 2023)
AI & LLM Engineering · Prompt Engineering · Card 044/081hard
Standard zero-shot chain-of-thought prompting (Kojima et al., 2022) elicits step-by-step reasoning by appending an explicit instruction such as "Let's think step by step" before greedily decoding the response. Wang & Zhou (2024), "Chain-of-Thought Reasoning Without Prompting," report finding reasoning paths in a pretrained model without adding any such instruction to the prompt at all. What do they change instead, and what do they observe as a result?
AInstead of adding any chain-of-thought instruction to the prompt, they change how the very first token of the response is decoded, inspecting the top-k alternative tokens rather than only the single highest-probability greedy token at that first step; branching down some of those alternative paths reveals chain-of-thought reasoning that was already latent in the pretrained model, and when such a path appears, the model tends to show higher confidence in its final answer
BThey fine-tune the pretrained model on a small set of chain-of-thought demonstrations, after which greedy decoding alone reproduces step-by-step reasoning without needing the "let's think step by step" instruction
CThey replace the pretrained model's tokenizer with one that segments numbers digit by digit, which the paper claims is solely responsible for eliciting latent reasoning at the decoding stage
DThey add a hidden system-level instruction that is invisible to the end user but functionally identical to Kojima et al.'s "let's think step by step" phrase, appended before decoding begins
Correct answer: .
Wang and Zhou leave the prompt completely untouched and instead intervene at decoding time: at the first decoding step, rather than following only the single highest-probability greedy token, they branch out and inspect the top-k alternative tokens available at that step, then continue decoding down several of those alternative branches to completion. They find that some of these alternative branches naturally contain step-by-step reasoning that greedy decoding alone would never surface, showing that this reasoning ability is already latent in the pretrained model rather than something an explicit instruction creates from nothing; they also observe that when a decoded path does contain such a reasoning chain, the model's own probability assigned to its final answer tends to be higher, giving a usable signal for picking good paths. The option describing fine-tuning is wrong because the paper's central claim is that no additional training is needed; the reasoning is uncovered from an already-pretrained model purely by changing the decoding procedure. The option describing a digit-by-digit tokenizer swap is wrong because the method makes no change to tokenization at all. The option describing a hidden instruction is wrong because the entire point of the paper is that reasoning appears without any instruction, hidden or otherwise, added to the prompt.
Source: Wang & Zhou, 'Chain-of-Thought Reasoning Without Prompting' (arXiv:2402.10200, 2024)
AI & LLM Engineering · Prompt Engineering · Card 045/081medium
Zou et al. (2023), "Universal and Transferable Adversarial Attacks on Aligned Language Models," introduce the Greedy Coordinate Gradient (GCG) method for finding a short string that, appended to a harmful request, causes safety-trained models to comply. How does GCG find this adversarial suffix, and what does "transferable" mean in the paper's results?
AGCG works by manually testing suffixes proposed by a human red team, ranking them only by how grammatically fluent they read, with no use of gradients or automated search at all
BGCG uses gradient information from one or more open-weight models to greedily search for, and iteratively replace, individual tokens in a candidate suffix so as to increase the likelihood that the target model begins its response by complying with the harmful request; "transferable" describes the finding that a suffix optimized against open-weight models such as Vicuna also induces objectionable output when tested against unrelated closed models it was never optimized against
CGCG requires direct write access to the target model's weights during the attack itself, so it cannot be used against a closed model served only through an API, and "transferable" refers only to porting the attack's code between programming languages
DGCG produces a suffix that is unique to a single specific harmful request and a single specific model, and "transferable" describes only how the resulting text can be copy-pasted between chat interfaces of the same model
Correct answer: .
GCG is an optimization-based attack: it uses gradients computed on one or more open-weight models to identify, at each position in a candidate suffix, which token substitutions would most increase the probability that the target model's response begins by complying with the harmful request, then greedily applies the best substitution found and repeats across many iterations until an effective suffix emerges. The paper's transferability result is that a suffix optimized this way against accessible open-weight models such as Vicuna-7B and 13B still induces objectionable output when later tested against entirely different, closed models it was never optimized against directly, including public chat interfaces the attacker never had gradient access to. The option describing manual human testing ranked by fluency is wrong because GCG is an automated, gradient-guided search, and the resulting suffixes are typically not fluent at all. The option claiming direct weight access to the target model is required is wrong because the transfer result specifically demonstrates success against models the attacker never had such access to. The option restricting transferability to copy-pasting within one model's interfaces is wrong because the paper's finding is that the suffix generalizes across different underlying models, not just across interfaces of the same model.
Source: Zou, Wang, Carlini, Nasr, Kolter & Fredrikson, 'Universal and Transferable Adversarial Attacks on Aligned Language Models' (arXiv:2307.15043, 2023)
AI & LLM Engineering · Prompt Engineering · Card 046/081medium
Tree of Thoughts (Yao et al., 2023) lets a model explore multiple reasoning branches and backtrack from dead ends, but each branch in a tree can only split further or be discarded; branches cannot be recombined with each other. Besta et al., "Graph of Thoughts: Solving Elaborate Problems with Large Language Models," propose a structure that goes beyond this constraint. What does Graph of Thoughts add on top of the tree structure, and what capability does that addition unlock?
AIt removes branching entirely and forces the model onto a single linear chain, trading the exploration benefits of Tree of Thoughts for a large reduction in the number of model calls needed
BIt adds a fixed, hand-written decision tree of if-then rules external to the language model that decides which existing branch to keep, replacing the model's own judgment about branch quality
CIt models the reasoning process as an arbitrary graph, where individual "thoughts" are vertices and dependencies between them are edges rather than being restricted to a single parent-to-child tree shape; this lets thoughts explored on separate branches be merged, aggregated, or fed back into each other through feedback loops, rather than only ever being extended or dropped
DIt requires training a separate small classifier model to score each branch numerically, since the underlying language model in this framework is never asked to judge or compare its own branches
Correct answer: .
Graph of Thoughts generalizes the tree structure of Tree of Thoughts into an arbitrary graph, treating each unit of information the model generates as a vertex and each dependency between units as an edge, rather than restricting every unit to exactly one parent in a strict tree. Because a graph allows edges between any vertices, not just parent-to-child links, this structure lets the framework combine thoughts that were developed independently on separate branches into a single synergistic result, distill a whole network of thoughts down to its essence, or route a thought's output back into an earlier point in the graph as feedback, none of which a tree shape permits once branches have diverged. The option describing removal of branching for a linear chain is wrong because Graph of Thoughts adds structure beyond a tree, it does not simplify below one. The option describing a fixed external decision tree of if-then rules is wrong because the graph's vertices and edges represent the language model's own generated thoughts and their dependencies, not a hand-written rule set replacing the model's judgment. The option describing a separately trained classifier is wrong because the framework still relies on the language model itself to generate, evaluate, and combine thoughts within the graph.
Source: Besta, Blach, Kubicek, Gerstenberger, Podstawski, Gianinazzi, Gajda, Lehmann, Niewiadomski, Nyczyk & Hoefler, 'Graph of Thoughts: Solving Elaborate Problems with Large Language Models' (arXiv:2308.09687, 2023/2024)
AI & LLM Engineering · Prompt Engineering · Card 047/081easy
Khattab et al. (2023), "DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines," argue that hand-written free-form prompt templates discovered by trial and error make LLM pipelines brittle and hard to reuse across models. What does DSPy have the programmer do instead, and what does the framework itself handle automatically?
AThe programmer writes the exact final prompt wording as before, and DSPy's only contribution is translating that wording into several different natural languages automatically
BThe programmer must manually rewrite every prompt for every new base model DSPy is pointed at, since the framework provides no automatic prompt generation or tuning of its own
CThe programmer specifies only the desired output token count, and DSPy pads or truncates the model's natural response to match that fixed length regardless of prompt wording
DThe programmer declares a "signature" describing, in a structured and model-agnostic way, what inputs a step needs and what outputs it should produce, and composes these signatures into modules forming a pipeline; DSPy's own compiler then automatically generates and tunes the actual natural-language prompt text, and can select or generate few-shot demonstrations, needed to make each declared step work
Correct answer: .
DSPy replaces hand-crafted free-form prompt strings with a declarative "signature," a structured, model-agnostic description of what inputs a pipeline step consumes and what outputs it should produce, and lets the programmer compose several such signatures into modules that form a full pipeline. Rather than the programmer discovering and hard-coding the exact wording that makes a given base model behave correctly, DSPy's own compiler takes over that job: it automatically generates and tunes the natural-language prompt text for each declared step, and can also select or generate the few-shot demonstrations used within that prompt, so the same declared pipeline can be recompiled for a different base model without the programmer rewriting prompt wording by hand. The option describing automatic translation into other natural languages is wrong because DSPy's compilation targets prompt effectiveness for a given model and task, not multilingual translation of a fixed wording. The option claiming DSPy provides no automatic prompt generation is wrong because automatic prompt generation and tuning by the compiler is the paper's central contribution. The option describing padding or truncating output to a fixed token count is wrong because DSPy's declared signatures describe input/output structure and intent, not a fixed output length enforced independently of prompt wording.
Source: Khattab, Singhvi, Maheshwari, Zhang, Santhanam, Vardhamanan, Haq, Sharma, Joshi, Moazam et al., 'DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines' (arXiv:2310.03714, 2023)
AI & LLM Engineering · Prompt Engineering · Card 048/081easy
An engineering team building a customer-support chatbot that can call internal tools has already learned OWASP's distinction between direct prompt injection (attacker text typed straight into the chat) and indirect prompt injection (malicious instructions hidden in a document, webpage, or tool result the model later reads). Separately from that classification, what defense-in-depth mitigations does the OWASP Top 10 for LLM Applications (2025) recommend for reducing the risk and blast radius of a successful prompt injection?
AA combination of layered controls: constraining model behavior and output format through the system prompt, segregating untrusted external content so it cannot be interpreted as an instruction, restricting the model's tools and permissions to the minimum needed for the task, and requiring human approval before any high-risk or irreversible action is carried out
BRelying on a single measure, training a dedicated classifier that scans every user message for injection attempts, which OWASP describes as sufficient on its own once deployed, without any additional tool-permission or output controls
CDisabling all tool use entirely for any model that might ever process text from an external source, since OWASP states this is the only mitigation that fully eliminates the risk
DEncrypting the model's system prompt so that it cannot be extracted through prompt leaking, which OWASP identifies as its primary recommended defense against prompt injection specifically
Correct answer: .
OWASP's guidance treats prompt injection as a risk that cannot be fully eliminated by any single control, so it recommends layering several mitigations together: constraining what the model is allowed to do and what output format it must follow through the system prompt, segregating content pulled in from untrusted external sources so the model does not treat it as an instruction, restricting the tools and permissions available to the model to only what a given task actually needs, and requiring a human to approve any action that would be high-risk or hard to reverse before it is carried out. The option describing a single injection-scanning classifier as sufficient on its own is wrong because OWASP explicitly frames defense in depth, combining several layers, as necessary rather than relying on one filter. The option describing disabling all tool use for any model touching external text is wrong because OWASP's recommendations focus on constraining and monitoring tool use, not eliminating it outright, and does not present blanket disabling as the only real mitigation. The option describing system-prompt encryption as OWASP's primary recommended defense is wrong because encrypting the system prompt addresses prompt leaking, a related but separate risk, and is not the mitigation OWASP centers for prompt injection itself.
Source: OWASP, 'OWASP Top 10 for LLM Applications 2025', LLM01: Prompt Injection
AI & LLM Engineering · Prompt Engineering · Card 049/081easy
Du et al. (2023), "Improving Factuality and Reasoning in Language Models through Multiagent Debate," test an alternative to having a single model instance revise its own answer alone, as in Self-Refine. What debate procedure do they propose, and how does it differ from a single model critiquing and refining its own output?
AA single model instance argues both sides of a debate against itself within one continuous response, then declares its own earlier argument the winner without ever comparing the two arguments against each other
BMultiple separate instances of a language model each independently produce an answer and its reasoning, then are shown one another's answers and reasoning over several rounds and asked to update their own response in light of the others' arguments, continuing until the instances converge on a shared final answer; this differs from Self-Refine, which uses one model instance for every role, generator, critic, and refiner, with no independent second party involved
COne model instance generates several candidate answers, and a much smaller, separately trained model is trained to grade which of those answers is factually best, with no back-and-forth exchange of arguments involved
DMultiple model instances are merged into a single set of weights via averaging before generating one answer, so no exchange of natural-language arguments occurs between separate active instances at inference time
Correct answer: .
Multiagent debate has multiple separate instances of a language model, or of several different models, each independently generate their own answer and reasoning for the same question, then exposes each instance to the other instances' answers and reasoning over multiple rounds, asking every instance to update its own response in light of what the others argued; the process continues across rounds until the instances converge on a shared final answer, which the paper reports improves factual accuracy and reduces hallucinated or fallacious answers compared to a single instance working alone. This differs from Self-Refine specifically because Self-Refine uses exactly one model instance to play every role, generating a draft, critiquing it, and refining it, with no second, independently reasoning party ever involved, whereas debate depends on genuinely separate instances that can disagree with each other. The option describing one instance arguing both sides internally is wrong because it never actually confronts an independent second perspective, unlike true multiagent debate. The option describing a separately trained grading model with no argument exchange is wrong because debate's core mechanism is the multi-round exchange of reasoning between instances, not a one-shot grading step. The option describing weight averaging before generation is wrong because debate operates entirely at inference time through natural-language exchange, never merging the models' parameters together.
Source: Du, Li, Torralba, Tenenbaum & Mordatch, 'Improving Factuality and Reasoning in Language Models through Multiagent Debate' (arXiv:2305.14325, 2023)
AI & LLM Engineering · Prompt Engineering · Card 050/081easy
A developer wants an LLM to extract structured data (name, date, amount) from unstructured invoice text. They try asking the model in plain English but get inconsistent output formats. What prompting technique would most reliably produce consistently structured output?
AProviding a concrete output schema or example in the prompt (few-shot prompting with structured examples), such as showing the model one or two completed extractions in the exact JSON format desired, so the model replicates the structure rather than inventing its own format each time
BIncreasing the temperature parameter to its maximum value, which encourages the model to explore more formatting options and eventually converge on the correct one
CAsking the model to explain its reasoning step by step before producing any output, which guarantees the output will be in valid JSON without needing to specify a schema
DSending the same prompt multiple times and averaging the responses, since the model's output will naturally stabilise into a consistent format with enough repetitions
Correct answer: .
Few-shot prompting with structured examples is the most reliable way to get consistent output formatting. By showing the model exactly what a correct extraction looks like — including the JSON keys, their order, and the value formats — the model learns to replicate that structure. Many APIs also offer structured-output or JSON-mode features that constrain the model's generation to valid JSON, which pairs with few-shot examples for even higher reliability. Raising temperature increases randomness, making output less consistent, not more. Chain-of-thought reasoning helps with complex logic but does not inherently enforce a specific output format. Sending the same prompt repeatedly wastes tokens and does not converge on a stable format because each generation is independent.
AI & LLM Engineering · Prompt Engineering · Card 051/081easy
A developer adds the instruction 'Think step by step before answering' to a prompt and notices the model's accuracy improves on a multi-step math problem. What is this technique called, and why does it help?
AChain-of-thought (CoT) prompting — it works because the intermediate reasoning tokens give the model more computation to decompose the problem, catch errors in sub-steps, and maintain state across a multi-step calculation rather than attempting to jump directly from the question to the final answer
BRetrieval-augmented generation (RAG) — the instruction causes the model to search an external database for the answer, which is why accuracy improves
CTemperature tuning — the phrase 'think step by step' automatically lowers the model's temperature parameter, making the output more deterministic and therefore more accurate
DPrompt injection — the instruction hijacks the model's control flow, which coincidentally produces correct answers for math problems but is considered a security vulnerability
Correct answer: .
Chain-of-thought prompting, popularised by Wei et al. (2022), improves performance on reasoning tasks by encouraging the model to generate intermediate steps. Each reasoning token effectively gives the model additional 'thinking time' — since transformers perform a fixed amount of computation per token, more tokens mean more opportunities to break down complex logic, track intermediate results, and self-correct. The instruction does not trigger any external retrieval (that would require an actual RAG system, not a prompt phrase). It does not change the temperature parameter, which is set via the API, not through prompt text. Calling it prompt injection misuses the term; prompt injection refers to adversarial inputs that override system instructions, not legitimate prompting techniques.
Source: Wei et al. (2022), 'Chain-of-Thought Prompting Elicits Reasoning in Large Language Models', arXiv:2201.11903
AI & LLM Engineering · Prompt Engineering · Card 052/081easy
A team notices that when a prompt includes an opinionated aside right before the actual question (for example, 'I really think the answer is X, but what do you think?'), an LLM's response is swayed toward agreeing with the aside rather than giving an independent, well-reasoned answer, even when the aside is unrelated to what is actually correct. Weston & Sukhbaatar (2023), 'System 2 Attention (is something you might need too),' propose a two-stage inference method to reduce this kind of susceptibility to irrelevant or biasing context. What does System 2 Attention (S2A) do?
AIt fine-tunes the model's attention weights on a dataset of biased versus unbiased prompts so the model learns, once and for all, to ignore opinionated asides in any future prompt
BIt has the LLM first regenerate the input context in an initial inference pass, rewriting it to strip out irrelevant or biasing content such as the aside, and then attends to and answers based on only that regenerated context in a second pass
CIt lowers the sampling temperature to 0 for the duration of the response, which is described as suppressing the model's tendency to mirror the emotional tone of the prompt
DIt appends a fixed disclaimer instructing the model to disregard any opinions in the prompt, relying on the model always giving the highest priority to instructions placed at the very end of a prompt
Correct answer: .
System 2 Attention is a two-stage, inference-only method: in the first pass, the model is prompted to regenerate the given context, rewriting it so that it retains only the information relevant to answering correctly and drops irrelevant or biasing material such as an opinionated aside; in the second pass, the model attends to and answers from that regenerated context rather than the original one. The paper reports this reduces sycophancy and increases factuality and objectivity on tasks containing opinion or irrelevant information. The option describing weight fine-tuning is wrong because S2A requires no training or weight updates at all — it works purely through additional inference-time prompting on a frozen model, which is the whole point of the technique being deployable on models a team cannot retrain. The option about lowering temperature to 0 is wrong because temperature only controls how deterministically the model samples from its output distribution; it does nothing to change what content the model attends to in its input, so a biasing aside earlier in the context would still be attended to and would still be able to sway a low-temperature response. The option about a fixed end-of-prompt disclaimer is wrong on two counts: S2A does not rely on a static instruction at all, and the premise that end-of-prompt instructions always receive the highest priority is not something the paper claims or relies on — instruction position effects like recency bias are a separate, unrelated phenomenon from S2A's actual regenerate-then-attend mechanism.
Source: Weston & Sukhbaatar, 'System 2 Attention (is something you might need too)' (2023), arXiv:2311.11829
AI & LLM Engineering · Prompt Engineering · Card 053/081medium
A team building few-shot chain-of-thought prompts for grade-school and competition math problems has a small pool of hand-written exemplars, some walking through only two reasoning steps and others walking through eight or more before reaching an answer. Fu et al. (2022), 'Complexity-Based Prompting for Multi-Step Reasoning,' study how the choice of exemplars, and how multiple sampled outputs are combined, affects accuracy on this kind of multi-step task. What does their complexity-based approach do?
AIt selects the shortest, simplest exemplars available for the few-shot prompt, on the theory that concise demonstrations reduce the chance the model copies an irrelevant reasoning pattern
BIt selects exemplars at random regardless of step count, then replaces greedy decoding with a single high-temperature sample to increase output diversity
CIt selects exemplars with the fewest reasoning steps, then applies a majority vote over multiple greedily decoded generations of the same prompt
DIt deliberately selects exemplars with more reasoning steps as the few-shot demonstrations, and when combining several sampled reasoning chains for a test question, favors a majority vote taken among the more complex generated chains rather than weighting every sampled chain equally
Correct answer: .
Fu et al. show that, for multi-step reasoning tasks, deliberately choosing few-shot exemplars with more reasoning steps ('complex' demonstrations) as the prompt outperforms choosing simpler ones, on the reasoning that a prompt should model the depth of reasoning the test problem is likely to require. They pair this with a complexity-based consistency step: rather than sampling several chain-of-thought outputs and taking a plain majority vote across all of them equally, they take the majority vote specifically among the generated reasoning chains that turned out to be the most complex (the most reasoning steps), on the finding that more complex generated chains are more likely to be correct. The option favoring the shortest exemplars is wrong because it is the opposite of the paper's core finding — simpler demonstrations under-prepare the model for problems that need several intermediate steps. The option using random exemplar selection with a single high-temperature sample is wrong on both counts: exemplar choice is not random in this method, and a single higher-temperature sample is not a substitute for sampling multiple chains and voting among them. The option combining 'fewest reasoning steps' exemplars with a plain majority vote is wrong because it inverts the exemplar-selection finding (fewest instead of most reasoning steps) and describes ordinary self-consistency's equal-weighted vote rather than the paper's complexity-weighted vote among the more complex sampled chains.
Source: Fu et al., 'Complexity-Based Prompting for Multi-Step Reasoning' (2022), arXiv:2210.00720
AI & LLM Engineering · Prompt Engineering · Card 054/081easy
A team has a large pool of thousands of labeled input-output pairs but can fit only a handful of them as few-shot exemplars into a single prompt due to context-length limits. Rather than picking the same fixed set of exemplars for every test input, Liu et al. (2021), 'What Makes Good In-Context Examples for GPT-3?,' propose retrieving a different set of exemplars for each new input at query time. How do they decide which exemplars from the pool to retrieve for a given input?
AFor each test input, they embed it and retrieve the exemplars from the candidate pool whose embeddings are most semantically similar (nearest neighbours) to that input, so a different exemplar set is used per query
BThey retrieve exemplars uniformly at random from the pool for every query, on the reasoning that random sampling prevents the model from overfitting to any single fixed prompt template
CThey select whichever exemplars have the shortest input text, since shorter exemplars leave more of the context window available for the model's own reasoning
DThey select whichever exemplars appear earliest in the training set's original ordering, since fixing exemplar order across every query improves output calibration
Correct answer: .
Liu et al. propose retrieval-based exemplar selection: for each new test input, an embedding model is used to find the exemplars in a large candidate pool that are semantically closest to that specific input, and those nearest-neighbour exemplars are used to build a custom few-shot prompt for that query, rather than reusing one static set of exemplars for every input. The paper reports this consistently outperforms random exemplar selection, with especially large gains on tasks like table-to-text generation and open-domain question answering, because semantically related examples give the model more relevant demonstrations of the specific mapping needed for that input and reduce noise from mismatched exemplars. The option describing random exemplar retrieval is wrong because it is presented in the paper as the weaker baseline the retrieval-based method is compared against and improves upon, not as the method itself. The option about picking the shortest exemplars is wrong because exemplar length plays no role in the selection criterion described in the paper; it optimizes for semantic relevance to the query, not brevity. The option about fixing exemplar order by the training set's original ordering is wrong because it describes a static, query-independent selection rule, which is exactly the kind of fixed-prompt approach the paper's per-query retrieval method is designed to improve on, and the paper's contribution concerns which exemplars to select, not their ordering within the prompt.
Source: Liu et al., 'What Makes Good In-Context Examples for GPT-3?' (2021), arXiv:2101.06804
AI & LLM Engineering · Prompt Engineering · Card 055/081medium
A team wants a single LLM deployment to handle a wide range of unrelated task types, such as coding, arithmetic word problems, and creative writing, without hand-crafting a separate task-specific prompt template for each one in advance, and without any additional fine-tuning. Suzgun & Kalai (2024), 'Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding,' propose using a single underlying model in two distinct roles at inference time. What is this structure?
AA single fixed instruction is prepended to every task type, and the model answers directly in one pass with no decomposition into subtasks at all
BTwo separately fine-tuned checkpoints of the same base model are created, one trained to plan and one trained to execute, and they pass messages to each other over an API
COne instance of the model acts as a high-level 'conductor' that breaks the task into subtasks and writes a tailored instruction for each, then dispatches each subtask to a separate 'expert' instance of the same underlying model, before integrating their outputs into one final response, all through prompting alone
DThe model is asked to generate several candidate prompts for the task, and a human reviewer manually selects the best one before any generation of the final response begins
Correct answer: .
Meta-prompting turns a single language model into a 'conductor' that receives the overall task, decomposes it into subtasks, and generates a tailored high-level instruction for each subtask; it then routes each subtask, with its own instruction, to a separate 'expert' instance of that same underlying model (a fresh conversation or role-prompted call, not a different model or checkpoint), collects the experts' outputs, and integrates them into the final response. Because this scaffolding is achieved entirely through prompting and orchestration of one frozen model, it is task-agnostic and needs no task-specific fine-tuning or hand-written template per task type. The option with a single fixed instruction and no decomposition is wrong because it describes the plain, undifferentiated prompting the paper is specifically trying to improve on, not meta-prompting's conductor-and-experts structure. The option describing two separately fine-tuned checkpoints is wrong because meta-prompting uses one frozen, already-trained model in two prompted roles, not two differently trained models; introducing separate fine-tuning would defeat the task-agnostic, no-training-required framing of the method. The option involving a human manually selecting among candidate prompts is wrong because meta-prompting's conductor-expert loop is fully automatic at inference time, with no human in the loop selecting or approving intermediate prompts.
Source: Suzgun & Kalai, 'Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding' (2024), arXiv:2401.12954
AI & LLM Engineering · Prompt Engineering · Card 056/081hard
After generating several candidate chain-of-thought answers to an arithmetic word problem, a team wants a way to score which candidate is most likely correct, using only the same LLM that generated the candidates, without access to any external calculator, code interpreter, or separately trained verifier model. Weng et al. (2022), 'Large Language Models are Better Reasoners with Self-Verification,' propose a backward-verification procedure for exactly this. How does it score a given candidate answer?
AIt asks the model to restate the candidate answer in different words and checks whether the restatement uses the same number of tokens as the original, treating matching length as a proxy for confidence
BIt takes the candidate's conclusion, uses it as a given condition to construct a new question that masks one of the original problem's stated conditions, asks the model to predict that masked original condition from the conclusion and the rest of the problem, and scores the candidate by how accurately the model recovers it
CIt has the model assign itself a numeric confidence score from 0 to 100 directly, based on how fluent and well-organized the candidate's prose reads, without referring back to the original problem statement at all
DIt counts how many of the candidate's chain-of-thought sentences use passive voice, on the reasoning that passive constructions indicate the model is unsure of a causal claim it is making
Correct answer: .
Weng et al.'s backward verification works by treating a candidate's conclusion as if it were a known fact: it builds a new, derived question in which the conclusion is supplied as a condition and one of the original problem's own stated conditions is masked out, then asks the same model to predict that masked condition using the conclusion and the remaining original conditions. If the model can accurately recover the masked condition this way, the candidate answer is scored as more likely correct, because a correct conclusion should let you reconstruct the facts that were used to reach it; an incorrect conclusion, by contrast, tends to produce inconsistent or wrong reconstructions. This gives an interpretable, self-contained validation score with no external calculator, code execution, or separately trained verifier involved. The option about matching restatement length is wrong because token count says nothing about logical consistency between a conclusion and the problem's original conditions, which is the actual mechanism being scored. The option about a direct numeric self-rating based on prose fluency is wrong because it never checks the conclusion against the problem's conditions at all, and confident-sounding but wrong answers are a well-known failure mode that this kind of unconditioned self-rating would not catch. The option about counting passive-voice sentences is wrong because grammatical voice has no established connection to reasoning correctness and plays no role in the paper's method.
Source: Weng et al., 'Large Language Models are Better Reasoners with Self-Verification' (2022), arXiv:2212.09561
AI & LLM Engineering · Prompt Engineering · Card 057/081easy
A team wants an LLM to double-check and, if needed, revise its own answer to a math word problem across several automatic rounds within the same session, stopping once the answer stops changing, without using any external tool, calculator, or separately trained verifier model, and without simply asking the model to critique its own reasoning in open-ended prose. Zheng et al. (2023), 'Progressive-Hint Prompting Improves Reasoning in Large Language Models' (PHP), propose a specific mechanism for this. What does PHP do?
AIt generates ten independent chain-of-thought answers to the question in a single pass and returns whichever answer appears most frequently among them, with no further rounds of interaction
BIt fine-tunes a small verifier model on a labeled dataset of correct and incorrect chain-of-thought traces, then uses that verifier to pick the best of several candidate answers
CIt rewrites the original question once, in a single pass, to remove ambiguous phrasing, then re-asks the rewritten version and returns that answer directly with no further rounds
DIt feeds the model's own most recently generated answer back into the prompt as a 'hint' alongside the original question and re-asks, repeating this hint-and-reask loop across further rounds until two consecutive rounds produce the same answer, at which point that stable answer is treated as final
Correct answer: .
Progressive-Hint Prompting automates multiple rounds of interaction between the user and the model by feeding the model's own previously generated answer back into the prompt as a hint alongside the original question, then re-asking; this hint-and-reask cycle repeats across further automatic rounds, using each round's answer as the next round's hint, until two consecutive rounds converge on the same answer, at which point that stable answer is returned as final. The paper frames this as orthogonal to, and combinable with, techniques like chain-of-thought and self-consistency, since it operates on top of whatever single-pass answer those techniques already produce. The option about generating ten independent answers and taking the most frequent one describes self-consistency's single-pass majority vote, not PHP's iterative hint-and-reask loop across multiple rounds. The option about fine-tuning a separate verifier model is wrong because PHP requires no training data, labeled traces, or separately trained model at all — it only re-prompts the same frozen LLM with its own prior answer. The option about a single rewrite-and-reask pass is wrong because it describes only one round with no hint from a previous answer and no repetition until convergence, missing PHP's defining progressive, multi-round hinting mechanism entirely.
Source: Zheng et al., 'Progressive-Hint Prompting Improves Reasoning in Large Language Models' (2023), arXiv:2304.09797
AI & LLM Engineering · Prompt Engineering · Card 058/081easy
A team wants a zero-shot chain-of-thought prompting method that produces highly structured, easy-to-parse intermediate reasoning rather than free-flowing prose, and that does not require any hand-written multi-step exemplars in the prompt. Jin & Lu (2023), 'Tab-CoT: Zero-shot Tabular Chain of Thought,' propose prompting the model to organize its reasoning in a specific format. What is it?
AThe model is prompted to lay out its reasoning as a markdown-style table, with a row per reasoning step and columns such as the step number, a subquestion for that step, the process or calculation used to answer it, and the resulting intermediate value, reaching the final answer only after the table is filled in
BThe model is prompted to output its reasoning as a single JSON object containing only one field, the final answer, with no intermediate reasoning captured anywhere in the output
CThe model is prompted to number each sentence of an otherwise free-form paragraph of reasoning sequentially, without imposing any column structure on the content
DThe model is first given several worked, table-formatted examples as few-shot exemplars, and only then asked to fill in a table of its own for the new problem
Correct answer: .
Tab-CoT prompts the model to express its reasoning as a table rather than as prose, with each row representing one reasoning step and columns capturing the step index, a subquestion the model is trying to answer at that step, the process or calculation it performs to answer that subquestion, and the resulting intermediate value; the model fills in the table row by row and only reaches its final answer once the table is complete. This structured format is generated by the model itself with a zero-shot instruction, so it removes the need for any hand-written multi-step exemplars while still making the reasoning process explicit and easy to parse programmatically. The option describing a bare JSON object with only a final answer is wrong because it captures no intermediate reasoning at all, which is the opposite of Tab-CoT's purpose of making the reasoning process explicit and structured. The option about numbering sentences in an otherwise free-form paragraph is wrong because it imposes no column structure and does not organize the reasoning into the step, subquestion, process, and result fields that define Tab-CoT's table format. The option requiring worked table-formatted few-shot exemplars is wrong because Tab-CoT is specifically presented as a zero-shot method: the model produces the tabular structure itself from a general instruction, without needing any hand-written table exemplars first, which is exactly what distinguishes it from a few-shot tabular approach.
Source: Jin & Lu, 'Tab-CoT: Zero-shot Tabular Chain of Thought' (2023), arXiv:2305.17812, Findings of ACL 2023
AI & LLM Engineering · Prompt Engineering · Card 059/081hard
A team has several different prompt templates for the same classification task, phrased slightly differently (for example, a yes/no question form versus an open-ended question form), and each template alone gives noisy, inconsistent predictions on held-out examples, with no single template reliably best across all inputs. Arora et al. (2022/2023), 'Ask Me Anything: A Simple Strategy for Prompting Language Models' (AMA), propose combining the noisy predictions from multiple such prompts into one final label. How do they do this, beyond a simple equally-weighted majority vote?
AThey discard every template except whichever single one scores highest on a large labeled validation set, and use only that one template at test time from then on
BThey train a large separate classifier on the LLM's internal hidden-state activations, and this classifier alone produces the final label, ignoring the individual prompts' answers entirely
CThey treat each prompt template as a noisy 'weak label source' over the unlabeled test examples and combine the templates' predictions using a weak-supervision aggregation method, which models and corrects for each prompt's individual reliability and for correlations between prompts, rather than counting every template's vote equally
DThey average the raw output token probabilities of every template's completion character by character, since averaging logits is mathematically equivalent to weak-supervision aggregation for a classification task
Correct answer: .
Ask Me Anything treats each of the several imperfect prompt templates as an independent noisy 'weak label source' over the unlabeled test set, similar to how weak supervision treats independent noisy labeling functions, and aggregates their predictions with a weak-supervision method that estimates each source's reliability and any correlations between sources from the pattern of agreement and disagreement across all templates and examples, then combines them into a final label weighted by that estimated reliability, rather than giving every template's vote equal weight as a plain majority vote would. The paper reports this weak-supervision aggregation outperforms simple majority voting on several benchmarks. The option that keeps only the single best-scoring template is wrong because it discards the multi-prompt aggregation entirely and requires a large labeled validation set to pick that one template, which defeats the paper's goal of combining several imperfect, cheaply-obtained prompts. The option describing a separate classifier trained on hidden-state activations is wrong because AMA works purely from the models' output predictions across prompts, with no probing of internal activations or separate classifier training. The option about averaging raw token probabilities is wrong because logit averaging is not mathematically equivalent to weak-supervision aggregation, which explicitly models per-source reliability and correlation rather than blending output probabilities uniformly.
Source: Arora et al., 'Ask Me Anything: A Simple Strategy for Prompting Language Models' (2022/2023, ICLR 2023), arXiv:2210.02441
AI & LLM Engineering · Prompt Engineering · Card 060/081easy
A team's prompts routinely include a lengthy few-shot exemplar set plus a long retrieved reference document, and they want to cut the number of tokens billed per request and reduce latency, without retraining or fine-tuning the target LLM, and while keeping task accuracy close to what the full, uncompressed prompt achieves. Jiang et al. (2023), 'LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models,' propose a method for this. What does it do?
AIt replaces every word in the prompt with its shortest available synonym from a fixed thesaurus, on the theory that shorter synonyms always exist and always preserve the original meaning
BIt uses a separate, small language model to estimate each token's perplexity given its context, then removes the lowest-information tokens under a target compression ratio, budgeting more aggressive compression for parts of the prompt such as exemplars while preserving instructions more conservatively
CIt has the target LLM itself summarize the entire prompt into a single sentence, then discards the original text and sends only that one-sentence summary to the target model
DIt strips only whitespace, line breaks, and punctuation from the prompt, since token count is otherwise fixed once the vocabulary of words in the prompt is decided
Correct answer: .
LLMLingua uses a small, separate language model to compute each token's perplexity in context, treating low-perplexity (highly predictable, low-information) tokens as safe to drop and higher-perplexity tokens as more essential to keep, and removes tokens under a target overall compression ratio; it also applies a dynamic budget across different parts of the prompt, allowing more aggressive compression on components like few-shot exemplars while compressing instructions more conservatively, since instructions tend to be more sensitive to information loss. This lets a team shrink a prompt's token count, and therefore its cost and latency, on an unmodified target LLM while the paper reports accuracy stays close to the uncompressed baseline on benchmarks like GSM8K and BBH. The option about replacing words with thesaurus synonyms is wrong because it is not the mechanism used, and synonym substitution does not reliably reduce token count or preserve conditional dependencies between the retained tokens the way perplexity-based removal does. The option about the target LLM summarizing its own prompt into one sentence is wrong because LLMLingua's compression is done by a separate small model operating on the original tokens, not by having the target model generate and then rely on a lossy natural-language summary of itself. The option about only stripping whitespace and punctuation is wrong because such formatting characters are a minor part of typical prompt token counts, and this approach would not achieve anything close to the large compression ratios LLMLingua reports.
Source: Jiang et al., 'LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models' (2023), arXiv:2310.05736
AI & LLM Engineering · Prompt Engineering · Card 061/081medium
Min et al. (2022) found that replacing the correct labels in few-shot demonstrations with random incorrect ones barely hurts a large language model's in-context learning accuracy on many tasks, a result that is hard to square with the intuitive idea that the model is learning the exact input-output mapping from the examples the way a small supervised classifier would. Xie et al. (2021), 'An Explanation of In-context Learning as Implicit Bayesian Inference,' propose a theoretical account of why few-shot demonstrations help at all, developed independently of that later empirical finding. What is their explanation?
AThe demonstrations work purely as a formatting cue: the model has already memorized the exact benchmark's answer key during pretraining, and the examples only signal which memorized answer set to output
BEach demonstration silently triggers a gradient-descent-like weight update inside the frozen model's forward computation, functionally equivalent to running a few steps of fine-tuning on those examples before generating the answer
CThe examples work only because the model recognizes their surface formatting, such as delimiters and layout, and pattern-matches new inputs against that formatting alone, independent of the examples' actual content
DBecause long, coherent pretraining documents implicitly require the model to infer a shared latent concept or topic connecting the text seen so far in order to predict the next token well, at inference time a prompt's demonstrations similarly signal a shared latent concept, and the model performs an implicit Bayesian inference over that latent concept from the examples and applies it to the query, even when the examples' exact input-output mapping is partly noisy
Correct answer: .
Xie et al. argue that in-context learning emerges from how language models are pretrained: long, coherent pretraining documents implicitly require the model to infer a shared latent document-level concept in order to predict upcoming tokens well throughout that document. At inference time, a prompt's demonstrations play an analogous role, implicitly signaling a shared latent concept or task that connects them; the model performs an implicit Bayesian inference over that latent concept given the demonstrations, and then applies the inferred concept to answer the new query. This account explains why in-context learning can work even when the exact mapping in the demonstrations is imperfect, since the model is inferring the underlying task rather than memorizing a precise function from inputs to outputs, distinct from and developed independently of the later empirical observation that scrambling demonstration labels often barely hurts accuracy. The option about memorized benchmark answer keys is wrong because it would not explain in-context learning on genuinely novel tasks or example sets the model could not have memorized, which the phenomenon generalizes to. The option describing an implicit gradient-descent update inside the forward pass is wrong because that describes a different, later theoretical account of in-context learning, not the implicit Bayesian inference over a latent concept that Xie et al. propose. The option about pattern-matching only on surface formatting is wrong because it would predict that scrambling demonstration content while keeping formatting fixed should have no effect on accuracy, which is a stronger and different claim than the concept-inference account this paper actually makes.
Source: Xie et al., 'An Explanation of In-context Learning as Implicit Bayesian Inference' (2021/2022, ICLR 2022), arXiv:2111.02080
AI & LLM Engineering · Prompt Engineering · Card 062/081medium
A team wants a single LLM to handle a whole benchmark of related but non-identical reasoning problems by selecting and combining general reasoning strategies (such as 'break the problem into subproblems' or 'think step by step') into a structure tailored to the task family, but without hand-designing that structure themselves and without paying the cost of an expensive multi-path search (such as exploring and scoring many branches) on every single problem instance. Zhou et al. (2024), 'Self-Discover: Large Language Models Self-Compose Reasoning Structures,' propose a way to do this. What is their two-stage approach?
AThe model runs a full Tree-of-Thoughts search independently on every problem instance, exploring and scoring several candidate reasoning branches per instance and keeping whichever branch scores highest for that instance
BThe model is fine-tuned once, using gradient updates on a labeled dataset of worked reasoning traces for the task family, to internalize a single fixed reasoning procedure it then applies to every instance
CIn a one-time 'discover' stage, run only over a small number of unlabeled example problems from the task, the model selects, adapts, and composes a handful of generic atomic reasoning modules (e.g., critical thinking, decomposition into subtasks) into one explicit, task-specific reasoning structure; that same discovered structure is then reused to solve every individual problem in the task at ordinary single-pass inference cost, without repeating the discovery step per instance
DA separate, larger 'teacher' model first solves every problem in the benchmark and writes out its reasoning, and the target model is then prompted with these solved examples as few-shot demonstrations for each new instance
Correct answer: .
Self-Discover operates in two distinct stages. In a one-time 'discover' stage, the model looks at a handful of unlabeled example problems from the task and, working purely through prompting, selects a subset of generic atomic reasoning modules from a fixed set (things like 'break the problem into subproblems' or 'use critical thinking'), adapts their wording to the specific task, and composes them into one explicit reasoning structure, essentially a task-specific reasoning template. This structure is discovered only once per task family. In the second stage, that same discovered structure is reused, unchanged, as the prompt scaffold for solving every individual problem in the task, at the same single-pass inference cost as an ordinary prompted answer, with no further search or repeated discovery per instance. The option describing per-instance Tree-of-Thoughts search is wrong because that describes exploring and scoring multiple reasoning branches for each individual problem, which is exactly the expensive, inference-heavy alternative the paper shows Self-Discover outperforms while using far less compute, not what Self-Discover itself does. The option describing gradient-based fine-tuning is wrong because Self-Discover never updates any model weights; the reasoning structure is composed entirely through prompting the frozen model. The option describing a separate teacher model whose solved examples become few-shot demonstrations is wrong because Self-Discover uses only the one target model to both discover the structure and solve every instance, and does not rely on any worked example solutions being fed in as demonstrations at inference time.
Source: Zhou et al., 'Self-Discover: Large Language Models Self-Compose Reasoning Structures' (NeurIPS 2024), arXiv:2402.03620
AI & LLM Engineering · Prompt Engineering · Card 063/081hard
A team wants to automatically improve both a task-prompt (the instruction actually shown to the LLM for solving a target task) and the process used to generate new candidate task-prompts, across many rounds, using only a training set of labeled examples to score fitness and without any gradient-based training of the LLM's weights. Fernando et al. (2023), 'Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution,' propose a method for this that goes beyond single-pass approaches such as Automatic Prompt Engineer (APE), which generates one batch of candidate instructions from example input-output pairs and picks the best-scoring one. What does Promptbreeder do differently?
AIt maintains an evolving population of task-prompts, using an LLM to repeatedly mutate them into new candidates across many generations and keeping variants that score best on a fitness set as in a genetic algorithm; crucially, the mutation-prompts that instruct the LLM how to mutate the task-prompts are themselves part of the evolving population and are improved over the same generations, making the whole process self-referential rather than a single-round generate-and-select step
BIt fine-tunes a smaller auxiliary model on the training set's labeled examples to directly predict, in one shot, which single task-prompt from a fixed candidate list will score highest, then discards all other candidates
CIt has a single frozen task-prompt debated word-by-word between two independent LLM instances until they reach unanimous agreement on its final wording, without ever generating new candidate prompts beyond edits proposed during the debate
DIt searches over prompts using gradient-based backpropagation directly through the LLM's embedding layer to find a continuous soft-prompt vector, which is then rounded to the nearest discrete tokens for the final task-prompt
Correct answer: .
Promptbreeder maintains an evolving population of task-prompts and uses an LLM to mutate them into new candidate task-prompts across many generations, keeping the fittest-scoring variants on a training set much like a genetic algorithm. What makes the method self-referential is that the mutation-prompts, the instructions that tell the LLM how to mutate a given task-prompt, are themselves part of what evolves: they are mutated and selected over the same generations as the task-prompts, so the system improves not just its candidate solutions but its own process for generating new candidates. This lets it outperform single-pass, non-iterative methods like APE as well as fixed prompting strategies such as Chain-of-Thought and Plan-and-Solve on reasoning benchmarks. The option describing a fine-tuned auxiliary model predicting a single best prompt in one shot is wrong because Promptbreeder never trains any separate model and never collapses to a single one-shot prediction; it iterates over many generations of an evolving population instead. The option describing two LLM instances debating a single frozen prompt to unanimous agreement is wrong because Promptbreeder does not use a debate protocol at all, and it maintains and mutates a whole population of diverse candidate prompts rather than converging edits onto one fixed prompt. The option describing gradient-based backpropagation into a continuous soft-prompt vector is wrong because that describes a different technique, prompt tuning over continuous embeddings, which requires access to and differentiation through the model's parameters; Promptbreeder instead operates entirely through discrete, LLM-generated prompt mutations with no gradients computed on the target model at all.
Source: Fernando et al., 'Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution' (2023), arXiv:2309.16797
AI & LLM Engineering · Prompt Engineering · Card 064/081easy
After a large language model generates an answer to an open-ended question, a team wants the same model, without any additional training and without access to any separately trained verifier, to estimate how likely it is that its own answer is actually correct. Kadavath et al. (2022), 'Language Models (Mostly) Know What They Know,' study prompting the model to do exactly this. What do they have the model do, and what do they find about larger models' resulting estimates?
AThe model is asked to restate its own answer using different wording several times; if the reworded versions are lexically identical to the original, the model is deemed confident, and if they differ at all, the answer is discarded regardless of correctness
BThe model is shown only the bare question again with no memory of its own prior answer and asked to guess whether a typical test-taker would find the question easy or hard, using that difficulty guess as a stand-in for confidence in its own specific answer
CThe model is fine-tuned on a labeled dataset of past correct and incorrect answers so that a separate output head learns to predict correctness directly from the question text alone, without ever seeing the model's actual generated answer
DThe model is prompted to look at the question together with its own proposed answer and output the probability, called P(True), that the proposed answer is correct; for larger models, this self-reported P(True) is reasonably well-calibrated against actual accuracy across many multiple-choice and true/false questions when the format is set up appropriately
Correct answer: .
Kadavath et al. have the model self-evaluate by showing it the question together with its own previously proposed answer and prompting it to output a single probability, which they call P(True), that this specific proposed answer is correct. Across diverse multiple-choice and true/false question sets, they find that larger models' self-reported P(True) values are reasonably well-calibrated against actual accuracy, meaning that answers the model rates as, say, 80% likely to be true are actually correct roughly 80% of the time, once the prompting format is set up appropriately (including showing the model its own sampled answer rather than asking in the abstract). The option about restating the answer in different words and checking for lexical identity is wrong because that describes a paraphrase-consistency heuristic, not the P(True) probability-elicitation method the paper studies, and exact rewording identity is not how the paper measures confidence. The option about guessing question difficulty for a generic test-taker while ignoring the model's own specific answer is wrong because the whole point of the method is to evaluate the model's own particular proposed answer, not to produce a generic difficulty estimate detached from that answer. The option describing a fine-tuned classifier head that predicts correctness from the question alone, without seeing the generated answer, is wrong because it both requires additional supervised training, which the setup explicitly avoids, and evaluates only the question rather than the model's own specific answer, unlike P(True), which conditions on that answer directly.
Source: Kadavath et al., 'Language Models (Mostly) Know What They Know' (2022), arXiv:2207.05221
AI & LLM Engineering · Prompt Engineering · Card 065/081medium
A team deploys an LLM-based chatbot and wants a lightweight, training-free filter to flag inputs that may contain a Greedy Coordinate Gradient (GCG)-style adversarial suffix, the kind of algorithmically optimized nonsense-looking string appended to a harmful request to try to force the model to comply. Alon & Kamfonas (2023), 'Detecting Language Model Attacks with Perplexity,' propose using what signal for this, and what do they find about GCG suffixes compared to ordinary natural-language text?
AThe cosine similarity between the embedding of the full input and the embeddings of a library of known jailbreak prompts; GCG suffixes are found to always land within a fixed similarity threshold of at least one known jailbreak
BThe perplexity of the input text under a language model, i.e., how surprised the model is by the text on a token-by-token basis; GCG-optimized adversarial suffixes are found to have dramatically higher perplexity than ordinary natural-language text, because the optimization process searches for tokens that manipulate the target model's internals rather than tokens that read as fluent language
CThe total character length of the input string alone; GCG suffixes are found to always exceed a fixed length threshold that ordinary user messages never reach
DWhether the input contains any token from a fixed manually curated blocklist of profanity and violence-related words; GCG suffixes are found to always include at least one such blocklisted token
Correct answer: .
Alon and Kamfonas propose using perplexity, a measure of how surprised a language model is by a piece of text on a token-by-token basis, as a training-free detection signal. They find that adversarial suffixes produced by Greedy Coordinate Gradient search have dramatically higher perplexity than ordinary natural-language text, because GCG's optimization process searches token-by-token for whichever tokens most effectively push the target model's internal computations toward compliance, with no objective favoring fluent or natural-sounding language; the result reads as gibberish to a language model, which is precisely what a perplexity filter can catch, including via a sliding-window variant for suffixes embedded within otherwise-normal text. The option about embedding similarity to a library of known jailbreak prompts is wrong because that describes a semantic-matching defense against known, previously catalogued jailbreak templates, not the perplexity-based approach this paper proposes, and it would not generalize to newly optimized suffixes that resemble no prior known jailbreak. The option about raw character length is wrong because length alone is not the signal the paper studies or relies on; GCG suffixes are not defined or reliably identified by a fixed length threshold. The option about a manually curated profanity or violence blocklist is wrong because GCG-optimized suffixes are largely nonsensical token sequences with no reason to contain any specific blocklisted word, which is exactly why keyword-based filters are easy for such attacks to slip past, unlike a perplexity-based filter that catches the suffix's inherent unnaturalness regardless of its specific vocabulary.
Source: Alon & Kamfonas, 'Detecting Language Model Attacks with Perplexity' (2023), arXiv:2308.14132
AI & LLM Engineering · Prompt Engineering · Card 066/081easy
A team wants GPT-4 to produce short, information-dense summaries of news articles, but plain single-pass summarization prompts tend to produce summaries that are fluent but omit many of the article's salient entities (names, numbers, organizations) in favor of generic language, while summaries that try to cram in every entity at once tend to read as unreadable lists. Adams et al. (2023), 'From Sparse to Dense: GPT-4 Summarization with Chain of Density Prompting,' propose a prompting procedure to balance this trade-off. What does Chain of Density (CoD) prompting do?
AIt has the model generate five independent summaries of different fixed lengths in parallel, from very short to very long, and a human reviewer then manually picks whichever one seems most balanced
BIt retrieves the five most similar previously-published human-written summaries from a database and averages their wording token-by-token to produce a new summary
CIt has the model iteratively rewrite the same fixed-length summary across several rounds, identifying one to three additional salient, previously-missing entities from the article each round and fusing them into the existing summary without increasing its overall length, producing a progressively denser final summary
DIt asks the model to list every named entity in the article first, then simply concatenates that raw entity list in front of an unrelated, separately generated generic summary
Correct answer: .
Chain of Density prompting has the model rewrite the same summary across several iterative rounds while holding its overall length roughly fixed. In each round, the model is prompted to identify one to three additional salient entities from the article that are missing from the current draft and to fuse them into the existing summary, typically by making the prose more compact or fusing clauses, rather than by simply appending more text. Repeating this for a handful of rounds produces a final summary that is progressively denser in genuinely salient content while staying the same length as the very first draft, which the paper shows humans generally prefer over both the sparse first draft and over summaries that are so dense they become an unreadable list. The option describing five independently generated summaries of different lengths reviewed manually by a human is wrong because CoD's rounds are not independent parallel drafts of varying length; they are a single sequential rewriting process at constant length, requiring no human intervention to pick among them. The option describing retrieval and averaging of previously published human summaries is wrong because CoD involves no external database of prior summaries at all; every draft is generated and revised by the model itself from the source article. The option describing an entity list simply concatenated in front of an unrelated generic summary is wrong because it describes exactly the disjointed, low-quality output CoD is designed to avoid; CoD fuses newly identified entities into the existing summary's prose rather than prepending a separate, disconnected list.
Source: Adams et al., 'From Sparse to Dense: GPT-4 Summarization with Chain of Density Prompting' (2023), arXiv:2309.04269
AI & LLM Engineering · Prompt Engineering · Card 067/081easy
A team building a closed-book question-answering system (no external documents, search engine, or vector database available at inference time) wants to improve factual accuracy on questions that hinge on a specific passage the model was likely exposed to during pretraining, without adding any retrieval system. Sun et al. (2022), 'Recitation-Augmented Language Models,' propose RECITE, a two-step generation procedure for this. What does RECITE have the model do, and how does this differ from retrieval-augmented generation?
AThe model first samples one or more relevant passages purely from its own parameters, by generating text it 'recites' as if quoting a source it was trained on, and then conditions on that self-generated recitation to produce its final answer; unlike retrieval-augmented generation, no external corpus or search index is queried at any point, since the 'retrieved' passage comes entirely from the model's own memory
BThe model queries an external search engine to retrieve the single most relevant Wikipedia passage, then quotes that retrieved passage verbatim as its final answer without any further generation
CThe model is fine-tuned to memorize the entire pretraining corpus word-for-word, so that at inference time it can always retrieve a verbatim passage through gradient-free nearest-neighbor lookup over its own weights
DThe model generates several candidate answers independently and then recites, i.e., repeats, whichever candidate answer appears most frequently among the samples, without generating or conditioning on any intermediate passage
Correct answer: .
RECITE works in two generation steps performed entirely by the same model, with no external system involved. First, the model samples one or more passages it 'recites,' generating text as if quoting a source from its own pretraining, drawing purely on its parameters rather than looking anything up. Second, the model conditions on that self-generated recitation as additional context to produce its final answer. This differs from retrieval-augmented generation specifically in where the supporting passage comes from: retrieval-augmented systems query an external corpus or search index to fetch a real passage, whereas RECITE's 'retrieved' passage is itself generated by the model from its own parametric memory, so no external corpus or index is ever queried. The option describing querying an external search engine and quoting a retrieved Wikipedia passage verbatim as the final answer is wrong because that is standard retrieval-augmented generation, the exact external-lookup approach RECITE is designed to avoid, and RECITE also still generates a separate final answer rather than outputting the passage itself as the answer. The option describing fine-tuning the model to memorize the corpus for gradient-free nearest-neighbor lookup over its weights is wrong because RECITE involves no additional fine-tuning step at all; it uses the model's existing pretrained knowledge through ordinary prompted generation. The option describing majority-vote repetition of the most common candidate final answer is wrong because that describes a self-consistency-style voting procedure over final answers directly, with no intermediate recited passage generated or conditioned on, which is a different technique from RECITE's explicit two-step recite-then-answer structure.
Source: Sun et al., 'Recitation-Augmented Language Models' (2022), arXiv:2210.01296
AI & LLM Engineering · Prompt Engineering · Card 068/081medium
A team has a large pool of thousands of unlabeled questions for a new reasoning task and a limited budget for human annotators to write chain-of-thought exemplars, so they can only afford to hand-annotate a small handful of questions as few-shot demonstrations. Rather than picking that handful at random, Diao et al. (2023), 'Active Prompting with Chain-of-Thought for Large Language Models,' propose a selection procedure. What do they do to decide which questions from the pool are most worth sending to a human annotator?
AThey select whichever questions are shortest in token count, reasoning that short questions are cheapest for human annotators to work through and annotate quickly
BThey randomly sample an equal number of questions from every topic category in the pool, ensuring balanced topical coverage regardless of how difficult any individual question is for the model
CThey fine-tune a separate small classifier model to predict each question's ground-truth answer directly, then send only the questions the classifier gets wrong to human annotators
DFor each unlabeled question, they have the target LLM generate several chain-of-thought answers via repeated sampling and compute an uncertainty metric (such as disagreement among the sampled answers) for that question; questions where the model's several sampled answers disagree most are judged most uncertain and are prioritized for human annotation, on the reasoning that resolving the model's biggest uncertainties yields the most useful new exemplars
Correct answer: .
Active Prompting has the target LLM sample several chain-of-thought answers for each unlabeled question in the pool and then computes an uncertainty metric from that set of sampled answers, such as how much they disagree with one another (the paper also studies related metrics like entropy and variance over the sampled answers). Questions where the model's own sampled answers disagree the most are judged the most uncertain, and these are the questions prioritized for human annotation, on the reasoning that annotating exemplars for the questions the model already finds easy contributes little, while annotating the questions where the model is most uncertain yields exemplars that most improve downstream performance. The option selecting the shortest questions by token count is wrong because token length has no established relationship to how uncertain or difficult a question is for the model, and the paper's selection criterion is based on measured model disagreement, not surface length. The option selecting an equal number of questions from every topic category is wrong because it ignores the model's actual per-question uncertainty entirely in favor of a fixed topical quota, which is the opposite of what Active Prompting measures and prioritizes. The option describing a separately fine-tuned classifier that predicts ground-truth answers directly is wrong both because no such separate classifier is trained in this method, and because the pool is explicitly unlabeled, so there is no ground truth available to check a classifier's predictions against in the first place; the uncertainty signal instead comes entirely from the target LLM's own sampled chain-of-thought outputs.
Source: Diao et al., 'Active Prompting with Chain-of-Thought for Large Language Models' (2023), arXiv:2302.12246
AI & LLM Engineering · Prompt Engineering · Card 069/081easy
A team has only a handful of input-output example pairs demonstrating some task (for instance, pairs of a word and its plural form) and, instead of writing a few-shot prompt with those examples for a model to imitate, wants the model to explicitly state, in a single natural-language sentence, what the underlying task instruction actually is. Honovich et al. (2022), 'Instruction Induction: From Few Examples to Natural Language Task Descriptions,' study prompting language models to do this directly. What do they have the model do, and how do they evaluate whether the induced instruction is any good?
AThey have the model produce more input-output pairs in the same style as the given examples, and judge the induction successful only if a human annotator cannot tell the generated pairs apart from the original ones
BThey fine-tune the model on a large labeled corpus of (examples, correct-instruction) pairs so it learns to map any set of examples directly onto a matching instruction from that fixed labeled set
CThey prompt the model with only the small set of input-output example pairs and ask it to generate, in natural language, the instruction that would produce those outputs from those inputs; they then evaluate a generated instruction by executing it, giving that instruction alone (with no examples) to a model and checking whether it correctly reproduces the original outputs on new inputs
DThey ask the model to classify the task into one of a small fixed list of predefined categories (such as 'translation' or 'sentiment') rather than producing any free-form natural-language description of the task
Correct answer: .
Honovich et al. prompt the model with nothing but the small set of input-output example pairs and ask it to generate, in ordinary natural language, the instruction that would explain how those outputs were produced from those inputs, a task they call instruction induction. To evaluate whether a generated instruction is actually good, rather than just plausible-sounding, they use an execution-based metric: they take the generated instruction alone, with no examples attached, give it to a model on new inputs, and check whether the model, following only that instruction, reproduces the correct outputs. An instruction only counts as good if executing it actually recovers the intended task behavior. The option about generating more input-output pairs and judging success by whether a human can tell them apart from the originals is wrong because the goal here is inducing an explicit natural-language instruction, not generating additional example pairs, and human indistinguishability is not the paper's evaluation metric. The option about fine-tuning on a large labeled corpus of examples-to-instruction pairs is wrong because the setup is a few-shot prompting task with no such training corpus or fine-tuning step; the model must induce the instruction purely from the handful of given examples at inference time. The option about classifying the task into a small fixed list of predefined categories is wrong because it restricts induction to a closed label set, whereas the paper's task is explicitly open-ended: the model must produce a free-form natural-language instruction, which the execution-based metric then verifies works.
Source: Honovich et al., 'Instruction Induction: From Few Examples to Natural Language Task Descriptions' (ACL 2023), arXiv:2205.10782
AI & LLM Engineering · Prompt Engineering · Card 070/081easy
A team wants an LLM's output to always be syntactically valid according to a specific formal grammar (for example, always a well-formed date in `YYYY-MM-DD` format, or always valid according to a custom domain-specific mini-language), with a hard guarantee rather than just a strong tendency, and without retraining the model or relying on the model to self-correct after the fact. Willard & Louf (2023), 'Efficient Guided Generation for Large Language Models,' propose a technique (implemented in the open-source Outlines library) for this. What does their approach do at each decoding step?
AIt generates the full response first with no restriction, then runs a separate regular-expression validator afterward and asks the model to regenerate from scratch if the output fails validation, repeating until a valid output happens to appear
BIt appends a natural-language instruction to the prompt describing the required grammar in words (such as 'please only output a valid date') and relies on the model's instruction-following ability alone to comply, with no mechanism enforcing the constraint at the token level
CIt fine-tunes the model's weights on a large synthetic dataset of only grammar-valid outputs until the model has implicitly learned to always emit valid text on its own
DIt builds an index, from a regular expression or context-free grammar, over the model's vocabulary that tracks which tokens are valid continuations at each position of the grammar (modeled as a finite-state machine or pushdown automaton); at every decoding step, this index is used to mask out the logits of any token that would make the output invalid according to the grammar, so only grammar-consistent tokens can ever be sampled, guaranteeing valid output structurally rather than by hope or after-the-fact correction
Correct answer: .
Willard and Louf's approach compiles a regular expression or context-free grammar into an index over the model's vocabulary, modeling valid strings as transitions of a finite-state machine (or, for context-free grammars, a pushdown automaton), which tracks exactly which tokens are legal continuations at each position in the grammar. At every decoding step, this index is consulted to mask out the logits of any vocabulary token that would make the generated string invalid according to the grammar, so sampling only ever chooses among grammar-consistent tokens. Because invalid tokens are structurally excluded before sampling rather than merely discouraged, the output is guaranteed to satisfy the grammar, not just likely to, and this holds without retraining the model or needing any generate-then-fix loop. The option describing generating unrestricted text first and regenerating from scratch on validation failure is wrong because that offers only a probabilistic tendency toward validity with no guarantee of eventual success or bounded cost, unlike token-level masking, which enforces validity at every single step. The option describing a natural-language instruction with no token-level enforcement is wrong because it relies entirely on the model's willingness to comply, which the paper explicitly moves beyond by enforcing the constraint mechanically at the logit level regardless of what the model would have otherwise generated. The option describing fine-tuning on a synthetic dataset of valid outputs is wrong because Willard and Louf's method requires no additional training of any kind; the vocabulary index can be applied to a frozen, already-trained model, and even a fine-tuned model given this option would still lack a hard guarantee, since fine-tuning only shifts a distribution rather than mechanically excluding invalid tokens.
Source: Willard & Louf, 'Efficient Guided Generation for Large Language Models' (2023), arXiv:2307.09702
AI & LLM Engineering · Prompt Engineering · Card 071/081hard
A team applies self-consistency (Wang et al., 2022) to a code-generation task by sampling several chain-of-thought solutions and picking the final answer that appears identically most often among the samples, and it works well. When they try the same majority-vote approach on an open-ended long-form summarization task, they find it breaks down, because no two of the several sampled full-length summaries are ever exactly identical, so a literal majority vote never finds any answer that 'wins.' Chen et al. (2023), 'Universal Self-Consistency for Large Language Model Generation,' propose Universal Self-Consistency (USC) to extend the technique to exactly this kind of task. What does USC do differently from the original self-consistency method's majority vote?
AIt shortens every sampled summary down to a single keyword before comparing them, so that exact-match voting becomes possible again on the reduced representations
BIt trains a separate learned reward model on human preference labels over pairs of summaries, and uses that reward model's score alone to pick the best sampled summary, without any involvement from the original LLM at this stage
CInstead of relying on exact-match voting over the final answers, it feeds all of the sampled candidate outputs back into the same LLM in a single prompt and asks that LLM itself to read across the candidates and select (or synthesize) whichever one is most consistent with the majority of the others, which works even when no two full free-form outputs are ever character-for-character identical
DIt discards self-consistency's sampling step entirely and instead asks the model to produce just one single output per problem, deterministically, using greedy decoding
Correct answer: .
Universal Self-Consistency keeps the original method's idea of sampling several candidate outputs, but replaces exact-match majority voting over final answers with an LLM-based selection step: all of the sampled candidates are placed together into a single prompt, and the same LLM is asked to read across them and pick out, or synthesize, whichever candidate is most consistent with the majority of the others. Because this comparison is made by the LLM reading the full content of each candidate rather than by checking for character-for-character identity, it works on free-form, long-form outputs such as summaries, open-ended answers, or code, where no two full outputs are ever likely to be exactly identical even when they agree in substance. The option describing shortening every summary to a single keyword before voting is wrong because that discards essentially all of the summary's content before comparison, and USC's actual mechanism instead has the LLM directly read and compare the full candidates rather than reducing them to a lossy representation. The option describing a separately trained reward model scoring summaries without LLM involvement is wrong because USC explicitly leverages the LLM itself, not any separate trained model, to perform the consistency judgment; no additional reward model is trained. The option describing dropping sampling entirely in favor of one deterministic greedy output is wrong because it eliminates self-consistency's core premise of sampling multiple candidates and comparing them; USC still samples multiple candidates exactly as before, and only changes how the best one is selected from among them.
Source: Chen et al., 'Universal Self-Consistency for Large Language Model Generation' (2023), arXiv:2311.17311
AI & LLM Engineering · Prompt Engineering · Card 072/081easy
A team already follows Anthropic's advice to give Claude a role by writing a short, fixed description such as 'You are a seasoned data scientist' into the system prompt, and reuses that same sentence for every request the application handles. They want to go further: automatically tailoring a distinct, detailed expert identity for each new instruction the system receives, rather than reusing one static description every time. Xu et al. (2023), 'ExpertPrompting: Instructing Large Language Models to be Distinguished Experts,' propose a method for exactly this. How does ExpertPrompting generate the expert identity used for a given instruction, and how does that differ from simply assigning the same fixed role on every call?
AExpertPrompting hand-writes a fixed library of a few dozen expert personas in advance (for example, 'senior software engineer,' 'tax attorney') and has the model pick the single closest-matching persona from that pre-written library for each new instruction, the same static-description approach as ordinary role prompting, just drawing from a longer list
BExpertPrompting uses in-context learning to have the model itself write a new, detailed description of a specific expert identity customized to the particular instruction at hand, including a plausible background for that expert, and then conditions its actual answer on that freshly generated description, so the persona is synthesized per instruction rather than reused verbatim from one fixed sentence
CExpertPrompting fine-tunes the base model's weights on a labeled dataset of expert personas so the model permanently role-plays as one single designated expert in every future conversation, with no further prompt engineering needed at inference time
DExpertPrompting removes any description of a role or identity from the prompt entirely and instead infers an appropriate expert persona purely from statistical patterns in the user's own writing style, with no expert-identity text ever appearing in the prompt itself
Correct answer: .
ExpertPrompting uses in-context learning (few-shot exemplars of instruction-to-expert-description pairs) to have the model automatically synthesize a detailed, customized description of a distinguished expert's identity and background for each specific instruction it receives, then has the model produce its actual answer conditioned on that freshly generated expert background, rather than reusing one hand-written role sentence for every request. The paper reports that GPT-4-based evaluation judged these expert-conditioned answers to be of significantly higher quality than vanilla answers produced without any persona at all, and that an open-source model instruction-tuned on this expert-augmented data (ExpertLLaMA) reached roughly 96% of the original ChatGPT's capability on their benchmark. The option describing a pre-written library of a few dozen fixed personas is wrong because that is just a larger version of ordinary static role prompting (Anthropic's documented 'give Claude a role' technique already tested elsewhere in this topic) and does not involve generating a new, instruction-specific description at all. The option describing fine-tuning the base model's weights to permanently adopt one persona is wrong because ExpertPrompting is a prompting-time technique applied per instruction, not a weight-update procedure, and it explicitly produces a different expert for different instructions rather than one fixed persona for all future conversations. The option describing inferring a persona purely from the user's writing style with no expert-identity text in the prompt at all is wrong because the method's entire mechanism is generating and then explicitly including a written expert-identity description in the prompt before answering, not omitting one.
Source: Xu, Yang, Lin, Wang, Zhou, Zhang & Mao, 'ExpertPrompting: Instructing Large Language Models to be Distinguished Experts' (arXiv:2305.14688, 2023)
AI & LLM Engineering · Prompt Engineering · Card 073/081easy
A team's LLM-based assistant reads an externally sourced document, such as fetched web content, as part of its context, and the team is worried that if the document contains attacker-planted text phrased as an instruction, the model might carry it out as though the user had asked for it directly. Separately from the tool-permission and human-approval controls already covered by OWASP's general defense-in-depth guidance, Hines et al. (2024), 'Defending Against Indirect Prompt Injection Attacks With Spotlighting,' propose a prompting-level technique aimed specifically at this. What does spotlighting do to the untrusted document text itself, and why does this make the model less likely to treat it as an instruction?
ASpotlighting scans the document and deletes any sentence phrased as an imperative instruction before the document ever reaches the model's context window, so the model never sees any instruction-like phrasing at all
BSpotlighting routes the document to a separate, smaller classifier model trained specifically to detect injection attempts, and only forwards the document into the main model's context if that classifier scores it as safe
CSpotlighting translates the document into a different natural language from the rest of the prompt, relying on the main model performing worse in that language to prevent it from successfully following any instruction embedded in the translated text
DSpotlighting visibly transforms the untrusted text itself, for example by marking it with distinctive delimiters, interleaving it with marker characters, or encoding it (such as in base64), paired with a system instruction telling the model to treat anything appearing in that transformed form only as data to read rather than as instructions to follow, which makes injected imperative-sounding text stand out as untrusted instead of blending in with the model's genuine instructions
Correct answer: .
Spotlighting works by visibly transforming the untrusted content before it enters the model's context, using one of several variants such as marking the text with distinctive delimiter tokens, interleaving it with marker characters, or encoding it (the paper finds base64-style encoding the strongest of the variants tested), and pairing that transformed text with an explicit system instruction telling the model that anything appearing in that transformed form is data to be read and reasoned about, never a command to be obeyed. This gives injected instruction-like phrasing a visibly different 'shape' from the model's genuine instructions, rather than letting it blend in verbatim. The technique is used in production as part of Microsoft's Prompt Shields in Azure AI Foundry, though the authors note it does not change which specific attacks succeed so much as reduce overall attack rates, and that it works largely by disrupting how injection triggers tokenize, which means it may not generalize as well to languages whose tokenization differs substantially from the languages it was evaluated on. This is a different, lower-level mechanism from the general OWASP defense-in-depth list (layered tool-permission and human-approval controls) and from simply classifying an attack as direct or indirect, both already covered elsewhere for this exam. The option describing deleting imperative sentences outright is wrong because spotlighting transforms and marks the text rather than removing content from it. The option describing a separate classifier model gating what gets forwarded is wrong because spotlighting does not filter content with an external classifier; it changes how the untrusted text itself is presented to the one model doing the reasoning. The option describing translation into a different language to rely on weaker performance there is wrong because spotlighting's documented variants are marking, interleaving, and encoding of the same content, not machine translation, and the paper does not rely on degraded cross-lingual performance as its mechanism.
Source: Hines, Lopez, Hall, Zarfati, Zunger & Kiciman (Microsoft), 'Defending Against Indirect Prompt Injection Attacks With Spotlighting' (arXiv:2403.14720, 2024)
AI & LLM Engineering · Prompt Engineering · Card 074/081medium
A team is building a few-shot chain-of-thought prompt for a new task and has a very large pool of unlabeled questions but only a small, fixed budget for human annotators to write out full worked exemplars by hand. They must decide, in advance and before any labels exist, which specific unlabeled questions are even worth sending to an annotator — a different problem from ranking already-labeled examples by similarity to one particular test input, or from ranking the model's own generated answers by uncertainty. Su et al. (2022), 'Selective Annotation Makes Language Models Better Few-Shot Learners,' propose a method called vote-k for exactly this unlabeled-pool selection step. How does vote-k decide which unlabeled examples get sent for annotation?
AVote-k builds a graph over the unlabeled pool based on embedding-similarity neighborhoods, scores each candidate by how well-connected it is to many other examples in the pool, and discounts candidates that sit too close to examples already selected, so the resulting annotated set is both broadly representative of the whole pool and spread out rather than clustered in one region of it
BVote-k sends every single example in the unlabeled pool to a human annotator up front, then has the language model vote on which of the resulting fully labeled examples to keep in the final few-shot prompt, discarding whichever receive the fewest votes
CVote-k ranks unlabeled examples purely by their embedding similarity to the one specific test input currently being answered, selecting only the handful closest to that particular input for annotation every time a new test input is processed
DVote-k measures the language model's output entropy across several answers it generates for each unlabeled example, without ever examining the pool's embedding structure, and sends only the highest-entropy examples off for annotation
Correct answer: .
Vote-k operates purely over the unlabeled pool, before any test input is seen and before any labels exist: it builds a graph connecting examples based on embedding-similarity neighborhoods, scores each unlabeled candidate by how central or well-connected it is within that graph (a rough proxy for how representative it is of the pool as a whole), and applies a discount to candidates that are too similar to examples already chosen for annotation, so the selection process actively favors spreading across different regions of the pool instead of clustering around whichever region scored highest first. The paper reports this selective-annotation step, combined with retrieving from the resulting annotated pool at test time, achieves performance close to much more expensive fully supervised fine-tuning while using far fewer annotated examples, and clearly outperforms selecting the same number of examples at random. This is a distinct problem from, and a different method than, two other exemplar-selection techniques already covered elsewhere in this topic: choosing which already-labeled examples to retrieve for one specific test input by similarity (a test-time retrieval problem assuming labels already exist), and choosing which unlabeled examples are most worth sending to a human annotator based on the model's own uncertainty over candidate answers it generates for them. The option describing sending the entire pool to annotators and then voting on which finished labels to keep is wrong because vote-k's entire purpose is reducing annotation cost by selecting a small subset before annotation happens, not annotating everything and filtering afterward. The option describing ranking by similarity to one specific test input is wrong because that describes a test-time retrieval approach operating over already-labeled data, not vote-k's pre-annotation selection from an unlabeled pool. The option describing ranking by the model's own output entropy on generated answers is wrong because that describes an uncertainty-based active-selection approach, not vote-k's graph-based diversity-and-representativeness method, which never examines model-generated answers at all.
Source: Su, Kasai, Wu, Shi, Wang, Xin, Zhang, Ostendorf, Zettlemoyer, Smith & Yu, 'Selective Annotation Makes Language Models Better Few-Shot Learners' (arXiv:2209.01975, ICLR 2023)
AI & LLM Engineering · Prompt Engineering · Card 075/081medium
A red-teaming exercise finds that a single, blunt request for clearly harmful content is refused outright, and so is a many-shot jailbreak attempt that stuffs hundreds of fabricated compliant-dialogue turns into one long prompt submitted all at once. But the team finds a different attack succeeds: across several separate, individually mild-looking conversation turns, each one explicitly referencing and building on the model's own prior reply, the conversation gradually arrives at the same harmful output that neither the one blunt request nor the single long fabricated-dialogue prompt could get. Russinovich, Salem & Eldan (2024), 'Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack,' describe this technique. How does Crescendo structure its attack across turns, and why does gradual escalation succeed where a single obviously harmful request does not?
ACrescendo submits one very long prompt, structurally identical to many-shot jailbreaking, except that the fabricated turns depict the assistant repeatedly refusing rather than complying, which the authors report paradoxically increases compliance on the model's real final turn
BCrescendo relies entirely on an algorithmically optimized, nonsense-looking suffix appended to a single harmful request, similar to the Greedy Coordinate Gradient method, rather than using multiple separate conversational turns at all
CCrescendo starts with a general, benign question related to the target topic and, across multiple separate turns, progressively escalates the request while explicitly referencing the model's own prior replies, adapting or substituting later turns depending on whether the model complied or refused at each step, so that by the final turn the model is nudged into continuing a trajectory it has effectively already been following rather than ever confronting one single, obviously harmful ask
DCrescendo has hundreds of different users each submit one small, innocuous-looking fragment of the harmful request independently in separate, unrelated conversations, then reassembles their separate outputs outside the model entirely, so no single conversation with the model ever contains more than one fragment
Correct answer: .
Crescendo begins with an innocuous, general question or framing related to the eventual target topic and then, over several distinct conversational turns, incrementally escalates the request while explicitly referencing and building on what the model said in its own prior replies, with the attacker adapting or substituting later turns based on whether the model complied or pushed back at a given step. By the time the final turn arrives, the model is being asked to continue a trajectory of its own prior outputs rather than being confronted with one single, obviously harmful request it could cleanly refuse, which the authors report lets Crescendo surpass other state-of-the-art jailbreak techniques, achieving notably higher attack success rates on models including GPT-4 and Gemini-Pro than the baselines they compared against, and that the technique also transfers to multimodal models. This is a distinct mechanism from two other jailbreak techniques already covered elsewhere in this topic: many-shot jailbreaking, which submits a single very long prompt packed with hundreds of fabricated dialogue turns all at once rather than engaging across genuinely separate turns, and Greedy Coordinate Gradient, which searches for an optimized adversarial token suffix rather than structuring a multi-turn conversation at all. The option describing one long prompt of fabricated refusals is wrong because Crescendo's turns are genuinely separate conversational exchanges with the actual model, not a single prompt simulating a fake dialogue history. The option describing an optimized nonsense suffix is wrong because that describes Greedy Coordinate Gradient's single-request mechanism, not Crescendo's multi-turn escalation. The option describing fragmenting the request across many different users' unrelated conversations is wrong because Crescendo is carried out by one attacker across one evolving conversation with continuity between turns, not fragments scattered across unrelated sessions that are reassembled outside the model.
Source: Russinovich, Salem & Eldan (Microsoft), 'Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack' (arXiv:2404.01833, 2024; USENIX Security 2025)
AI & LLM Engineering · Prompt Engineering · Card 076/081hard
A team already uses Gao et al.'s PAL (Program-Aided Language Models) approach: the model writes a complete Python program for a word problem, and the entire program is handed to an external interpreter to execute, which reliably fixes arithmetic slips but only works when the whole problem can be expressed as code the interpreter can actually run end to end. They now hit tasks with sub-steps that have no well-defined executable semantics, such as a step inside an otherwise code-like procedure that says to check whether a sentence sounds sarcastic, where handing the full program to an external interpreter alone breaks down because the interpreter has no way to execute an undefined operation like that. Li et al. (2023), 'Chain of Code: Reasoning with a Language Model-Augmented Code Emulator,' propose an approach for exactly this. What does Chain of Code have the system do differently from handing a complete program to an external interpreter alone?
AChain of Code abandons writing any code the moment it detects a step without well-defined executable semantics, falling back to standard natural-language chain-of-thought prompting for the entire problem instead of writing a program at all
BChain of Code has the model write flexible pseudocode that freely mixes genuinely executable code with semantic sub-steps left loosely defined in natural language, then runs this through an interpreter that executes whichever parts it can and, whenever it reaches a sub-step it cannot resolve, hands that specific step to the language model itself to simulate the expected output as an 'LMulator,' before the interpreter resumes executing the rest of the program using that simulated result
CChain of Code trains a dedicated external interpreter via supervised fine-tuning to recognize and directly execute informal natural-language instructions as if they were valid code, removing the language model from the execution loop entirely once that training is complete
DChain of Code requires every sub-step in the program to be rewritten as strictly valid, fully executable code before any execution begins, rejecting any pseudocode step with undefined behavior rather than letting the interpreter hand such a step to the model
Correct answer: .
Chain of Code broadens code-augmented reasoning by having the model write pseudocode that freely interleaves genuinely executable code with sub-steps left semantically loose and written in natural language wherever the task calls for an operation that has no well-defined executable form, such as judging tone or sentiment. This pseudocode is then run through an interpreter that executes whatever parts of it are valid code as normal, and whenever execution reaches a sub-step it cannot resolve, it hands that specific step off to the language model, acting as an 'LMulator,' to simulate the expected output for that step, after which the interpreter resumes executing the remainder of the program using the simulated result exactly as if it had been computed normally. This lets the overall approach broaden the scope of questions a code-driven method can correctly answer well beyond PAL, which only works when the entire problem can be expressed as a complete program an external interpreter can run from start to finish with no undefined steps. The option describing falling back entirely to natural-language chain-of-thought the moment an unresolvable step appears is wrong because Chain of Code's whole point is continuing to use code structure around the unresolvable step, handing only that specific step to the model rather than abandoning code for the entire problem. The option describing fine-tuning a dedicated external interpreter to execute natural language directly is wrong because Chain of Code uses the existing language model itself, with no additional training, to simulate the unresolvable steps, not a separately trained execution engine. The option describing rejecting any step with undefined behavior and requiring fully valid code throughout is wrong because that describes the rigid, PAL-like constraint Chain of Code is specifically designed to relax by allowing the interpreter to hand such steps to the model instead of rejecting them.
Source: Li, Liang, Zeng, Chen, Hausman, Sadigh, Levine, Fei-Fei, Xia & Ichter, 'Chain of Code: Reasoning with a Language Model-Augmented Code Emulator' (arXiv:2312.04474, 2023)
AI & LLM Engineering · Prompt Engineering · Card 077/081easy
A team wants GPT-4V to answer questions that require pointing to one specific, precise region of an image (for example, 'what color is the object in the upper-left corner, not the one in the center') rather than just describing the image as a whole, and they want this to work in a zero-shot setting with no fine-tuning of the model. Yang et al. (2023), 'Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V,' propose a prompting method for this. What does Set-of-Mark (SoM) prompting do to the image before it is given to the model?
AIt converts the image into a text caption produced by a separate captioning model and sends only that caption text to GPT-4V, since the paper's method assumes the model cannot process raw pixel regions directly
BIt crops the image into a fixed grid of equally sized tiles and numbers each tile in reading order, regardless of what objects happen to fall inside any given tile
CIt uses off-the-shelf interactive segmentation models to partition the image into regions at different granularities, then overlays visual marks, such as numbers, letters, masks, or boxes, on those regions, so the model can refer to a specific object by its mark instead of a vague spatial description
DIt fine-tunes a lightweight adapter on top of GPT-4V's vision encoder using labeled bounding-box data so the model learns to ground references internally, without changing how the image itself is presented to it
Correct answer: .
Set-of-Mark (SoM) prompting is a purely inference-time, prompting-level technique: it first runs an off-the-shelf interactive segmentation model (such as SAM or similar panoptic/instance segmenters) over the input image to partition it into regions at multiple granularities, then overlays a distinct visual mark, such as a number, letter, mask outline, or bounding box, on each region before the marked-up image is sent to GPT-4V alongside the text prompt. Because every region now carries its own visible, unambiguous label, the model can be asked a question that refers to 'the object marked 7' or similar instead of relying on vague spatial language like 'the thing in the upper-left,' and the paper reports that GPT-4V with SoM in a zero-shot setting outperforms specialized, fully fine-tuned referring-expression-comprehension and segmentation models on benchmarks such as RefCOCOg. The option describing captioning the image and discarding the pixels is wrong because SoM still sends the image itself, with marks overlaid on it, rather than replacing it with text. The option describing a fixed, content-blind tiling grid is wrong because SoM's regions come from segmentation models that respect actual object boundaries, not arbitrary equal-sized tiles. The option describing fine-tuning an adapter on the vision encoder is wrong because SoM requires no training or weight changes at all; it is applied entirely by preprocessing the image before it ever reaches the frozen model.
Source: Yang, Zhang, Li, Zou, Li & Gao, 'Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V' (arXiv:2310.11441, 2023)
AI & LLM Engineering · Prompt Engineering · Card 078/081medium
A team is building a multi-step reasoning system that needs to pause partway through a problem to call an external tool, such as a calculator or a search API, and then continue reasoning using the tool's output, similar to what Yao et al.'s ReAct framework enables. Rather than hand-writing a new set of task-specific, interleaved reasoning-and-tool-use demonstrations for every new task the way ReAct's prompt author must, the team wants suitable demonstrations selected automatically based on tasks solved before. Paranjape et al. (2023), 'ART: Automatic Reasoning and Tool-use,' propose a framework for this. According to the paper, how does ART construct and run its prompt for a new task, and what specifically does it automate that ReAct leaves to a human?
AGiven a new task, ART retrieves demonstrations of similar multi-step reasoning and tool use from a task library rather than requiring a human to hand-craft them, generates the reasoning-and-tool-call program for the new input, pauses generation whenever a tool is called, and resumes once the tool's output has been inserted
BART eliminates the need for any demonstrations whatsoever, generating tool calls purely by looking up the task's name in a fixed table, whereas ReAct still requires the model to interleave reasoning and actions inside its own generated text
CART fine-tunes a separate, small classifier model to predict which tool should be called at each reasoning step, replacing ReAct's prompting-only interleaving with a trained decision model
DART requires an engineer to manually verify and hand-edit every retrieved demonstration before each run, trading ReAct's fully human-written demonstrations for human-reviewed ones instead of removing human involvement
Correct answer: .
ART (Automatic Reasoning and Tool-use) targets exactly the manual-authoring burden that ReAct still carries: prior work on chain-of-thought and tool-use prompting, including ReAct, typically requires hand-crafting task-specific demonstrations and carefully scripting how model generation interleaves with tool calls for every new task. ART instead maintains a library of demonstrations of multi-step reasoning and tool use drawn from related tasks solved previously, and for a new task it automatically retrieves suitable demonstrations from that library rather than having a human write new ones. Using a frozen LLM, it then generates the reasoning-and-tool-call program for the new input, seamlessly pausing generation whenever the program calls an external tool, and resumes generation once that tool's output has been integrated back into the context. The option describing eliminating demonstrations entirely via a name-based lookup table is wrong because ART still relies on retrieved few-shot demonstrations of reasoning and tool use; it automates their selection, not their existence. The option describing a separately fine-tuned classifier choosing tools is wrong because ART uses the same frozen LLM to generate the full reasoning-and-tool-call program itself, with no additional trained model in the loop. The option describing mandatory human review of every retrieved demonstration before each run is wrong because the paper's point is removing that per-task human involvement, not merely shifting it from writing demonstrations to reviewing them.
Source: Paranjape, Lundberg, Singh, Hajishirzi, Zettlemoyer & Ribeiro, 'ART: Automatic Multi-step Reasoning and Tool-use for Large Language Models' (arXiv:2303.09014, 2023)
AI & LLM Engineering · Prompt Engineering · Card 079/081medium
A team's task is complex enough that Zhou et al.'s Least-to-Most prompting, which lists a sequence of simpler subquestions and answers them in order within one running prompt, still struggles on some of the sub-steps, particularly ones that are themselves hard to answer reliably in a single pass or that need to recurse on a much smaller version of the same overall problem. Khot et al. (2022), 'Decomposed Prompting: A Modular Approach for Solving Complex Tasks,' propose DECOMP for this. How does DECOMP structure its solution, and how does this differ from Least-to-Most's approach?
ADECOMP requires every subquestion to be generated and answered inside the same single prompt used by Least-to-Most, differing only in the order in which the subquestions are listed
BDECOMP generates every subquestion and its answer in one uninterrupted pass with no ability to call back into the overall decomposition procedure, whereas Least-to-Most allows the model to revisit and revise earlier subquestions
CDECOMP replaces prompting entirely with task-specific fine-tuning, training a separate specialized model for each sub-task instead of prompting a shared language model for every step, unlike Least-to-Most's purely prompting-based approach
DDECOMP delegates each sub-task to its own dedicated prompting-based handler drawn from a library of such handlers, rather than solving every subquestion inside one shared prompt the way Least-to-Most does, and a sub-task whose difficulty comes from a large input can recursively decompose further by invoking the same DECOMP procedure again on a smaller version of itself
Correct answer: .
Decomposed Prompting (DECOMP) addresses complex tasks by breaking them, via prompting, into simpler sub-tasks that are each delegated to a separate prompting-based handler drawn from a library of such handlers, rather than requiring one shared prompt to generate and answer every subquestion itself the way Least-to-Most does. This modularity means a sub-task can be handled by whichever specialized handler is best suited to it, and, critically, when a sub-task's difficulty comes from having a very large input rather than from needing a fundamentally different skill, that sub-task can recursively invoke the same DECOMP procedure again on a smaller version of the same problem, something a single flat prompt answering a fixed list of subquestions in sequence has no mechanism to do. The option describing DECOMP as using the same single prompt as Least-to-Most, differing only in subquestion order, is wrong because it ignores DECOMP's defining feature of delegating to separate modular handlers rather than keeping everything inside one prompt. The option claiming DECOMP cannot call back into its own procedure while Least-to-Most can revise earlier subquestions is wrong because it reverses the actual capabilities: DECOMP's recursive, modular structure is exactly what enables calling back into the procedure on smaller inputs, which a single linear Least-to-Most prompt does not support. The option describing DECOMP as replacing prompting with task-specific fine-tuning is wrong because DECOMP's handlers are themselves prompting-based language-model calls, not separately trained models, keeping the method fully in the prompting paradigm.
Source: Khot, Trivedi, Finlayson, Fu, Richardson, Clark & Sabharwal, 'Decomposed Prompting: A Modular Approach for Solving Complex Tasks' (arXiv:2210.02406, ICLR 2023)
AI & LLM Engineering · Prompt Engineering · Card 080/081hard
A team wants to automatically improve a task prompt using only a training set of labeled examples and an LLM API, without any gradient-based training of the model's weights. They are already aware of Yang et al.'s OPRO, which treats an LLM as an optimizer weighing a running trajectory of past prompt-and-score pairs to propose better prompts, and of Fernando et al.'s Promptbreeder, which evolves a population of prompts through LLM-driven mutation and crossover operators scored against a training set. Pryzant et al. (2023), 'Automatic Prompt Optimization with “Gradient Descent” and Beam Search,' propose ProTeGi, a different mechanism for the same kind of problem. What does ProTeGi do?
AProTeGi numerically differentiates the prompt's token embeddings with respect to a loss computed on the training set, taking literal gradient-descent steps in embedding space before decoding the result back into text
BProTeGi runs the current prompt over a minibatch of training examples, has the LLM generate natural-language 'textual gradients' that criticize the specific ways the prompt's outputs failed, edits the prompt in the semantic direction opposite that criticism to produce candidate rewrites, and selects among the growing set of candidates using beam search combined with a bandit-style selection procedure
CProTeGi asks the LLM to propose a single best replacement prompt directly from the training examples in one generation step, with no iterative criticism-and-edit loop and no search over multiple candidates
DProTeGi mutates and recombines pairs of candidate prompts drawn from an evolving population using crossover and mutation operators scored against a fitness function, the same mechanism Promptbreeder uses
Correct answer: .
ProTeGi is inspired by numerical gradient descent but operates entirely in natural language rather than on model weights or embeddings: it evaluates the current prompt on a minibatch of training examples, then has the LLM itself produce a 'textual gradient,' a natural-language critique describing specifically how and why the prompt's outputs failed on that minibatch, much as a numerical gradient points toward the direction of greatest error. That textual gradient is then 'propagated' back into the prompt by having the LLM edit the prompt in the semantic direction opposite the criticism, producing one or more candidate rewrites; because this process can generate many candidates across iterations, ProTeGi uses beam search together with a bandit-style selection procedure to decide which candidates are worth expanding further and which to discard. The option describing literal numerical differentiation of token embeddings is wrong because ProTeGi never touches embeddings or computes an actual mathematical gradient; its 'gradient' is a natural-language critique used only as an analogy. The option describing a single one-shot best-prompt generation with no criticism-and-edit loop or search is wrong because ProTeGi's defining mechanism is precisely its iterative critique-edit-search loop, which that option omits entirely. The option describing population-based crossover and mutation is wrong because that describes Promptbreeder's evolutionary mechanism, which the question already identifies as a separate, contrasting method, not ProTeGi's gradient-and-beam-search approach.
Source: Pryzant, Iter, Li, Lee, Zhu & Zeng, 'Automatic Prompt Optimization with "Gradient Descent" and Beam Search' (arXiv:2305.03495, EMNLP 2023)
AI & LLM Engineering · Prompt Engineering · Card 081/081easy
A team deploying a ChatGPT-based assistant wants a lightweight, purely prompting-level defense against jailbreak attempts, one that does not require architecture-level controls such as tool-permission scoping or human approval steps, and does not involve any additional training of the model. Xie et al. (2023), 'Defending ChatGPT against jailbreak attack via self-reminders,' published in Nature Machine Intelligence, propose a technique drawing on the psychological concept of a self-reminder. What does their system-mode self-reminder technique do, and what effect on jailbreak success rate did their experiments report?
AIt inspects each user message for keywords associated with known jailbreak prompt templates and, if any are found, replaces the flagged span with a hidden untrusted-text marker before the model ever reads it, the same marking mechanism later generalized by spotlighting
BIt fine-tunes the underlying model on a dataset of jailbreak attempts paired with refusals, so the defense is built into the model's weights rather than applied at prompting time
CIt wraps the user's query within a system-level reminder, placed before and after the query, asking the model to respond responsibly and in accordance with guidelines; across their experiments this reduced the jailbreak success rate from 67.21% to 19.34%
DIt sends the user's query to a separate, independently trained classifier model that scores the query for jailbreak intent before the primary model ever sees it, automatically blocking any query that scores above a fixed threshold
Correct answer: .
Drawing on the psychological concept of a self-reminder, the authors propose a 'system-mode self-reminder' that wraps the user's query between reminder text placed both before and after it, prompting the model to remember that it should respond responsibly and in accordance with its guidelines even as it processes whatever the enclosed query asks. This is purely a prompting-time intervention, requiring no change to the model's weights and no external architecture such as tool scoping or human-in-the-loop approval, which makes it attractive as a cheap, easily deployable mitigation. Across their experiments using a jailbreak-prompt dataset the authors constructed, this technique reduced the jailbreak attack success rate from 67.21% down to 19.34%, a substantial but not complete reduction. The option describing keyword-based detection and hidden untrusted-text markers is wrong because that describes a different, later technique (spotlighting) aimed at indirect prompt injection from untrusted documents, not the reminder-wrapping mechanism this paper proposes for jailbreak prompts typed directly by a user. The option describing fine-tuning on jailbreak-and-refusal pairs is wrong because the self-reminder technique is applied entirely at inference time through the prompt, with no additional training involved. The option describing a separately trained classifier scoring and blocking queries is wrong because the paper's mechanism never introduces an external classifier model; it relies solely on reminder text added around the query for the same model to read.