passdrill

Prompt Engineering

81 cards · AI & LLM Engineering · answer each one, then read the explanation. Your score tallies below. Looking for Which prompting technique to use: CoT vs self-consistency vs ReAct vs ToT? Read the explainer.

0 / 81 answered · 0 correct

AI & LLM Engineering · Prompt Engineering · Card 001/081 easy

According to the original GPT-3 paper, "Language Models are Few-Shot Learners" (Brown et al., 2020), which best describes the difference between zero-shot and few-shot prompting?

  1. Zero-shot and few-shot both require gradient updates to the model's weights; they differ only in how many examples are used per update
  2. Zero-shot requires the model to be fine-tuned on the target task first, while few-shot requires no training at all
  3. Few-shot prompting means the model is shown zero examples but asked to solve the task in fewer than five reasoning steps
  4. Zero-shot provides no task examples in the prompt, relying only on a natural language instruction, while few-shot includes a small number of input-output examples in the prompt before the actual query
AI & LLM Engineering · Prompt Engineering · Card 002/081 easy

Under the technique introduced by Wei et al. (2022) in "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models," what does chain-of-thought prompting add to a standard few-shot prompt?

  1. It instructs the model to output only the final answer with no explanation, in order to reduce token usage and cost
  2. It includes intermediate reasoning steps leading to the final answer within the few-shot exemplars, rather than showing only the input and final answer
  3. It replaces the task examples with a single, more detailed instruction and removes all examples from the prompt
  4. It requires retraining the model on a dataset of step-by-step solutions before it can be used
AI & LLM Engineering · Prompt Engineering · Card 003/081 easy

Anthropic's prompt engineering documentation for Claude recommends "giving Claude a role." Where does it say a role description should be written, and why?

  1. In the final line of the user's message, because Claude only applies role instructions that appear immediately before the question being asked
  2. In a separate fine-tuning dataset, because role behavior can only be changed by retraining the model on role-labeled examples
  3. In the system prompt, because even a single sentence establishing a role there focuses Claude's behavior and tone for the specific use case
  4. In the assistant's prior turn, because Claude infers its role only from how it phrased earlier responses in the conversation
AI & LLM Engineering · Prompt Engineering · Card 004/081 easy

In LLM inference APIs such as Anthropic's Messages API and OpenAI's Chat Completions API, what effect does lowering the `temperature` sampling parameter toward 0 have on generated text?

  1. It sharpens the probability distribution over next tokens so the model more consistently picks the highest-probability token, producing more deterministic, less varied output
  2. It disables sampling entirely and forces the model to retrieve an exact quote from its training data
  3. It increases the number of tokens the model is allowed to generate in a single response
  4. It reduces the model's context window, so fewer previous tokens are considered when generating each new token
AI & LLM Engineering · Prompt Engineering · Card 005/081 easy

Per the OWASP Top 10 for LLM Applications (2025), which best distinguishes "indirect" prompt injection from "direct" prompt injection?

  1. Direct prompt injection means malicious instructions are typed straight into the model's input by the user; indirect prompt injection means the malicious instructions are hidden in external content, such as a webpage or document, that the LLM later ingests and follows
  2. Direct prompt injection only affects open-source models, while indirect prompt injection only affects closed, API-based models
  3. Indirect prompt injection requires physical access to the server running the model, while direct prompt injection can be performed remotely over the network
  4. Direct prompt injection is a purely theoretical risk with no real-world examples, while indirect prompt injection has already been fixed in all major LLM products
AI & LLM Engineering · Prompt Engineering · Card 006/081 easy

According to Anthropic's prompt engineering documentation, what is the stated benefit of wrapping different parts of a prompt (instructions, context, examples, input) in distinct XML tags such as `<instructions>` and `<context>`?

  1. It compresses the prompt so that it consumes fewer tokens than the equivalent plain-text prompt
  2. It is required syntax without which the Messages API will reject the request with a formatting error
  3. It automatically translates the tagged sections into a different language before the model processes them
  4. It helps Claude parse complex prompts unambiguously by clearly separating different types of content, reducing the chance the model misinterprets what is instruction versus example versus input
AI & LLM Engineering · Prompt Engineering · Card 007/081 medium

Zhao et al. (2021), "Calibrate Before Use: Improving Few-Shot Performance of Language Models," identify "recency bias" as one cause of instability in few-shot prompting. What does recency bias describe?

  1. The tendency of a model's accuracy to decline over time as new versions of the model are released
  2. The model favoring the most recently published research papers when asked to cite sources
  3. The model's tendency to disproportionately predict whichever label appeared in the example placed nearest the end of the few-shot prompt, regardless of the true input
  4. The tendency to give more weight to the very first example in a few-shot prompt while ignoring later examples
AI & LLM Engineering · Prompt Engineering · Card 008/081 medium

Wang et al. (2022), "Self-Consistency Improves Chain of Thought Reasoning in Language Models," propose replacing greedy decoding with what alternative decoding strategy for chain-of-thought prompts?

  1. Discard chain-of-thought reasoning entirely and instead retrieve the answer from an external search engine
  2. Sample multiple diverse reasoning paths for the same question at a nonzero temperature, then take a majority vote over the final answers each path arrives at
  3. Always generate exactly one reasoning path deterministically, then ask a separate human reviewer to check it before accepting the answer
  4. Fine-tune the model on the correct chain-of-thought path found by brute-force search over the entire training set
AI & LLM Engineering · Prompt Engineering · Card 009/081 medium

Zhou et al. (2022), "Least-to-Most Prompting Enables Complex Reasoning in Large Language Models," describe a two-stage strategy for solving problems harder than those shown in the prompt's examples. What is that strategy?

  1. First prompt the model to decompose the problem into a sequence of simpler subproblems, then sequentially prompt it to solve each subproblem in order, feeding each prior subproblem's answer into the context used for the next
  2. First run the same prompt through several different LLMs from different vendors, then pick whichever vendor's answer appears most frequently
  3. First ask the model to guess the final answer directly, then ask it to generate a chain-of-thought justification for that already-chosen answer after the fact
  4. First fine-tune the model on the hardest available examples, then evaluate it zero-shot on easier examples to measure generalization downward
AI & LLM Engineering · Prompt Engineering · Card 010/081 medium

OpenAI's documentation distinguishes "Structured Outputs" from the older "JSON mode" feature of its Chat Completions / Responses APIs. According to that documentation, what is the key difference between the two?

  1. The two features are functionally identical; "Structured Outputs" is simply a rebranding of "JSON mode" with no change in behavior
  2. Structured Outputs can only be used with image inputs, while JSON mode is restricted to text-only prompts
  3. JSON mode guarantees schema conformance, while Structured Outputs only guarantees syntactically valid JSON without any schema checking
  4. Both guarantee syntactically valid JSON, but only Structured Outputs also guarantees the output conforms to the caller's supplied JSON Schema (for example, required keys and enum values); JSON mode guarantees valid JSON syntax only
AI & LLM Engineering · Prompt Engineering · Card 011/081 hard

Turpin et al. (2023), "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting," demonstrate unfaithfulness by manipulating what feature of a few-shot prompt, and observing what result?

  1. They removed all chain-of-thought reasoning from the few-shot examples entirely and found the model refused to answer at all without it
  2. They increased the number of few-shot examples from 2 to 200 and found chain-of-thought accuracy improved with no change in faithfulness concerns
  3. They reordered the multiple-choice options in the few-shot examples so the correct answer was biased to always fall on a particular letter (e.g., always "(A)"); the model's chain-of-thought then rationalized picking that biased letter while never mentioning the answer ordering as its real reason
  4. They translated the few-shot examples into a different natural language and found the model's final answers became random regardless of the question
AI & LLM Engineering · Prompt Engineering · Card 012/081 hard

Sclar et al. (2023), "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design," measure how much purely cosmetic formatting choices (such as separators and spacing) in a few-shot prompt affect accuracy, holding the semantic content constant. What did they find?

  1. Accuracy differences from formatting disappeared entirely once few-shot examples were replaced with zero-shot instructions
  2. Purely formatting-level changes to a semantically identical few-shot prompt caused accuracy swings of up to tens of accuracy points (as much as 76 points for one open-source model tested), and this sensitivity persisted even with larger models, more few-shot examples, and instruction tuning
  3. Formatting only matters for image-based prompts and has no measurable effect on plain-text few-shot prompts
  4. Formatting choices had no measurable effect on accuracy once a model exceeded roughly one billion parameters, fully resolving the issue at modern model scales
AI & LLM Engineering · Prompt Engineering · Card 013/081 easy

Kojima et al. (2022), "Large Language Models are Zero-Shot Reasoners," show that a specific technique substantially improves LLM performance on multi-step reasoning benchmarks without any task-specific worked examples in the prompt. What is this zero-shot chain-of-thought technique, and how does it differ from the few-shot chain-of-thought prompting of Wei et al. (2022)?

  1. Providing the correct final answer to the model up front and asking it to work backward to justify it, whereas Wei et al.'s method asks for the answer with no justification
  2. Appending a task-agnostic trigger phrase such as "Let's think step by step" before the answer, with no worked examples in the prompt at all, whereas Wei et al.'s method requires several few-shot exemplars that each include their own written-out reasoning steps
  3. Replacing the multiple-choice options with open-ended text so the model can no longer guess from answer choices, a change unrelated to Wei et al.'s few-shot method
  4. Fine-tuning the model on a small labeled set of step-by-step solutions before inference, whereas Wei et al.'s method requires no training at all
AI & LLM Engineering · Prompt Engineering · Card 014/081 medium

Zhou et al. (2022/2023), "Large Language Models Are Human-Level Prompt Engineers," propose Automatic Prompt Engineer (APE). How does APE generate and select an effective task instruction, according to the paper?

  1. It requires a human panel to write hundreds of candidate instructions, which the LLM then simply ranks by fluency without regard to task performance
  2. It fine-tunes the model's weights on many instruction/response pairs and treats the resulting model checkpoint itself as the "instruction"
  3. It works only for classification tasks with a fixed label set and cannot generate free-form instructions for open-ended tasks
  4. It treats the instruction itself as a "program": an LLM proposes a pool of candidate instructions from a handful of input-output demonstrations, and each candidate is scored (for example, by how well it reproduces the demonstrations when used as a prompt) so the highest-scoring instruction can be selected or further refined
AI & LLM Engineering · Prompt Engineering · Card 015/081 hard

Min et al. (2022), "Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?," test in-context learning by randomly replacing the labels in few-shot demonstrations with incorrect ones. What did they find, and what does it suggest about why few-shot demonstrations help?

  1. Replacing the demonstrations' labels with random, often-incorrect ones barely hurt accuracy across a range of classification and multiple-choice tasks, suggesting that the correctness of the input-label mapping matters far less than other aspects of the demonstrations, such as the label space, the input distribution, and the overall format
  2. The experiment could not be run, because in-context learning requires every demonstration label to be verified against a held-out validation set before each query
  3. Replacing the labels with random ones improved accuracy beyond using correct labels, showing that few-shot demonstrations actively mislead the model and should generally be avoided
  4. Replacing the labels with random ones caused accuracy to collapse to chance level on every task tested, confirming that the model learns the exact input-label mapping shown in the demonstrations the way a supervised classifier would
AI & LLM Engineering · Prompt Engineering · Card 016/081 hard

Lu et al. (2022), "Fantastically Ordered Prompts and Where to Find Them," study how the order of few-shot examples within an otherwise-identical prompt affects accuracy. What did they find, and what method did they propose to pick a good order without a labeled validation set?

  1. They found the best fix was to always sort examples alphabetically by their label text, which fully eliminated order sensitivity across all models and tasks tested
  2. They found order sensitivity only affects models under one billion parameters, and it disappears automatically once a model is scaled past that size
  3. The order of the same few-shot examples can swing accuracy from near state-of-the-art to close to random guessing; they proposed generating an artificial, unlabeled "probing" set from the language model itself and selecting the ordering whose predicted-label distribution has favorable entropy statistics on that set, without needing any labeled dev data
  4. Example order has no measurable effect on accuracy once the examples themselves are held constant, so the paper concludes ordering can safely be ignored
AI & LLM Engineering · Prompt Engineering · Card 017/081 easy

Liu et al. (2022), "Generated Knowledge Prompting for Commonsense Reasoning," propose a two-stage prompting pipeline for commonsense question answering. What are the two stages?

  1. First, prompt a language model to generate several relevant knowledge statements about the question's topic; second, provide those generated knowledge statements as additional context alongside the original question in a separate prompt that produces the final answer
  2. First, ask the model to answer the question directly with no context; second, ask a different model to translate that answer into another language to check consistency
  3. First, retrieve documents from a fixed external knowledge base such as an encyclopedia; second, fine-tune the model's weights on those retrieved documents before it answers the question
  4. First, generate several candidate final answers; second, average their token probabilities together to produce a single blended answer string
AI & LLM Engineering · Prompt Engineering · Card 018/081 medium

Wang et al. (2023), "Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models," target error categories that plague Kojima et al.'s zero-shot "Let's think step by step" prompting, such as missing reasoning steps and calculation errors. What does Plan-and-Solve prompting add to address this, without using any worked examples?

  1. It replaces natural-language reasoning entirely with a formal, executable program that a separate code interpreter runs to produce the final answer
  2. It asks the model to skip straight to the final numeric answer without showing any intermediate reasoning, in order to avoid calculation errors introduced by long chains of text
  3. It requires collecting several hundred worked examples with expert-annotated plans and fine-tuning the model on them before it can perform any reasoning
  4. It instructs the model, in a single zero-shot prompt, to first devise a plan that divides the overall task into smaller subtasks, and then to carry out that plan step by step before giving the final answer
AI & LLM Engineering · Prompt Engineering · Card 019/081 medium

Press et al. (2022), "Measuring and Narrowing the Compositionality Gap in Language Models," introduce the "self-ask" prompting method to address multi-hop questions whose sub-answers the model already knows individually but fails to combine correctly. How does self-ask structure the model's output?

  1. The compositionality gap is closed simply by scaling up model size, and self-ask is presented only as an unrelated historical baseline that the paper argues should be discarded
  2. The prompt's few-shot exemplars teach the model to explicitly decide whether a follow-up question is needed, then pose and answer that follow-up question itself within the same generation, repeating as needed before stating the final answer -- a structure into which an external search engine can optionally be plugged to answer the follow-up questions instead
  3. Two separate, independently prompted models debate each other's answers over several rounds until they converge on the same final answer, which is the paper's central proposed method
  4. The model is asked to answer the multi-hop question directly in a single token with no intermediate text of any kind, which the paper shows eliminates the compositionality gap entirely
AI & LLM Engineering · Prompt Engineering · Card 020/081 easy

Zhang et al. (2022), "Automatic Chain of Thought Prompting in Large Language Models," propose Auto-CoT to avoid the manual effort, and potential for hand-written mistakes, involved in writing few-shot chain-of-thought exemplars by hand. What two-step procedure does Auto-CoT use to build its demonstrations automatically?

  1. First, fine-tune the model on a large labeled reasoning dataset; second, discard the few-shot examples entirely, since the fine-tuned model no longer needs any demonstrations
  2. First, ask human annotators to hand-write a reasoning chain for every single question in the dataset; second, use the model to pick which of those human-written chains looks most fluent
  3. First, cluster the dataset's questions by similarity and pick one representative question from each cluster; second, generate a reasoning chain for each representative question automatically using zero-shot chain-of-thought (for example, "Let's think step by step"), assembling the resulting diverse set of question-plus-generated-chain pairs into the few-shot demonstrations
  4. First, generate one reasoning chain for the very first question in the dataset; second, reuse that exact same chain, unmodified, as the only demonstration for every other question regardless of topic
AI & LLM Engineering · Prompt Engineering · Card 021/081 easy

According to Anthropic's current prompt engineering documentation for Claude, roughly how many examples should a multishot (few-shot) prompt include for best results, and what qualities should those examples have?

  1. At least 50 examples are required before Claude can reliably follow the demonstrated pattern at all
  2. About 3 to 5 examples, each relevant to the actual use case and diverse enough (including edge cases) that Claude does not pick up unintended patterns from the examples
  3. The examples must all be drawn from the exact same edge case repeated with different wording, since variety between examples confuses Claude about which pattern to follow
  4. Exactly 1 example is optimal, and using more than 1 measurably degrades Claude's performance on every task
AI & LLM Engineering · Prompt Engineering · Card 022/081 medium

Anthropic's current documentation on long-context prompting recommends where to place long documents or data-rich inputs (roughly 20,000+ tokens) relative to the query and instructions. What does it recommend, and what improvement does it cite?

  1. Split the long document into many short, separate API calls of under 500 tokens each, since Claude cannot process more than roughly 20,000 tokens of context in a single request
  2. Interleave single sentences of the long document between each instruction sentence so that context and instructions alternate line by line throughout the prompt
  3. Place the long documents and other inputs near the top of the prompt, above the query, instructions, and examples; the documentation notes that putting queries at the end can improve response quality by up to 30 percent in tests, especially for complex, multidocument inputs
  4. Place the long documents at the very end of the prompt, after the query and instructions, because Claude always weighs the last few hundred tokens of a prompt most heavily regardless of content length
AI & LLM Engineering · Prompt Engineering · Card 023/081 easy

According to Anthropic's current documentation on chaining complex prompts for Claude, what is described as the most common prompt-chaining pattern, and how is it structured across separate API calls?

  1. Load balancing: send the exact same prompt to several different Claude API keys simultaneously and return whichever response arrives first, discarding the rest
  2. Cache warming: repeatedly resend an unchanged prompt purely to keep it available in the prompt cache, with no reviewing or refining step involved
  3. Weight merging: average the model weights used to generate two separate draft responses into a single hybrid model before producing the final answer
  4. Self-correction: generate a draft response in one API call, have Claude review that draft against stated criteria in a second call, then have Claude refine the draft based on that review in a third call, so each step is a separate call that can be logged, evaluated, or branched on
AI & LLM Engineering · Prompt Engineering · Card 024/081 easy

Anthropic's prompt caching documentation describes marking a "cache breakpoint" with the `cache_control` parameter. For a request that mixes stable content (tool definitions, a system prompt, a large reference document) with content that changes on every call (the current user turn), where should that breakpoint be placed, and why?

  1. On the last content block whose prefix is identical across requests -- that is, after the stable tool definitions, system prompt, and document, and before the changing per-request content -- because the cache only stores what comes before the breakpoint, and placing it on content that changes every request would make the cached prefix's hash change each time, producing no cache hits
  2. Nowhere -- Anthropic's API caches every request identically by default with no `cache_control` parameter or configuration needed
  3. On the changing user message itself, because caching is described as useful specifically for content that differs on every single call
  4. On the very first token of the entire request, including the tool definitions, so that almost nothing in the request ends up covered by the cached prefix
AI & LLM Engineering · Prompt Engineering · Card 025/081 medium

An engineer is building an agent that must look up information from a search API partway through solving a multi-step question and adjust its plan based on what the search returns. According to Yao et al. (2022), "ReAct: Synergizing Reasoning and Acting in Language Models," what does the ReAct prompting framework do to make this possible?

  1. Generating only the sequence of actions to take, such as API calls, without ever producing any intermediate natural-language reasoning
  2. Interleaving natural-language reasoning traces with task-specific actions and the observations those actions return, within a single prompted trajectory, so reasoning can decide the next action and each new observation can update the reasoning that follows
  3. Training a separate reasoning model and a separate acting model with reinforcement learning and combining their outputs after each has finished running independently
  4. Producing a chain-of-thought explanation for a problem but never issuing any call to an external tool or API
AI & LLM Engineering · Prompt Engineering · Card 026/081 hard

A researcher is applying a language model to a puzzle-like task, such as the Game of 24, where an early move can turn out to be a dead end many steps later, and simply extending a single left-to-right chain of thought performs poorly. According to Yao et al. (2023), "Tree of Thoughts: Deliberate Problem Solving with Large Language Models," what does the Tree of Thoughts framework add on top of chain-of-thought prompting to address this?

  1. Producing one single, linear sequence of reasoning steps from the problem to the final answer, exactly as in standard chain-of-thought prompting
  2. Sampling many complete, independent chain-of-thought reasoning paths for the whole problem and picking the final answer that the largest number of them agree on
  3. Framing problem solving as a search over a tree whose nodes are intermediate "thoughts": the model generates and self-evaluates multiple candidate next thoughts at each step, and the search can look ahead or backtrack using strategies such as breadth-first or depth-first search, at substantially higher inference-time compute cost than a single reasoning chain
  4. Training a separate value function offline to score candidate solutions, then using that fixed value function alone to pick a solution without the language model exploring or backtracking at inference time
AI & LLM Engineering · Prompt Engineering · Card 027/081 easy

A developer is building an application on OpenAI's Chat Completions API and needs to send the model both a high-level instruction that should shape its behavior for the whole conversation and the end user's own question. Per OpenAI's API documentation, how should these two pieces of context be structured in the request?

  1. Both pieces of context are sent as a single combined string, since the Chat Completions API has no way to distinguish who authored which part of the input
  2. Only the end user's question can be sent to the model; there is no supported way to give it standing instructions that apply to the whole conversation
  3. The high-level instruction must be repeated inside every individual user message, since a message with a system role is discarded by the API after the first turn
  4. Per OpenAI's API documentation, the request is built as an array of role-tagged messages: the high-level instruction is sent as a system message placed at the start of the array, the end user's question is sent as a user message, and any of the model's own prior responses are represented back to it as assistant messages
AI & LLM Engineering · Prompt Engineering · Card 028/081 easy

A team building a math word-problem solver notices that a large language model using standard chain-of-thought prompting writes out a correct step-by-step plan in natural language but still makes an arithmetic slip when computing the final number, producing a wrong answer despite sound reasoning. Gao et al. (2022), "PAL: Program-Aided Language Models," propose an alternative prompting approach specifically to fix this class of error. What does PAL have the model do differently?

  1. PAL asks the model to write out the same natural-language reasoning chain twice independently and takes whichever of the two final numeric answers appears first
  2. PAL replaces every arithmetic step with a request for the model to look up the answer in a retrieved external knowledge base of pre-solved problems
  3. PAL prompts the model to translate the problem into intermediate steps expressed as runnable code (for example Python statements) and hands that generated program to an external interpreter to execute, using the interpreter's output as the final answer instead of having the model compute it itself
  4. PAL fine-tunes the underlying model on millions of additional arithmetic examples so it memorizes correct computations for common operation patterns
AI & LLM Engineering · Prompt Engineering · Card 029/081 easy

A developer notices that a long-form LLM answer (for example, 'list and explain five approaches to X') is slow to generate because the model must produce the entire response as one long sequential stream of tokens, point by point, in order. Ning et al. (2023), "Skeleton-of-Thought: Large Language Models Can Do Parallel Decoding," propose a prompting-and-decoding approach aimed specifically at cutting this latency without changing the model's weights. What does Skeleton-of-Thought do?

  1. It first prompts the model to generate a brief skeleton of the answer's main points, then issues separate parallel calls (or batched decoding) to expand each point concurrently, before combining the expanded points into the final answer
  2. It shortens the answer by instructing the model to omit the elaboration for each point, trading completeness for speed
  3. It routes the request to a smaller, faster model for a first draft, then upgrades to the original model only for a final proofreading pass
  4. It caches the model's response to the same prompt across users, so only the first request incurs the full sequential generation cost
AI & LLM Engineering · Prompt Engineering · Card 030/081 easy

Deng et al. (2023), "Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves," start from the observation that a human's phrasing of a question often carries ambiguity or missing context that a language model reads differently than the human intended. What does the Rephrase and Respond (RaR) method have the model do about this, and how does the paper describe its relationship to chain-of-thought prompting?

  1. RaR trains a separate classifier to detect ambiguous questions and reroutes only those questions to a human reviewer before any model response is generated
  2. RaR has the model rephrase and expand the given question in its own words, adding clarifying detail, before answering, which the paper shows is complementary to chain-of-thought rather than a replacement for it, and combining the two performs better than either alone
  3. RaR instructs the model to translate the question into a different natural language first, on the theory that translation removes ambiguity, and this fully replaces the need for chain-of-thought
  4. RaR skips rephrasing entirely and instead asks the model to answer the same question multiple times, then rephrases only the final selected answer for readability
AI & LLM Engineering · Prompt Engineering · Card 031/081 easy

A developer calling Anthropic's Messages API wants Claude's response to always begin with a valid JSON object and never with an introductory sentence like 'Here is the JSON you requested:'. According to Anthropic's documentation on prefilling Claude's response, how is this achieved, and through what channel is the technique available?

  1. By adding a `response_format: json` parameter to the request, which is documented as the only supported way to constrain where Claude's output can start
  2. By appending a stop sequence equal to the word 'Here' so that Claude is blocked from generating that word at the start of its answer
  3. By fine-tuning a custom version of Claude on examples that start directly with `{`, since the base model cannot be steered to skip a preamble through prompting alone
  4. By including the desired starting text (for example an opening `{`) as the beginning of the assistant turn itself in the API call, so Claude's generated response continues directly from that point; the documentation notes this prefill technique is API-only and is not supported on the newest Claude models
AI & LLM Engineering · Prompt Engineering · Card 032/081 easy

A developer who has used Anthropic's Messages API is used to marking a `cache_control` breakpoint on the content she wants cached (such as a large system prompt or reference document), so that stable content is reused across calls and only new content is billed at full price. Moving the same workload to OpenAI's API, she wants to know whether she must add an equivalent manual marker to get a caching discount. Per OpenAI's own documentation on prompt caching, what should she expect?

  1. No manual marker is required for basic caching: OpenAI's prompt caching is applied automatically, matching and reusing the longest previously-seen prefix of a prompt once the prompt is long enough to qualify, without any `cache_control`-style parameter needed to opt in
  2. Yes, the requirement is identical: she must annotate the exact same content block with a `cache_control` field, using the same syntax as Anthropic's API, or no caching discount will ever apply
  3. No, because OpenAI's API does not offer any form of prompt caching at all, regardless of prompt length or repetition
  4. Yes, but only through a separate paid add-on subscription that must be purchased before any caching discount becomes available on cached tokens
AI & LLM Engineering · Prompt Engineering · Card 033/081 easy

Madaan et al. (2023), "Self-Refine: Iterative Refinement with Self-Feedback," propose a prompting loop that improves a language model's own output without any additional training data, fine-tuning, or external verifier model. What is the loop, and which single model performs every role in it?

  1. A larger 'teacher' model grades the outputs of a separate, smaller 'student' model and returns a numeric score that the student uses to pick from several candidate answers
  2. A retrieval system feeds the model documents relevant to its own previous answer, and the model simply copies the most relevant retrieved sentence into its final response
  3. The same pretrained model first produces an initial answer, then critiques that answer as feedback, then uses its own critique to produce a refined answer, repeating this generate-feedback-refine loop for multiple rounds; no separate model or extra training is involved at any stage
  4. A human reviewer reads the model's first answer and writes the feedback, which is then pasted back into a second prompt for the model to revise
AI & LLM Engineering · Prompt Engineering · Card 034/081 medium

A model is asked a specific, detail-heavy physics question and, despite reasoning step by step, applies the wrong underlying formula because it dives straight into the specific numbers without first recalling which general principle governs the situation. Zheng et al. (2023), "Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models," propose Step-Back Prompting to address exactly this failure mode. How does the technique change what the model is prompted to do before answering?

  1. It has the model answer the specific question first, and only afterward asks it to state which general principle it implicitly used, purely as a post-hoc explanation with no effect on the answer
  2. It has the model retrieve the original specific question from a database of similar past questions and copy the closest match's stored answer
  3. It has the model break the specific question into the smallest possible sub-questions and answer each sub-question independently before summing the sub-answers
  4. It first prompts the model to derive a higher-level, more abstract question or general principle from the specific instance, then has the model reason from that abstraction to answer the original specific question, rather than reasoning directly from the details alone
AI & LLM Engineering · Prompt Engineering · Card 035/081 medium

Chia et al. (2023), "Contrastive Chain-of-Thought Prompting," start from a surprising finding in prior work: giving a model chain-of-thought demonstrations with deliberately invalid reasoning had only a small negative effect compared to fully valid demonstrations, suggesting standard CoT does not clearly teach the model what mistakes to avoid. What do the authors propose instead, and what is its rationale?

  1. Removing chain-of-thought demonstrations altogether and relying solely on zero-shot instructions, on the theory that any demonstration risks teaching a spurious pattern
  2. Providing both valid and invalid reasoning demonstrations side by side for the same or paired problems, together with an automatic method for constructing such contrastive demonstrations, so the model is explicitly shown examples of flawed reasoning alongside correct reasoning rather than only ever seeing correct chains
  3. Increasing the number of valid chain-of-thought demonstrations from a handful to several dozen, so that sheer repetition overwhelms any influence from invalid reasoning
  4. Replacing human-written demonstrations with demonstrations generated entirely by a separate, larger teacher model, without regard to whether the reasoning in them is valid or invalid
AI & LLM Engineering · Prompt Engineering · Card 036/081 medium

A team wants the reasoning-boosting benefit of few-shot chain-of-thought prompting on a new task, but has no labeled exemplars available and does not want to spend engineering effort hand-writing or retrieving demonstrations for every new problem type. Yasunaga et al. (2023), "Large Language Models as Analogical Reasoners," propose Analogical Prompting to address this. What does the model do under this method, and what does it avoid needing?

  1. The model retrieves the single most similar labeled example from a large curated exemplar database using a nearest-neighbor embedding search, then copies that example's reasoning structure exactly
  2. The model is fine-tuned on a small set of analogous problems drawn from a related domain before being asked to solve the target problem
  3. The model asks a human expert, via an interactive clarification step, to supply one worked analogy before attempting to solve the problem
  4. The model is prompted to recall or self-generate relevant exemplars and, where useful, relevant background knowledge related to the given problem before solving it, entirely on its own; the method avoids needing any labeled exemplars to be hand-written or retrieved from an external source
AI & LLM Engineering · Prompt Engineering · Card 037/081 medium

A model asked to write a short biographical paragraph about a real historical figure produces a mostly accurate response but invents two incorrect dates. Dhuliawala et al. (2023), "Chain-of-Verification Reduces Hallucination in Large Language Models," propose a four-step prompting procedure (CoVe) meant to catch this kind of error before the response is delivered to the user. What are the four steps, in order?

  1. Draft an answer, then immediately ask the model to rate its own confidence on a 1-10 scale, then deliver whichever draft scores highest without any further steps
  2. Draft an initial response, plan a set of targeted verification questions that would fact-check specific claims in that draft, answer each verification question independently so those answers are not biased by the original draft, then use the verification answers to produce a final, revised response
  3. Draft an answer, translate it into a different language and back, compare the two versions for discrepancies, and keep whichever version is shorter
  4. Draft an answer, retrieve external documents matching every named entity in the draft, replace every named entity with whatever the retrieved documents say without any further verification step, and stop
AI & LLM Engineering · Prompt Engineering · Card 038/081 hard

A team wants to steer a large black-box LLM they cannot fine-tune (no access to weights) toward producing summaries that reliably include certain keywords, without hand-crafting a new instruction for every single input document. Li et al. (2023), "Guiding Large Language Models via Directional Stimulus Prompting," propose a framework for exactly this setting. What mechanism does Directional Stimulus Prompting use, and how is that mechanism trained?

  1. It trains a separate, small, tunable policy model (such as a T5-sized model) to generate an instance-specific 'directional stimulus,' a short hint or set of keywords tailored to each input, which is inserted into the prompt sent to the black-box LLM; the policy model itself is optimized via supervised fine-tuning on labeled data and/or reinforcement learning using rewards derived from the black-box LLM's own output quality
  2. It fine-tunes the black-box LLM's own weights directly on a small labeled dataset of ideal summaries, despite the team's lack of weight access, by using a gradient-free zeroth-order optimization method against the model's API
  3. It hard-codes the same fixed list of keywords into every prompt regardless of the input document, relying on the black-box LLM to decide on its own which of the fixed keywords are actually relevant
  4. It relies entirely on the black-box LLM's built-in retrieval-augmented generation feature to pull in relevant keywords from a search index, with no additional model or training step involved
AI & LLM Engineering · Prompt Engineering · Card 039/081 hard

A team building on OpenAI's newer reasoning models (the o-series and later) via the Chat Completions or Responses API wants to know how the older 'system' role relates to the newer 'developer' role, and where each level of instruction sits in OpenAI's documented instruction-following hierarchy. Per OpenAI's Model Spec and API guidance, which best describes this?

  1. The 'system' and 'developer' roles are two independent, co-equal channels with no defined precedence between them, so a conflict between a system message and a developer message is resolved arbitrarily by the model at random
  2. The 'developer' role sits below the 'user' role in authority, meaning any instruction a user provides during the conversation can always override an instruction given via the developer role
  3. For the newer models, the 'developer' role takes over the position the 'system' role used to occupy for API callers, sitting above the 'user' role in OpenAI's documented chain-of-command (platform-level instructions rank above developer instructions, which in turn rank above user instructions, with assistant and tool messages carrying no independent authority)
  4. The 'system' role was retired entirely with no replacement, and any instruction that used to go in a system message must now be passed only through the per-request `instructions` parameter, with no message-role equivalent available at all
AI & LLM Engineering · Prompt Engineering · Card 040/081 medium

A red-teaming exercise finds that stuffing an LLM's prompt with a few dozen fabricated dialogue turns showing a compliant assistant answering harmful requests has little effect on the model's safety behavior, but scaling the same fabricated dialogue up to hundreds of turns reliably breaks it. Per Anthropic's 2024 research on this technique, called "many-shot jailbreaking," what mechanism explains why effectiveness increases so sharply with the number of turns, and which models are most exposed?

  1. It exploits in-context learning: the growing number of faux dialogue turns showing harmful compliance functions like few-shot demonstrations that steer the model's behavior toward the pattern shown, with attack success following a power-law increase as the number of turns grows; models with the largest context windows are most exposed because they can fit enough turns for the effect to take hold
  2. It works by repeatedly asking the same harmful question in slightly reworded form until the model's output-length limit forces it to answer directly instead of refusing
  3. It exploits a training-data leakage bug specific to one vendor's models, where feeding back memorized fragments of the safety-training dataset itself disables the safety filter
  4. It works by encoding the harmful request in a language underrepresented in the safety-training data, exhausting a fixed per-request translation budget before any safety classifier runs
AI & LLM Engineering · Prompt Engineering · Card 041/081 easy

A developer needs an LLM to classify 600 short product reviews as positive or negative, paying per input-and-output token through a hosted API. Sending each review as its own separate call means the same lengthy classification instructions and few-shot examples are repeated in every single request. Cheng et al. (2023), "Batch Prompting: Efficient Inference with Large Language Model APIs," propose an alternative. What does batch prompting do to cut this cost?

  1. It fine-tunes a smaller specialized model on the same classification examples so the large model is no longer needed once training completes
  2. It groups several independent input samples together into a single prompt, sharing one copy of the instructions and in-context examples across the whole group, and has the model return the answers for all of the grouped samples in one response instead of issuing a separate call per sample
  3. It reduces cost by having the model skip the few-shot examples entirely and rely purely on zero-shot instructions for every request
  4. It caches every previous request-response pair, so a later request only costs tokens if its text is not an exact character-for-character match to an earlier one
AI & LLM Engineering · Prompt Engineering · Card 042/081 hard

A team wants an LLM to write and keep improving its own instruction for a reasoning task, given only a way to score any candidate instruction against a small held-out set (for example, accuracy on GSM8K). Zhou et al.'s Automatic Prompt Engineer (APE) generates a batch of candidate instructions in one pass from example input-output pairs and selects whichever scores best. Yang et al. (2023), "Large Language Models as Optimizers," propose OPRO for the same kind of problem. How does OPRO's procedure differ from APE's one-pass generate-and-select approach?

  1. OPRO trains a separate small neural network purely on the numeric scores to predict the best wording, without the LLM ever seeing the scores itself
  2. OPRO removes scoring from the loop entirely and instead has the LLM vote on which of several candidate instructions sounds most natural to a human reader
  3. OPRO runs an iterative optimization loop: at each step it feeds the model a "meta-prompt" containing the trajectory of previously tried instructions together with their scores, and asks the model to propose a new instruction intended to score higher than the ones already tried, so each new attempt is conditioned on the full history of earlier attempts rather than everything being generated in one pass
  4. OPRO requires updating the LLM's own weights via gradient descent on the scoring function, making it inapplicable to closed-weight models accessed only through an API
AI & LLM Engineering · Prompt Engineering · Card 043/081 easy

Li et al. (2023), "Large Language Models Understand and Can Be Enhanced by Emotional Stimuli," test appending short psychologically-motivated phrases, such as "This is very important for my career" or "You'd better be sure and think carefully," to the end of otherwise-unchanged task instructions, an approach the paper calls EmotionPrompt. What did the paper report as the effect of adding these phrases, and did the technique require retraining the model?

  1. The phrases had no measurable effect on any tested model, confirming that language models are insensitive to wording intended to convey urgency or stakes
  2. The phrases improved performance only after the model was fine-tuned on a dataset of emotionally-annotated examples paired with correct answers, so the effect required retraining rather than being a pure prompting technique
  3. The phrases decreased benchmark performance by making the model overly cautious and more likely to refuse to answer, though human raters still preferred the more cautious tone
  4. Adding these appended emotional-stimulus phrases to otherwise-unchanged instructions produced measurable improvements in benchmark performance and in human ratings of the resulting text across multiple LLMs, and the technique required no retraining or fine-tuning of any kind, since it works purely by changing the wording of the prompt
AI & LLM Engineering · Prompt Engineering · Card 044/081 hard

Standard zero-shot chain-of-thought prompting (Kojima et al., 2022) elicits step-by-step reasoning by appending an explicit instruction such as "Let's think step by step" before greedily decoding the response. Wang & Zhou (2024), "Chain-of-Thought Reasoning Without Prompting," report finding reasoning paths in a pretrained model without adding any such instruction to the prompt at all. What do they change instead, and what do they observe as a result?

  1. Instead of adding any chain-of-thought instruction to the prompt, they change how the very first token of the response is decoded, inspecting the top-k alternative tokens rather than only the single highest-probability greedy token at that first step; branching down some of those alternative paths reveals chain-of-thought reasoning that was already latent in the pretrained model, and when such a path appears, the model tends to show higher confidence in its final answer
  2. They fine-tune the pretrained model on a small set of chain-of-thought demonstrations, after which greedy decoding alone reproduces step-by-step reasoning without needing the "let's think step by step" instruction
  3. They replace the pretrained model's tokenizer with one that segments numbers digit by digit, which the paper claims is solely responsible for eliciting latent reasoning at the decoding stage
  4. They add a hidden system-level instruction that is invisible to the end user but functionally identical to Kojima et al.'s "let's think step by step" phrase, appended before decoding begins
AI & LLM Engineering · Prompt Engineering · Card 045/081 medium

Zou et al. (2023), "Universal and Transferable Adversarial Attacks on Aligned Language Models," introduce the Greedy Coordinate Gradient (GCG) method for finding a short string that, appended to a harmful request, causes safety-trained models to comply. How does GCG find this adversarial suffix, and what does "transferable" mean in the paper's results?

  1. GCG works by manually testing suffixes proposed by a human red team, ranking them only by how grammatically fluent they read, with no use of gradients or automated search at all
  2. GCG uses gradient information from one or more open-weight models to greedily search for, and iteratively replace, individual tokens in a candidate suffix so as to increase the likelihood that the target model begins its response by complying with the harmful request; "transferable" describes the finding that a suffix optimized against open-weight models such as Vicuna also induces objectionable output when tested against unrelated closed models it was never optimized against
  3. GCG requires direct write access to the target model's weights during the attack itself, so it cannot be used against a closed model served only through an API, and "transferable" refers only to porting the attack's code between programming languages
  4. GCG produces a suffix that is unique to a single specific harmful request and a single specific model, and "transferable" describes only how the resulting text can be copy-pasted between chat interfaces of the same model
AI & LLM Engineering · Prompt Engineering · Card 046/081 medium

Tree of Thoughts (Yao et al., 2023) lets a model explore multiple reasoning branches and backtrack from dead ends, but each branch in a tree can only split further or be discarded; branches cannot be recombined with each other. Besta et al., "Graph of Thoughts: Solving Elaborate Problems with Large Language Models," propose a structure that goes beyond this constraint. What does Graph of Thoughts add on top of the tree structure, and what capability does that addition unlock?

  1. It removes branching entirely and forces the model onto a single linear chain, trading the exploration benefits of Tree of Thoughts for a large reduction in the number of model calls needed
  2. It adds a fixed, hand-written decision tree of if-then rules external to the language model that decides which existing branch to keep, replacing the model's own judgment about branch quality
  3. It models the reasoning process as an arbitrary graph, where individual "thoughts" are vertices and dependencies between them are edges rather than being restricted to a single parent-to-child tree shape; this lets thoughts explored on separate branches be merged, aggregated, or fed back into each other through feedback loops, rather than only ever being extended or dropped
  4. It requires training a separate small classifier model to score each branch numerically, since the underlying language model in this framework is never asked to judge or compare its own branches
AI & LLM Engineering · Prompt Engineering · Card 047/081 easy

Khattab et al. (2023), "DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines," argue that hand-written free-form prompt templates discovered by trial and error make LLM pipelines brittle and hard to reuse across models. What does DSPy have the programmer do instead, and what does the framework itself handle automatically?

  1. The programmer writes the exact final prompt wording as before, and DSPy's only contribution is translating that wording into several different natural languages automatically
  2. The programmer must manually rewrite every prompt for every new base model DSPy is pointed at, since the framework provides no automatic prompt generation or tuning of its own
  3. The programmer specifies only the desired output token count, and DSPy pads or truncates the model's natural response to match that fixed length regardless of prompt wording
  4. The programmer declares a "signature" describing, in a structured and model-agnostic way, what inputs a step needs and what outputs it should produce, and composes these signatures into modules forming a pipeline; DSPy's own compiler then automatically generates and tunes the actual natural-language prompt text, and can select or generate few-shot demonstrations, needed to make each declared step work
AI & LLM Engineering · Prompt Engineering · Card 048/081 easy

An engineering team building a customer-support chatbot that can call internal tools has already learned OWASP's distinction between direct prompt injection (attacker text typed straight into the chat) and indirect prompt injection (malicious instructions hidden in a document, webpage, or tool result the model later reads). Separately from that classification, what defense-in-depth mitigations does the OWASP Top 10 for LLM Applications (2025) recommend for reducing the risk and blast radius of a successful prompt injection?

  1. A combination of layered controls: constraining model behavior and output format through the system prompt, segregating untrusted external content so it cannot be interpreted as an instruction, restricting the model's tools and permissions to the minimum needed for the task, and requiring human approval before any high-risk or irreversible action is carried out
  2. Relying on a single measure, training a dedicated classifier that scans every user message for injection attempts, which OWASP describes as sufficient on its own once deployed, without any additional tool-permission or output controls
  3. Disabling all tool use entirely for any model that might ever process text from an external source, since OWASP states this is the only mitigation that fully eliminates the risk
  4. Encrypting the model's system prompt so that it cannot be extracted through prompt leaking, which OWASP identifies as its primary recommended defense against prompt injection specifically
AI & LLM Engineering · Prompt Engineering · Card 049/081 easy

Du et al. (2023), "Improving Factuality and Reasoning in Language Models through Multiagent Debate," test an alternative to having a single model instance revise its own answer alone, as in Self-Refine. What debate procedure do they propose, and how does it differ from a single model critiquing and refining its own output?

  1. A single model instance argues both sides of a debate against itself within one continuous response, then declares its own earlier argument the winner without ever comparing the two arguments against each other
  2. Multiple separate instances of a language model each independently produce an answer and its reasoning, then are shown one another's answers and reasoning over several rounds and asked to update their own response in light of the others' arguments, continuing until the instances converge on a shared final answer; this differs from Self-Refine, which uses one model instance for every role, generator, critic, and refiner, with no independent second party involved
  3. One model instance generates several candidate answers, and a much smaller, separately trained model is trained to grade which of those answers is factually best, with no back-and-forth exchange of arguments involved
  4. Multiple model instances are merged into a single set of weights via averaging before generating one answer, so no exchange of natural-language arguments occurs between separate active instances at inference time
AI & LLM Engineering · Prompt Engineering · Card 050/081 easy

A developer wants an LLM to extract structured data (name, date, amount) from unstructured invoice text. They try asking the model in plain English but get inconsistent output formats. What prompting technique would most reliably produce consistently structured output?

  1. Providing a concrete output schema or example in the prompt (few-shot prompting with structured examples), such as showing the model one or two completed extractions in the exact JSON format desired, so the model replicates the structure rather than inventing its own format each time
  2. Increasing the temperature parameter to its maximum value, which encourages the model to explore more formatting options and eventually converge on the correct one
  3. Asking the model to explain its reasoning step by step before producing any output, which guarantees the output will be in valid JSON without needing to specify a schema
  4. Sending the same prompt multiple times and averaging the responses, since the model's output will naturally stabilise into a consistent format with enough repetitions
AI & LLM Engineering · Prompt Engineering · Card 051/081 easy

A developer adds the instruction 'Think step by step before answering' to a prompt and notices the model's accuracy improves on a multi-step math problem. What is this technique called, and why does it help?

  1. Chain-of-thought (CoT) prompting — it works because the intermediate reasoning tokens give the model more computation to decompose the problem, catch errors in sub-steps, and maintain state across a multi-step calculation rather than attempting to jump directly from the question to the final answer
  2. Retrieval-augmented generation (RAG) — the instruction causes the model to search an external database for the answer, which is why accuracy improves
  3. Temperature tuning — the phrase 'think step by step' automatically lowers the model's temperature parameter, making the output more deterministic and therefore more accurate
  4. Prompt injection — the instruction hijacks the model's control flow, which coincidentally produces correct answers for math problems but is considered a security vulnerability
AI & LLM Engineering · Prompt Engineering · Card 052/081 easy

A team notices that when a prompt includes an opinionated aside right before the actual question (for example, 'I really think the answer is X, but what do you think?'), an LLM's response is swayed toward agreeing with the aside rather than giving an independent, well-reasoned answer, even when the aside is unrelated to what is actually correct. Weston & Sukhbaatar (2023), 'System 2 Attention (is something you might need too),' propose a two-stage inference method to reduce this kind of susceptibility to irrelevant or biasing context. What does System 2 Attention (S2A) do?

  1. It fine-tunes the model's attention weights on a dataset of biased versus unbiased prompts so the model learns, once and for all, to ignore opinionated asides in any future prompt
  2. It has the LLM first regenerate the input context in an initial inference pass, rewriting it to strip out irrelevant or biasing content such as the aside, and then attends to and answers based on only that regenerated context in a second pass
  3. It lowers the sampling temperature to 0 for the duration of the response, which is described as suppressing the model's tendency to mirror the emotional tone of the prompt
  4. It appends a fixed disclaimer instructing the model to disregard any opinions in the prompt, relying on the model always giving the highest priority to instructions placed at the very end of a prompt
AI & LLM Engineering · Prompt Engineering · Card 053/081 medium

A team building few-shot chain-of-thought prompts for grade-school and competition math problems has a small pool of hand-written exemplars, some walking through only two reasoning steps and others walking through eight or more before reaching an answer. Fu et al. (2022), 'Complexity-Based Prompting for Multi-Step Reasoning,' study how the choice of exemplars, and how multiple sampled outputs are combined, affects accuracy on this kind of multi-step task. What does their complexity-based approach do?

  1. It selects the shortest, simplest exemplars available for the few-shot prompt, on the theory that concise demonstrations reduce the chance the model copies an irrelevant reasoning pattern
  2. It selects exemplars at random regardless of step count, then replaces greedy decoding with a single high-temperature sample to increase output diversity
  3. It selects exemplars with the fewest reasoning steps, then applies a majority vote over multiple greedily decoded generations of the same prompt
  4. It deliberately selects exemplars with more reasoning steps as the few-shot demonstrations, and when combining several sampled reasoning chains for a test question, favors a majority vote taken among the more complex generated chains rather than weighting every sampled chain equally
AI & LLM Engineering · Prompt Engineering · Card 054/081 easy

A team has a large pool of thousands of labeled input-output pairs but can fit only a handful of them as few-shot exemplars into a single prompt due to context-length limits. Rather than picking the same fixed set of exemplars for every test input, Liu et al. (2021), 'What Makes Good In-Context Examples for GPT-3?,' propose retrieving a different set of exemplars for each new input at query time. How do they decide which exemplars from the pool to retrieve for a given input?

  1. For each test input, they embed it and retrieve the exemplars from the candidate pool whose embeddings are most semantically similar (nearest neighbours) to that input, so a different exemplar set is used per query
  2. They retrieve exemplars uniformly at random from the pool for every query, on the reasoning that random sampling prevents the model from overfitting to any single fixed prompt template
  3. They select whichever exemplars have the shortest input text, since shorter exemplars leave more of the context window available for the model's own reasoning
  4. They select whichever exemplars appear earliest in the training set's original ordering, since fixing exemplar order across every query improves output calibration
AI & LLM Engineering · Prompt Engineering · Card 055/081 medium

A team wants a single LLM deployment to handle a wide range of unrelated task types, such as coding, arithmetic word problems, and creative writing, without hand-crafting a separate task-specific prompt template for each one in advance, and without any additional fine-tuning. Suzgun & Kalai (2024), 'Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding,' propose using a single underlying model in two distinct roles at inference time. What is this structure?

  1. A single fixed instruction is prepended to every task type, and the model answers directly in one pass with no decomposition into subtasks at all
  2. Two separately fine-tuned checkpoints of the same base model are created, one trained to plan and one trained to execute, and they pass messages to each other over an API
  3. One instance of the model acts as a high-level 'conductor' that breaks the task into subtasks and writes a tailored instruction for each, then dispatches each subtask to a separate 'expert' instance of the same underlying model, before integrating their outputs into one final response, all through prompting alone
  4. The model is asked to generate several candidate prompts for the task, and a human reviewer manually selects the best one before any generation of the final response begins
AI & LLM Engineering · Prompt Engineering · Card 056/081 hard

After generating several candidate chain-of-thought answers to an arithmetic word problem, a team wants a way to score which candidate is most likely correct, using only the same LLM that generated the candidates, without access to any external calculator, code interpreter, or separately trained verifier model. Weng et al. (2022), 'Large Language Models are Better Reasoners with Self-Verification,' propose a backward-verification procedure for exactly this. How does it score a given candidate answer?

  1. It asks the model to restate the candidate answer in different words and checks whether the restatement uses the same number of tokens as the original, treating matching length as a proxy for confidence
  2. It takes the candidate's conclusion, uses it as a given condition to construct a new question that masks one of the original problem's stated conditions, asks the model to predict that masked original condition from the conclusion and the rest of the problem, and scores the candidate by how accurately the model recovers it
  3. It has the model assign itself a numeric confidence score from 0 to 100 directly, based on how fluent and well-organized the candidate's prose reads, without referring back to the original problem statement at all
  4. It counts how many of the candidate's chain-of-thought sentences use passive voice, on the reasoning that passive constructions indicate the model is unsure of a causal claim it is making
AI & LLM Engineering · Prompt Engineering · Card 057/081 easy

A team wants an LLM to double-check and, if needed, revise its own answer to a math word problem across several automatic rounds within the same session, stopping once the answer stops changing, without using any external tool, calculator, or separately trained verifier model, and without simply asking the model to critique its own reasoning in open-ended prose. Zheng et al. (2023), 'Progressive-Hint Prompting Improves Reasoning in Large Language Models' (PHP), propose a specific mechanism for this. What does PHP do?

  1. It generates ten independent chain-of-thought answers to the question in a single pass and returns whichever answer appears most frequently among them, with no further rounds of interaction
  2. It fine-tunes a small verifier model on a labeled dataset of correct and incorrect chain-of-thought traces, then uses that verifier to pick the best of several candidate answers
  3. It rewrites the original question once, in a single pass, to remove ambiguous phrasing, then re-asks the rewritten version and returns that answer directly with no further rounds
  4. It feeds the model's own most recently generated answer back into the prompt as a 'hint' alongside the original question and re-asks, repeating this hint-and-reask loop across further rounds until two consecutive rounds produce the same answer, at which point that stable answer is treated as final
AI & LLM Engineering · Prompt Engineering · Card 058/081 easy

A team wants a zero-shot chain-of-thought prompting method that produces highly structured, easy-to-parse intermediate reasoning rather than free-flowing prose, and that does not require any hand-written multi-step exemplars in the prompt. Jin & Lu (2023), 'Tab-CoT: Zero-shot Tabular Chain of Thought,' propose prompting the model to organize its reasoning in a specific format. What is it?

  1. The model is prompted to lay out its reasoning as a markdown-style table, with a row per reasoning step and columns such as the step number, a subquestion for that step, the process or calculation used to answer it, and the resulting intermediate value, reaching the final answer only after the table is filled in
  2. The model is prompted to output its reasoning as a single JSON object containing only one field, the final answer, with no intermediate reasoning captured anywhere in the output
  3. The model is prompted to number each sentence of an otherwise free-form paragraph of reasoning sequentially, without imposing any column structure on the content
  4. The model is first given several worked, table-formatted examples as few-shot exemplars, and only then asked to fill in a table of its own for the new problem
AI & LLM Engineering · Prompt Engineering · Card 059/081 hard

A team has several different prompt templates for the same classification task, phrased slightly differently (for example, a yes/no question form versus an open-ended question form), and each template alone gives noisy, inconsistent predictions on held-out examples, with no single template reliably best across all inputs. Arora et al. (2022/2023), 'Ask Me Anything: A Simple Strategy for Prompting Language Models' (AMA), propose combining the noisy predictions from multiple such prompts into one final label. How do they do this, beyond a simple equally-weighted majority vote?

  1. They discard every template except whichever single one scores highest on a large labeled validation set, and use only that one template at test time from then on
  2. They train a large separate classifier on the LLM's internal hidden-state activations, and this classifier alone produces the final label, ignoring the individual prompts' answers entirely
  3. They treat each prompt template as a noisy 'weak label source' over the unlabeled test examples and combine the templates' predictions using a weak-supervision aggregation method, which models and corrects for each prompt's individual reliability and for correlations between prompts, rather than counting every template's vote equally
  4. They average the raw output token probabilities of every template's completion character by character, since averaging logits is mathematically equivalent to weak-supervision aggregation for a classification task
AI & LLM Engineering · Prompt Engineering · Card 060/081 easy

A team's prompts routinely include a lengthy few-shot exemplar set plus a long retrieved reference document, and they want to cut the number of tokens billed per request and reduce latency, without retraining or fine-tuning the target LLM, and while keeping task accuracy close to what the full, uncompressed prompt achieves. Jiang et al. (2023), 'LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models,' propose a method for this. What does it do?

  1. It replaces every word in the prompt with its shortest available synonym from a fixed thesaurus, on the theory that shorter synonyms always exist and always preserve the original meaning
  2. It uses a separate, small language model to estimate each token's perplexity given its context, then removes the lowest-information tokens under a target compression ratio, budgeting more aggressive compression for parts of the prompt such as exemplars while preserving instructions more conservatively
  3. It has the target LLM itself summarize the entire prompt into a single sentence, then discards the original text and sends only that one-sentence summary to the target model
  4. It strips only whitespace, line breaks, and punctuation from the prompt, since token count is otherwise fixed once the vocabulary of words in the prompt is decided
AI & LLM Engineering · Prompt Engineering · Card 061/081 medium

Min et al. (2022) found that replacing the correct labels in few-shot demonstrations with random incorrect ones barely hurts a large language model's in-context learning accuracy on many tasks, a result that is hard to square with the intuitive idea that the model is learning the exact input-output mapping from the examples the way a small supervised classifier would. Xie et al. (2021), 'An Explanation of In-context Learning as Implicit Bayesian Inference,' propose a theoretical account of why few-shot demonstrations help at all, developed independently of that later empirical finding. What is their explanation?

  1. The demonstrations work purely as a formatting cue: the model has already memorized the exact benchmark's answer key during pretraining, and the examples only signal which memorized answer set to output
  2. Each demonstration silently triggers a gradient-descent-like weight update inside the frozen model's forward computation, functionally equivalent to running a few steps of fine-tuning on those examples before generating the answer
  3. The examples work only because the model recognizes their surface formatting, such as delimiters and layout, and pattern-matches new inputs against that formatting alone, independent of the examples' actual content
  4. Because long, coherent pretraining documents implicitly require the model to infer a shared latent concept or topic connecting the text seen so far in order to predict the next token well, at inference time a prompt's demonstrations similarly signal a shared latent concept, and the model performs an implicit Bayesian inference over that latent concept from the examples and applies it to the query, even when the examples' exact input-output mapping is partly noisy
AI & LLM Engineering · Prompt Engineering · Card 062/081 medium

A team wants a single LLM to handle a whole benchmark of related but non-identical reasoning problems by selecting and combining general reasoning strategies (such as 'break the problem into subproblems' or 'think step by step') into a structure tailored to the task family, but without hand-designing that structure themselves and without paying the cost of an expensive multi-path search (such as exploring and scoring many branches) on every single problem instance. Zhou et al. (2024), 'Self-Discover: Large Language Models Self-Compose Reasoning Structures,' propose a way to do this. What is their two-stage approach?

  1. The model runs a full Tree-of-Thoughts search independently on every problem instance, exploring and scoring several candidate reasoning branches per instance and keeping whichever branch scores highest for that instance
  2. The model is fine-tuned once, using gradient updates on a labeled dataset of worked reasoning traces for the task family, to internalize a single fixed reasoning procedure it then applies to every instance
  3. In a one-time 'discover' stage, run only over a small number of unlabeled example problems from the task, the model selects, adapts, and composes a handful of generic atomic reasoning modules (e.g., critical thinking, decomposition into subtasks) into one explicit, task-specific reasoning structure; that same discovered structure is then reused to solve every individual problem in the task at ordinary single-pass inference cost, without repeating the discovery step per instance
  4. A separate, larger 'teacher' model first solves every problem in the benchmark and writes out its reasoning, and the target model is then prompted with these solved examples as few-shot demonstrations for each new instance
AI & LLM Engineering · Prompt Engineering · Card 063/081 hard

A team wants to automatically improve both a task-prompt (the instruction actually shown to the LLM for solving a target task) and the process used to generate new candidate task-prompts, across many rounds, using only a training set of labeled examples to score fitness and without any gradient-based training of the LLM's weights. Fernando et al. (2023), 'Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution,' propose a method for this that goes beyond single-pass approaches such as Automatic Prompt Engineer (APE), which generates one batch of candidate instructions from example input-output pairs and picks the best-scoring one. What does Promptbreeder do differently?

  1. It maintains an evolving population of task-prompts, using an LLM to repeatedly mutate them into new candidates across many generations and keeping variants that score best on a fitness set as in a genetic algorithm; crucially, the mutation-prompts that instruct the LLM how to mutate the task-prompts are themselves part of the evolving population and are improved over the same generations, making the whole process self-referential rather than a single-round generate-and-select step
  2. It fine-tunes a smaller auxiliary model on the training set's labeled examples to directly predict, in one shot, which single task-prompt from a fixed candidate list will score highest, then discards all other candidates
  3. It has a single frozen task-prompt debated word-by-word between two independent LLM instances until they reach unanimous agreement on its final wording, without ever generating new candidate prompts beyond edits proposed during the debate
  4. It searches over prompts using gradient-based backpropagation directly through the LLM's embedding layer to find a continuous soft-prompt vector, which is then rounded to the nearest discrete tokens for the final task-prompt
AI & LLM Engineering · Prompt Engineering · Card 064/081 easy

After a large language model generates an answer to an open-ended question, a team wants the same model, without any additional training and without access to any separately trained verifier, to estimate how likely it is that its own answer is actually correct. Kadavath et al. (2022), 'Language Models (Mostly) Know What They Know,' study prompting the model to do exactly this. What do they have the model do, and what do they find about larger models' resulting estimates?

  1. The model is asked to restate its own answer using different wording several times; if the reworded versions are lexically identical to the original, the model is deemed confident, and if they differ at all, the answer is discarded regardless of correctness
  2. The model is shown only the bare question again with no memory of its own prior answer and asked to guess whether a typical test-taker would find the question easy or hard, using that difficulty guess as a stand-in for confidence in its own specific answer
  3. The model is fine-tuned on a labeled dataset of past correct and incorrect answers so that a separate output head learns to predict correctness directly from the question text alone, without ever seeing the model's actual generated answer
  4. The model is prompted to look at the question together with its own proposed answer and output the probability, called P(True), that the proposed answer is correct; for larger models, this self-reported P(True) is reasonably well-calibrated against actual accuracy across many multiple-choice and true/false questions when the format is set up appropriately
AI & LLM Engineering · Prompt Engineering · Card 065/081 medium

A team deploys an LLM-based chatbot and wants a lightweight, training-free filter to flag inputs that may contain a Greedy Coordinate Gradient (GCG)-style adversarial suffix, the kind of algorithmically optimized nonsense-looking string appended to a harmful request to try to force the model to comply. Alon & Kamfonas (2023), 'Detecting Language Model Attacks with Perplexity,' propose using what signal for this, and what do they find about GCG suffixes compared to ordinary natural-language text?

  1. The cosine similarity between the embedding of the full input and the embeddings of a library of known jailbreak prompts; GCG suffixes are found to always land within a fixed similarity threshold of at least one known jailbreak
  2. The perplexity of the input text under a language model, i.e., how surprised the model is by the text on a token-by-token basis; GCG-optimized adversarial suffixes are found to have dramatically higher perplexity than ordinary natural-language text, because the optimization process searches for tokens that manipulate the target model's internals rather than tokens that read as fluent language
  3. The total character length of the input string alone; GCG suffixes are found to always exceed a fixed length threshold that ordinary user messages never reach
  4. Whether the input contains any token from a fixed manually curated blocklist of profanity and violence-related words; GCG suffixes are found to always include at least one such blocklisted token
AI & LLM Engineering · Prompt Engineering · Card 066/081 easy

A team wants GPT-4 to produce short, information-dense summaries of news articles, but plain single-pass summarization prompts tend to produce summaries that are fluent but omit many of the article's salient entities (names, numbers, organizations) in favor of generic language, while summaries that try to cram in every entity at once tend to read as unreadable lists. Adams et al. (2023), 'From Sparse to Dense: GPT-4 Summarization with Chain of Density Prompting,' propose a prompting procedure to balance this trade-off. What does Chain of Density (CoD) prompting do?

  1. It has the model generate five independent summaries of different fixed lengths in parallel, from very short to very long, and a human reviewer then manually picks whichever one seems most balanced
  2. It retrieves the five most similar previously-published human-written summaries from a database and averages their wording token-by-token to produce a new summary
  3. It has the model iteratively rewrite the same fixed-length summary across several rounds, identifying one to three additional salient, previously-missing entities from the article each round and fusing them into the existing summary without increasing its overall length, producing a progressively denser final summary
  4. It asks the model to list every named entity in the article first, then simply concatenates that raw entity list in front of an unrelated, separately generated generic summary
AI & LLM Engineering · Prompt Engineering · Card 067/081 easy

A team building a closed-book question-answering system (no external documents, search engine, or vector database available at inference time) wants to improve factual accuracy on questions that hinge on a specific passage the model was likely exposed to during pretraining, without adding any retrieval system. Sun et al. (2022), 'Recitation-Augmented Language Models,' propose RECITE, a two-step generation procedure for this. What does RECITE have the model do, and how does this differ from retrieval-augmented generation?

  1. The model first samples one or more relevant passages purely from its own parameters, by generating text it 'recites' as if quoting a source it was trained on, and then conditions on that self-generated recitation to produce its final answer; unlike retrieval-augmented generation, no external corpus or search index is queried at any point, since the 'retrieved' passage comes entirely from the model's own memory
  2. The model queries an external search engine to retrieve the single most relevant Wikipedia passage, then quotes that retrieved passage verbatim as its final answer without any further generation
  3. The model is fine-tuned to memorize the entire pretraining corpus word-for-word, so that at inference time it can always retrieve a verbatim passage through gradient-free nearest-neighbor lookup over its own weights
  4. The model generates several candidate answers independently and then recites, i.e., repeats, whichever candidate answer appears most frequently among the samples, without generating or conditioning on any intermediate passage
AI & LLM Engineering · Prompt Engineering · Card 068/081 medium

A team has a large pool of thousands of unlabeled questions for a new reasoning task and a limited budget for human annotators to write chain-of-thought exemplars, so they can only afford to hand-annotate a small handful of questions as few-shot demonstrations. Rather than picking that handful at random, Diao et al. (2023), 'Active Prompting with Chain-of-Thought for Large Language Models,' propose a selection procedure. What do they do to decide which questions from the pool are most worth sending to a human annotator?

  1. They select whichever questions are shortest in token count, reasoning that short questions are cheapest for human annotators to work through and annotate quickly
  2. They randomly sample an equal number of questions from every topic category in the pool, ensuring balanced topical coverage regardless of how difficult any individual question is for the model
  3. They fine-tune a separate small classifier model to predict each question's ground-truth answer directly, then send only the questions the classifier gets wrong to human annotators
  4. For each unlabeled question, they have the target LLM generate several chain-of-thought answers via repeated sampling and compute an uncertainty metric (such as disagreement among the sampled answers) for that question; questions where the model's several sampled answers disagree most are judged most uncertain and are prioritized for human annotation, on the reasoning that resolving the model's biggest uncertainties yields the most useful new exemplars
AI & LLM Engineering · Prompt Engineering · Card 069/081 easy

A team has only a handful of input-output example pairs demonstrating some task (for instance, pairs of a word and its plural form) and, instead of writing a few-shot prompt with those examples for a model to imitate, wants the model to explicitly state, in a single natural-language sentence, what the underlying task instruction actually is. Honovich et al. (2022), 'Instruction Induction: From Few Examples to Natural Language Task Descriptions,' study prompting language models to do this directly. What do they have the model do, and how do they evaluate whether the induced instruction is any good?

  1. They have the model produce more input-output pairs in the same style as the given examples, and judge the induction successful only if a human annotator cannot tell the generated pairs apart from the original ones
  2. They fine-tune the model on a large labeled corpus of (examples, correct-instruction) pairs so it learns to map any set of examples directly onto a matching instruction from that fixed labeled set
  3. They prompt the model with only the small set of input-output example pairs and ask it to generate, in natural language, the instruction that would produce those outputs from those inputs; they then evaluate a generated instruction by executing it, giving that instruction alone (with no examples) to a model and checking whether it correctly reproduces the original outputs on new inputs
  4. They ask the model to classify the task into one of a small fixed list of predefined categories (such as 'translation' or 'sentiment') rather than producing any free-form natural-language description of the task
AI & LLM Engineering · Prompt Engineering · Card 070/081 easy

A team wants an LLM's output to always be syntactically valid according to a specific formal grammar (for example, always a well-formed date in `YYYY-MM-DD` format, or always valid according to a custom domain-specific mini-language), with a hard guarantee rather than just a strong tendency, and without retraining the model or relying on the model to self-correct after the fact. Willard & Louf (2023), 'Efficient Guided Generation for Large Language Models,' propose a technique (implemented in the open-source Outlines library) for this. What does their approach do at each decoding step?

  1. It generates the full response first with no restriction, then runs a separate regular-expression validator afterward and asks the model to regenerate from scratch if the output fails validation, repeating until a valid output happens to appear
  2. It appends a natural-language instruction to the prompt describing the required grammar in words (such as 'please only output a valid date') and relies on the model's instruction-following ability alone to comply, with no mechanism enforcing the constraint at the token level
  3. It fine-tunes the model's weights on a large synthetic dataset of only grammar-valid outputs until the model has implicitly learned to always emit valid text on its own
  4. It builds an index, from a regular expression or context-free grammar, over the model's vocabulary that tracks which tokens are valid continuations at each position of the grammar (modeled as a finite-state machine or pushdown automaton); at every decoding step, this index is used to mask out the logits of any token that would make the output invalid according to the grammar, so only grammar-consistent tokens can ever be sampled, guaranteeing valid output structurally rather than by hope or after-the-fact correction
AI & LLM Engineering · Prompt Engineering · Card 071/081 hard

A team applies self-consistency (Wang et al., 2022) to a code-generation task by sampling several chain-of-thought solutions and picking the final answer that appears identically most often among the samples, and it works well. When they try the same majority-vote approach on an open-ended long-form summarization task, they find it breaks down, because no two of the several sampled full-length summaries are ever exactly identical, so a literal majority vote never finds any answer that 'wins.' Chen et al. (2023), 'Universal Self-Consistency for Large Language Model Generation,' propose Universal Self-Consistency (USC) to extend the technique to exactly this kind of task. What does USC do differently from the original self-consistency method's majority vote?

  1. It shortens every sampled summary down to a single keyword before comparing them, so that exact-match voting becomes possible again on the reduced representations
  2. It trains a separate learned reward model on human preference labels over pairs of summaries, and uses that reward model's score alone to pick the best sampled summary, without any involvement from the original LLM at this stage
  3. Instead of relying on exact-match voting over the final answers, it feeds all of the sampled candidate outputs back into the same LLM in a single prompt and asks that LLM itself to read across the candidates and select (or synthesize) whichever one is most consistent with the majority of the others, which works even when no two full free-form outputs are ever character-for-character identical
  4. It discards self-consistency's sampling step entirely and instead asks the model to produce just one single output per problem, deterministically, using greedy decoding
AI & LLM Engineering · Prompt Engineering · Card 072/081 easy

A team already follows Anthropic's advice to give Claude a role by writing a short, fixed description such as 'You are a seasoned data scientist' into the system prompt, and reuses that same sentence for every request the application handles. They want to go further: automatically tailoring a distinct, detailed expert identity for each new instruction the system receives, rather than reusing one static description every time. Xu et al. (2023), 'ExpertPrompting: Instructing Large Language Models to be Distinguished Experts,' propose a method for exactly this. How does ExpertPrompting generate the expert identity used for a given instruction, and how does that differ from simply assigning the same fixed role on every call?

  1. ExpertPrompting hand-writes a fixed library of a few dozen expert personas in advance (for example, 'senior software engineer,' 'tax attorney') and has the model pick the single closest-matching persona from that pre-written library for each new instruction, the same static-description approach as ordinary role prompting, just drawing from a longer list
  2. ExpertPrompting uses in-context learning to have the model itself write a new, detailed description of a specific expert identity customized to the particular instruction at hand, including a plausible background for that expert, and then conditions its actual answer on that freshly generated description, so the persona is synthesized per instruction rather than reused verbatim from one fixed sentence
  3. ExpertPrompting fine-tunes the base model's weights on a labeled dataset of expert personas so the model permanently role-plays as one single designated expert in every future conversation, with no further prompt engineering needed at inference time
  4. ExpertPrompting removes any description of a role or identity from the prompt entirely and instead infers an appropriate expert persona purely from statistical patterns in the user's own writing style, with no expert-identity text ever appearing in the prompt itself
AI & LLM Engineering · Prompt Engineering · Card 073/081 easy

A team's LLM-based assistant reads an externally sourced document, such as fetched web content, as part of its context, and the team is worried that if the document contains attacker-planted text phrased as an instruction, the model might carry it out as though the user had asked for it directly. Separately from the tool-permission and human-approval controls already covered by OWASP's general defense-in-depth guidance, Hines et al. (2024), 'Defending Against Indirect Prompt Injection Attacks With Spotlighting,' propose a prompting-level technique aimed specifically at this. What does spotlighting do to the untrusted document text itself, and why does this make the model less likely to treat it as an instruction?

  1. Spotlighting scans the document and deletes any sentence phrased as an imperative instruction before the document ever reaches the model's context window, so the model never sees any instruction-like phrasing at all
  2. Spotlighting routes the document to a separate, smaller classifier model trained specifically to detect injection attempts, and only forwards the document into the main model's context if that classifier scores it as safe
  3. Spotlighting translates the document into a different natural language from the rest of the prompt, relying on the main model performing worse in that language to prevent it from successfully following any instruction embedded in the translated text
  4. Spotlighting visibly transforms the untrusted text itself, for example by marking it with distinctive delimiters, interleaving it with marker characters, or encoding it (such as in base64), paired with a system instruction telling the model to treat anything appearing in that transformed form only as data to read rather than as instructions to follow, which makes injected imperative-sounding text stand out as untrusted instead of blending in with the model's genuine instructions
AI & LLM Engineering · Prompt Engineering · Card 074/081 medium

A team is building a few-shot chain-of-thought prompt for a new task and has a very large pool of unlabeled questions but only a small, fixed budget for human annotators to write out full worked exemplars by hand. They must decide, in advance and before any labels exist, which specific unlabeled questions are even worth sending to an annotator — a different problem from ranking already-labeled examples by similarity to one particular test input, or from ranking the model's own generated answers by uncertainty. Su et al. (2022), 'Selective Annotation Makes Language Models Better Few-Shot Learners,' propose a method called vote-k for exactly this unlabeled-pool selection step. How does vote-k decide which unlabeled examples get sent for annotation?

  1. Vote-k builds a graph over the unlabeled pool based on embedding-similarity neighborhoods, scores each candidate by how well-connected it is to many other examples in the pool, and discounts candidates that sit too close to examples already selected, so the resulting annotated set is both broadly representative of the whole pool and spread out rather than clustered in one region of it
  2. Vote-k sends every single example in the unlabeled pool to a human annotator up front, then has the language model vote on which of the resulting fully labeled examples to keep in the final few-shot prompt, discarding whichever receive the fewest votes
  3. Vote-k ranks unlabeled examples purely by their embedding similarity to the one specific test input currently being answered, selecting only the handful closest to that particular input for annotation every time a new test input is processed
  4. Vote-k measures the language model's output entropy across several answers it generates for each unlabeled example, without ever examining the pool's embedding structure, and sends only the highest-entropy examples off for annotation
AI & LLM Engineering · Prompt Engineering · Card 075/081 medium

A red-teaming exercise finds that a single, blunt request for clearly harmful content is refused outright, and so is a many-shot jailbreak attempt that stuffs hundreds of fabricated compliant-dialogue turns into one long prompt submitted all at once. But the team finds a different attack succeeds: across several separate, individually mild-looking conversation turns, each one explicitly referencing and building on the model's own prior reply, the conversation gradually arrives at the same harmful output that neither the one blunt request nor the single long fabricated-dialogue prompt could get. Russinovich, Salem & Eldan (2024), 'Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack,' describe this technique. How does Crescendo structure its attack across turns, and why does gradual escalation succeed where a single obviously harmful request does not?

  1. Crescendo submits one very long prompt, structurally identical to many-shot jailbreaking, except that the fabricated turns depict the assistant repeatedly refusing rather than complying, which the authors report paradoxically increases compliance on the model's real final turn
  2. Crescendo relies entirely on an algorithmically optimized, nonsense-looking suffix appended to a single harmful request, similar to the Greedy Coordinate Gradient method, rather than using multiple separate conversational turns at all
  3. Crescendo starts with a general, benign question related to the target topic and, across multiple separate turns, progressively escalates the request while explicitly referencing the model's own prior replies, adapting or substituting later turns depending on whether the model complied or refused at each step, so that by the final turn the model is nudged into continuing a trajectory it has effectively already been following rather than ever confronting one single, obviously harmful ask
  4. Crescendo has hundreds of different users each submit one small, innocuous-looking fragment of the harmful request independently in separate, unrelated conversations, then reassembles their separate outputs outside the model entirely, so no single conversation with the model ever contains more than one fragment
AI & LLM Engineering · Prompt Engineering · Card 076/081 hard

A team already uses Gao et al.'s PAL (Program-Aided Language Models) approach: the model writes a complete Python program for a word problem, and the entire program is handed to an external interpreter to execute, which reliably fixes arithmetic slips but only works when the whole problem can be expressed as code the interpreter can actually run end to end. They now hit tasks with sub-steps that have no well-defined executable semantics, such as a step inside an otherwise code-like procedure that says to check whether a sentence sounds sarcastic, where handing the full program to an external interpreter alone breaks down because the interpreter has no way to execute an undefined operation like that. Li et al. (2023), 'Chain of Code: Reasoning with a Language Model-Augmented Code Emulator,' propose an approach for exactly this. What does Chain of Code have the system do differently from handing a complete program to an external interpreter alone?

  1. Chain of Code abandons writing any code the moment it detects a step without well-defined executable semantics, falling back to standard natural-language chain-of-thought prompting for the entire problem instead of writing a program at all
  2. Chain of Code has the model write flexible pseudocode that freely mixes genuinely executable code with semantic sub-steps left loosely defined in natural language, then runs this through an interpreter that executes whichever parts it can and, whenever it reaches a sub-step it cannot resolve, hands that specific step to the language model itself to simulate the expected output as an 'LMulator,' before the interpreter resumes executing the rest of the program using that simulated result
  3. Chain of Code trains a dedicated external interpreter via supervised fine-tuning to recognize and directly execute informal natural-language instructions as if they were valid code, removing the language model from the execution loop entirely once that training is complete
  4. Chain of Code requires every sub-step in the program to be rewritten as strictly valid, fully executable code before any execution begins, rejecting any pseudocode step with undefined behavior rather than letting the interpreter hand such a step to the model
AI & LLM Engineering · Prompt Engineering · Card 077/081 easy

A team wants GPT-4V to answer questions that require pointing to one specific, precise region of an image (for example, 'what color is the object in the upper-left corner, not the one in the center') rather than just describing the image as a whole, and they want this to work in a zero-shot setting with no fine-tuning of the model. Yang et al. (2023), 'Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V,' propose a prompting method for this. What does Set-of-Mark (SoM) prompting do to the image before it is given to the model?

  1. It converts the image into a text caption produced by a separate captioning model and sends only that caption text to GPT-4V, since the paper's method assumes the model cannot process raw pixel regions directly
  2. It crops the image into a fixed grid of equally sized tiles and numbers each tile in reading order, regardless of what objects happen to fall inside any given tile
  3. It uses off-the-shelf interactive segmentation models to partition the image into regions at different granularities, then overlays visual marks, such as numbers, letters, masks, or boxes, on those regions, so the model can refer to a specific object by its mark instead of a vague spatial description
  4. It fine-tunes a lightweight adapter on top of GPT-4V's vision encoder using labeled bounding-box data so the model learns to ground references internally, without changing how the image itself is presented to it
AI & LLM Engineering · Prompt Engineering · Card 078/081 medium

A team is building a multi-step reasoning system that needs to pause partway through a problem to call an external tool, such as a calculator or a search API, and then continue reasoning using the tool's output, similar to what Yao et al.'s ReAct framework enables. Rather than hand-writing a new set of task-specific, interleaved reasoning-and-tool-use demonstrations for every new task the way ReAct's prompt author must, the team wants suitable demonstrations selected automatically based on tasks solved before. Paranjape et al. (2023), 'ART: Automatic Reasoning and Tool-use,' propose a framework for this. According to the paper, how does ART construct and run its prompt for a new task, and what specifically does it automate that ReAct leaves to a human?

  1. Given a new task, ART retrieves demonstrations of similar multi-step reasoning and tool use from a task library rather than requiring a human to hand-craft them, generates the reasoning-and-tool-call program for the new input, pauses generation whenever a tool is called, and resumes once the tool's output has been inserted
  2. ART eliminates the need for any demonstrations whatsoever, generating tool calls purely by looking up the task's name in a fixed table, whereas ReAct still requires the model to interleave reasoning and actions inside its own generated text
  3. ART fine-tunes a separate, small classifier model to predict which tool should be called at each reasoning step, replacing ReAct's prompting-only interleaving with a trained decision model
  4. ART requires an engineer to manually verify and hand-edit every retrieved demonstration before each run, trading ReAct's fully human-written demonstrations for human-reviewed ones instead of removing human involvement
AI & LLM Engineering · Prompt Engineering · Card 079/081 medium

A team's task is complex enough that Zhou et al.'s Least-to-Most prompting, which lists a sequence of simpler subquestions and answers them in order within one running prompt, still struggles on some of the sub-steps, particularly ones that are themselves hard to answer reliably in a single pass or that need to recurse on a much smaller version of the same overall problem. Khot et al. (2022), 'Decomposed Prompting: A Modular Approach for Solving Complex Tasks,' propose DECOMP for this. How does DECOMP structure its solution, and how does this differ from Least-to-Most's approach?

  1. DECOMP requires every subquestion to be generated and answered inside the same single prompt used by Least-to-Most, differing only in the order in which the subquestions are listed
  2. DECOMP generates every subquestion and its answer in one uninterrupted pass with no ability to call back into the overall decomposition procedure, whereas Least-to-Most allows the model to revisit and revise earlier subquestions
  3. DECOMP replaces prompting entirely with task-specific fine-tuning, training a separate specialized model for each sub-task instead of prompting a shared language model for every step, unlike Least-to-Most's purely prompting-based approach
  4. DECOMP delegates each sub-task to its own dedicated prompting-based handler drawn from a library of such handlers, rather than solving every subquestion inside one shared prompt the way Least-to-Most does, and a sub-task whose difficulty comes from a large input can recursively decompose further by invoking the same DECOMP procedure again on a smaller version of itself
AI & LLM Engineering · Prompt Engineering · Card 080/081 hard

A team wants to automatically improve a task prompt using only a training set of labeled examples and an LLM API, without any gradient-based training of the model's weights. They are already aware of Yang et al.'s OPRO, which treats an LLM as an optimizer weighing a running trajectory of past prompt-and-score pairs to propose better prompts, and of Fernando et al.'s Promptbreeder, which evolves a population of prompts through LLM-driven mutation and crossover operators scored against a training set. Pryzant et al. (2023), 'Automatic Prompt Optimization with “Gradient Descent” and Beam Search,' propose ProTeGi, a different mechanism for the same kind of problem. What does ProTeGi do?

  1. ProTeGi numerically differentiates the prompt's token embeddings with respect to a loss computed on the training set, taking literal gradient-descent steps in embedding space before decoding the result back into text
  2. ProTeGi runs the current prompt over a minibatch of training examples, has the LLM generate natural-language 'textual gradients' that criticize the specific ways the prompt's outputs failed, edits the prompt in the semantic direction opposite that criticism to produce candidate rewrites, and selects among the growing set of candidates using beam search combined with a bandit-style selection procedure
  3. ProTeGi asks the LLM to propose a single best replacement prompt directly from the training examples in one generation step, with no iterative criticism-and-edit loop and no search over multiple candidates
  4. ProTeGi mutates and recombines pairs of candidate prompts drawn from an evolving population using crossover and mutation operators scored against a fitness function, the same mechanism Promptbreeder uses
AI & LLM Engineering · Prompt Engineering · Card 081/081 easy

A team deploying a ChatGPT-based assistant wants a lightweight, purely prompting-level defense against jailbreak attempts, one that does not require architecture-level controls such as tool-permission scoping or human approval steps, and does not involve any additional training of the model. Xie et al. (2023), 'Defending ChatGPT against jailbreak attack via self-reminders,' published in Nature Machine Intelligence, propose a technique drawing on the psychological concept of a self-reminder. What does their system-mode self-reminder technique do, and what effect on jailbreak success rate did their experiments report?

  1. It inspects each user message for keywords associated with known jailbreak prompt templates and, if any are found, replaces the flagged span with a hidden untrusted-text marker before the model ever reads it, the same marking mechanism later generalized by spotlighting
  2. It fine-tunes the underlying model on a dataset of jailbreak attempts paired with refusals, so the defense is built into the model's weights rather than applied at prompting time
  3. It wraps the user's query within a system-level reminder, placed before and after the query, asking the model to respond responsibly and in accordance with guidelines; across their experiments this reduced the jailbreak success rate from 67.21% to 19.34%
  4. It sends the user's query to a separate, independently trained classifier model that scores the query for jailbreak intent before the primary model ever sees it, automatically blocking any query that scores above a fixed threshold