64 cards · AI & LLM Engineering · answer each one, then read the explanation. Your score tallies below. Looking for Temperature vs top-p vs top-k: how they combine (worked example)? Read the explainer.
0 / 64 answered · 0 correct
Link copied — send it to a friend!
AI & LLM Engineering · Building with LLM APIs · Card 001/064easy
When an LLM API's "function calling" (tool use) feature returns a function call in its response, what actually happens next in a typical integration?
AThe API executes the function directly on the vendor's infrastructure and returns only the final answer text, with no involvement from the calling application
BThe response contains a structured description of the function name and the arguments the model wants to pass; the calling application must run the actual function itself and send the result back in a follow-up request
CThe model runs the function inside its own weights using an internal code interpreter, producing the function's return value without any external execution step
DThe request fails with an error unless the function has already been executed and its output included as part of the original prompt
Correct answer: .
This is correct because standard function/tool calling returns a structured block naming the tool and the arguments the model wants to pass; the model itself has no way to execute arbitrary code or reach external systems, so the calling application is responsible for actually running the corresponding function and returning its result in a subsequent request so the model can incorporate it. The option describing vendor-side execution confuses this with built-in "server tools" (such as a hosted web-search tool) that some vendors run on their own infrastructure -- that is the exception, not the general mechanic of user-defined function calling. The option about the model running the function internally is wrong because a language model's weights encode learned associations, not a general-purpose interpreter capable of arbitrary code execution. The option requiring the function's output before the first request is backwards: the whole point of the feature is that the model requests the call first, before any result exists.
Source: Anthropic, 'Tool use with Claude' overview, https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview; OpenAI, 'Function calling' guide, https://platform.openai.com/docs/guides/function-calling
AI & LLM Engineering · Building with LLM APIs · Card 002/064easy
Many LLM provider APIs offer a "streaming" mode for text generation, typically implemented over Server-Sent Events (SSE). What does enabling streaming change about how a client receives the model's output?
AThe client receives the complete response in a single payload, but compressed with gzip to reduce total bandwidth compared to a non-streaming request
BThe model generates its answer in a single internal pass regardless of the setting; streaming only changes how the response is logged on the provider's servers, not how the client receives it
CThe client receives the response as a sequence of incremental chunks over an open connection as tokens are generated, rather than waiting for generation to finish before anything is returned
DThe client must poll a separate status endpoint at fixed intervals to check whether generation has completed, since no data is returned until the full response is ready
Correct answer: .
Streaming delivers partial output incrementally, as a sequence of small chunks sent as events over an open HTTP connection while the model is still generating, letting a client start displaying tokens (for example, in a chat UI) well before the full response completes, which lowers perceived latency for long responses. Gzip compression of a single complete payload is unrelated to streaming -- it is an orthogonal transport-level optimization that applies equally to non-streaming responses and does not change how many payloads are sent. Streaming very much changes what the client itself receives over the wire, not just how the provider logs the response server-side; the client's connection stays open and receives multiple discrete events in order. Polling a separate status endpoint describes an asynchronous batch-style pattern, which is a different integration model from token-by-token streaming and is typically used for long-running bulk jobs rather than conversational responses.
Source: OpenAI, API reference on streaming responses, https://platform.openai.com/docs/api-reference/streaming; Anthropic, Messages API streaming documentation, https://platform.claude.com/docs/en/build-with-claude/streaming
AI & LLM Engineering · Building with LLM APIs · Card 003/064easy
An LLM API's "context window" limit of, say, 200,000 tokens applies to what, specifically, during a single API call?
AThe combined total of the input tokens sent in the request (system prompt, conversation history, and any injected documents) plus the output tokens the model generates in response
BOnly the tokens in the user's most recent message, since earlier turns in the conversation are automatically summarized and don't count against the limit
COnly the tokens the model generates in its response, since input tokens are processed by a separate, effectively unlimited ingestion pipeline
DThe number of separate API requests a client can make per minute before being rate-limited
Correct answer: .
The context window is a hard ceiling on the total number of tokens the model can attend to in one call, and that total includes everything fed in as input (system prompt, prior conversation turns, retrieved documents, tool definitions) plus every token the model produces as output; if the combined total would exceed the limit, the request is rejected or there is no room left for the model to generate. Treating only the latest message as counted is wrong because unless a client deliberately truncates or summarizes prior turns itself before sending the request, full conversation history sent as input still counts toward the same shared budget -- there is no automatic summarization performed by the API. Treating only output tokens as counted is wrong because input processing consumes the same limited context, not a separate unlimited channel. Confusing this with a per-minute request cap conflates the context window with rate limiting, a completely separate constraint governing how many calls can be made over time, not how much content fits within one call.
AI & LLM Engineering · Building with LLM APIs · Card 004/064medium
A client application calling an LLM API in a tight loop starts receiving HTTP 429 responses. What is the generally recommended way to handle this, per common LLM provider API documentation?
AImmediately retry the exact same request as fast as possible in a loop, since 429 responses are transient and will resolve within milliseconds if retried aggressively
BSwitch to a different, unrelated API endpoint entirely, since a 429 on one endpoint indicates that the provider's entire platform is unavailable
CReduce the `max_tokens` parameter on the failing request, since 429 responses indicate the requested output would be too long to generate
DBack off and retry after a delay that increases with each subsequent failure (exponential backoff), typically with some added random jitter, rather than retrying immediately or at a fixed short interval
Correct answer: .
A 429 status code signals that the client has exceeded a rate or usage limit (such as requests per minute or tokens per minute), and the standard remediation across REST-style LLM provider APIs is exponential backoff: wait progressively longer between retries as failures continue, with random jitter added so that many concurrent clients don't all retry at the exact same moment and cause a new burst of failures. Retrying immediately and aggressively is the opposite of what is recommended, since it worsens the very condition causing the 429s rather than relieving it. A 429 on one endpoint reflects that specific endpoint's or account's usage limit, not a platform-wide outage, so switching to an unrelated endpoint does not address the underlying limit at all. Reducing `max_tokens` addresses a different failure mode entirely -- a request that would generate too much output or exceed a context-length limit -- not a rate-limit error, which is about request or token volume over time rather than the size of any single request.
AI & LLM Engineering · Building with LLM APIs · Card 005/064easy
How does a typical LLM provider's embeddings endpoint differ from its text-completion/chat-generation endpoint?
AThe embeddings endpoint is simply a faster version of the generation endpoint that returns shorter natural-language answers to save on output tokens
BThe embeddings endpoint takes text as input and returns a fixed-length numeric vector representing that text's meaning, rather than generating new natural-language text
CThe embeddings endpoint only works on images, while the generation endpoint only works on text, so the two cannot be used on the same type of content
DThe embeddings endpoint requires fine-tuning a custom model first, while the generation endpoint works with any base model out of the box
Correct answer: .
An embeddings endpoint converts input text into a fixed-length numeric vector that captures its semantic content, positioning similar meanings closer together in vector space; this is used for tasks like semantic search, clustering, and retrieval rather than producing readable prose, which is fundamentally different from a generation endpoint that autoregressively produces new natural-language tokens. Describing it as merely a "faster" generation endpoint is wrong because it does not generate text at all -- it returns numbers, not sentences. Restricting embeddings to images only is wrong, since provider embeddings endpoints primarily operate on text (some also support images), while the generation endpoint's core case is likewise text, so the two are not split along an image-versus-text line. Requiring fine-tuning first is wrong because base embeddings models are available for immediate use via the API just like base generation models, with fine-tuning being an optional enhancement for either endpoint type rather than a prerequisite for either one.
AI & LLM Engineering · Building with LLM APIs · Card 006/064easy
What does the `stop` (or "stop sequences") parameter available in most LLM generation APIs do?
AIt tells the API to stop generating further tokens as soon as any one of a specified list of strings appears in the output, and to exclude that string from the returned text
BIt sets a maximum wall-clock time limit in seconds after which the API forcibly terminates the connection regardless of how much text has been generated
CIt specifies a list of words the model is never permitted to generate anywhere in its response, causing an error if any of them would otherwise be produced
DIt pauses generation partway through and waits for the client to send an approval signal before continuing to generate the remainder of the response
Correct answer: .
Stop sequences let a caller specify one or more strings that, if generated, cause the API to immediately halt generation at that point and return the output produced so far, with the matched stop string itself typically omitted from the returned text; this is commonly used to cut a response off at a natural boundary, such as before a model would start role-playing the next turn of a conversation. A wall-clock timeout is a distinct, separate control unrelated to matching specific text content, and would cut a response off regardless of what was actually generated. A blanket prohibition on ever producing certain words describes content filtering or constrained decoding, not what stop sequences do, since stop sequences only trigger termination once a match completes rather than preventing the string from ever being produced. Pausing for manual approval mid-stream describes an entirely different, human-in-the-loop interaction pattern that generation APIs do not implement via this parameter.
Source: OpenAI, Chat Completions API reference, `stop` parameter, https://platform.openai.com/docs/api-reference/chat/create; Anthropic, Messages API reference, `stop_sequences` parameter, https://platform.claude.com/docs/en/api/messages
AI & LLM Engineering · Building with LLM APIs · Card 007/064hard
A developer is using Anthropic's Messages API with two tools defined, `get_weather` and `send_email`, and wants to guarantee that Claude calls `get_weather` specifically on this turn rather than answering in plain text or calling `send_email`. According to Anthropic's tool-use documentation, how is this accomplished?
ABy removing `send_email` from the `tools` array entirely for this request, since `tool_choice` can only force "any" tool use, not a specific named tool
BBy setting `tool_choice` to `{"type": "any"}`, which restricts the model to only the first tool listed in the `tools` array
CBy setting `tool_choice` to `{"type": "tool", "name": "get_weather"}`, which forces the model to call that specific named tool rather than deciding on its own or calling a different one
DBy adding the instruction "you must call get_weather" only to the system prompt, since `tool_choice` itself has no mechanism for naming a specific tool
Correct answer: .
Anthropic's Messages API exposes a `tool_choice` parameter with four possible values: `auto` (the default, letting the model decide whether and which tool to call), `any` (forcing some tool call but not a particular one), `none` (preventing tool use), and `tool` -- and setting `tool_choice` to `{"type": "tool", "name": "get_weather"}` specifically forces that named tool to be called on this turn, which is exactly the guarantee the developer wants. Removing the other tool from the array would only prevent it from being offered at all, and does not actually force a call, since the model could still decline to call the only remaining tool under the default `auto` behavior. The `any` value forces some tool call, but not a specific one -- it does not default to "the first tool listed"; the model still chooses among whichever tools are available to it. Relying only on a system-prompt instruction is a weaker, unenforced nudge that Anthropic's own documentation contrasts with the `tool_choice` mechanism, which structurally guarantees the behavior rather than merely encouraging it through prompting.
AI & LLM Engineering · Building with LLM APIs · Card 008/064medium
A developer building a multi-turn conversational app compares OpenAI's Chat Completions API to its Responses API. According to OpenAI's documentation, what is a key difference in how each manages conversation state across turns?
AChat Completions automatically stores and threads every conversation server-side with no client involvement, while the Responses API requires the client to resend the entire message history on every call
BBoth APIs require the exact same manual approach: the client must always reconstruct and resend the full list of prior user and assistant messages with every request, with no built-in alternative in either API
CThe Responses API has no way to maintain multi-turn context at all, and is only suitable for single-turn, stateless requests unrelated to any previous exchange
DWith Chat Completions, the client must append prior turns into the message array and resend the full history each call, while the Responses API can instead reference a prior turn via a `previous_response_id` parameter so the server carries the context forward
Correct answer: .
OpenAI's documentation describes Chat Completions as requiring the client to manually reconstruct conversation state by including the model's previous outputs as part of the input and appending that to each new request's message array, whereas the Responses API offers a `previous_response_id` parameter as a way to chain a new request to an earlier response so the server-side context is carried forward without the client re-sending everything itself. The first option reverses which API requires manual resending -- it is Chat Completions, not the Responses API, that lacks this shortcut. The option claiming both APIs are purely manual with no alternative ignores the Responses API's `previous_response_id` mechanism entirely. The option claiming the Responses API cannot maintain multi-turn context at all is wrong because chaining via `previous_response_id`, along with the related Conversations API for longer-running threads, is specifically built for exactly that purpose.
AI & LLM Engineering · Building with LLM APIs · Card 009/064medium
An LLM-powered agent has a tool that charges a customer's payment method, and the agent's HTTP call to your backend times out after the charge has actually already been processed. The agent's retry logic then calls the same tool again with the same arguments. What design choice prevents this from resulting in a duplicate charge?
AHaving the tool accept a unique idempotency key per logical operation, so the backend can recognize a retried call with the same key and return the original result instead of processing the charge a second time
BIncreasing the model's `max_tokens` limit, so the agent has enough space to reason more carefully about whether a retry is safe before calling the tool again
CLowering the model's `temperature` to 0, so the agent always generates the exact same tool call arguments and therefore never issues an unintended duplicate request
DDisabling the agent's ability to call tools more than once per conversation, so any repeated call is rejected outright regardless of what happened to the first one
Correct answer: .
The generally correct fix is at the tool and backend layer, not the model layer: attaching a unique idempotency key to each logical operation lets the backend detect that a retried request represents the same intended operation and short-circuit to returning the original result rather than executing the side effect again, which is the standard pattern for exactly this kind of network-timeout-then-retry scenario in payment and other side-effecting APIs, and is the recommended orchestration-layer practice for agent tool calls that can be retried. Increasing `max_tokens` only affects how much text the model can generate and has no bearing on whether a backend call is executed twice. Lowering temperature to 0 only makes the model's chosen arguments more deterministic across similar prompts; it does nothing to prevent the exact same request from actually being re-sent and re-executed after a timeout, since the retry happens regardless of how deterministic the arguments were. Blanket-disabling repeat calls to any tool would also break legitimate cases where a tool is meant to be called multiple times with different arguments in the same conversation, making it too blunt a fix for this specific failure mode.
Source: OpenAI, 'In production' guide for Agentic Commerce (idempotency keys for retried tool/side-effect calls), https://developers.openai.com/commerce/guides/production
AI & LLM Engineering · Building with LLM APIs · Card 010/064easy
Some LLM APIs support returning `logprobs` (log probabilities) alongside generated text. What do these values represent, and what are they typically used for?
AThey report how many milliseconds the API took to generate each token, and are used purely for latency monitoring and performance debugging
BThey report, for each generated token, the log probability the model assigned to it (and often to alternative candidate tokens), and are typically used to gauge the model's confidence or build classifiers from token likelihoods
CThey report the total dollar cost billed for each individual token, broken out on a per-token basis for detailed cost accounting
DThey report which of several fine-tuned model versions actually generated each token, for use in auditing which checkpoint produced a given response
Correct answer: .
Log probabilities express, on a log scale, how likely the model considered each token it actually output (and optionally some number of alternative top candidate tokens at that position), which is useful for gauging the model's confidence in its own output, flagging likely hallucinations where confidence is unusually low, or building lightweight classifiers on top of a generation task by comparing the relative log probabilities of a small set of candidate answers. They have nothing to do with latency, which providers report separately, if at all, as timing metadata rather than as a probability value. They also are not a per-token cost breakdown; billing is typically reported as aggregate input and output token counts rather than a probability-shaped field attached to each token. And they do not identify which fine-tuned checkpoint generated a token -- that information, if available at all, would be reported as separate model or version metadata on the response, not encoded in a probability value.
Source: OpenAI, Chat Completions API reference, `logprobs`/`top_logprobs` parameters, https://platform.openai.com/docs/api-reference/chat/create
AI & LLM Engineering · Building with LLM APIs · Card 011/064medium
An LLM generation API exposes both a `top_p` (nucleus sampling) parameter and, in some APIs, a `top_k` parameter, in addition to `temperature`. How does nucleus sampling with `top_p` differ from `top_k` sampling in how it restricts the model's next-token choices?
A`top_p` and `top_k` are two names for the exact same mechanism, differing only in whether the cutoff value is expressed as a percentage or as a raw integer
B`top_k` restricts choices based on cumulative probability mass, so its cutoff point moves depending on how confident the model is at each step, while `top_p` always keeps exactly the same fixed number of candidates
C`top_p` restricts sampling to the smallest set of most-probable tokens whose cumulative probability reaches the threshold `p`, so the number of candidates varies step to step, while `top_k` always keeps a fixed number of the highest-probability tokens regardless of how the probability mass is distributed
DBoth parameters only take effect when `temperature` is set to exactly 0, and have no effect on generation at any other temperature value
Correct answer: .
Nucleus sampling (`top_p`) builds a candidate pool by adding tokens in order of probability until their cumulative probability reaches the threshold `p`, so a peaked distribution yields a small pool and a flatter distribution yields a larger one -- the pool size adapts to the model's confidence at each generation step. `top_k` sampling instead always keeps exactly the same fixed number, `k`, of the highest-probability tokens as candidates regardless of how concentrated or spread out the probability mass is, which can include unlikely tokens when the true distribution is sharply peaked, or exclude reasonable ones when it is flat. The two are not the same mechanism expressed in different units, since one produces a variable-size candidate pool and the other a fixed-size one. The option describing `top_k` as the variable, confidence-dependent one and `top_p` as fixed reverses their actual behavior. Both parameters shape the candidate pool independently of `temperature` and take effect at any temperature setting, not only at exactly 0, where sampling becomes fully deterministic and such cutoffs become moot only in the trivial sense that a single token already dominates the distribution.
Source: OpenAI, API reference `top_p` parameter documentation, https://platform.openai.com/docs/api-reference/chat/create; Holtzman et al. (2020), 'The Curious Case of Neural Text Degeneration' (arXiv:1904.09751), introducing nucleus sampling
AI & LLM Engineering · Building with LLM APIs · Card 012/064hard
A team needs to run sentiment classification over 40,000 archived support tickets and is not latency-sensitive, but wants to minimize per-request cost. According to OpenAI's Batch API documentation, what tradeoff does using the Batch API (instead of the standard synchronous chat completions endpoint) involve?
AThe Batch API charges the same per-token price as the synchronous API but guarantees results within 60 seconds regardless of batch size
BThe Batch API is free of charge for any volume of requests, but results are only available after a mandatory 7-day waiting period
CThe Batch API only accepts a single request per batch, so the team would still need to submit 40,000 separate batch jobs to process all the tickets
DThe Batch API offers roughly a 50% cost discount compared to the synchronous API, but processes the submitted batch of requests asynchronously with results typically available within 24 hours rather than immediately, and draws from a separate rate-limit pool
Correct answer: .
OpenAI documents the Batch API as accepting a file of many requests, each with its own identifier for matching results back to inputs, that is processed asynchronously against a separate rate-limit pool, at roughly a 50% cost discount relative to the equivalent synchronous requests, with completion typically well within the fixed 24-hour completion window rather than returned immediately -- a tradeoff of latency for cost that fits exactly the non-latency-sensitive bulk classification use case described. The claim of same price but a guaranteed 60-second turnaround is wrong on both counts: the price is discounted relative to the synchronous endpoint, and results are not guaranteed within any such short window. The claim that it is entirely free with a mandatory week-long wait invents figures the documentation does not state; the actual discount is roughly half price, not zero, and the completion window is 24 hours, not 7 days. The claim of a one-request-per-batch limit is backwards: a single batch file is specifically built to bundle a large number of requests, documented up to tens of thousands, into one submission, which is the entire point of the feature for a use case like this one.
Source: OpenAI, Batch API guide, https://developers.openai.com/api/docs/guides/batch
AI & LLM Engineering · Building with LLM APIs · Card 013/064medium
A developer building a chat feature needs the model's reply to reliably parse as a JSON object matching a specific schema, with no missing required fields and no invalid enum values. According to OpenAI's documentation, how do its `json_object` response format and its `json_schema` Structured Outputs mode (with `strict` set to true) differ in what they actually guarantee?
A`response_format` `json_object` mode and `json_schema` mode with `strict` enabled provide the exact same guarantee: both ensure the output matches a developer-supplied schema exactly, including all required fields and valid enum values
B`json_object` mode guarantees the output matches the developer's schema exactly, while `json_schema` mode with `strict` enabled guarantees only that the output is syntactically valid JSON, with no schema enforcement
C`json_object` mode guarantees only that the output is syntactically valid JSON, with no guarantee it matches any particular structure, while `json_schema` mode with `strict` set to true constrains generation so the model reliably includes every required field and only valid enum values as defined in the developer's supplied schema
Dboth modes only take effect when `temperature` is also set to 0; without that, neither has any influence on whether the output is valid JSON or matches a schema
Correct answer: .
The option correctly describing json_object mode is right because OpenAI's documentation distinguishes legacy JSON mode, which only ensures the output parses as syntactically valid JSON without validating its shape against any schema, from strict json_schema Structured Outputs, which constrains the model's generation process itself so that a required key is never omitted and an enum field never receives a value outside the allowed set. The option claiming both modes provide the identical guarantee ignores this distinction and would leave a developer unprotected against missing fields if they mistakenly relied on plain JSON mode for schema conformance. The option that reverses which mode enforces the schema gets the two backwards: it is the strict json_schema mode, not plain JSON mode, that enforces structure. The option tying both modes to a specific temperature setting invents a dependency that does not exist -- output format and schema enforcement operate independently of temperature, which only affects sampling randomness among otherwise valid completions.
AI & LLM Engineering · Building with LLM APIs · Card 014/064easy
A developer comparing Anthropic's Messages API to some other chat-style LLM APIs notices a structural difference in how the system prompt is supplied. How does Anthropic's Messages API handle the system prompt?
AAnthropic's Messages API accepts a top-level `system` parameter, kept separate from the `messages` array, to carry the system prompt, whereas some other chat-style APIs instead embed the system prompt as a message with a system or developer role placed inside the same messages list
BAnthropic requires the system prompt to be the first entry in the `messages` array with a role of `system`, exactly matching how some other chat-style APIs structure a request
CAnthropic's Messages API has no mechanism for supplying a system prompt at all; instructions can only be embedded inside the first user message
Dthe `system` parameter selects which dated model snapshot handles the request, and has nothing to do with supplying instructions to the model
Correct answer: .
Anthropic's Messages API takes a dedicated top-level `system` parameter that sits alongside, rather than inside, the `messages` array, which itself alternates only `user` and `assistant` roles; this is a genuine structural difference from some other chat-style APIs, whose request format instead includes the system prompt as just another entry in a single flat list of role-tagged messages. The option describing a `messages`-array system entry describes that other pattern, not Anthropic's actual `system`/`messages` split, so it is wrong for the API being asked about. The option claiming no system-prompt mechanism exists is wrong because the dedicated `system` parameter is exactly that mechanism, and is extensively documented and supported. The option describing `system` as a model-selection parameter confuses it with the separate `model` field in the request body, which is what actually selects the model version, not `system`.
Source: Anthropic, Messages API reference, https://platform.claude.com/docs/en/api/messages
AI & LLM Engineering · Building with LLM APIs · Card 015/064hard
A developer calls a reasoning model through OpenAI's Responses API with `max_output_tokens` set, and gets back a response whose `output_text` is empty even though the usage data shows a large number of billed output tokens. According to OpenAI's documentation, what is most likely happening, and how should the application detect it?
Athe response's `status` field will be `error` with an `error.type` of `server_error`, indicating an internal fault on OpenAI's infrastructure that is unrelated to the token limit set on the request
Bthis cannot happen when `max_output_tokens` is set, since OpenAI's documentation guarantees a visible answer is always produced in full before any internal reasoning tokens are counted against the limit
Cthe response's `output` array will always still contain a text message item in this situation, so an application never needs to check anything beyond `output_text` before using the result
Dthe response's `status` field will be `incomplete`, with `incomplete_details.reason` set to `max_output_tokens`, because a reasoning model's internal reasoning tokens can consume the entire token budget before any visible answer is produced; an application should check `status` and `incomplete_details` rather than assuming `output_text` is populated, and should reserve enough budget for both reasoning and the visible answer
Correct answer: .
OpenAI's documentation on reasoning models describes exactly this failure mode: because a reasoning model generates internal reasoning tokens in addition to the visible answer, and both draw from the same `max_output_tokens` budget, the model can exhaust the entire budget on reasoning before producing any visible text, leaving `output_text` empty despite non-zero billed usage; the documented signal for this is a response `status` of `incomplete` together with `incomplete_details.reason` equal to `max_output_tokens`, and the fix is reserving a larger token budget. Describing this as a `server_error` is wrong because nothing about the infrastructure has failed -- the request completed, just without reaching a visible answer, which is a documented, expected outcome rather than a fault. The option claiming this can never happen contradicts the documented behavior directly, since the guarantee it describes does not exist. The option asserting a text item is always present is wrong because the whole documented failure mode is that no message item, and therefore no `output_text`, may be emitted when reasoning consumes the full budget.
AI & LLM Engineering · Building with LLM APIs · Card 016/064easy
LLM generation APIs commonly expose a `temperature` parameter alongside sampling-restriction parameters like `top_p` or `top_k`. Mechanically, what does `temperature` itself do to the model's next-token choice?
Atemperature removes the least probable tokens from consideration entirely, functioning identically to `top_k` sampling with the temperature value acting as the value of k
Btemperature scales the model's output logits before they are converted into probabilities via softmax; lower values sharpen the resulting distribution toward the highest-probability tokens, while higher values flatten it, making comparatively less probable tokens more likely to be sampled
Ctemperature is applied only after a token has already been sampled, triggering a re-roll whenever a separate filter model flags the chosen token as low quality
Dtemperature has no effect on which token is sampled and instead controls only how many separate completions the API returns for a single request
Correct answer: .
Temperature is a scaling factor applied to the model's raw output logits before the softmax function converts them into a probability distribution over the vocabulary; dividing logits by a small temperature value sharpens the distribution so the highest-probability tokens dominate even further, approaching a near-deterministic choice as temperature approaches zero, while dividing by a larger value flattens the distribution, giving lower-probability tokens a comparatively higher chance of being sampled. Describing temperature as functionally identical to `top_k` is wrong because `top_k` restricts the candidate pool to a fixed count of the highest-probability tokens rather than reshaping the probability distribution itself, which is a distinct mechanism. Describing temperature as a post-hoc quality filter invents a re-rolling mechanism that generation APIs do not implement as part of this parameter. Describing temperature as controlling the number of returned completions confuses it with an entirely separate parameter, such as a count of choices to generate, that some APIs expose independently.
Source: OpenAI, Chat Completions API reference, `temperature` parameter, https://platform.openai.com/docs/api-reference/chat/create
AI & LLM Engineering · Building with LLM APIs · Card 017/064easy
A developer wants to send an image alongside a text question to a model through OpenAI's Responses API. According to OpenAI's current documentation, how is the image included in the request?
Athe request's `input` array includes a content item with `type` set to `input_image`, whose `image_url` field can be either a fully qualified URL pointing to an image or a base64-encoded data URL, placed alongside a separate `input_text` item carrying the text portion of the same message
Bimages can only be referenced by a fully qualified URL; base64-encoded image data placed directly in the request body is rejected by the endpoint
Cimages must be attached as raw binary data outside the JSON request body, since the `input` array can only ever contain plain text strings
Dimages are supplied through the `messages` array using a content type of `image_url` only, with no `input_image` type and no base64 option available
Correct answer: .
OpenAI's current documentation for the Responses API shows a content item of type `input_image` inside the `input` array's message content, alongside an `input_text` item for the accompanying question, and the `image_url` field on that item accepts either a plain URL pointing to a hosted image or a base64-encoded data URL, giving developers a choice depending on whether the image is already hosted somewhere accessible. Restricting the field to URLs only is wrong because the documented base64 data-URL form is explicitly supported as an alternative. Claiming images can only be sent as raw binary data outside the JSON body is wrong because the documented mechanism keeps the image reference inside the JSON `input` array like any other content item. Describing the content type as `image_url` inside a `messages` array describes an older, different API shape and gets both the array name and the content type's `type` string wrong for the API actually being asked about.
Source: OpenAI, Images and vision guide, https://developers.openai.com/api/docs/guides/images-vision
AI & LLM Engineering · Building with LLM APIs · Card 018/064hard
A developer marks a `cache_control` breakpoint on a large, static system prompt sent with every request to Anthropic's Messages API. According to Anthropic's documentation, how does this affect billing across the first request and later requests that reuse the same cached prefix?
Athe first request's cache-write tokens are billed at a discount below the base input-token rate, while every later request that hits the same cache is billed at a premium above the base rate, since maintaining a cache costs more than reading fresh input
Bcache writes and cache hits are always billed at exactly the same per-token rate as ordinary, uncached input tokens, so prompt caching only ever provides a latency benefit and never a cost benefit
Cthe first request that establishes the cache is billed at a premium above the base input-token rate for the cached portion, since writing a new cache entry costs more, while later requests whose prefix hash matches that cached entry are billed at a steep discount for the reused portion, which is where the overall cost savings come from
Dprompt caching only ever reduces the cost of output tokens, never input tokens, because it works by shortening how much text the model generates in its response rather than by reusing previously processed input
Correct answer: .
Anthropic's documented pricing structure charges a premium multiplier, above the base input-token rate, for tokens written into a new cache entry on a cache miss, and a steep discount, well below the base rate, for tokens read from a matching cache entry on a subsequent cache hit -- so the first request that establishes the cache actually costs more than an equivalent uncached request, and the savings only materialize on later requests whose prefix hash matches what was already cached. The option reversing which side is discounted and which is a premium gets the direction backwards relative to the documented pricing tiers. The option claiming no price difference exists ignores the documented multiplier structure entirely and would leave a developer unable to explain why caching is recommended as a cost-saving feature at all. The option tying caching to output-token cost is wrong because prompt caching operates entirely on the input side, reusing previously processed prompt content, and has no mechanism for shortening generated output.
AI & LLM Engineering · Building with LLM APIs · Card 019/064medium
A developer needs Claude to analyze both the text and the visual layout, charts, and images on each page of a PDF report, not just extract raw text. According to Anthropic's documentation, how is a PDF supplied to the Messages API?
APDFs cannot be sent as input to the Messages API at all; the developer's own application must first convert the PDF to plain text before any part of it can be included in a request
Ba message's content array can include a block of type `document`, whose source can be a base64-encoded PDF, a URL pointing to a hosted PDF, or a `file_id` from the Files API, letting the model reason over both the extracted text and each page's visual layout, charts, and images
CPDFs are supported only via a separate proprietary file format that the developer must first produce from the PDF using a vendor-provided offline conversion tool before uploading
Dthe `document` content block only extracts and returns the PDF's raw text back to the developer as a standalone response; it cannot be used as part of a prompt for the model to reason over
Correct answer: .
Anthropic's documentation describes a `document` content block that accepts a PDF either as base64-encoded data, as a URL referencing a hosted PDF, or as a `file_id` obtained from the Files API, and states that because PDF support relies on Claude's vision capabilities, each page is processed as both text and image, so the model can reason over charts, tables, and layout in addition to extracted text rather than text alone. The option claiming PDFs cannot be sent at all is directly contradicted by this documented feature. The option requiring a proprietary offline conversion format invents a step that the documented base64, URL, and Files API options make unnecessary. The option describing the block as only returning raw text back to the developer misdescribes it as an extraction utility rather than as prompt content the model itself reasons over within the same request.
Source: Anthropic, PDF support documentation, https://platform.claude.com/docs/en/build-with-claude/pdf-support
AI & LLM Engineering · Building with LLM APIs · Card 020/064medium
An agent built on OpenAI's Responses API needs to look up two independent pieces of information in a single turn, using the same tool twice with different arguments. According to OpenAI's documentation, how does the API support this, and what must the client do in response?
Aa single assistant turn can never include more than one function call; the model must wait for the client to answer the first call before it is permitted to request the second one
Bthe client may return one combined `function_call_output` message covering every function call made in that turn, as long as the message lists all of the relevant call IDs together
C`parallel_tool_calls` can only ever be set to true, since the API provides no way to restrict a response to at most one function call per turn
Da single assistant turn can include multiple function calls, each carrying its own `call_id`, and the client must send back one `function_call_output` item per `call_id` containing that call's result; setting `parallel_tool_calls` to false instead restricts the model to at most one function call per turn
Correct answer: .
OpenAI's documentation states that the model may call multiple functions in a single turn, with the response's output array containing a separate entry, each with its own `call_id`, for every distinct function invocation, and that the client must process each one independently and return a separate `function_call_output` message per `call_id` rather than merging them; it also documents `parallel_tool_calls` as a request parameter that, when set to false, restricts the model to exactly zero or one function call per turn. The option denying multiple calls per turn contradicts the documented multi-call output shape directly. The option allowing a single combined output message is wrong because the documented pattern matches results back to individual call IDs one at a time, not as a bundled response. The option claiming `parallel_tool_calls` cannot be disabled is wrong because the documented false setting exists specifically to enforce a single-call-per-turn restriction.
Source: OpenAI, Function calling guide, https://developers.openai.com/api/docs/guides/function-calling
AI & LLM Engineering · Building with LLM APIs · Card 021/064easy
Every request to Anthropic's Messages API must include an `anthropic-version` header, such as `anthropic-version: 2023-06-01`. According to Anthropic's documentation, what does this header actually do?
Athe header is purely optional and informational; omitting it, or sending an arbitrary unrecognized string, has no effect on how the request is processed
Bthe header selects which specific model snapshot answers the request, serving as an alternative to setting the `model` field in the request body
Cthe header is required on every request and pins that request to a documented API version; Anthropic's versioning policy preserves a given version's existing input and output parameters while allowing additive changes, such as new optional inputs or new output values, so pinning a version protects an integration from breaking changes introduced by later versions
Dthe header's value must change to a new date every single day, since each calendar day's API responses use an incompatible request and response format from the previous day
Correct answer: .
Anthropic's documentation states that the `anthropic-version` header is required on requests and that, for a given version, Anthropic preserves existing input and output parameters while reserving the right to make additive changes such as new optional inputs, new output values, or new enum-like variants -- so an integration that pins a specific version string is protected from having its existing request and response handling broken by later additive changes. Describing the header as optional and inert is wrong because it is a required header with defined semantics, not a no-op. Describing it as a model-selection mechanism confuses it with the separate `model` field, which is what actually determines which model answers the request. Claiming the version string must change daily invents a versioning cadence that contradicts the documented version history, where a single version string like `2023-06-01` remains valid and stable over long periods.
Source: Anthropic, API versioning documentation, https://platform.claude.com/docs/en/api/versioning
AI & LLM Engineering · Building with LLM APIs · Card 022/064easy
Anthropic's Messages API can return either a 429 status with a `rate_limit_error` type or a 529 status with an `overloaded_error` type. According to Anthropic's documentation, what is the practical difference between these two conditions?
Aa 429 with a `rate_limit_error` type means the calling organization's own limits, such as a requests-per-minute cap or a configured spend limit, have been exceeded, while a 529 with an `overloaded_error` type means the API itself is temporarily overloaded across all users, a capacity condition unrelated to that organization's own usage
B429 and 529 both describe exactly the same underlying condition, and application error-handling code does not need to distinguish between them
Ca 529 status means the calling application's API key has been revoked, while a 429 status means the request body was malformed and rejected by validation
Da 429 status can only ever be returned by the Batch API, while a 529 status can only ever be returned by the standard synchronous Messages API, so the two are distinguished solely by which endpoint returned them
Correct answer: .
Anthropic's documented error reference ties `rate_limit_error` specifically to the 429 status, describing it as the organization having hit its own rate limit, usage-tier spend cap, or a workspace-specific spend limit, while it ties `overloaded_error` specifically to the 529 status and explains that this happens when the API experiences high traffic across all users, a shared-capacity condition rather than anything the calling organization did. Treating the two as interchangeable is wrong because they point to different root causes with different appropriate responses, even though both can warrant retrying with backoff. Describing 529 as a revoked key and 429 as malformed input is wrong because those conditions are documented separately as 401 `authentication_error` and 400 `invalid_request_error` respectively, not as 429 or 529. Restricting each code to a single specific endpoint is wrong because both status codes are documented as general API-wide conditions, not endpoint-specific ones.
Source: Anthropic, API errors documentation, https://platform.claude.com/docs/en/api/errors
AI & LLM Engineering · Building with LLM APIs · Card 023/064medium
Anthropic's Messages API exposes a separate `/v1/messages/count_tokens` endpoint. According to Anthropic's documentation, what does this endpoint let a developer do, and what does calling it cost?
Ait returns the exact monetary cost, in the organization's billing currency, of a hypothetical request, in addition to a token count
Bcalling it consumes the same input-token allowance and counts fully against the same per-minute rate limit as an ordinary message request, because it internally runs the full generation pipeline
Cit can count tokens only in plain text messages, and cannot account for the additional tokens that tool definitions or images would add to an actual request
Dit lets a developer estimate how many tokens a prospective request would consume, including messages, tool definitions, and images, without actually creating a message or generating any output, which is useful for pre-flight cost estimation and for checking that a request will fit inside the model's context window
Correct answer: .
Anthropic's documentation describes the Token Count API as a way to count the tokens a Message would consume, including tools, images, and documents, without actually creating that Message -- meaning no output is generated and the call is meant specifically for pre-flight estimation, such as confirming a request stays within a context-window budget before sending it for real. Claiming it returns an exact monetary cost is wrong because the documented response is a token count, not a currency-denominated price, even though a developer could separately multiply that count by a known per-token rate. Claiming it costs the same as a full generation request is wrong because the entire documented purpose of the endpoint is to estimate usage cheaply without running generation. Claiming it ignores tools and images is wrong because the documentation explicitly states that tools, images, and documents are included in the count, not just plain text.
Source: Anthropic, Count tokens API reference, https://platform.claude.com/docs/en/api/messages-count-tokens
AI & LLM Engineering · Building with LLM APIs · Card 024/064easy
OpenAI offers both a plain model name (an alias, such as a model's short identifier) and a dated snapshot identifier for the same underlying model family. According to OpenAI's documentation, how do these two kinds of identifiers differ, and why might a production application choose one over the other?
Aa dated snapshot identifier automatically updates over time to point at the newest underlying model, while a plain alias always stays pinned to whichever exact version existed when the application first started calling it
Ba plain alias name is a moving pointer that OpenAI can repoint to a newer underlying snapshot over time, while a dated snapshot identifier stays fixed to one specific model version indefinitely; production applications often pin to a dated snapshot specifically to avoid the model's behavior changing underneath them without warning
Caliases and dated snapshots are billed at different per-token prices for otherwise identical model capability, with the alias name always priced lower than any dated snapshot of the same model
Ddated snapshot identifiers are available only to enterprise customers under a custom contract, while every other developer can only ever call the plain alias name
Correct answer: .
OpenAI documents plain model names as aliases that can be repointed to a newer underlying snapshot as the model family is updated, while a dated snapshot identifier, which includes a specific release date in its name, stays fixed to that one version indefinitely; this is why production applications that need stable, reproducible behavior commonly pin to a dated snapshot rather than the alias, so an unannounced model update does not silently change outputs, latency, or capabilities the application depends on. The option reversing which identifier is the moving one gets the mechanism backwards relative to what is documented. The option inventing a pricing difference between an alias and a dated snapshot of the same underlying model is not supported by documented pricing, which is tied to the model family and capability tier rather than to whether the identifier is dated. The option restricting dated snapshots to enterprise contracts is wrong because dated snapshots are documented as available to any developer calling the API, not gated behind a special contract tier.
AI & LLM Engineering · Building with LLM APIs · Card 025/064medium
A developer building an OpenAI Responses API integration wants the model to call some function on this turn — any of the tools it has available — but does not want to pin the choice to one specific named function, and also does not want the model skipping tool use entirely and answering in plain text instead. According to OpenAI's documentation, which tool_choice setting achieves this?
ASetting tool_choice to "required", which forces the model to call at least one of the available functions on this turn without pinning the choice to any particular named function
BSetting tool_choice to "auto", since forcing the model toward some unspecified tool without naming one is not actually possible with any documented tool_choice value
CSetting tool_choice to an object naming one specific function, since OpenAI provides no way to require some tool call without also naming which tool must be called
DSetting tool_choice to "none", since "none" only prevents the model from producing a free-form text answer while still allowing it to call a tool if it chooses to
Correct answer: .
OpenAI's documentation describes "required" as forcing the model to invoke at least one of the functions supplied in the request, without constraining which specific function it picks — exactly the behavior of guaranteeing a tool call while leaving the choice of tool open. The option claiming this outcome is unreachable with any documented setting is wrong because "required" is documented for precisely this purpose; "auto" is a real value but it leaves the model free to skip tool use and answer in plain text, which is the behavior the developer is trying to avoid. Naming one specific function in tool_choice does let a developer force a particular tool, but that over-constrains the request here, since the developer wants any available tool called, not one designated in advance. "none" does the opposite of what its name might suggest a reader could assume: it disables tool calls outright and forces a plain-text answer, rather than merely discouraging a text-only response while still permitting a tool call.
Source: OpenAI, Function calling guide, https://developers.openai.com/api/docs/guides/function-calling
AI & LLM Engineering · Building with LLM APIs · Card 026/064hard
A developer using Anthropic's Messages API applies a cache_control breakpoint to a large, static system prompt, and is deciding between the default ephemeral cache lifetime and the optional extended lifetime enabled by adding "ttl": "1h" to that breakpoint. According to Anthropic's documentation, how do the two lifetimes differ in duration and in what the request that writes to the cache is charged, and does either lifetime change what a later cache hit costs?
AThe default cache lasts 1 hour, the extended option shortens it to 5 minutes, and both a 5-minute and a 1-hour cache write are billed at the same 1.25x multiplier over the base input token price
BBoth lifetimes cost the same to write to the cache, but choosing the 1-hour option charges double the normal input price on every later cache read for as long as that entry stays cached
CThe default cache lasts 5 minutes and the extended option lasts 1 hour; writing to a 5-minute cache costs 1.25x the base input token price while writing to a 1-hour cache costs 2x that price, but a later cache read is billed at the same reduced fraction of the base price regardless of which TTL wrote the entry
DThe default cache lasts 5 minutes and the extended option lasts 1 hour, but the TTL only changes how long the entry survives — Anthropic charges the identical price for a cache write regardless of which TTL is chosen, and cache reads are billed at the full base input token price
Correct answer: .
Anthropic documents the default ephemeral cache lifetime as 5 minutes, measured from the start of the request that writes or reads the entry, with an optional 1-hour lifetime available by setting a ttl field on the cache_control breakpoint. Writing to the cache costs more than an ordinary input token at either lifetime, but the multiplier itself differs: 1.25x the base input price for a 5-minute write versus 2x the base input price for a 1-hour write, reflecting the longer-lived entry's added cost. Reading from the cache, however, is documented at the same steeply discounted fraction of the base price no matter which TTL originally wrote that entry — the discount applies to any cache hit, not just hits against the shorter-lived tier. The option that swaps which duration is the default and which is the extended option has the two lifetimes backwards. The option inventing a doubled read price for the 1-hour tier, and the option claiming writes cost the same regardless of TTL while reads lose their discount entirely, both misstate the documented pricing structure, which ties the price difference specifically to the write operation, not to reads.
AI & LLM Engineering · Building with LLM APIs · Card 027/064medium
A developer sends a high-resolution photograph to a vision-capable OpenAI model with the image detail setting left at high (or auto), rather than set to low. According to OpenAI's documentation, how is that image converted into billable tokens?
AThe image is billed purely by its file size in kilobytes, independent of its pixel dimensions or the requested detail setting
BThe image is first scaled down to fit within a maximum pixel dimension, then divided into a grid of fixed-size tiles, and the final token count is the model's fixed base token cost plus a per-tile token cost multiplied by the number of tiles needed to cover the scaled image
CAt high or auto detail the image always costs exactly the same fixed number of tokens it would cost at low detail, since the detail setting only affects output quality rather than billed input tokens
DThe image is billed at one flat per-image token cost that is identical across every vision-capable model regardless of resolution or detail setting
Correct answer: .
OpenAI documents a two-part calculation for high or auto detail: the image is scaled down to fit within a maximum bounding dimension while preserving its aspect ratio, and if its shorter side still exceeds a smaller threshold that side is reduced further, then the scaled image is covered by a grid of fixed-size square tiles. The final token count is the model's fixed base token amount plus a per-tile token cost multiplied by however many tiles are needed to cover the resized image, so a larger or more detailed image costs more tokens than a small one. The option billing purely by file-size in kilobytes invents a metric OpenAI does not document; actual byte size is unrelated to the pixel-and-tile calculation. The option claiming high and low detail always cost the same wrongly generalizes low detail's flat, resolution-independent base cost to the high/auto path, when the tiling calculation specifically scales with resolution. The flat, model-agnostic per-image price is also incorrect, since the documented base and per-tile token amounts are specific to each model rather than uniform across every vision-capable model.
Source: OpenAI, Images and vision guide, https://developers.openai.com/api/docs/guides/images
AI & LLM Engineering · Building with LLM APIs · Card 028/064easy
A developer using OpenAI's text-embedding-3-large model wants shorter vectors to cut vector-database storage costs, so they call the embeddings endpoint with the dimensions parameter set below the model's native output size. According to OpenAI's documentation, why can the resulting shortened embedding still remain highly useful for retrieval instead of degrading arbitrarily once truncated?
AThe dimensions parameter runs a separate, smaller embedding model that was trained from scratch specifically to output that exact requested size
BShortening only ever removes dimensions that carry formatting metadata rather than semantic content, so no meaningful information is discarded no matter how small a size is requested
CThe API automatically applies a separate dimensionality-reduction algorithm, such as PCA, to the full-size embedding after generation to compress it down to the requested size
DThe embedding models are trained using Matryoshka Representation Learning, a technique that concentrates the most important concept-representing information toward the earlier dimensions of the vector, so truncating the trailing dimensions preserves most of the useful signal instead of degrading arbitrarily
Correct answer: .
OpenAI documents its text-embedding-3 models as trained with Matryoshka Representation Learning, a technique that arranges the most important, concept-representing information toward the earlier dimensions of the output vector, so that simply truncating the later dimensions preserves most of the embedding's useful signal rather than corrupting it unpredictably; OpenAI's own example notes a text-embedding-3-large embedding shortened to 256 dimensions still outperforms an unshortened, larger legacy embedding. The option describing a separately trained smaller model is wrong because the dimensions parameter truncates one model's native output rather than invoking a distinct model trained at that size. The option claiming only metadata is discarded overstates the guarantee; some quality is traded away as the vector shrinks, it simply degrades gracefully rather than not at all. The option describing an automatic post-hoc algorithm such as PCA is also incorrect, since the property comes from how the model was trained in the first place, not from a compression step applied afterward.
AI & LLM Engineering · Building with LLM APIs · Card 029/064hard
A developer enables Anthropic's citations feature on a large source document passed to the Messages API, so that the model's response includes text blocks with citations pointing back to exact passages in that document. According to Anthropic's documentation, how does the cited_text field returned in each citation affect billed tokens, and how does a citation's location reference differ between a plain text document and a PDF document?
AThe cited_text field does not count toward billed output tokens, nor toward input tokens if it is passed back into a later turn, and citation locations are given as character index ranges for a plain text document but as a page number range for a PDF
BThe cited_text field is billed as ordinary output tokens exactly like the rest of the response text, and citation locations use the same character index ranges regardless of whether the source was a plain text document or a PDF
CEnabling citations always incurs a separate flat per-request surcharge on top of ordinary token pricing, and citation locations are returned as an arbitrary chunk identifier that carries no positional information within the document
DThe cited_text field does not count toward billed output tokens, but citation locations are identical in format between a plain text document and a PDF, both expressed as character index ranges
Correct answer: .
Anthropic documents that the cited_text field is provided for convenience and is explicitly excluded from billed output tokens when generated, and is likewise excluded from billed input tokens if that same cited_text is passed back into a later conversation turn, which makes the citations feature efficient even for lengthy quoted passages. The location reference is also documented as depending on the source document's type: a plain text document is chunked into sentences and cited using character index ranges, while a PDF is cited using a page number range, and a custom content document instead uses block index ranges. The option billing cited_text as ordinary output tokens contradicts the documented token exclusion. The option inventing a flat per-request surcharge and a non-informative chunk identifier is not supported by the documentation, which ties citation locations to specific, meaningful positional ranges rather than opaque identifiers. The option that correctly exempts cited_text from output billing but claims plain text and PDF citations share the same character-index format is also wrong, since PDFs are specifically cited by page range rather than by character offset.
AI & LLM Engineering · Building with LLM APIs · Card 030/064easy
A developer building a high-volume integration with Anthropic's Messages API wants to throttle their own request rate proactively, before ever triggering a 429 rate-limit error, by tracking how much of the per-minute request and token allowance remains. According to Anthropic's documentation, how can the application learn this without first waiting for a 429 response?
AIt cannot be done proactively; the only way to learn how close the account is to its rate limit is to keep sending requests until a 429 response is returned
BThe Claude Console dashboard is the only place this information is exposed; it is not returned in any HTTP response header on ordinary, successful API calls
CEvery Messages API response, not only a 429, includes headers such as anthropic-ratelimit-requests-remaining and anthropic-ratelimit-tokens-remaining along with a corresponding reset time, so the application can read these headers on successful responses and slow down before it is actually rate limited
DThe remaining-capacity information is included only in the JSON response body of a successful request, never in response headers, so the application must parse the body of every response to extract it
Correct answer: .
Anthropic documents a set of response headers, including anthropic-ratelimit-requests-remaining, anthropic-ratelimit-tokens-remaining, and matching reset-time headers, that are returned on every Messages API response rather than only on a 429, letting an application read its remaining capacity on ordinary successful calls and back off before it actually gets rate limited. The option claiming this can only be discovered reactively by triggering 429s ignores that these headers accompany every response. The option restricting this visibility to the console dashboard is also wrong, since the documentation specifically lists these as HTTP response headers available programmatically on each API call, not only as a human-facing dashboard view. The option claiming the information lives only in the JSON response body rather than in headers reverses the actual mechanism — the values are exposed as headers precisely so an application does not need to parse or depend on the body's shape to find them.
AI & LLM Engineering · Building with LLM APIs · Card 031/064easy
A developer calls OpenAI's Chat Completions API with the n parameter set to 3 instead of leaving it at its default. According to OpenAI's documentation, what does this change, and how does it affect billed usage?
AIt sends the same request in parallel to three different underlying model versions and returns whichever response arrives first, at no extra token cost since only one response is kept
BIt generates three independent chat completion choices for the same input message in a single request, and the developer is billed for the output tokens generated across all three choices combined, not just one
CIt repeats the user's input message three times within the same prompt before generating a single completion, which increases input token billing but leaves output token billing unchanged
DIt sets the maximum number of follow-up turns allowed in a single conversation thread to three before the API automatically ends the conversation
Correct answer: .
OpenAI documents the n parameter as controlling how many chat completion choices the model generates for a single input message within one request, and explicitly notes that billing reflects the tokens generated across all of the resulting choices combined, recommending developers keep n at its default of 1 to minimize cost unless they specifically need multiple alternative completions. The option describing a race across different model versions invents behavior that is not documented; n does not select or vary the underlying model, and there is no free extra completion returned at no cost. The option describing repeated input text is also incorrect, since n does not duplicate the prompt itself; it duplicates the generation process for the same single prompt, which affects output token billing, not input token billing. The option describing a turn-limit setting confuses n with an unrelated concept, since n governs completions per request rather than how many turns a conversation may contain.
Source: OpenAI, Chat Completions API reference, https://developers.openai.com/api/docs/api-reference/chat/create
AI & LLM Engineering · Building with LLM APIs · Card 032/064easy
An application reads the stop_reason field on a Claude Messages API response before deciding whether to continue a multi-turn agent loop. According to Anthropic's documentation, which of the following correctly matches a stop_reason value to what it indicates happened?
AA stop_reason of "stop_sequence" means the response was truncated because it reached the max_tokens limit specified in the request
BA stop_reason of "tool_use" means the model refused to answer the request for policy reasons and produced no other usable content
CA stop_reason of "max_tokens" means the model matched one of the custom strings supplied in the stop_sequences parameter
DA stop_reason of "tool_use" means the model generated one or more tool_use content blocks and is waiting for the application to run those tools and return results before it continues
Correct answer: .
Anthropic documents "tool_use" as the stop_reason returned when the model has generated one or more tool_use content blocks and is pausing its turn so the calling application can execute those tools and send back tool_result blocks before the conversation continues — this is the signal an agent loop watches for to know it needs to act before calling the API again. A dedicated "refusal" value, not "tool_use", is what Anthropic documents for a policy-based decline, so pairing tool_use with a refusal meaning swaps two distinct, separately documented values. The two options that swap "stop_sequence" and "max_tokens" also mismatch documented meanings: "max_tokens" indicates the response was cut off for hitting the token limit, while "stop_sequence" indicates the model's output matched one of the custom strings supplied in the stop_sequences parameter, with the matched string reported separately in the response's own stop_sequence field — describing either value with the other's documented meaning is incorrect.
Source: Anthropic, Messages API reference, https://platform.claude.com/docs/en/api/messages
AI & LLM Engineering · Building with LLM APIs · Card 033/064medium
A developer who has only built agent loops against Anthropic's Messages API is porting the same tool-calling logic to OpenAI's Responses API, and needs to submit the output of a tool the model just called back into the conversation. According to each vendor's documentation, how does the mechanism for submitting a tool result differ between the two APIs?
AAnthropic's Messages API has no separate role for tool results — the result is sent as an ordinary user-role message containing a tool_result content block that references the original call's tool_use_id, whereas OpenAI's Responses API instead uses a dedicated function_call_output item carrying the original call_id and an output value, without wrapping it in a user-role message
BBoth APIs require the result to be submitted using a dedicated "tool" role message type that behaves identically between the two vendors, so the same request body can be reused unmodified
CAnthropic's Messages API requires a dedicated "tool" role message referencing tool_use_id, while OpenAI's Responses API instead expects the tool result to be appended as plain user-role text with no structured reference to which call it answers
DNeither API has any structured way to link a tool result back to the specific tool call it answers; correlation is inferred purely from the order in which messages appear in the conversation
Correct answer: .
Anthropic's own documentation contrasts its approach with "APIs that separate tool use or use special roles like tool or function": the Claude Messages API integrates tool results directly into ordinary user-role messages as a tool_result content block that references the id of the original tool_use block. OpenAI's Responses API instead documents a dedicated function_call_output item type carrying the original call's call_id and an output value, submitted as its own item rather than nested inside a user-role message. The option claiming both vendors share one identical "tool" role is wrong precisely because Anthropic's documentation states it deliberately avoids a separate role, unlike some other APIs. The option that swaps which vendor uses a dedicated role and which uses unstructured text reverses the actual mechanisms documented by each vendor. The option claiming neither API structurally links a result to its originating call is also incorrect, since both tool_use_id and call_id are documented fields that exist specifically to provide that correlation.
AI & LLM Engineering · Building with LLM APIs · Card 034/064easy
A team is building an AI agent that needs to connect to several external systems — a calendar, a database, and a search engine — and wants to avoid writing a separate bespoke integration for each one against each LLM vendor's own proprietary function-calling schema. What does adopting the Model Context Protocol (MCP) let them do instead?
AMCP is a proprietary Anthropic-only feature of the Messages API that cannot be used with any other model provider's product
BMCP replaces an LLM's own function-calling or tool-use mechanism entirely, so a model connected through MCP no longer needs any tools parameter or tool-call-style content blocks at all
CMCP is an open, vendor-neutral client-server protocol that standardizes how an AI application connects to external tools, data sources, and prompts, so a single MCP server implementation for a given system can be reused across different compliant AI applications instead of writing a separate bespoke connector per vendor
DMCP is a data format for packaging training data used to fine-tune a model on a company's internal documents, rather than something used at inference time
Correct answer: .
MCP is documented as an open standard, described by its own documentation as being like a standardized port for connecting AI applications to external data sources, tools, and prompt workflows, and it is explicitly supported across multiple AI applications and vendors, including Claude, ChatGPT, and development tools such as VS Code and Cursor — meaning one MCP server built for a given system can be reused by any compliant client rather than reimplemented per vendor. The option describing MCP as Anthropic-exclusive contradicts this documented cross-vendor support. The option claiming MCP eliminates the need for tool-use mechanics entirely misunderstands the layering: an MCP server still surfaces its capabilities as tools the model calls through the model's own tool-use mechanism, with MCP standardizing the connector beneath that layer rather than replacing it. The option recasting MCP as a training-data format is also incorrect, since MCP operates at inference time as a live connection protocol between a running AI application and external systems, not as a dataset used to fine-tune a model.
Source: Model Context Protocol documentation, https://modelcontextprotocol.io/introduction
AI & LLM Engineering · Building with LLM APIs · Card 035/064medium
A developer calling Anthropic's Messages API gives Claude two independent tools, and on a single turn Claude decides both are needed to answer the user's question. According to Anthropic's documentation, how does the API surface this, and what must the client do in its next request?
AThe assistant response can contain multiple tool_use content blocks in the same turn, one per tool call, each with its own id; the client must run each tool and reply with a separate tool_result block carrying the matching tool_use_id for each call, with all tool_result blocks placed before any other content in that user message
BClaude can only request one tool per turn, so it deliberately picks the single most useful tool and defers the second lookup to a follow-up turn after receiving the first tool_result
CThe API automatically merges both tool calls into a single tool_use block with a combined input object, and the client returns one tool_result covering both calls at once
DParallel tool calls are only available through the separate Message Batches API, so a normal synchronous Messages API request always forces Claude to pick a single tool per turn
Correct answer: .
Anthropic's documentation describes a single assistant turn as capable of returning several tool_use content blocks at once, each carrying its own unique id used to match it to a result. The client must execute every tool call and respond with one tool_result block per tool_use_id, and the documentation is explicit that these tool_result blocks must immediately follow their corresponding tool_use blocks and must all come before any other text in that user message — sending text ahead of a tool_result triggers an error. The option describing Claude as limited to one tool per turn contradicts this documented parallel capability and would waste an unnecessary round trip. The option describing the two calls as merged into a single block is wrong because each tool call keeps a distinct id specifically so its result can be individually matched; a merged block would make that matching impossible. The option tying parallel tool calls to the separate Batches API confuses two unrelated features: parallel tool use is a synchronous Messages API mechanic, while the Batches API is an entirely separate asynchronous bulk-processing endpoint.
AI & LLM Engineering · Building with LLM APIs · Card 036/064easy
Anthropic's Messages API supports both user-defined client tools (like a custom get_weather function) and built-in server tools (like web_search or code_execution). According to Anthropic's documentation, what is the key operational difference between the two?
AServer tools can only be used through the Message Batches API, while client tools work only with standard synchronous Messages API requests
BA client tool's call must be executed by the developer's own application code, with the result sent back in a follow-up request, whereas a server tool runs on Anthropic's own infrastructure and its result is included directly in that same response, without the client executing anything or making another API call
CServer tools give Claude the ability to rewrite its own system prompt mid-conversation, while client tools cannot alter the conversation's configuration at all
DClient tool results are limited to plain text, while server tools are the only tool type whose tool_result can contain images or documents
Correct answer: .
Anthropic's documentation splits tools by where their code executes: for a client tool, the response carries a tool_use block that the developer's own application must act on, running the corresponding function and sending its outcome back as a tool_result in a subsequent request — an agent loop the developer's code drives. A server tool such as web_search, web_fetch, or code_execution instead runs on Anthropic's own infrastructure; Claude calls it and its result is folded directly into the same response, so the developer sees the outcome without executing anything or making a second call, and no agent loop is needed for that call. Restricting server tools to the Batches API is incorrect — both tool types operate within the ordinary synchronous Messages API, and the Batches API is an unrelated asynchronous bulk-processing feature. Server tools do not grant Claude the ability to rewrite its own system prompt; that configuration remains entirely under the developer's control. Tool results are not divided this way either: a client tool's tool_result can carry text, image, or document content, so the content-type restriction attributed to client tools does not hold.
Source: Anthropic, Server tools, https://platform.claude.com/docs/en/agents-and-tools/tool-use/server-tools
AI & LLM Engineering · Building with LLM APIs · Card 037/064medium
A developer needs to run overnight sentiment analysis on a large batch of saved support transcripts using Claude, with no need for an immediate reply to any single request. According to Anthropic's documentation, how does the Message Batches API structure this workload, and what does it cost relative to the standard Messages API?
AEvery request must be sent one at a time over a persistent WebSocket connection that stays open until all results have streamed back
BEach request is submitted individually to the standard /v1/messages endpoint but tagged with a batch_id header, and Anthropic bills these at the same per-token rate as any other synchronous request
CThe developer submits a single batch request containing an inline array of individual Messages requests, each tagged with its own custom_id so its result can be matched back afterward since result order is not guaranteed; the batch processes asynchronously (most batches finish in under an hour) and is billed at a 50% discount off standard per-token pricing once results are retrieved
DBatches replace token-based pricing entirely with one flat fee per batch, regardless of how many requests or tokens it contains
Correct answer: .
Anthropic's Message Batches API takes a bundled array of individual Messages requests, supplied inline in the batch-creation call rather than as an uploaded file, with each entry carrying its own custom_id and the same parameters (model, max_tokens, messages) used in an ordinary request. Because results are not guaranteed to come back in the same order they were submitted, the custom_id is what lets the developer reassociate each result with its original request. The system then processes the whole batch asynchronously, with most batches completing in under an hour, and bills the tokens involved at a 50% discount versus standard synchronous pricing. Requiring one-at-a-time submission over a held-open WebSocket describes neither this feature nor any Anthropic streaming mechanism, which uses one-way server-sent events per request rather than a batch-spanning socket. Tagging individual synchronous calls with a batch_id header isn't how batching works and would not produce any cost discount, since standard per-token pricing applies to ordinary requests. Batches do not replace token-based billing with a flat fee — the discount is explicitly a percentage off the same per-token rates.
AI & LLM Engineering · Building with LLM APIs · Card 038/064hard
A developer sends Claude a 1000x1000 pixel JPEG image (within the standard resolution tier, no downscaling triggered) alongside a text question through the Messages API. According to Anthropic's documentation, how is the number of image tokens this costs actually calculated?
AAnthropic calculates image tokens the same way as OpenAI: the image is divided into fixed 512x512-pixel tiles, with each tile plus a fixed base cost added to the total
BAnthropic charges a flat token fee per image regardless of resolution, so a 1000x1000 image costs exactly the same number of tokens as a 200x200 image
CThe image's base64-encoded string is billed directly as though its character count were ordinary text input tokens
DClaude views an image as a grid of 28x28-pixel patches, with each patch counted as one visual token, giving a cost of ceil(width / 28) x ceil(height / 28); for a 1000x1000 image this works out to 1,296 visual tokens
Correct answer: .
Anthropic's documentation states that Claude views images in patches rather than pixels, with each patch a 28x28-pixel block counted as one visual token, so an image costs ceil(width / 28) x ceil(height / 28) visual tokens. For a 1000x1000 image, ceil(1000/28) is 36 on each dimension, giving 36 x 36 = 1,296 visual tokens, matching the exact figure Anthropic's own cost table lists for a 1-megapixel image at the standard resolution tier. Attributing OpenAI's fixed-tile approach to Anthropic mixes up the two vendors' distinct and separately documented mechanisms — Anthropic's own patch-grid method is the one that applies here. A flat per-image fee is inconsistent with Anthropic's own table, which shows token cost scaling with resolution (a 200x200 image costs 64 tokens, far fewer than a 1000x1000 image's 1,296). Billing the raw base64 character count as text tokens misdescribes vision entirely: Claude's image tokenization is patch-based and entirely separate from how its text tokenizer counts characters in an encoded string.
AI & LLM Engineering · Building with LLM APIs · Card 039/064easy
A developer's agent repeatedly includes the same large reference PDF in every turn of a long multi-turn conversation with Claude through the Messages API. According to Anthropic's documentation, how does the Files API help avoid the overhead of resending that PDF's bytes on every request?
AThe Files API automatically detects duplicate attachments already present in the conversation history and silently deduplicates them, with no change needed to how the file is referenced in the request
BThe developer uploads the PDF once to get back a file_id; later Messages API requests reference that file_id in a document content block instead of re-embedding the file's base64 bytes each turn, and the upload, download, list, retrieve, and delete operations themselves are free — only the file's content actually used within a Messages request is billed, as input tokens
CFiles uploaded through the Files API are cached for exactly five minutes and must be re-uploaded after that window before they can be referenced again
DThe Files API only accepts image files, so a PDF must still be base64-encoded and resent in full on every request regardless of its size
Correct answer: .
Anthropic's Files API documentation describes a create-once, use-many-times approach: a file is uploaded a single time and returns a unique file_id, which subsequent Messages requests can reference inside a document (or image) content block in place of resending the file's raw bytes. The documentation is explicit that uploading, downloading, listing, retrieving metadata for, and deleting files are all free operations, and that only file content actually drawn on within a Messages request gets billed, as ordinary input tokens. There is no documented automatic deduplication of repeated attachments already present in history — the developer must deliberately switch to referencing the file_id to get this benefit. The five-minute figure describes a different feature entirely (a prompt-caching TTL setting), not how uploaded files persist, since Files API uploads remain available until deleted or until any expires_in_seconds the developer set elapses. The Files API is not image-only; it explicitly supports PDFs and plain text as document content blocks alongside images.
AI & LLM Engineering · Building with LLM APIs · Card 040/064easy
A developer sets the same seed integer across repeated calls to OpenAI's Chat Completions API, keeping every other parameter identical, hoping for reproducible outputs. According to OpenAI's documentation, what guarantee does this actually provide, and what does the system_fingerprint field returned with each response help the developer detect?
AThe seed parameter makes the system attempt deterministic sampling on a best-effort basis rather than offering a strict guarantee, so repeated requests are only mostly identical; system_fingerprint identifies the current backend model and infrastructure configuration, so a change in its value between calls signals a backend change that can explain why outputs stopped matching
BThe seed parameter is a strict, guaranteed source of bit-for-bit identical output on every call, and system_fingerprint is simply a random request-tracing identifier unrelated to reproducibility
CThe seed value determines which fine-tuned model snapshot handles the request, while system_fingerprint reports the end user's device or browser fingerprint for analytics purposes
DSetting a seed disables sampling entirely and forces greedy decoding, so the temperature and top_p parameters no longer have any effect on the response
Correct answer: .
OpenAI's documentation describes the seed parameter as making the system sample deterministically on a best-effort basis, explicitly stating that determinism is not guaranteed even when the same seed and other parameters are reused. The system_fingerprint value returned alongside each response identifies the particular combination of model weights, infrastructure, and configuration currently serving the request; comparing it across calls lets a developer tell whether an unexpected change in output stems from OpenAI updating its backend rather than from anything the developer changed. Calling seed a strict, guaranteed determinism mechanism overstates what the documentation promises, and dismissing system_fingerprint as an unrelated tracing ID ignores its documented purpose of monitoring backend changes. The seed parameter has no role in choosing a fine-tuned snapshot, and system_fingerprint reports server-side configuration, not anything about the end user's device or browser. Setting a seed also does not disable sampling parameters like temperature or top_p — it works alongside them to make sampling more reproducible, not replace it with greedy decoding.
AI & LLM Engineering · Building with LLM APIs · Card 041/064easy
A team wants to build a natural, low-latency spoken-voice agent that can be interrupted mid-response ('barge-in') and hear the user talking while it's still speaking. According to OpenAI's documentation, why is the Realtime API a better fit for this than streaming text responses from the Chat Completions or Responses API over server-sent events?
AServer-sent event streaming already supports full-duplex audio in both directions, so the Realtime API differs only in offering a wider range of voices
BThe Realtime API works by converting speech to text, calling the standard Chat Completions endpoint in a fast loop, and converting the reply back to speech, which is functionally the only difference from building this pipeline yourself
CThe Chat Completions API's server-sent event streaming already maintains a persistent connection and conversational state across turns identical to the Realtime API, so switching would only change pricing
DThe Realtime API keeps a persistent, stateful WebSocket or WebRTC session that streams audio natively in both directions, supporting low first-audio latency and barge-in, whereas standard chat APIs follow a request/response model where server-sent events stream text tokens in one direction per call, with no native audio support
Correct answer: .
OpenAI's documentation describes the Realtime API as connecting over WebRTC in the browser or WebSocket on the server, establishing a persistent, stateful session in which audio streams natively in both directions at once, purpose-built for speech-to-speech voice agents that support barge-in and low first-audio latency. Standard chat APIs, by contrast, are a request/response model: each call is a discrete request, and server-sent-event streaming within that call sends text tokens one way, from server to client, with no native audio handling at all. Claiming SSE already supports full-duplex audio reverses this distinction, since SSE is inherently a one-directional text stream. Describing the Realtime API as merely orchestrating separate speech-to-text, chat-completion, and text-to-speech calls misses that it processes audio natively within one stateful session rather than chaining independent client-managed calls. Chat Completions' SSE streaming is also not equivalent to a persistent session — a new request is required for each turn, unlike the single ongoing connection the Realtime API maintains.
Source: OpenAI, Realtime API guide, https://developers.openai.com/api/docs/guides/realtime
AI & LLM Engineering · Building with LLM APIs · Card 042/064easy
A developer wants to screen user-submitted text and images for policy violations (hate speech, self-harm content, violence, and similar categories) before passing them to a chat model, without paying for a full model generation just to get a safety judgment. According to OpenAI's documentation, what does the Moderation API provide for this?
AThe Moderation API is a paid feature billed at the same per-token rate as chat completions, since it must run a full generation to judge the content
BThere is no separate moderation endpoint; classification is only available by prompting a regular chat model with a custom system prompt asking it to judge the content
CA separate, free-to-use moderation endpoint that classifies text and/or image input and returns a flagged boolean along with a dictionary of per-category violation flags and per-category confidence scores, without generating any chat response
DThe Moderation API only accepts and classifies images, so text content must be screened through a different, paid endpoint instead
Correct answer: .
OpenAI's documentation describes the moderation endpoint as a dedicated resource that classifies text or image inputs without generating a model response, explicitly stating that it is free to use. Its output includes a flagged boolean, a dictionary of per-category violation flags covering categories like hate, self-harm, and violence, and per-category confidence scores between 0 and 1, with the omni-moderation-latest model able to handle both text and image inputs. Billing it like a chat completion misdescribes the feature, since no generation occurs and no per-token charge applies. Claiming no dedicated endpoint exists ignores that OpenAI documents this as its own moderations resource, separate from prompting a chat model with a custom safety instruction — an unreliable, undocumented workaround by comparison. The claim that it is image-only is also incorrect, since the documented model explicitly classifies text and image input together, not one to the exclusion of the other.
AI & LLM Engineering · Building with LLM APIs · Card 043/064hard
A team building a customer-support triage bot on an OpenAI reasoning model finds responses too slow and expensive for high-volume simple classification, while a separate internal tool doing complex multi-step debugging analysis on the same model family needs much more thorough internal reasoning. According to OpenAI's documentation, how does the reasoning-effort setting (reasoning.effort / reasoning_effort) address both cases, and how are the tokens spent on that internal reasoning billed?
AEffort is a tunable knob, ranging from low through high depending on the model, that trades response latency and cost against how thoroughly the model reasons before answering — lower effort suits fast, simple tasks like triage, higher effort suits complex, less latency-sensitive analysis — and the tokens spent on that internal reasoning are billed as output tokens even though the reasoning content itself is not shown back to the caller
BReasoning effort only controls the model's writing tone, becoming more formal at higher settings, and has no effect on latency, cost, or how much internal reasoning the model performs
CReasoning tokens spent at any effort level are provided completely free of charge, since they are considered part of the model's internal process rather than part of the visible response
DEffort must be set identically across every request in an organization; it cannot be varied per request based on that request's task complexity
Correct answer: .
OpenAI's documentation frames the reasoning-effort setting as a tuning knob rather than the primary lever for output quality: it trades latency and cost against how thoroughly the model reasons internally before producing a final answer, with lower settings suited to fast, high-volume, low-complexity work like triage or customer support, and higher settings reserved for complex, less latency-sensitive tasks like debugging or research analysis. This setting is applied per request, so the triage bot and the debugging tool can each choose the effort level appropriate to their own workload on the same model family. The documentation also states plainly that reasoning tokens, while not visible via the API, still occupy context and are billed as output tokens — so the internal reasoning driving a high-effort request has a real cost even though its content is hidden. Describing effort as controlling only writing tone misattributes a stylistic effect to a parameter that actually governs reasoning depth. Claiming reasoning tokens are free contradicts the documented output-token billing. Claiming effort is fixed organization-wide misdescribes it as a request-level parameter that can be tuned per call to match each task's complexity.
AI & LLM Engineering · Building with LLM APIs · Card 044/064medium
A developer used to Anthropic's Messages API, where a cache_control breakpoint must be explicitly added to a prompt to enable prompt caching, moves the same kind of application onto OpenAI's API. According to OpenAI's documentation, what does the developer need to do to get prompt caching benefits there?
AExactly the same as Anthropic: an explicit cache breakpoint marker must be added to the prompt, or caching never activates on OpenAI's models
BNothing extra in general: prompt caching is enabled by default on supported OpenAI models with no special parameter required, automatically discounting the portion of a repeated prompt prefix that qualifies once it reaches the model's minimum cacheable-length threshold; only on newer model generations can the developer optionally choose between implicit and explicit cache breakpoints via a caching-options parameter
COpenAI does not offer any form of prompt caching; only Anthropic's Messages API supports this feature
DPrompt caching on OpenAI's API only applies to the embeddings endpoint and has no effect on chat completion requests
Correct answer: .
OpenAI's documentation describes prompt caching as enabled by default for supported models, requiring no special parameter to activate: once a prompt's repeated prefix reaches the model's minimum cacheable length (a threshold that depends on the model and, for older generations, on factors like tools and images in the request), the qualifying portion is automatically discounted on a cache hit. Only on newer model generations does the documentation add an optional prompt_cache_options.mode setting letting a developer choose between implicit and explicit cache breakpoints, in contrast to Anthropic's Messages API, where a cache_control breakpoint must be explicitly placed in the prompt for caching to happen at all. Claiming OpenAI requires the identical explicit-breakpoint mechanism misses this default-on, no-configuration-needed design for the common case. Claiming OpenAI has no caching feature at all contradicts its documented automatic caching. Restricting the benefit to the embeddings endpoint is also incorrect, since the documented caching mechanism discounts repeated prompt prefixes on chat/completion-style requests, not embeddings calls.
AI & LLM Engineering · Building with LLM APIs · Card 045/064medium
A developer building on Claude's Messages API enables extended thinking with `thinking: {"type": "enabled", "budget_tokens": 4000}` alongside `max_tokens: 8000`. According to Anthropic's documentation, what does `budget_tokens` actually control, and how is the model's internal reasoning billed?
A`budget_tokens` is a hard ceiling Claude's internal reasoning can never cross, and thinking tokens are billed at a separate, discounted rate that does not count toward the response's output-token total
B`budget_tokens` sets a target for how many tokens Claude's internal reasoning may use, but Claude can finish well under that target, `max_tokens` remains the actual hard ceiling on the turn's combined reasoning-plus-answer output, and the tokens spent reasoning are billed as ordinary output tokens, reported separately in the response's usage data
C`budget_tokens` only controls how much of Claude's reasoning is shown back to the developer in the response; the full reasoning is always generated and billed at the same length no matter what value is configured
D`budget_tokens` is a free reasoning allowance that Anthropic does not bill at all, since only the final text answer Claude produces counts toward billed output tokens
Correct answer: .
budget_tokens sets a target token budget for Claude's internal reasoning before it starts its final answer, but the documentation is explicit that this is a target rather than a strict cap: Claude may stop reasoning well before using the full amount, and the real hard ceiling on the turn's total output remains max_tokens, which the configured budget must stay under. Whatever tokens Claude actually spends reasoning are billed as part of the response's output tokens, not as a separate discounted category, and the exact count spent this way is reported in the response's usage.output_tokens_details.thinking_tokens field so a developer can track what a given budget actually cost. The option describing the budget as an inescapable hard ceiling on reasoning reverses which parameter is the real ceiling and invents a discounted billing category that does not exist. The option claiming the budget only limits what is displayed misunderstands the mechanism entirely: the budget shapes how much reasoning Claude actually performs, not just how much of it is surfaced afterward. The option claiming reasoning is free contradicts the documented billing treatment directly, since thinking tokens are counted as billed output tokens like any other generated content.
AI & LLM Engineering · Building with LLM APIs · Card 046/064medium
A developer adds Anthropic's code execution tool to a Messages API request so Claude can run Python against an uploaded CSV file. The same request also defines a client tool, `send_slack_message`, that Claude has used before. According to Anthropic's documentation, how does fulfilling a code execution call differ architecturally from fulfilling a call to the client-defined `send_slack_message` tool?
AThere is no architectural difference: both are fulfilled the same way, with the API returning a request naming what to run and the developer's own code responsible for actually running it and sending back a result
BThe code execution tool can be invoked at most once per conversation for the entire lifetime of that conversation, while `send_slack_message` may be called an unlimited number of times
CCode execution calls are always queued and answered only after every other tool call in the same turn has been resolved, regardless of the order Claude requested them in
DCode execution runs entirely server-side: Anthropic's API executes the command itself inside a sandboxed container and returns the result within the same response, so the developer never runs anything or sends back a tool_result for it themselves, unlike `send_slack_message`, whose tool_use block the developer's own code must execute and answer; the one exception is that if Claude calls both tools in the same turn, the code execution result is withheld until after the developer returns the client tool's result
Correct answer: .
Anthropic's documentation describes code execution as a server tool: when Claude requests a command, the API runs it itself inside a secure sandboxed container and returns stdout, stderr, and any generated files directly in the same response, so the developer never writes execution logic or sends a tool_result block for it, in sharp contrast to a client-defined tool like a Slack-sending function, whose tool_use request the developer's own backend must actually carry out before replying with the outcome. The documented exception is that when Claude calls a server tool like code execution alongside a client tool in the same turn, the API withholds the server tool's result and returns only the client tool's call, waiting for the developer to supply that client tool's result before it runs and returns the code execution outcome in a later response. The option claiming no architectural difference exists ignores this documented client-versus-server execution split entirely. The option inventing a single-use-per-conversation limit for code execution and the option claiming code execution is always queued behind every other tool call both invent constraints the documentation never states.
AI & LLM Engineering · Building with LLM APIs · Card 047/064easy
A developer adds Anthropic's web_search tool to a Messages API request, sets `allowed_domains` to a short list of trusted sites, and Claude performs three searches over the course of the conversation to answer the user's question. According to Anthropic's documentation, how is this billed, and what rule governs the domain-filtering parameters?
AEach search is billed per use ($10 per 1,000 searches) in addition to the standard token cost of the search-generated content that enters the context, and a request may set only one of `allowed_domains` or `blocked_domains`, never both together
BWeb search itself is always free; only the tokens in the final answer are billed, and `allowed_domains` and `blocked_domains` must always be supplied together so the API can cross-check them against each other
CEach search is billed per use, but the content returned by the search is entirely free and never counts as input tokens no matter how much text is retrieved
DThere is no per-search charge at all; billing is based solely on how many domains appear in `allowed_domains`, with more domains costing more regardless of how many searches Claude actually performs
Correct answer: .
Anthropic's documentation prices web search at $10 per 1,000 searches, charged in addition to the standard input-token cost of whatever search-result content ends up in Claude's context, and each search counts as one use regardless of how many results it returns. On the domain-filtering side, the documentation states that a request must supply only one of allowed_domains or blocked_domains, never both, and returns a 400 error if both are present in the same request. The option claiming search is entirely free contradicts the documented per-search charge, and inventing a requirement that both domain lists be supplied together reverses the actual mutual-exclusivity rule. The option claiming retrieved search content never counts as input tokens is wrong because the documentation explicitly describes search-result content as counted in input tokens as it enters the conversation; only certain citation fields like the cited text snippet, title, and URL are carved out as not counted. The option pricing based on domain-list length rather than search count invents a billing dimension nowhere in the documented pricing structure.
Source: Anthropic, Web search tool documentation, https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool
AI & LLM Engineering · Building with LLM APIs · Card 048/064medium
An agent built on Claude's Messages API runs a research task for dozens of turns, calling tools repeatedly, and the conversation's growing tool-result history is approaching the model's context window limit even though the task isn't finished. The developer configures Anthropic's context_management feature with a `clear_tool_uses_20250919` edit. According to Anthropic's documentation, what does this feature do, and where does it run?
AIt permanently deletes the oldest messages from the developer's own stored conversation history in their database, so the developer must reconstruct any cleared turns from their own logs if they are ever needed again
BIt has Claude itself decide, mid-turn, which of its own past tool calls to omit from its next response, based on which ones it judges are no longer relevant to the current step
CIt runs server-side, before the prompt reaches Claude: once the conversation's input tokens cross a configured trigger threshold, it automatically clears the content of the oldest tool results in chronological order, replacing each with placeholder text, while keeping a configurable number of the most recent tool uses intact, all without altering the full conversation history the developer's own client still holds
DIt compresses every tool result in the conversation into a shorter paraphrase using a separate summarization model call, rather than clearing any of them outright, and this paraphrasing counts as an extra billed request each time it runs
Correct answer: .
Anthropic's context editing feature, configured through the context_management parameter, applies the clear_tool_uses_20250919 edit server-side, before the request is processed: once the conversation's input tokens exceed a configured trigger value, it clears the content of the oldest tool results first, working forward in chronological order, and replaces each cleared result with placeholder text so Claude can see that something was removed rather than silently losing context. A configurable keep setting preserves a chosen number of the most recent tool uses in full, and the developer's own client-side copy of the conversation remains completely unmodified, since the editing only affects what is sent to the model on that request. The option describing permanent deletion from the developer's own database misattributes what is really a per-request, model-facing transformation to the developer's own storage layer. The option describing Claude itself as the one deciding what to omit inverts where this logic lives; it is the API's server-side context management step, not a judgment call the model makes mid-turn. The option describing summarization by a separate model call invents a mechanism the documentation does not describe; the feature clears content by replacing it with a placeholder, not by paraphrasing it through an additional billed call.
AI & LLM Engineering · Building with LLM APIs · Card 049/064easy
A developer calling OpenAI's Chat Completions API wants to strongly discourage the model from ever producing a specific token, without removing it from the model's vocabulary entirely or rewriting the prompt. They use the `logit_bias` parameter. According to OpenAI's API reference, what does this parameter actually do?
AIt accepts a list of literal banned words as plain text strings, which the API matches against the generated text after decoding and deletes if found, before the response is returned to the developer
BIt changes the sampling temperature applied only to the specific tokens named, leaving every other token's probability governed by the request's own top-level temperature setting
CIt reorders the model's vocabulary so that the named tokens are permanently removed from consideration on every subsequent request made with the same API key, not just the current one
DIt accepts a map from a token's numeric ID in the model's tokenizer to a bias value between -100 and 100, which is added directly to that token's logit before sampling, with values near -100 effectively banning the token and values near 100 effectively forcing its selection
Correct answer: .
OpenAI's API reference documents logit_bias as a JSON object mapping token IDs, not literal words, to bias values constrained to the range -100 to 100, and these bias values are added directly to the corresponding token's logit before the sampling step that picks the next token, so a value near -100 makes that token's selection essentially impossible and a value near 100 makes it dominate the choice. The option describing word-level string matching and after-the-fact deletion misdescribes the mechanism as post-processing rather than as an adjustment to the model's own token probabilities during generation. The option describing a token-specific temperature override is wrong because temperature is a single global sampling-randomness setting applied uniformly to the whole distribution, not something logit_bias touches per token. The option describing a permanent, cross-request vocabulary change is wrong because the bias values are scoped to the single request that includes them, with no persistent effect on later calls or on other API keys.
Source: OpenAI, Chat Completions API reference (logit_bias parameter), https://developers.openai.com/api/docs/api-reference/chat/create
AI & LLM Engineering · Building with LLM APIs · Card 050/064hard
A developer defines a JSON schema for OpenAI's Structured Outputs feature with `strict` set to true, and wants one field, `middle_name`, to be genuinely optional so that some responses can simply omit it. According to OpenAI's documentation, why doesn't setting `middle_name` as optional the normal JSON Schema way (leaving it out of the `required` array) work here, and what must the developer do instead?
ALeaving a field out of `required` works exactly as in ordinary JSON Schema; strict mode places no additional constraint on the `required` array beyond what standard JSON Schema already allows
BStrict mode requires every property defined in the schema to be listed in `required`, so a field can never be truly absent from the output; to model an optional-feeling field, the developer instead types it as a union that allows null (alongside `additionalProperties: false`) and lets the model return null when there is no value, while still including the field's key in the response
CStrict mode ignores the `required` array entirely and instead infers which fields are mandatory from the order they appear in the `properties` object, with earlier fields treated as mandatory and later ones as optional
DStrict mode requires the developer to submit two entirely separate schemas, one listing only the mandatory fields and one listing only the optional fields, and to specify in the request which of the two schemas applies to the current call
Correct answer: .
OpenAI's Structured Outputs documentation states that strict mode requires every property declared under a schema's properties to also appear in that schema's required array, which means there is no way to make a field truly absent from the response the way a plain JSON Schema optional field would be. To model a field that can be meaningfully empty, the documented pattern is to widen its type to a union that includes null, alongside setting additionalProperties to false on the schema, and let the model emit null for that key rather than omitting the key altogether; the key itself is always present in the output, only its value can be empty. The option claiming ordinary optional-field behavior works unchanged contradicts this documented required-everything constraint directly. The option describing required-ness as inferred from property order invents a mechanism the documentation does not describe; required status comes only from explicit membership in the required array. The option describing two separate submitted schemas invents an entirely fictitious submission format; a single schema object is what strict mode actually validates against.
AI & LLM Engineering · Building with LLM APIs · Card 051/064easy
A team building a voice agent on OpenAI's Realtime API compares the per-token price of audio input tokens against the per-token price of text input tokens on the same model. According to OpenAI's published pricing, what should they expect?
AAudio and text tokens are always priced identically per token on Realtime models, since the API converts audio to an internal text representation before counting tokens
BAudio tokens are priced lower than text tokens, since audio is a lossy, more compressed signal that costs the model less to process per token than dense text
CAudio tokens are priced substantially higher than text tokens on the same Realtime model, with separate published per-million-token rates for text input, audio input, and audio output, rather than a single shared token rate
DRealtime models bill only by wall-clock session duration in minutes, with no separate per-token rate for audio or text at all
Correct answer: .
OpenAI's published Realtime API pricing lists separate per-million-token rates for text input, audio input, and audio output on the same model, and the audio rates are substantially higher than the text rate -- on one current Realtime model, audio input is priced several times higher than text input per token. This reflects that processing and generating audio tokens costs more than processing dense text tokens of the same count, the opposite of the option claiming audio is cheaper because it is a compressed signal. The option claiming audio and text share one identical per-token rate ignores the documented separate pricing categories entirely. The option claiming billing is based purely on session minutes with no token-level rate at all contradicts the documented per-token pricing structure; even though a rough dollars-per-minute figure can be derived from typical token-per-minute audio rates as a convenience, the underlying billing unit is still the token.
Source: OpenAI, API pricing, https://developers.openai.com/api/docs/pricing
AI & LLM Engineering · Building with LLM APIs · Card 052/064easy
A developer's application calls OpenAI's Responses API for a task expected to take several minutes of model processing, and wants to avoid the request failing due to a client-side or gateway connection timeout while waiting synchronously for the full response. According to OpenAI's documentation, how does background mode address this?
ASetting `background` to true on the request lets the model process the task asynchronously; the initial call returns immediately with an in-progress response object, and the application polls a GET endpoint using that response's ID, checking again while its status is queued or in_progress, until it reaches a terminal state and the output can be read
BBackground mode requires the developer to first split the task into several smaller requests themselves, since the API itself has no way to run a single request for longer than its normal synchronous timeout
CBackground mode automatically retries the entire request from scratch every time the connection drops, silently discarding any progress the model had already made before the drop
DBackground mode only changes how the response is formatted once it's ready; the request is still processed synchronously and the client connection must stay open for the task's entire duration either way
Correct answer: .
OpenAI's documentation describes background mode as an asynchronous processing option enabled by setting background to true on a Responses API request: rather than holding the connection open until the model finishes, the initial call returns right away with a response object whose status is queued or in_progress, and the application is expected to poll a GET endpoint for that same response ID, continuing to poll while the status remains queued or in_progress, until the response reaches a terminal state and its output can be retrieved. This is specifically documented as the fix for long-running tasks that would otherwise risk hitting timeouts or connectivity issues on a held-open synchronous connection. The option requiring the developer to manually split the task into smaller requests describes a workaround for the problem background mode is specifically meant to eliminate, not what background mode itself does. The option describing automatic full retries from scratch on every disconnect is wrong because the documented mechanism is status polling against a persisted response object, not repeated resubmission. The option claiming the connection must still stay open throughout directly contradicts the documented point of the feature, which is to let the client disconnect and reconnect later via polling.
AI & LLM Engineering · Building with LLM APIs · Card 053/064easy
A developer using an OpenAI GPT-5-family model wants shorter, terser answers for simple code-lookup questions and longer, more thoroughly explained answers for complex debugging questions, without changing how much internal reasoning the model performs on either kind of request. According to OpenAI's documentation, which parameter is designed for this, and how does it differ from the model's reasoning-effort setting?
AThe `max_tokens` parameter, since lowering it for simple questions and raising it for complex ones is the documented way to control how much explanation the model includes, and it has always worked this way for every OpenAI model
BThe `verbosity` parameter, set to low, medium, or high, which is documented as controlling the length and level of detail in the model's final answer independently of `reasoning_effort`, which instead controls how much internal reasoning the model performs before answering
CThe `reasoning_effort` parameter itself, which OpenAI's documentation describes as jointly controlling both how much the model reasons internally and how long its final answer is, with no separate setting for output length
DThe `temperature` parameter, since raising temperature is documented as producing systematically longer, more detailed answers, while lowering it produces systematically terser ones
Correct answer: .
OpenAI's documentation introduces a verbosity parameter on GPT-5-family models, with low, medium, and high settings, specifically to control how long and detailed the model's final answer is -- low verbosity favors terse, efficient answers well suited to something like a quick code snippet, while high verbosity favors a more thorough, explanation-rich answer -- and this is documented as a separate control from reasoning_effort, which governs how much internal reasoning the model performs before producing that answer, not how long the answer itself reads. The option pointing to max_tokens is wrong because that parameter only truncates output at a hard length ceiling; it does not shape the model's own judgment about how much detail to include, and cutting it off mid-explanation is a blunt truncation rather than a deliberate terseness choice. The option describing reasoning_effort as also controlling output length conflates the two documented parameters that were introduced specifically to be independent of each other. The option pointing to temperature is wrong because temperature governs sampling randomness among token choices, not the deliberate length or detail level of the response.
AI & LLM Engineering · Building with LLM APIs · Card 054/064hard
Across nearly every major LLM provider's published pricing, generating one output token from a model costs more than processing one input (prompt) token of the same model, often by a factor of three to five times. What is the primary technical reason for this asymmetry, rather than it being an arbitrary business markup?
AOutput tokens are billed higher purely because providers want to discourage overly long responses for environmental reasons, with no meaningful difference in the underlying computation cost between processing an input token and generating an output token
BInput tokens require more GPU memory to store than output tokens, since the entire prompt must remain in memory throughout generation, while each output token can be discarded immediately after it is produced
COutput tokens cost more because they are always run through a separate, larger safety-filtering model before being returned to the developer, adding a second full model pass on top of generation
DProcessing the input prompt happens in a single parallel forward pass over all prompt tokens at once (prefill), which uses the available compute efficiently, while generating output happens one token at a time in a sequential, autoregressive process where each new token depends on every token generated before it, requiring roughly as many sequential forward passes as there are output tokens and using the hardware less efficiently per token
Correct answer: .
The prefill stage of LLM inference processes an entire prompt in one parallelized forward pass, since every input token's representation can be computed simultaneously, which uses GPU compute efficiently and amortizes cost well across many tokens at once. Generation, by contrast, is autoregressive: each output token depends on every token produced before it, so the model must run a fresh forward pass for every single new token, one after another, which underutilizes the same hardware per token compared to prefill's batched efficiency, and this computational asymmetry, not an arbitrary policy choice, is why major providers consistently price output tokens several times higher than input tokens. The option attributing the price gap to environmental discouragement invents a motive with no technical basis and ignores the real compute-cost driver. The option about input tokens needing more memory reverses the actual memory pressure, since it is the growing key-value cache accumulated during sequential generation, not simply holding the prompt, that dominates memory use as output grows. The option describing a mandatory second full model pass for safety filtering invents a step that is not what drives this specific, universal input-versus-output price gap across providers; content moderation, where used, is a separate concern from the base generation cost asymmetry being asked about here.
Source: NVIDIA, 'LLM Inference Benchmarking: Fundamental Concepts' (prefill vs. decode phases); OpenAI and Anthropic published API pricing pages showing the consistent output-over-input price ratio
AI & LLM Engineering · Building with LLM APIs · Card 055/064easy
A developer wants their support bot to answer questions by searching a folder of about 200 internal policy PDFs. Rather than building a chunking pipeline, choosing an embedding model, and standing up a vector database themselves, they want a single API tool to handle retrieval so they can get a first version working quickly. According to OpenAI's documentation, what does the Responses API's `file_search` tool provide?
AA hosted vector store the developer can upload files into, plus a built-in tool that automatically chunks, embeds, and semantically searches those files for relevant passages during a model turn, without the developer running their own embedding pipeline or vector database
BA local, on-device retrieval engine that runs entirely inside the developer's own infrastructure, with OpenAI providing only the embedding model weights for the developer to host and query
CA tool that can only match on exact keyword overlap between the user's question and the uploaded documents, since no embedding or semantic search is involved
DA tool that requires every uploaded document to already be pre-chunked and pre-embedded by the developer before upload, with the tool only handling the final similarity-ranking step
Correct answer: .
OpenAI's file_search tool, available through the Responses API, lets a developer upload files into a hosted vector store; OpenAI's infrastructure handles chunking, embedding, and storage, and the tool automatically performs semantic search over that store when the model decides a lookup is needed, returning relevant passages inline during the turn -- exactly the 'skip building RAG infrastructure' capability that fits a single bot answering questions over a folder of PDFs. The option describing an on-device engine with only weights supplied inverts this: the store and search run on OpenAI's hosted infrastructure, not the developer's own servers. The option limiting the tool to exact keyword overlap is wrong because the tool performs semantic, embedding-based search, not literal string matching, so it can surface relevant passages that don't share exact wording with the query. The option requiring the developer to pre-chunk and pre-embed documents before upload is wrong because chunking and embedding are exactly the steps the hosted tool performs automatically once raw files are uploaded, which is the point of using it instead of a self-built pipeline.
Source: OpenAI, Responses API tools guide, https://developers.openai.com/api/docs/guides/tools-file-search
AI & LLM Engineering · Building with LLM APIs · Card 056/064medium
A developer adds a `cache_control` breakpoint to a system prompt sent with every call to Anthropic's Messages API, but the prompt is shorter than the model's documented minimum cacheable length. According to Anthropic's documentation, what happens on this and later requests, and how can the developer confirm from the response itself whether caching actually took effect?
AThe request fails with an error identifying which minimum-length requirement was not met, so the developer must catch and handle this specific error before retrying
BThe prompt is still cached, but only at a partial discount scaled to how far short of the minimum it falls, and the response's usage fields report that partial discount directly
CThe prompt below the minimum length is processed without caching and without an error, and the developer can confirm this by checking the response's usage fields: if both `cache_creation_input_tokens` and `cache_read_input_tokens` report zero, the prompt was not cached
DAnthropic's API automatically pads the prompt up to the minimum length behind the scenes so that caching still applies, and the padding is reflected as extra input tokens in the usage response
Correct answer: .
Anthropic's documentation states that prompts shorter than a model's minimum cacheable length cannot be cached even if marked with `cache_control`: such requests are simply processed without caching, and no error is returned. To confirm whether caching actually took effect on a given request, the documentation points developers to the response's usage fields -- if both `cache_creation_input_tokens` and `cache_read_input_tokens` come back as zero, the prompt was not cached, most likely because it fell short of the minimum length for that model. The option describing an identifiable error is wrong because no error is raised in this case; the request still succeeds, just without caching. The option describing a partial, scaled discount is wrong because caching is all-or-nothing per Anthropic's pricing -- there is no partial-credit mechanism for prompts that fall short of the minimum. The option describing automatic padding is wrong because Anthropic does not silently insert extra tokens into a request to force it over the caching threshold; the prompt is processed exactly as sent.
AI & LLM Engineering · Building with LLM APIs · Card 057/064hard
A developer is building a coding assistant that asks GPT-4o to make small, targeted edits to an existing file -- most of the file's content stays identical, and only a few lines change each time. Regenerating the entire file token-by-token from scratch is slow given how little actually changes on each call. According to OpenAI's documentation, what does the Predicted Outputs feature (the `prediction` parameter) do to help?
AIt lets the developer supply the existing, mostly-unchanged text as a predicted version of the output; the API then checks the predicted tokens against what it would generate and can skip ahead through matching stretches instead of generating every token sequentially, reducing latency when most of the output matches the prediction
BIt caches the entire previous completion server-side and simply returns it unchanged on the next call unless the prompt hash changes, so no new generation happens at all for near-identical requests
CIt fine-tunes a smaller distilled model on the fly from the developer's prediction text, then uses that distilled model to generate the response instead of the original GPT-4o model
DIt pre-validates the prediction text against the developer's JSON schema before generation starts, rejecting the API call early if the prediction itself would not satisfy the schema
Correct answer: .
OpenAI's Predicted Outputs feature lets a developer pass a `prediction` parameter containing an expected version of the output -- such as the mostly-unchanged existing file -- and the model uses that prediction to speed up generation: rather than sequentially generating every token from scratch, it can verify and skip ahead through stretches where the actual output matches the predicted text, only truly generating the parts that differ, which is exactly why it fits a coding-assistant workload where most tokens don't change between edits. The option describing server-side caching that returns the same completion unchanged is wrong because Predicted Outputs still runs real generation and can diverge from the prediction wherever the true output differs; it isn't a cache-hit mechanism. The option describing on-the-fly distillation to a smaller model is wrong because no separate distilled model is created or swapped in; the same target model still produces the output. The option describing early schema pre-validation of the prediction is wrong because Predicted Outputs is a latency and generation-speed optimization, not a schema-validation step performed before generation begins.
AI & LLM Engineering · Building with LLM APIs · Card 058/064easy
A developer's existing integration against OpenAI's Chat Completions API sets `max_tokens` to cap response length. When they point the same integration at an o-series reasoning model instead of a standard chat model, the call fails or behaves unexpectedly. According to OpenAI's documentation, what is the relevant difference for reasoning models here?
AReasoning models ignore any token-limit parameter entirely and always generate until they reach their absolute maximum context length, regardless of what the developer sets
BOpenAI's Chat Completions API deprecated `max_tokens` in favor of `max_completion_tokens` for use with reasoning models, and this newer parameter caps the combined total of the model's internal reasoning tokens plus its visible output tokens, not just the visible output alone
CReasoning models require the token limit to be set as a request header rather than a body parameter, since reasoning tokens are billed through a separate metering system
D`max_tokens` still works identically for reasoning models, and the failure described must be caused by an unrelated authentication or network issue
Correct answer: .
OpenAI's documentation states that `max_tokens` was deprecated in favor of `max_completion_tokens` for Chat Completions requests to reasoning models (the o-series), because reasoning models spend part of their output-token budget on internal reasoning tokens that are never returned as visible text, and `max_completion_tokens` caps that combined total of hidden reasoning tokens plus visible output tokens together, whereas the older `max_tokens` parameter is not accepted the same way for these models. The option claiming reasoning models ignore any limit and always run to the context maximum is wrong; a completion-token cap is still respected, it just needs the newer parameter name. The option describing a request header is wrong because the limit is still a body parameter, just renamed, not moved to a header or a separate metering channel. The option dismissing the failure as unrelated is wrong because the parameter mismatch is precisely a documented, reproducible cause of failures when porting a standard-model integration to a reasoning model.
AI & LLM Engineering · Building with LLM APIs · Card 059/064easy
A developer uses OpenAI's Structured Outputs feature with a strict JSON schema requiring specific fields, and sends a request the model judges unsafe to fulfill. Forcing a refusal into the exact JSON shape the developer's schema demands would not make sense for a decline message. According to OpenAI's documentation, how does the API handle this case?
AIt silently returns an empty JSON object matching the schema's field names but with every value set to null, requiring the developer to infer a refusal from the all-null pattern
BIt raises an HTTP error status, forcing the developer to catch an exception rather than reading a normal successful response body
CIt returns a normal successful response, but with a separate `refusal` field containing the model's refusal message, distinct from the structured `content` field, so the developer can detect and handle a decline without it needing to conform to the requested schema
DIt force-fits the refusal text into the requested schema's first required string field, leaving all other fields empty, so the response technically still validates against the schema
Correct answer: .
OpenAI's Structured Outputs documentation describes a dedicated `refusal` field on the response message: when the model declines a request under strict mode, the API returns this separate field carrying the refusal text instead of trying to force that refusal into the developer's requested JSON schema, letting the developer check for a refusal explicitly rather than parsing a schema-shaped decline. The option describing an all-null JSON object is wrong because a refusal is not represented as null-valued schema fields; it is a distinct, clearly labeled field outside the structured content. The option describing an HTTP error is wrong because a refusal is still delivered as part of a normal successful response object, not as a thrown exception. The option describing force-fitting refusal text into the schema's first field is wrong because that is exactly the awkward outcome Structured Outputs' dedicated refusal field was designed to avoid, by keeping refusal text separate from schema-conforming content entirely.
AI & LLM Engineering · Building with LLM APIs · Card 060/064easy
A company already runs its infrastructure on AWS and wants to use Claude models while keeping billing, IAM permissions, and networking inside its existing AWS account, rather than managing a separate API key and vendor relationship directly with Anthropic. According to Anthropic's documentation, how can this be done?
AIt cannot be done; Claude models are only available directly through Anthropic's own API, with no first-party availability through any major cloud provider's platform
BThe company must first export the Claude model weights from Anthropic and self-host them on their own AWS EC2 instances, since Anthropic does not offer inference as a managed cloud service
CThe company can only use Claude on AWS by routing every request through Anthropic's own API using an Anthropic API key, with AWS providing nothing beyond generic network transit
DClaude models are available as a first-party option through Amazon Bedrock, using largely the same Messages API request and response shape as Anthropic's direct API, but authenticated and billed through the company's existing AWS account (AWS credentials and IAM) instead of an Anthropic API key
Correct answer: .
Anthropic's documentation confirms Claude models are offered as a first-party option through Amazon Bedrock (and similarly through Google Cloud Vertex AI), where the request and response shapes closely mirror the Messages API used on Anthropic's direct API, but authentication, billing, and IAM permissions flow through the company's existing cloud account rather than a separate Anthropic API key -- exactly the setup this AWS-based company wants. The option claiming Claude is unavailable through any cloud provider is wrong because Bedrock (and Vertex AI) availability is a documented, current offering. The option describing self-hosting exported model weights is wrong because Anthropic does not distribute Claude's weights for self-hosting; Bedrock provides managed inference, not exported weights. The option claiming AWS provides nothing beyond generic network transit and requests must still use an Anthropic API key is wrong because Bedrock access is authenticated through AWS credentials and IAM, not an Anthropic-issued key, which is precisely the integration benefit being asked about.
Source: Anthropic, Claude on Amazon Bedrock documentation, https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock
AI & LLM Engineering · Building with LLM APIs · Card 061/064hard
A team is building two separate products on OpenAI's Realtime API: a browser-based voice assistant that runs directly against a user's laptop microphone and speakers, and a backend telephony integration that connects to a third-party phone system server-to-server with no browser involved. According to OpenAI's documentation, which transport should each choose, and why?
AThe browser-based client should use WebRTC, since it is designed for client-side, real-time media use and handles network jitter, audio device capture, and variable connection quality automatically; the server-to-server telephony integration should use WebSocket, which is better suited to a backend-to-backend connection without a browser's media stack involved
BBoth should use WebSocket, since WebRTC is only supported for text-only Realtime API sessions and cannot carry audio at all
CThe browser-based client should use WebSocket, since browsers cannot establish WebRTC connections without a third-party plugin; the backend integration should use WebRTC, since it offers lower latency for server-to-server traffic
DBoth should use WebRTC, since OpenAI's documentation states WebSocket support for the Realtime API has been fully removed and is no longer a valid transport option
Correct answer: .
OpenAI's Realtime API documentation recommends WebRTC for client-side applications such as browser or mobile voice apps, since WebRTC is built for real-time media and automatically handles network jitter, echo cancellation, and audio device capture that a raw client would otherwise have to implement itself, while it recommends WebSocket for server-to-server integrations, such as a backend connecting to a telephony provider, where there is no browser media stack involved and a simpler persistent socket connection is more appropriate. The option claiming WebRTC is text-only and cannot carry audio is wrong, since WebRTC's entire purpose in this API is real-time audio streaming. The option reversing the recommendation, sending the browser client over WebSocket and the backend over WebRTC, is wrong because browsers support WebRTC natively without plugins, and WebRTC's benefits are specifically about client-side media handling that a backend integration doesn't need. The option claiming WebSocket support was fully removed is wrong because both transports remain documented, valid options for different use cases, not a deprecated-versus-current pairing.
Source: OpenAI, Realtime API guide, https://developers.openai.com/api/docs/guides/realtime
AI & LLM Engineering · Building with LLM APIs · Card 062/064medium
A developer building against the Responses API wants to estimate, before sending a request, exactly how many input tokens a prompt will consume once images, uploaded files, and tool/function schemas are all counted in, so they can trim it if it risks exceeding the model's context window. Running the request's text through `tiktoken` locally only approximates plain text and cannot account for the images, files, or tool schemas. According to OpenAI's documentation, how can the developer get an exact count instead?
AIt cannot be done exactly for anything beyond plain text; images, files, and tool schemas can only ever be approximated, never counted precisely, before a request is actually sent
BThe developer must send the full request to the standard Responses endpoint with generation enabled and then read the exact input token count back from the completed response's usage data, since no pre-send option exists
CThe developer can call a dedicated input token counting endpoint that accepts the same input format as the Responses API (including images, files, and tool definitions) and returns the exact input token count without performing any model generation
DThe developer must call the embeddings endpoint on the request's text content first, since the returned embedding vector's dimensionality directly reveals the exact token count of the full multimodal input
Correct answer: .
OpenAI's documentation describes a dedicated input token counting endpoint that accepts the same input format as the Responses API, so a developer can submit the same text, images, files, and tool definitions they intend to send for real and get back an exact input token count, including formatting tokens for message roles and boundaries, without triggering any model generation. This is precisely what closes the gap that a local tool like `tiktoken` cannot: `tiktoken` only approximates plain text and has no way to account for image token costs, file processing, or tool-schema overhead. The option claiming exact counts are impossible for anything beyond plain text is wrong because this endpoint is documented to return exact counts for images and files, not just approximations. The option requiring a full real request with generation enabled is wrong because the whole point of the dedicated counting endpoint is to get the count without paying for or waiting on generation. The option describing embedding-vector dimensionality as revealing token count is wrong because embedding dimensionality is a fixed property of the embedding model's output size, unrelated to how many tokens the input contained, and would not account for images, files, or tools at all.
AI & LLM Engineering · Building with LLM APIs · Card 063/064easy
A startup's rate limits on OpenAI's API were quite low when they first signed up, but several months later, after steady usage and payment history, they notice their per-minute request and token limits have increased automatically without contacting support. According to OpenAI's documentation, what explains this?
ARate limits are fixed per API key for its entire lifetime and cannot change; the startup must be mistaken, or a separate new key with higher limits was silently issued
BOpenAI places organizations into usage tiers that automatically raise associated rate limits as the organization's cumulative spend and payment history grow, with no manual request required for standard tier progression
CRate limits only increase if a developer manually opens a support ticket requesting a specific new limit, and support silently approved and applied the change without a reply email
DRate limits are recalculated fully at random on a rolling basis to load-balance traffic across OpenAI's infrastructure, unrelated to any individual organization's usage history
Correct answer: .
OpenAI's documentation describes a system of usage tiers: organizations are automatically placed into a tier based on cumulative spend and payment history, and each tier carries higher associated rate limits, so as an organization's usage and billing history accumulate, its limits rise automatically without needing a manual support request for standard tier progression. The option claiming limits are permanently fixed per key is wrong because tier-based increases are a documented, expected behavior over an account's lifetime. The option requiring a manual support ticket is wrong because standard tier increases happen automatically based on usage and payment history, not solely through a support request. The option describing random rolling recalculation for load-balancing is wrong because tier assignment is tied to an individual organization's own usage and billing history, not to overall infrastructure load-balancing.
AI & LLM Engineering · Building with LLM APIs · Card 064/064medium
A team building a debugging-assistant UI on an OpenAI o-series reasoning model wants to show users some indication of how the model reasoned through a problem, to build trust in its answer, but OpenAI's documentation is clear that the model's raw, full chain-of-thought reasoning tokens are not exposed to developers. According to OpenAI's documentation, how can the team still surface something about the model's reasoning process to users?
AIt cannot be done at all; no information about the model's internal reasoning is ever available through the API in any form for o-series models
BThe team should set `logprobs` on the response, since log probabilities of the output tokens directly encode the full internal reasoning trace in numerical form
CThe team should fine-tune a separate, smaller model specifically to imitate and re-articulate the reasoning model's likely internal reasoning, since this is the only officially documented workaround
DThe team can request a reasoning summary via the API (for example, setting a summary option on the reasoning configuration), which returns a natural-language, abbreviated summary of the model's internal reasoning without exposing the full raw chain-of-thought tokens
Correct answer: .
OpenAI's documentation describes a reasoning-summary option available for o-series reasoning models through the Responses API's reasoning configuration, which returns a natural-language, abbreviated summary of the model's internal reasoning process without exposing the full raw chain-of-thought tokens themselves, giving developers something legitimate to surface in a UI without conflicting with OpenAI's stated policy of keeping raw reasoning traces hidden. The option claiming no reasoning information is ever available is wrong because the summary option is exactly the documented, officially supported partial view into reasoning. The option describing logprobs as encoding the full reasoning trace is wrong because logprobs describe the probability distribution over generated output tokens, not the model's internal, hidden reasoning process. The option describing fine-tuning a separate imitator model is wrong because it invents an unofficial workaround where a much simpler, directly documented API option already exists.