passdrill

Building with LLM APIs

64 cards · AI & LLM Engineering · answer each one, then read the explanation. Your score tallies below. Looking for Temperature vs top-p vs top-k: how they combine (worked example)? Read the explainer.

0 / 64 answered · 0 correct

AI & LLM Engineering · Building with LLM APIs · Card 001/064 easy

When an LLM API's "function calling" (tool use) feature returns a function call in its response, what actually happens next in a typical integration?

  1. The API executes the function directly on the vendor's infrastructure and returns only the final answer text, with no involvement from the calling application
  2. The response contains a structured description of the function name and the arguments the model wants to pass; the calling application must run the actual function itself and send the result back in a follow-up request
  3. The model runs the function inside its own weights using an internal code interpreter, producing the function's return value without any external execution step
  4. The request fails with an error unless the function has already been executed and its output included as part of the original prompt
AI & LLM Engineering · Building with LLM APIs · Card 002/064 easy

Many LLM provider APIs offer a "streaming" mode for text generation, typically implemented over Server-Sent Events (SSE). What does enabling streaming change about how a client receives the model's output?

  1. The client receives the complete response in a single payload, but compressed with gzip to reduce total bandwidth compared to a non-streaming request
  2. The model generates its answer in a single internal pass regardless of the setting; streaming only changes how the response is logged on the provider's servers, not how the client receives it
  3. The client receives the response as a sequence of incremental chunks over an open connection as tokens are generated, rather than waiting for generation to finish before anything is returned
  4. The client must poll a separate status endpoint at fixed intervals to check whether generation has completed, since no data is returned until the full response is ready
AI & LLM Engineering · Building with LLM APIs · Card 003/064 easy

An LLM API's "context window" limit of, say, 200,000 tokens applies to what, specifically, during a single API call?

  1. The combined total of the input tokens sent in the request (system prompt, conversation history, and any injected documents) plus the output tokens the model generates in response
  2. Only the tokens in the user's most recent message, since earlier turns in the conversation are automatically summarized and don't count against the limit
  3. Only the tokens the model generates in its response, since input tokens are processed by a separate, effectively unlimited ingestion pipeline
  4. The number of separate API requests a client can make per minute before being rate-limited
AI & LLM Engineering · Building with LLM APIs · Card 004/064 medium

A client application calling an LLM API in a tight loop starts receiving HTTP 429 responses. What is the generally recommended way to handle this, per common LLM provider API documentation?

  1. Immediately retry the exact same request as fast as possible in a loop, since 429 responses are transient and will resolve within milliseconds if retried aggressively
  2. Switch to a different, unrelated API endpoint entirely, since a 429 on one endpoint indicates that the provider's entire platform is unavailable
  3. Reduce the `max_tokens` parameter on the failing request, since 429 responses indicate the requested output would be too long to generate
  4. Back off and retry after a delay that increases with each subsequent failure (exponential backoff), typically with some added random jitter, rather than retrying immediately or at a fixed short interval
AI & LLM Engineering · Building with LLM APIs · Card 005/064 easy

How does a typical LLM provider's embeddings endpoint differ from its text-completion/chat-generation endpoint?

  1. The embeddings endpoint is simply a faster version of the generation endpoint that returns shorter natural-language answers to save on output tokens
  2. The embeddings endpoint takes text as input and returns a fixed-length numeric vector representing that text's meaning, rather than generating new natural-language text
  3. The embeddings endpoint only works on images, while the generation endpoint only works on text, so the two cannot be used on the same type of content
  4. The embeddings endpoint requires fine-tuning a custom model first, while the generation endpoint works with any base model out of the box
AI & LLM Engineering · Building with LLM APIs · Card 006/064 easy

What does the `stop` (or "stop sequences") parameter available in most LLM generation APIs do?

  1. It tells the API to stop generating further tokens as soon as any one of a specified list of strings appears in the output, and to exclude that string from the returned text
  2. It sets a maximum wall-clock time limit in seconds after which the API forcibly terminates the connection regardless of how much text has been generated
  3. It specifies a list of words the model is never permitted to generate anywhere in its response, causing an error if any of them would otherwise be produced
  4. It pauses generation partway through and waits for the client to send an approval signal before continuing to generate the remainder of the response
AI & LLM Engineering · Building with LLM APIs · Card 007/064 hard

A developer is using Anthropic's Messages API with two tools defined, `get_weather` and `send_email`, and wants to guarantee that Claude calls `get_weather` specifically on this turn rather than answering in plain text or calling `send_email`. According to Anthropic's tool-use documentation, how is this accomplished?

  1. By removing `send_email` from the `tools` array entirely for this request, since `tool_choice` can only force "any" tool use, not a specific named tool
  2. By setting `tool_choice` to `{"type": "any"}`, which restricts the model to only the first tool listed in the `tools` array
  3. By setting `tool_choice` to `{"type": "tool", "name": "get_weather"}`, which forces the model to call that specific named tool rather than deciding on its own or calling a different one
  4. By adding the instruction "you must call get_weather" only to the system prompt, since `tool_choice` itself has no mechanism for naming a specific tool
AI & LLM Engineering · Building with LLM APIs · Card 008/064 medium

A developer building a multi-turn conversational app compares OpenAI's Chat Completions API to its Responses API. According to OpenAI's documentation, what is a key difference in how each manages conversation state across turns?

  1. Chat Completions automatically stores and threads every conversation server-side with no client involvement, while the Responses API requires the client to resend the entire message history on every call
  2. Both APIs require the exact same manual approach: the client must always reconstruct and resend the full list of prior user and assistant messages with every request, with no built-in alternative in either API
  3. The Responses API has no way to maintain multi-turn context at all, and is only suitable for single-turn, stateless requests unrelated to any previous exchange
  4. With Chat Completions, the client must append prior turns into the message array and resend the full history each call, while the Responses API can instead reference a prior turn via a `previous_response_id` parameter so the server carries the context forward
AI & LLM Engineering · Building with LLM APIs · Card 009/064 medium

An LLM-powered agent has a tool that charges a customer's payment method, and the agent's HTTP call to your backend times out after the charge has actually already been processed. The agent's retry logic then calls the same tool again with the same arguments. What design choice prevents this from resulting in a duplicate charge?

  1. Having the tool accept a unique idempotency key per logical operation, so the backend can recognize a retried call with the same key and return the original result instead of processing the charge a second time
  2. Increasing the model's `max_tokens` limit, so the agent has enough space to reason more carefully about whether a retry is safe before calling the tool again
  3. Lowering the model's `temperature` to 0, so the agent always generates the exact same tool call arguments and therefore never issues an unintended duplicate request
  4. Disabling the agent's ability to call tools more than once per conversation, so any repeated call is rejected outright regardless of what happened to the first one
AI & LLM Engineering · Building with LLM APIs · Card 010/064 easy

Some LLM APIs support returning `logprobs` (log probabilities) alongside generated text. What do these values represent, and what are they typically used for?

  1. They report how many milliseconds the API took to generate each token, and are used purely for latency monitoring and performance debugging
  2. They report, for each generated token, the log probability the model assigned to it (and often to alternative candidate tokens), and are typically used to gauge the model's confidence or build classifiers from token likelihoods
  3. They report the total dollar cost billed for each individual token, broken out on a per-token basis for detailed cost accounting
  4. They report which of several fine-tuned model versions actually generated each token, for use in auditing which checkpoint produced a given response
AI & LLM Engineering · Building with LLM APIs · Card 011/064 medium

An LLM generation API exposes both a `top_p` (nucleus sampling) parameter and, in some APIs, a `top_k` parameter, in addition to `temperature`. How does nucleus sampling with `top_p` differ from `top_k` sampling in how it restricts the model's next-token choices?

  1. `top_p` and `top_k` are two names for the exact same mechanism, differing only in whether the cutoff value is expressed as a percentage or as a raw integer
  2. `top_k` restricts choices based on cumulative probability mass, so its cutoff point moves depending on how confident the model is at each step, while `top_p` always keeps exactly the same fixed number of candidates
  3. `top_p` restricts sampling to the smallest set of most-probable tokens whose cumulative probability reaches the threshold `p`, so the number of candidates varies step to step, while `top_k` always keeps a fixed number of the highest-probability tokens regardless of how the probability mass is distributed
  4. Both parameters only take effect when `temperature` is set to exactly 0, and have no effect on generation at any other temperature value
AI & LLM Engineering · Building with LLM APIs · Card 012/064 hard

A team needs to run sentiment classification over 40,000 archived support tickets and is not latency-sensitive, but wants to minimize per-request cost. According to OpenAI's Batch API documentation, what tradeoff does using the Batch API (instead of the standard synchronous chat completions endpoint) involve?

  1. The Batch API charges the same per-token price as the synchronous API but guarantees results within 60 seconds regardless of batch size
  2. The Batch API is free of charge for any volume of requests, but results are only available after a mandatory 7-day waiting period
  3. The Batch API only accepts a single request per batch, so the team would still need to submit 40,000 separate batch jobs to process all the tickets
  4. The Batch API offers roughly a 50% cost discount compared to the synchronous API, but processes the submitted batch of requests asynchronously with results typically available within 24 hours rather than immediately, and draws from a separate rate-limit pool
AI & LLM Engineering · Building with LLM APIs · Card 013/064 medium

A developer building a chat feature needs the model's reply to reliably parse as a JSON object matching a specific schema, with no missing required fields and no invalid enum values. According to OpenAI's documentation, how do its `json_object` response format and its `json_schema` Structured Outputs mode (with `strict` set to true) differ in what they actually guarantee?

  1. `response_format` `json_object` mode and `json_schema` mode with `strict` enabled provide the exact same guarantee: both ensure the output matches a developer-supplied schema exactly, including all required fields and valid enum values
  2. `json_object` mode guarantees the output matches the developer's schema exactly, while `json_schema` mode with `strict` enabled guarantees only that the output is syntactically valid JSON, with no schema enforcement
  3. `json_object` mode guarantees only that the output is syntactically valid JSON, with no guarantee it matches any particular structure, while `json_schema` mode with `strict` set to true constrains generation so the model reliably includes every required field and only valid enum values as defined in the developer's supplied schema
  4. both modes only take effect when `temperature` is also set to 0; without that, neither has any influence on whether the output is valid JSON or matches a schema
AI & LLM Engineering · Building with LLM APIs · Card 014/064 easy

A developer comparing Anthropic's Messages API to some other chat-style LLM APIs notices a structural difference in how the system prompt is supplied. How does Anthropic's Messages API handle the system prompt?

  1. Anthropic's Messages API accepts a top-level `system` parameter, kept separate from the `messages` array, to carry the system prompt, whereas some other chat-style APIs instead embed the system prompt as a message with a system or developer role placed inside the same messages list
  2. Anthropic requires the system prompt to be the first entry in the `messages` array with a role of `system`, exactly matching how some other chat-style APIs structure a request
  3. Anthropic's Messages API has no mechanism for supplying a system prompt at all; instructions can only be embedded inside the first user message
  4. the `system` parameter selects which dated model snapshot handles the request, and has nothing to do with supplying instructions to the model
AI & LLM Engineering · Building with LLM APIs · Card 015/064 hard

A developer calls a reasoning model through OpenAI's Responses API with `max_output_tokens` set, and gets back a response whose `output_text` is empty even though the usage data shows a large number of billed output tokens. According to OpenAI's documentation, what is most likely happening, and how should the application detect it?

  1. the response's `status` field will be `error` with an `error.type` of `server_error`, indicating an internal fault on OpenAI's infrastructure that is unrelated to the token limit set on the request
  2. this cannot happen when `max_output_tokens` is set, since OpenAI's documentation guarantees a visible answer is always produced in full before any internal reasoning tokens are counted against the limit
  3. the response's `output` array will always still contain a text message item in this situation, so an application never needs to check anything beyond `output_text` before using the result
  4. the response's `status` field will be `incomplete`, with `incomplete_details.reason` set to `max_output_tokens`, because a reasoning model's internal reasoning tokens can consume the entire token budget before any visible answer is produced; an application should check `status` and `incomplete_details` rather than assuming `output_text` is populated, and should reserve enough budget for both reasoning and the visible answer
AI & LLM Engineering · Building with LLM APIs · Card 016/064 easy

LLM generation APIs commonly expose a `temperature` parameter alongside sampling-restriction parameters like `top_p` or `top_k`. Mechanically, what does `temperature` itself do to the model's next-token choice?

  1. temperature removes the least probable tokens from consideration entirely, functioning identically to `top_k` sampling with the temperature value acting as the value of k
  2. temperature scales the model's output logits before they are converted into probabilities via softmax; lower values sharpen the resulting distribution toward the highest-probability tokens, while higher values flatten it, making comparatively less probable tokens more likely to be sampled
  3. temperature is applied only after a token has already been sampled, triggering a re-roll whenever a separate filter model flags the chosen token as low quality
  4. temperature has no effect on which token is sampled and instead controls only how many separate completions the API returns for a single request
AI & LLM Engineering · Building with LLM APIs · Card 017/064 easy

A developer wants to send an image alongside a text question to a model through OpenAI's Responses API. According to OpenAI's current documentation, how is the image included in the request?

  1. the request's `input` array includes a content item with `type` set to `input_image`, whose `image_url` field can be either a fully qualified URL pointing to an image or a base64-encoded data URL, placed alongside a separate `input_text` item carrying the text portion of the same message
  2. images can only be referenced by a fully qualified URL; base64-encoded image data placed directly in the request body is rejected by the endpoint
  3. images must be attached as raw binary data outside the JSON request body, since the `input` array can only ever contain plain text strings
  4. images are supplied through the `messages` array using a content type of `image_url` only, with no `input_image` type and no base64 option available
AI & LLM Engineering · Building with LLM APIs · Card 018/064 hard

A developer marks a `cache_control` breakpoint on a large, static system prompt sent with every request to Anthropic's Messages API. According to Anthropic's documentation, how does this affect billing across the first request and later requests that reuse the same cached prefix?

  1. the first request's cache-write tokens are billed at a discount below the base input-token rate, while every later request that hits the same cache is billed at a premium above the base rate, since maintaining a cache costs more than reading fresh input
  2. cache writes and cache hits are always billed at exactly the same per-token rate as ordinary, uncached input tokens, so prompt caching only ever provides a latency benefit and never a cost benefit
  3. the first request that establishes the cache is billed at a premium above the base input-token rate for the cached portion, since writing a new cache entry costs more, while later requests whose prefix hash matches that cached entry are billed at a steep discount for the reused portion, which is where the overall cost savings come from
  4. prompt caching only ever reduces the cost of output tokens, never input tokens, because it works by shortening how much text the model generates in its response rather than by reusing previously processed input
AI & LLM Engineering · Building with LLM APIs · Card 019/064 medium

A developer needs Claude to analyze both the text and the visual layout, charts, and images on each page of a PDF report, not just extract raw text. According to Anthropic's documentation, how is a PDF supplied to the Messages API?

  1. PDFs cannot be sent as input to the Messages API at all; the developer's own application must first convert the PDF to plain text before any part of it can be included in a request
  2. a message's content array can include a block of type `document`, whose source can be a base64-encoded PDF, a URL pointing to a hosted PDF, or a `file_id` from the Files API, letting the model reason over both the extracted text and each page's visual layout, charts, and images
  3. PDFs are supported only via a separate proprietary file format that the developer must first produce from the PDF using a vendor-provided offline conversion tool before uploading
  4. the `document` content block only extracts and returns the PDF's raw text back to the developer as a standalone response; it cannot be used as part of a prompt for the model to reason over
AI & LLM Engineering · Building with LLM APIs · Card 020/064 medium

An agent built on OpenAI's Responses API needs to look up two independent pieces of information in a single turn, using the same tool twice with different arguments. According to OpenAI's documentation, how does the API support this, and what must the client do in response?

  1. a single assistant turn can never include more than one function call; the model must wait for the client to answer the first call before it is permitted to request the second one
  2. the client may return one combined `function_call_output` message covering every function call made in that turn, as long as the message lists all of the relevant call IDs together
  3. `parallel_tool_calls` can only ever be set to true, since the API provides no way to restrict a response to at most one function call per turn
  4. a single assistant turn can include multiple function calls, each carrying its own `call_id`, and the client must send back one `function_call_output` item per `call_id` containing that call's result; setting `parallel_tool_calls` to false instead restricts the model to at most one function call per turn
AI & LLM Engineering · Building with LLM APIs · Card 021/064 easy

Every request to Anthropic's Messages API must include an `anthropic-version` header, such as `anthropic-version: 2023-06-01`. According to Anthropic's documentation, what does this header actually do?

  1. the header is purely optional and informational; omitting it, or sending an arbitrary unrecognized string, has no effect on how the request is processed
  2. the header selects which specific model snapshot answers the request, serving as an alternative to setting the `model` field in the request body
  3. the header is required on every request and pins that request to a documented API version; Anthropic's versioning policy preserves a given version's existing input and output parameters while allowing additive changes, such as new optional inputs or new output values, so pinning a version protects an integration from breaking changes introduced by later versions
  4. the header's value must change to a new date every single day, since each calendar day's API responses use an incompatible request and response format from the previous day
AI & LLM Engineering · Building with LLM APIs · Card 022/064 easy

Anthropic's Messages API can return either a 429 status with a `rate_limit_error` type or a 529 status with an `overloaded_error` type. According to Anthropic's documentation, what is the practical difference between these two conditions?

  1. a 429 with a `rate_limit_error` type means the calling organization's own limits, such as a requests-per-minute cap or a configured spend limit, have been exceeded, while a 529 with an `overloaded_error` type means the API itself is temporarily overloaded across all users, a capacity condition unrelated to that organization's own usage
  2. 429 and 529 both describe exactly the same underlying condition, and application error-handling code does not need to distinguish between them
  3. a 529 status means the calling application's API key has been revoked, while a 429 status means the request body was malformed and rejected by validation
  4. a 429 status can only ever be returned by the Batch API, while a 529 status can only ever be returned by the standard synchronous Messages API, so the two are distinguished solely by which endpoint returned them
AI & LLM Engineering · Building with LLM APIs · Card 023/064 medium

Anthropic's Messages API exposes a separate `/v1/messages/count_tokens` endpoint. According to Anthropic's documentation, what does this endpoint let a developer do, and what does calling it cost?

  1. it returns the exact monetary cost, in the organization's billing currency, of a hypothetical request, in addition to a token count
  2. calling it consumes the same input-token allowance and counts fully against the same per-minute rate limit as an ordinary message request, because it internally runs the full generation pipeline
  3. it can count tokens only in plain text messages, and cannot account for the additional tokens that tool definitions or images would add to an actual request
  4. it lets a developer estimate how many tokens a prospective request would consume, including messages, tool definitions, and images, without actually creating a message or generating any output, which is useful for pre-flight cost estimation and for checking that a request will fit inside the model's context window
AI & LLM Engineering · Building with LLM APIs · Card 024/064 easy

OpenAI offers both a plain model name (an alias, such as a model's short identifier) and a dated snapshot identifier for the same underlying model family. According to OpenAI's documentation, how do these two kinds of identifiers differ, and why might a production application choose one over the other?

  1. a dated snapshot identifier automatically updates over time to point at the newest underlying model, while a plain alias always stays pinned to whichever exact version existed when the application first started calling it
  2. a plain alias name is a moving pointer that OpenAI can repoint to a newer underlying snapshot over time, while a dated snapshot identifier stays fixed to one specific model version indefinitely; production applications often pin to a dated snapshot specifically to avoid the model's behavior changing underneath them without warning
  3. aliases and dated snapshots are billed at different per-token prices for otherwise identical model capability, with the alias name always priced lower than any dated snapshot of the same model
  4. dated snapshot identifiers are available only to enterprise customers under a custom contract, while every other developer can only ever call the plain alias name
AI & LLM Engineering · Building with LLM APIs · Card 025/064 medium

A developer building an OpenAI Responses API integration wants the model to call some function on this turn — any of the tools it has available — but does not want to pin the choice to one specific named function, and also does not want the model skipping tool use entirely and answering in plain text instead. According to OpenAI's documentation, which tool_choice setting achieves this?

  1. Setting tool_choice to "required", which forces the model to call at least one of the available functions on this turn without pinning the choice to any particular named function
  2. Setting tool_choice to "auto", since forcing the model toward some unspecified tool without naming one is not actually possible with any documented tool_choice value
  3. Setting tool_choice to an object naming one specific function, since OpenAI provides no way to require some tool call without also naming which tool must be called
  4. Setting tool_choice to "none", since "none" only prevents the model from producing a free-form text answer while still allowing it to call a tool if it chooses to
AI & LLM Engineering · Building with LLM APIs · Card 026/064 hard

A developer using Anthropic's Messages API applies a cache_control breakpoint to a large, static system prompt, and is deciding between the default ephemeral cache lifetime and the optional extended lifetime enabled by adding "ttl": "1h" to that breakpoint. According to Anthropic's documentation, how do the two lifetimes differ in duration and in what the request that writes to the cache is charged, and does either lifetime change what a later cache hit costs?

  1. The default cache lasts 1 hour, the extended option shortens it to 5 minutes, and both a 5-minute and a 1-hour cache write are billed at the same 1.25x multiplier over the base input token price
  2. Both lifetimes cost the same to write to the cache, but choosing the 1-hour option charges double the normal input price on every later cache read for as long as that entry stays cached
  3. The default cache lasts 5 minutes and the extended option lasts 1 hour; writing to a 5-minute cache costs 1.25x the base input token price while writing to a 1-hour cache costs 2x that price, but a later cache read is billed at the same reduced fraction of the base price regardless of which TTL wrote the entry
  4. The default cache lasts 5 minutes and the extended option lasts 1 hour, but the TTL only changes how long the entry survives — Anthropic charges the identical price for a cache write regardless of which TTL is chosen, and cache reads are billed at the full base input token price
AI & LLM Engineering · Building with LLM APIs · Card 027/064 medium

A developer sends a high-resolution photograph to a vision-capable OpenAI model with the image detail setting left at high (or auto), rather than set to low. According to OpenAI's documentation, how is that image converted into billable tokens?

  1. The image is billed purely by its file size in kilobytes, independent of its pixel dimensions or the requested detail setting
  2. The image is first scaled down to fit within a maximum pixel dimension, then divided into a grid of fixed-size tiles, and the final token count is the model's fixed base token cost plus a per-tile token cost multiplied by the number of tiles needed to cover the scaled image
  3. At high or auto detail the image always costs exactly the same fixed number of tokens it would cost at low detail, since the detail setting only affects output quality rather than billed input tokens
  4. The image is billed at one flat per-image token cost that is identical across every vision-capable model regardless of resolution or detail setting
AI & LLM Engineering · Building with LLM APIs · Card 028/064 easy

A developer using OpenAI's text-embedding-3-large model wants shorter vectors to cut vector-database storage costs, so they call the embeddings endpoint with the dimensions parameter set below the model's native output size. According to OpenAI's documentation, why can the resulting shortened embedding still remain highly useful for retrieval instead of degrading arbitrarily once truncated?

  1. The dimensions parameter runs a separate, smaller embedding model that was trained from scratch specifically to output that exact requested size
  2. Shortening only ever removes dimensions that carry formatting metadata rather than semantic content, so no meaningful information is discarded no matter how small a size is requested
  3. The API automatically applies a separate dimensionality-reduction algorithm, such as PCA, to the full-size embedding after generation to compress it down to the requested size
  4. The embedding models are trained using Matryoshka Representation Learning, a technique that concentrates the most important concept-representing information toward the earlier dimensions of the vector, so truncating the trailing dimensions preserves most of the useful signal instead of degrading arbitrarily
AI & LLM Engineering · Building with LLM APIs · Card 029/064 hard

A developer enables Anthropic's citations feature on a large source document passed to the Messages API, so that the model's response includes text blocks with citations pointing back to exact passages in that document. According to Anthropic's documentation, how does the cited_text field returned in each citation affect billed tokens, and how does a citation's location reference differ between a plain text document and a PDF document?

  1. The cited_text field does not count toward billed output tokens, nor toward input tokens if it is passed back into a later turn, and citation locations are given as character index ranges for a plain text document but as a page number range for a PDF
  2. The cited_text field is billed as ordinary output tokens exactly like the rest of the response text, and citation locations use the same character index ranges regardless of whether the source was a plain text document or a PDF
  3. Enabling citations always incurs a separate flat per-request surcharge on top of ordinary token pricing, and citation locations are returned as an arbitrary chunk identifier that carries no positional information within the document
  4. The cited_text field does not count toward billed output tokens, but citation locations are identical in format between a plain text document and a PDF, both expressed as character index ranges
AI & LLM Engineering · Building with LLM APIs · Card 030/064 easy

A developer building a high-volume integration with Anthropic's Messages API wants to throttle their own request rate proactively, before ever triggering a 429 rate-limit error, by tracking how much of the per-minute request and token allowance remains. According to Anthropic's documentation, how can the application learn this without first waiting for a 429 response?

  1. It cannot be done proactively; the only way to learn how close the account is to its rate limit is to keep sending requests until a 429 response is returned
  2. The Claude Console dashboard is the only place this information is exposed; it is not returned in any HTTP response header on ordinary, successful API calls
  3. Every Messages API response, not only a 429, includes headers such as anthropic-ratelimit-requests-remaining and anthropic-ratelimit-tokens-remaining along with a corresponding reset time, so the application can read these headers on successful responses and slow down before it is actually rate limited
  4. The remaining-capacity information is included only in the JSON response body of a successful request, never in response headers, so the application must parse the body of every response to extract it
AI & LLM Engineering · Building with LLM APIs · Card 031/064 easy

A developer calls OpenAI's Chat Completions API with the n parameter set to 3 instead of leaving it at its default. According to OpenAI's documentation, what does this change, and how does it affect billed usage?

  1. It sends the same request in parallel to three different underlying model versions and returns whichever response arrives first, at no extra token cost since only one response is kept
  2. It generates three independent chat completion choices for the same input message in a single request, and the developer is billed for the output tokens generated across all three choices combined, not just one
  3. It repeats the user's input message three times within the same prompt before generating a single completion, which increases input token billing but leaves output token billing unchanged
  4. It sets the maximum number of follow-up turns allowed in a single conversation thread to three before the API automatically ends the conversation
AI & LLM Engineering · Building with LLM APIs · Card 032/064 easy

An application reads the stop_reason field on a Claude Messages API response before deciding whether to continue a multi-turn agent loop. According to Anthropic's documentation, which of the following correctly matches a stop_reason value to what it indicates happened?

  1. A stop_reason of "stop_sequence" means the response was truncated because it reached the max_tokens limit specified in the request
  2. A stop_reason of "tool_use" means the model refused to answer the request for policy reasons and produced no other usable content
  3. A stop_reason of "max_tokens" means the model matched one of the custom strings supplied in the stop_sequences parameter
  4. A stop_reason of "tool_use" means the model generated one or more tool_use content blocks and is waiting for the application to run those tools and return results before it continues
AI & LLM Engineering · Building with LLM APIs · Card 033/064 medium

A developer who has only built agent loops against Anthropic's Messages API is porting the same tool-calling logic to OpenAI's Responses API, and needs to submit the output of a tool the model just called back into the conversation. According to each vendor's documentation, how does the mechanism for submitting a tool result differ between the two APIs?

  1. Anthropic's Messages API has no separate role for tool results — the result is sent as an ordinary user-role message containing a tool_result content block that references the original call's tool_use_id, whereas OpenAI's Responses API instead uses a dedicated function_call_output item carrying the original call_id and an output value, without wrapping it in a user-role message
  2. Both APIs require the result to be submitted using a dedicated "tool" role message type that behaves identically between the two vendors, so the same request body can be reused unmodified
  3. Anthropic's Messages API requires a dedicated "tool" role message referencing tool_use_id, while OpenAI's Responses API instead expects the tool result to be appended as plain user-role text with no structured reference to which call it answers
  4. Neither API has any structured way to link a tool result back to the specific tool call it answers; correlation is inferred purely from the order in which messages appear in the conversation
AI & LLM Engineering · Building with LLM APIs · Card 034/064 easy

A team is building an AI agent that needs to connect to several external systems — a calendar, a database, and a search engine — and wants to avoid writing a separate bespoke integration for each one against each LLM vendor's own proprietary function-calling schema. What does adopting the Model Context Protocol (MCP) let them do instead?

  1. MCP is a proprietary Anthropic-only feature of the Messages API that cannot be used with any other model provider's product
  2. MCP replaces an LLM's own function-calling or tool-use mechanism entirely, so a model connected through MCP no longer needs any tools parameter or tool-call-style content blocks at all
  3. MCP is an open, vendor-neutral client-server protocol that standardizes how an AI application connects to external tools, data sources, and prompts, so a single MCP server implementation for a given system can be reused across different compliant AI applications instead of writing a separate bespoke connector per vendor
  4. MCP is a data format for packaging training data used to fine-tune a model on a company's internal documents, rather than something used at inference time
AI & LLM Engineering · Building with LLM APIs · Card 035/064 medium

A developer calling Anthropic's Messages API gives Claude two independent tools, and on a single turn Claude decides both are needed to answer the user's question. According to Anthropic's documentation, how does the API surface this, and what must the client do in its next request?

  1. The assistant response can contain multiple tool_use content blocks in the same turn, one per tool call, each with its own id; the client must run each tool and reply with a separate tool_result block carrying the matching tool_use_id for each call, with all tool_result blocks placed before any other content in that user message
  2. Claude can only request one tool per turn, so it deliberately picks the single most useful tool and defers the second lookup to a follow-up turn after receiving the first tool_result
  3. The API automatically merges both tool calls into a single tool_use block with a combined input object, and the client returns one tool_result covering both calls at once
  4. Parallel tool calls are only available through the separate Message Batches API, so a normal synchronous Messages API request always forces Claude to pick a single tool per turn
AI & LLM Engineering · Building with LLM APIs · Card 036/064 easy

Anthropic's Messages API supports both user-defined client tools (like a custom get_weather function) and built-in server tools (like web_search or code_execution). According to Anthropic's documentation, what is the key operational difference between the two?

  1. Server tools can only be used through the Message Batches API, while client tools work only with standard synchronous Messages API requests
  2. A client tool's call must be executed by the developer's own application code, with the result sent back in a follow-up request, whereas a server tool runs on Anthropic's own infrastructure and its result is included directly in that same response, without the client executing anything or making another API call
  3. Server tools give Claude the ability to rewrite its own system prompt mid-conversation, while client tools cannot alter the conversation's configuration at all
  4. Client tool results are limited to plain text, while server tools are the only tool type whose tool_result can contain images or documents
AI & LLM Engineering · Building with LLM APIs · Card 037/064 medium

A developer needs to run overnight sentiment analysis on a large batch of saved support transcripts using Claude, with no need for an immediate reply to any single request. According to Anthropic's documentation, how does the Message Batches API structure this workload, and what does it cost relative to the standard Messages API?

  1. Every request must be sent one at a time over a persistent WebSocket connection that stays open until all results have streamed back
  2. Each request is submitted individually to the standard /v1/messages endpoint but tagged with a batch_id header, and Anthropic bills these at the same per-token rate as any other synchronous request
  3. The developer submits a single batch request containing an inline array of individual Messages requests, each tagged with its own custom_id so its result can be matched back afterward since result order is not guaranteed; the batch processes asynchronously (most batches finish in under an hour) and is billed at a 50% discount off standard per-token pricing once results are retrieved
  4. Batches replace token-based pricing entirely with one flat fee per batch, regardless of how many requests or tokens it contains
AI & LLM Engineering · Building with LLM APIs · Card 038/064 hard

A developer sends Claude a 1000x1000 pixel JPEG image (within the standard resolution tier, no downscaling triggered) alongside a text question through the Messages API. According to Anthropic's documentation, how is the number of image tokens this costs actually calculated?

  1. Anthropic calculates image tokens the same way as OpenAI: the image is divided into fixed 512x512-pixel tiles, with each tile plus a fixed base cost added to the total
  2. Anthropic charges a flat token fee per image regardless of resolution, so a 1000x1000 image costs exactly the same number of tokens as a 200x200 image
  3. The image's base64-encoded string is billed directly as though its character count were ordinary text input tokens
  4. Claude views an image as a grid of 28x28-pixel patches, with each patch counted as one visual token, giving a cost of ceil(width / 28) x ceil(height / 28); for a 1000x1000 image this works out to 1,296 visual tokens
AI & LLM Engineering · Building with LLM APIs · Card 039/064 easy

A developer's agent repeatedly includes the same large reference PDF in every turn of a long multi-turn conversation with Claude through the Messages API. According to Anthropic's documentation, how does the Files API help avoid the overhead of resending that PDF's bytes on every request?

  1. The Files API automatically detects duplicate attachments already present in the conversation history and silently deduplicates them, with no change needed to how the file is referenced in the request
  2. The developer uploads the PDF once to get back a file_id; later Messages API requests reference that file_id in a document content block instead of re-embedding the file's base64 bytes each turn, and the upload, download, list, retrieve, and delete operations themselves are free — only the file's content actually used within a Messages request is billed, as input tokens
  3. Files uploaded through the Files API are cached for exactly five minutes and must be re-uploaded after that window before they can be referenced again
  4. The Files API only accepts image files, so a PDF must still be base64-encoded and resent in full on every request regardless of its size
AI & LLM Engineering · Building with LLM APIs · Card 040/064 easy

A developer sets the same seed integer across repeated calls to OpenAI's Chat Completions API, keeping every other parameter identical, hoping for reproducible outputs. According to OpenAI's documentation, what guarantee does this actually provide, and what does the system_fingerprint field returned with each response help the developer detect?

  1. The seed parameter makes the system attempt deterministic sampling on a best-effort basis rather than offering a strict guarantee, so repeated requests are only mostly identical; system_fingerprint identifies the current backend model and infrastructure configuration, so a change in its value between calls signals a backend change that can explain why outputs stopped matching
  2. The seed parameter is a strict, guaranteed source of bit-for-bit identical output on every call, and system_fingerprint is simply a random request-tracing identifier unrelated to reproducibility
  3. The seed value determines which fine-tuned model snapshot handles the request, while system_fingerprint reports the end user's device or browser fingerprint for analytics purposes
  4. Setting a seed disables sampling entirely and forces greedy decoding, so the temperature and top_p parameters no longer have any effect on the response
AI & LLM Engineering · Building with LLM APIs · Card 041/064 easy

A team wants to build a natural, low-latency spoken-voice agent that can be interrupted mid-response ('barge-in') and hear the user talking while it's still speaking. According to OpenAI's documentation, why is the Realtime API a better fit for this than streaming text responses from the Chat Completions or Responses API over server-sent events?

  1. Server-sent event streaming already supports full-duplex audio in both directions, so the Realtime API differs only in offering a wider range of voices
  2. The Realtime API works by converting speech to text, calling the standard Chat Completions endpoint in a fast loop, and converting the reply back to speech, which is functionally the only difference from building this pipeline yourself
  3. The Chat Completions API's server-sent event streaming already maintains a persistent connection and conversational state across turns identical to the Realtime API, so switching would only change pricing
  4. The Realtime API keeps a persistent, stateful WebSocket or WebRTC session that streams audio natively in both directions, supporting low first-audio latency and barge-in, whereas standard chat APIs follow a request/response model where server-sent events stream text tokens in one direction per call, with no native audio support
AI & LLM Engineering · Building with LLM APIs · Card 042/064 easy

A developer wants to screen user-submitted text and images for policy violations (hate speech, self-harm content, violence, and similar categories) before passing them to a chat model, without paying for a full model generation just to get a safety judgment. According to OpenAI's documentation, what does the Moderation API provide for this?

  1. The Moderation API is a paid feature billed at the same per-token rate as chat completions, since it must run a full generation to judge the content
  2. There is no separate moderation endpoint; classification is only available by prompting a regular chat model with a custom system prompt asking it to judge the content
  3. A separate, free-to-use moderation endpoint that classifies text and/or image input and returns a flagged boolean along with a dictionary of per-category violation flags and per-category confidence scores, without generating any chat response
  4. The Moderation API only accepts and classifies images, so text content must be screened through a different, paid endpoint instead
AI & LLM Engineering · Building with LLM APIs · Card 043/064 hard

A team building a customer-support triage bot on an OpenAI reasoning model finds responses too slow and expensive for high-volume simple classification, while a separate internal tool doing complex multi-step debugging analysis on the same model family needs much more thorough internal reasoning. According to OpenAI's documentation, how does the reasoning-effort setting (reasoning.effort / reasoning_effort) address both cases, and how are the tokens spent on that internal reasoning billed?

  1. Effort is a tunable knob, ranging from low through high depending on the model, that trades response latency and cost against how thoroughly the model reasons before answering — lower effort suits fast, simple tasks like triage, higher effort suits complex, less latency-sensitive analysis — and the tokens spent on that internal reasoning are billed as output tokens even though the reasoning content itself is not shown back to the caller
  2. Reasoning effort only controls the model's writing tone, becoming more formal at higher settings, and has no effect on latency, cost, or how much internal reasoning the model performs
  3. Reasoning tokens spent at any effort level are provided completely free of charge, since they are considered part of the model's internal process rather than part of the visible response
  4. Effort must be set identically across every request in an organization; it cannot be varied per request based on that request's task complexity
AI & LLM Engineering · Building with LLM APIs · Card 044/064 medium

A developer used to Anthropic's Messages API, where a cache_control breakpoint must be explicitly added to a prompt to enable prompt caching, moves the same kind of application onto OpenAI's API. According to OpenAI's documentation, what does the developer need to do to get prompt caching benefits there?

  1. Exactly the same as Anthropic: an explicit cache breakpoint marker must be added to the prompt, or caching never activates on OpenAI's models
  2. Nothing extra in general: prompt caching is enabled by default on supported OpenAI models with no special parameter required, automatically discounting the portion of a repeated prompt prefix that qualifies once it reaches the model's minimum cacheable-length threshold; only on newer model generations can the developer optionally choose between implicit and explicit cache breakpoints via a caching-options parameter
  3. OpenAI does not offer any form of prompt caching; only Anthropic's Messages API supports this feature
  4. Prompt caching on OpenAI's API only applies to the embeddings endpoint and has no effect on chat completion requests
AI & LLM Engineering · Building with LLM APIs · Card 045/064 medium

A developer building on Claude's Messages API enables extended thinking with `thinking: {"type": "enabled", "budget_tokens": 4000}` alongside `max_tokens: 8000`. According to Anthropic's documentation, what does `budget_tokens` actually control, and how is the model's internal reasoning billed?

  1. `budget_tokens` is a hard ceiling Claude's internal reasoning can never cross, and thinking tokens are billed at a separate, discounted rate that does not count toward the response's output-token total
  2. `budget_tokens` sets a target for how many tokens Claude's internal reasoning may use, but Claude can finish well under that target, `max_tokens` remains the actual hard ceiling on the turn's combined reasoning-plus-answer output, and the tokens spent reasoning are billed as ordinary output tokens, reported separately in the response's usage data
  3. `budget_tokens` only controls how much of Claude's reasoning is shown back to the developer in the response; the full reasoning is always generated and billed at the same length no matter what value is configured
  4. `budget_tokens` is a free reasoning allowance that Anthropic does not bill at all, since only the final text answer Claude produces counts toward billed output tokens
AI & LLM Engineering · Building with LLM APIs · Card 046/064 medium

A developer adds Anthropic's code execution tool to a Messages API request so Claude can run Python against an uploaded CSV file. The same request also defines a client tool, `send_slack_message`, that Claude has used before. According to Anthropic's documentation, how does fulfilling a code execution call differ architecturally from fulfilling a call to the client-defined `send_slack_message` tool?

  1. There is no architectural difference: both are fulfilled the same way, with the API returning a request naming what to run and the developer's own code responsible for actually running it and sending back a result
  2. The code execution tool can be invoked at most once per conversation for the entire lifetime of that conversation, while `send_slack_message` may be called an unlimited number of times
  3. Code execution calls are always queued and answered only after every other tool call in the same turn has been resolved, regardless of the order Claude requested them in
  4. Code execution runs entirely server-side: Anthropic's API executes the command itself inside a sandboxed container and returns the result within the same response, so the developer never runs anything or sends back a tool_result for it themselves, unlike `send_slack_message`, whose tool_use block the developer's own code must execute and answer; the one exception is that if Claude calls both tools in the same turn, the code execution result is withheld until after the developer returns the client tool's result
AI & LLM Engineering · Building with LLM APIs · Card 047/064 easy

A developer adds Anthropic's web_search tool to a Messages API request, sets `allowed_domains` to a short list of trusted sites, and Claude performs three searches over the course of the conversation to answer the user's question. According to Anthropic's documentation, how is this billed, and what rule governs the domain-filtering parameters?

  1. Each search is billed per use ($10 per 1,000 searches) in addition to the standard token cost of the search-generated content that enters the context, and a request may set only one of `allowed_domains` or `blocked_domains`, never both together
  2. Web search itself is always free; only the tokens in the final answer are billed, and `allowed_domains` and `blocked_domains` must always be supplied together so the API can cross-check them against each other
  3. Each search is billed per use, but the content returned by the search is entirely free and never counts as input tokens no matter how much text is retrieved
  4. There is no per-search charge at all; billing is based solely on how many domains appear in `allowed_domains`, with more domains costing more regardless of how many searches Claude actually performs
AI & LLM Engineering · Building with LLM APIs · Card 048/064 medium

An agent built on Claude's Messages API runs a research task for dozens of turns, calling tools repeatedly, and the conversation's growing tool-result history is approaching the model's context window limit even though the task isn't finished. The developer configures Anthropic's context_management feature with a `clear_tool_uses_20250919` edit. According to Anthropic's documentation, what does this feature do, and where does it run?

  1. It permanently deletes the oldest messages from the developer's own stored conversation history in their database, so the developer must reconstruct any cleared turns from their own logs if they are ever needed again
  2. It has Claude itself decide, mid-turn, which of its own past tool calls to omit from its next response, based on which ones it judges are no longer relevant to the current step
  3. It runs server-side, before the prompt reaches Claude: once the conversation's input tokens cross a configured trigger threshold, it automatically clears the content of the oldest tool results in chronological order, replacing each with placeholder text, while keeping a configurable number of the most recent tool uses intact, all without altering the full conversation history the developer's own client still holds
  4. It compresses every tool result in the conversation into a shorter paraphrase using a separate summarization model call, rather than clearing any of them outright, and this paraphrasing counts as an extra billed request each time it runs
AI & LLM Engineering · Building with LLM APIs · Card 049/064 easy

A developer calling OpenAI's Chat Completions API wants to strongly discourage the model from ever producing a specific token, without removing it from the model's vocabulary entirely or rewriting the prompt. They use the `logit_bias` parameter. According to OpenAI's API reference, what does this parameter actually do?

  1. It accepts a list of literal banned words as plain text strings, which the API matches against the generated text after decoding and deletes if found, before the response is returned to the developer
  2. It changes the sampling temperature applied only to the specific tokens named, leaving every other token's probability governed by the request's own top-level temperature setting
  3. It reorders the model's vocabulary so that the named tokens are permanently removed from consideration on every subsequent request made with the same API key, not just the current one
  4. It accepts a map from a token's numeric ID in the model's tokenizer to a bias value between -100 and 100, which is added directly to that token's logit before sampling, with values near -100 effectively banning the token and values near 100 effectively forcing its selection
AI & LLM Engineering · Building with LLM APIs · Card 050/064 hard

A developer defines a JSON schema for OpenAI's Structured Outputs feature with `strict` set to true, and wants one field, `middle_name`, to be genuinely optional so that some responses can simply omit it. According to OpenAI's documentation, why doesn't setting `middle_name` as optional the normal JSON Schema way (leaving it out of the `required` array) work here, and what must the developer do instead?

  1. Leaving a field out of `required` works exactly as in ordinary JSON Schema; strict mode places no additional constraint on the `required` array beyond what standard JSON Schema already allows
  2. Strict mode requires every property defined in the schema to be listed in `required`, so a field can never be truly absent from the output; to model an optional-feeling field, the developer instead types it as a union that allows null (alongside `additionalProperties: false`) and lets the model return null when there is no value, while still including the field's key in the response
  3. Strict mode ignores the `required` array entirely and instead infers which fields are mandatory from the order they appear in the `properties` object, with earlier fields treated as mandatory and later ones as optional
  4. Strict mode requires the developer to submit two entirely separate schemas, one listing only the mandatory fields and one listing only the optional fields, and to specify in the request which of the two schemas applies to the current call
AI & LLM Engineering · Building with LLM APIs · Card 051/064 easy

A team building a voice agent on OpenAI's Realtime API compares the per-token price of audio input tokens against the per-token price of text input tokens on the same model. According to OpenAI's published pricing, what should they expect?

  1. Audio and text tokens are always priced identically per token on Realtime models, since the API converts audio to an internal text representation before counting tokens
  2. Audio tokens are priced lower than text tokens, since audio is a lossy, more compressed signal that costs the model less to process per token than dense text
  3. Audio tokens are priced substantially higher than text tokens on the same Realtime model, with separate published per-million-token rates for text input, audio input, and audio output, rather than a single shared token rate
  4. Realtime models bill only by wall-clock session duration in minutes, with no separate per-token rate for audio or text at all
AI & LLM Engineering · Building with LLM APIs · Card 052/064 easy

A developer's application calls OpenAI's Responses API for a task expected to take several minutes of model processing, and wants to avoid the request failing due to a client-side or gateway connection timeout while waiting synchronously for the full response. According to OpenAI's documentation, how does background mode address this?

  1. Setting `background` to true on the request lets the model process the task asynchronously; the initial call returns immediately with an in-progress response object, and the application polls a GET endpoint using that response's ID, checking again while its status is queued or in_progress, until it reaches a terminal state and the output can be read
  2. Background mode requires the developer to first split the task into several smaller requests themselves, since the API itself has no way to run a single request for longer than its normal synchronous timeout
  3. Background mode automatically retries the entire request from scratch every time the connection drops, silently discarding any progress the model had already made before the drop
  4. Background mode only changes how the response is formatted once it's ready; the request is still processed synchronously and the client connection must stay open for the task's entire duration either way
AI & LLM Engineering · Building with LLM APIs · Card 053/064 easy

A developer using an OpenAI GPT-5-family model wants shorter, terser answers for simple code-lookup questions and longer, more thoroughly explained answers for complex debugging questions, without changing how much internal reasoning the model performs on either kind of request. According to OpenAI's documentation, which parameter is designed for this, and how does it differ from the model's reasoning-effort setting?

  1. The `max_tokens` parameter, since lowering it for simple questions and raising it for complex ones is the documented way to control how much explanation the model includes, and it has always worked this way for every OpenAI model
  2. The `verbosity` parameter, set to low, medium, or high, which is documented as controlling the length and level of detail in the model's final answer independently of `reasoning_effort`, which instead controls how much internal reasoning the model performs before answering
  3. The `reasoning_effort` parameter itself, which OpenAI's documentation describes as jointly controlling both how much the model reasons internally and how long its final answer is, with no separate setting for output length
  4. The `temperature` parameter, since raising temperature is documented as producing systematically longer, more detailed answers, while lowering it produces systematically terser ones
AI & LLM Engineering · Building with LLM APIs · Card 054/064 hard

Across nearly every major LLM provider's published pricing, generating one output token from a model costs more than processing one input (prompt) token of the same model, often by a factor of three to five times. What is the primary technical reason for this asymmetry, rather than it being an arbitrary business markup?

  1. Output tokens are billed higher purely because providers want to discourage overly long responses for environmental reasons, with no meaningful difference in the underlying computation cost between processing an input token and generating an output token
  2. Input tokens require more GPU memory to store than output tokens, since the entire prompt must remain in memory throughout generation, while each output token can be discarded immediately after it is produced
  3. Output tokens cost more because they are always run through a separate, larger safety-filtering model before being returned to the developer, adding a second full model pass on top of generation
  4. Processing the input prompt happens in a single parallel forward pass over all prompt tokens at once (prefill), which uses the available compute efficiently, while generating output happens one token at a time in a sequential, autoregressive process where each new token depends on every token generated before it, requiring roughly as many sequential forward passes as there are output tokens and using the hardware less efficiently per token
AI & LLM Engineering · Building with LLM APIs · Card 055/064 easy

A developer wants their support bot to answer questions by searching a folder of about 200 internal policy PDFs. Rather than building a chunking pipeline, choosing an embedding model, and standing up a vector database themselves, they want a single API tool to handle retrieval so they can get a first version working quickly. According to OpenAI's documentation, what does the Responses API's `file_search` tool provide?

  1. A hosted vector store the developer can upload files into, plus a built-in tool that automatically chunks, embeds, and semantically searches those files for relevant passages during a model turn, without the developer running their own embedding pipeline or vector database
  2. A local, on-device retrieval engine that runs entirely inside the developer's own infrastructure, with OpenAI providing only the embedding model weights for the developer to host and query
  3. A tool that can only match on exact keyword overlap between the user's question and the uploaded documents, since no embedding or semantic search is involved
  4. A tool that requires every uploaded document to already be pre-chunked and pre-embedded by the developer before upload, with the tool only handling the final similarity-ranking step
AI & LLM Engineering · Building with LLM APIs · Card 056/064 medium

A developer adds a `cache_control` breakpoint to a system prompt sent with every call to Anthropic's Messages API, but the prompt is shorter than the model's documented minimum cacheable length. According to Anthropic's documentation, what happens on this and later requests, and how can the developer confirm from the response itself whether caching actually took effect?

  1. The request fails with an error identifying which minimum-length requirement was not met, so the developer must catch and handle this specific error before retrying
  2. The prompt is still cached, but only at a partial discount scaled to how far short of the minimum it falls, and the response's usage fields report that partial discount directly
  3. The prompt below the minimum length is processed without caching and without an error, and the developer can confirm this by checking the response's usage fields: if both `cache_creation_input_tokens` and `cache_read_input_tokens` report zero, the prompt was not cached
  4. Anthropic's API automatically pads the prompt up to the minimum length behind the scenes so that caching still applies, and the padding is reflected as extra input tokens in the usage response
AI & LLM Engineering · Building with LLM APIs · Card 057/064 hard

A developer is building a coding assistant that asks GPT-4o to make small, targeted edits to an existing file -- most of the file's content stays identical, and only a few lines change each time. Regenerating the entire file token-by-token from scratch is slow given how little actually changes on each call. According to OpenAI's documentation, what does the Predicted Outputs feature (the `prediction` parameter) do to help?

  1. It lets the developer supply the existing, mostly-unchanged text as a predicted version of the output; the API then checks the predicted tokens against what it would generate and can skip ahead through matching stretches instead of generating every token sequentially, reducing latency when most of the output matches the prediction
  2. It caches the entire previous completion server-side and simply returns it unchanged on the next call unless the prompt hash changes, so no new generation happens at all for near-identical requests
  3. It fine-tunes a smaller distilled model on the fly from the developer's prediction text, then uses that distilled model to generate the response instead of the original GPT-4o model
  4. It pre-validates the prediction text against the developer's JSON schema before generation starts, rejecting the API call early if the prediction itself would not satisfy the schema
AI & LLM Engineering · Building with LLM APIs · Card 058/064 easy

A developer's existing integration against OpenAI's Chat Completions API sets `max_tokens` to cap response length. When they point the same integration at an o-series reasoning model instead of a standard chat model, the call fails or behaves unexpectedly. According to OpenAI's documentation, what is the relevant difference for reasoning models here?

  1. Reasoning models ignore any token-limit parameter entirely and always generate until they reach their absolute maximum context length, regardless of what the developer sets
  2. OpenAI's Chat Completions API deprecated `max_tokens` in favor of `max_completion_tokens` for use with reasoning models, and this newer parameter caps the combined total of the model's internal reasoning tokens plus its visible output tokens, not just the visible output alone
  3. Reasoning models require the token limit to be set as a request header rather than a body parameter, since reasoning tokens are billed through a separate metering system
  4. `max_tokens` still works identically for reasoning models, and the failure described must be caused by an unrelated authentication or network issue
AI & LLM Engineering · Building with LLM APIs · Card 059/064 easy

A developer uses OpenAI's Structured Outputs feature with a strict JSON schema requiring specific fields, and sends a request the model judges unsafe to fulfill. Forcing a refusal into the exact JSON shape the developer's schema demands would not make sense for a decline message. According to OpenAI's documentation, how does the API handle this case?

  1. It silently returns an empty JSON object matching the schema's field names but with every value set to null, requiring the developer to infer a refusal from the all-null pattern
  2. It raises an HTTP error status, forcing the developer to catch an exception rather than reading a normal successful response body
  3. It returns a normal successful response, but with a separate `refusal` field containing the model's refusal message, distinct from the structured `content` field, so the developer can detect and handle a decline without it needing to conform to the requested schema
  4. It force-fits the refusal text into the requested schema's first required string field, leaving all other fields empty, so the response technically still validates against the schema
AI & LLM Engineering · Building with LLM APIs · Card 060/064 easy

A company already runs its infrastructure on AWS and wants to use Claude models while keeping billing, IAM permissions, and networking inside its existing AWS account, rather than managing a separate API key and vendor relationship directly with Anthropic. According to Anthropic's documentation, how can this be done?

  1. It cannot be done; Claude models are only available directly through Anthropic's own API, with no first-party availability through any major cloud provider's platform
  2. The company must first export the Claude model weights from Anthropic and self-host them on their own AWS EC2 instances, since Anthropic does not offer inference as a managed cloud service
  3. The company can only use Claude on AWS by routing every request through Anthropic's own API using an Anthropic API key, with AWS providing nothing beyond generic network transit
  4. Claude models are available as a first-party option through Amazon Bedrock, using largely the same Messages API request and response shape as Anthropic's direct API, but authenticated and billed through the company's existing AWS account (AWS credentials and IAM) instead of an Anthropic API key
AI & LLM Engineering · Building with LLM APIs · Card 061/064 hard

A team is building two separate products on OpenAI's Realtime API: a browser-based voice assistant that runs directly against a user's laptop microphone and speakers, and a backend telephony integration that connects to a third-party phone system server-to-server with no browser involved. According to OpenAI's documentation, which transport should each choose, and why?

  1. The browser-based client should use WebRTC, since it is designed for client-side, real-time media use and handles network jitter, audio device capture, and variable connection quality automatically; the server-to-server telephony integration should use WebSocket, which is better suited to a backend-to-backend connection without a browser's media stack involved
  2. Both should use WebSocket, since WebRTC is only supported for text-only Realtime API sessions and cannot carry audio at all
  3. The browser-based client should use WebSocket, since browsers cannot establish WebRTC connections without a third-party plugin; the backend integration should use WebRTC, since it offers lower latency for server-to-server traffic
  4. Both should use WebRTC, since OpenAI's documentation states WebSocket support for the Realtime API has been fully removed and is no longer a valid transport option
AI & LLM Engineering · Building with LLM APIs · Card 062/064 medium

A developer building against the Responses API wants to estimate, before sending a request, exactly how many input tokens a prompt will consume once images, uploaded files, and tool/function schemas are all counted in, so they can trim it if it risks exceeding the model's context window. Running the request's text through `tiktoken` locally only approximates plain text and cannot account for the images, files, or tool schemas. According to OpenAI's documentation, how can the developer get an exact count instead?

  1. It cannot be done exactly for anything beyond plain text; images, files, and tool schemas can only ever be approximated, never counted precisely, before a request is actually sent
  2. The developer must send the full request to the standard Responses endpoint with generation enabled and then read the exact input token count back from the completed response's usage data, since no pre-send option exists
  3. The developer can call a dedicated input token counting endpoint that accepts the same input format as the Responses API (including images, files, and tool definitions) and returns the exact input token count without performing any model generation
  4. The developer must call the embeddings endpoint on the request's text content first, since the returned embedding vector's dimensionality directly reveals the exact token count of the full multimodal input
AI & LLM Engineering · Building with LLM APIs · Card 063/064 easy

A startup's rate limits on OpenAI's API were quite low when they first signed up, but several months later, after steady usage and payment history, they notice their per-minute request and token limits have increased automatically without contacting support. According to OpenAI's documentation, what explains this?

  1. Rate limits are fixed per API key for its entire lifetime and cannot change; the startup must be mistaken, or a separate new key with higher limits was silently issued
  2. OpenAI places organizations into usage tiers that automatically raise associated rate limits as the organization's cumulative spend and payment history grow, with no manual request required for standard tier progression
  3. Rate limits only increase if a developer manually opens a support ticket requesting a specific new limit, and support silently approved and applied the change without a reply email
  4. Rate limits are recalculated fully at random on a rolling basis to load-balance traffic across OpenAI's infrastructure, unrelated to any individual organization's usage history
AI & LLM Engineering · Building with LLM APIs · Card 064/064 medium

A team building a debugging-assistant UI on an OpenAI o-series reasoning model wants to show users some indication of how the model reasoned through a problem, to build trust in its answer, but OpenAI's documentation is clear that the model's raw, full chain-of-thought reasoning tokens are not exposed to developers. According to OpenAI's documentation, how can the team still surface something about the model's reasoning process to users?

  1. It cannot be done at all; no information about the model's internal reasoning is ever available through the API in any form for o-series models
  2. The team should set `logprobs` on the response, since log probabilities of the output tokens directly encode the full internal reasoning trace in numerical form
  3. The team should fine-tune a separate, smaller model specifically to imitate and re-articulate the reasoning model's likely internal reasoning, since this is the only officially documented workaround
  4. The team can request a reasoning summary via the API (for example, setting a summary option on the reasoning configuration), which returns a natural-language, abbreviated summary of the model's internal reasoning without exposing the full raw chain-of-thought tokens