20 cards · AWS SAA-C03 · answer each one, then read the explanation. Your score tallies below. Looking for AWS disaster recovery strategies: RTO/RPO worked example? Read the explainer.
0 / 20 answered · 0 correct
Link copied — send it to a friend!
AWS SAA-C03 · High Availability & DR · Card 001/020easy
A company runs a production database on a Multi-AZ RDS DB instance deployment (the single-standby kind, not a Multi-AZ DB cluster deployment). During a failover, which statement correctly describes what happens?
AThe standby instance has been serving read queries all along, so the failover simply stops routing writes to the old primary.
BRDS automatically promotes the synchronously-replicated standby and repoints the same DB instance endpoint to it, without the application needing a new connection string.
CThe application must be reconfigured with a new endpoint address, because the standby instance has a different DNS name from the primary.
DRDS asynchronously copies data to the standby, so some in-flight transactions are typically lost during the failover.
Correct answer: .
In a Multi-AZ RDS DB instance deployment (the single-standby kind, as opposed to a Multi-AZ DB cluster deployment), synchronous replication keeps a standby instance's storage in lockstep with the primary at all times, so a failover carries essentially no data loss — ruling out the idea that replication is asynchronous or lossy. That standby exists purely for failover, not for serving application reads, so it cannot have already been fielding read traffic before the switch. Because the deployment exposes a single, stable DB instance endpoint, RDS handles a failover by updating the underlying DNS record for that same endpoint to point at the newly promoted instance; the application keeps its existing connection string and simply reconnects, rather than needing a new endpoint address hardcoded anywhere.
Source: AWS RDS User Guide: Configuring and managing a Multi-AZ deployment (Multi-AZ DB instance deployments)
AWS SAA-C03 · High Availability & DR · Card 002/020medium
How does a Multi-AZ DB cluster deployment for RDS differ from a traditional Multi-AZ DB instance deployment (single standby)?
AA DB cluster deployment has only one standby instance, identical to a DB instance deployment, but replicates asynchronously instead of synchronously.
BA DB cluster deployment eliminates standby instances entirely and instead relies on cross-Region read replicas for failover.
CA DB cluster deployment adds a second standby instance that also cannot serve read traffic, purely to add redundancy against a second simultaneous failure.
DA DB cluster deployment spans three Availability Zones with two reader-capable standby instances, so those standbys can serve read traffic in addition to providing failover support.
Correct answer: .
A traditional Multi-AZ DB instance deployment has exactly one standby that exists solely for failover and never serves read traffic, so an option describing a second, still read-incapable standby only adds redundancy without changing that fundamental limitation, which isn't what a DB cluster deployment does. A Multi-AZ DB cluster deployment is a distinct architecture that spans three Availability Zones with a writer instance and two reader-capable standby instances; because those standbys can be queried directly, they add usable read capacity on top of failover protection, and their replication is synchronous, not asynchronous — so describing the cluster as merely swapping to async replication is backwards. Cross-Region read replicas are a separate feature entirely and aren't what enables faster, quorum-based failover within a Multi-AZ DB cluster.
Source: AWS RDS User Guide: Configuring and managing a Multi-AZ deployment (Multi-AZ DB cluster deployments)
AWS SAA-C03 · High Availability & DR · Card 003/020medium
A company wants to add a secondary AWS Region to an existing Aurora Global Database for disaster recovery, and later needs to relocate the primary Region there for planned maintenance with no data loss. Which statement is accurate?
AAurora can support the primary cluster plus up to ten read-only secondary clusters in other Regions, replicating writes to secondaries with typical latency under a second, and a planned relocation of the primary should use the switchover (previously called managed planned failover) operation rather than an unplanned failover.
BAurora Global Database limits a primary cluster to exactly one secondary Region, and any relocation of the primary — planned or unplanned — is called a failover.
CSecondary Aurora clusters in a global database are write-capable by default, so no special switchover operation is needed to relocate the primary.
DCross-Region replication in an Aurora Global Database uses the database engine's binlog-based replication, the same mechanism used by standard read replicas, rather than dedicated storage-level replication.
Correct answer: .
Aurora Global Database is built for exactly this pattern: one primary cluster can replicate to up to ten read-only secondary clusters in other Regions, using the underlying storage layer rather than the database engine's own replication mechanism, which is why replication lag is typically under a second rather than being bound by binlog-style replication delays — this also rules out any option describing engine-level binlog replication. Secondary clusters are read-only by design and cannot accept writes on their own, so no relocation can happen without an explicit operation. For a planned relocation with no data loss, AWS provides the switchover operation (the current name for what used to be called managed planned failover); a separate, unplanned failover operation is reserved for recovering from an actual outage in the primary Region.
Source: Amazon Aurora User Guide: Using Amazon Aurora Global Database
AWS SAA-C03 · High Availability & DR · Card 004/020hard
A company's compliance team sets a recovery objective allowing at most a few minutes of downtime and a few minutes of data loss after a full Regional outage, but cost is a secondary concern next to reliability. Per the AWS Well-Architected Reliability Pillar's disaster recovery strategies, which strategy best fits, and why do the cheaper options not fit?
ABackup and restore, because restoring from backups after any outage always meets a few-minutes recovery objective regardless of data size.
BPilot light, because keeping only core data services running with everything else fully scaled down and de-provisioned reliably yields minute-level recovery for any workload.
CWarm standby (or multi-site active/active if the budget allows), because pilot light and backup-and-restore both require provisioning or scaling up infrastructure before traffic can be served, which typically pushes recovery time well past a few minutes.
DAny of the four strategies works equally well, since AWS Regions fail over to each other automatically regardless of which DR strategy a workload's architecture implements.
Correct answer: .
AWS Regions never fail over to each other automatically — every one of the four Well-Architected disaster recovery strategies requires the workload's own architecture to make that happen, so treating them as interchangeable ignores how each is actually built. Backup and restore has the highest RTO and RPO of the four because it must provision infrastructure and restore data from scratch, which routinely takes far longer than a few minutes. Pilot light keeps only core data services running while everything else stays fully scaled down, so it still needs to provision and scale up compute before serving traffic — that startup time typically lands in the tens-of-minutes-to-hours range, not a reliable few minutes for any workload. Warm standby keeps a smaller but fully running, already-scaled copy of the environment, so it only needs to scale up rather than build from nothing, which is what makes minute-level recovery achievable without paying for full multi-site active/active.
AWS SAA-C03 · High Availability & DR · Card 005/020hard
A team configures a Route 53 health check on an application endpoint with the default consecutive-failure threshold, then observes that a single health checker briefly failing to reach the endpoint does not immediately mark it unhealthy in the Route 53 console. Which statement explains this correctly?
AThis is a bug; a single consecutive-failure count from any one health checker location always overrides Route 53's overall assessment of the endpoint.
BRoute 53 runs health checkers from many locations worldwide, and only marks an endpoint unhealthy once 18% or fewer of the reporting health checkers currently judge it healthy — a single checker's transient failure gets outvoted by the rest.
CRoute 53 disables DNS failover entirely until the endpoint has failed every health checker location at least once, then never re-enables it automatically.
DThe consecutive-failure threshold only applies to CloudWatch-alarm-based health checks, not to health checks that monitor an endpoint directly.
Correct answer: .
Route 53 doesn't rely on any single health checker's opinion: it operates checkers from many locations worldwide, and each one independently evaluates the endpoint using the consecutive-failure threshold that was configured. Route 53 then aggregates all of those independent opinions and only marks the endpoint unhealthy once 18% or fewer of the reporting checkers currently consider it healthy, precisely so that one checker's transient network blip can't flip the overall status on its own — that also means describing a single checker's failure count as automatically overriding the whole assessment gets the mechanism backwards. The consecutive-failure threshold isn't limited to CloudWatch-alarm-based health checks; it's the core mechanism for health checks that monitor an endpoint directly, too. And Route 53 doesn't require every single global location to fail before disabling failover, nor does it ever permanently disable failover once triggered — health status keeps being reevaluated continuously.
Source: Amazon Route 53 Developer Guide: How Amazon Route 53 determines whether a health check is healthy
AWS SAA-C03 · High Availability & DR · Card 006/020easy
A team enables S3 Cross-Region Replication on a bucket that already holds two years of existing objects, expecting all of that history to appear in the destination bucket automatically. What actually happens, and what should they do instead?
ACross-Region Replication automatically backfills every existing object the moment replication is enabled, so no extra action is needed.
BCross-Region Replication requires S3 Transfer Acceleration to be enabled first, and only then will pre-existing objects replicate automatically.
CCross-Region Replication silently deletes objects created before replication was enabled, since it only tracks objects going forward.
DCross-Region Replication (live replication) only replicates new and updated objects going forward; to replicate the pre-existing two years of objects, they need to run an S3 Batch Replication job.
Correct answer: .
Cross-Region Replication is a live replication feature, and live replication in S3 only ever copies new and updated objects from the moment the replication rule is enabled onward — it was never designed to look backward at a bucket's existing contents. That means objects added before replication was turned on need a separate, explicit action: an S3 Batch Replication job, which is specifically built to replicate existing objects on demand. Nothing about enabling Cross-Region Replication deletes older objects or touches them at all; they simply stay unreplicated until Batch Replication (or some other explicit copy) is run. Cross-Region Replication also has no dependency on S3 Transfer Acceleration, which is an unrelated feature for speeding up uploads over long distances, not a prerequisite for replication of any kind.
Source: Amazon S3 User Guide: Replicating objects within and across Regions
AWS SAA-C03 · High Availability & DR · Card 007/020easy
A company migrating its DR plan away from self-managed replication tooling wants continuous, near-real-time replication of its on-premises servers into AWS, with the ability to launch fully booted, native EC2 recovery instances within minutes of a declared disaster. Which AWS service matches this, and how does it keep steady-state costs down?
AAWS Elastic Disaster Recovery (DRS), which continuously replicates source servers at the block level into a low-cost staging area (minimal compute and storage) and only converts data into fully booted EC2 instances at the moment of a drill or actual recovery.
BAmazon S3 Cross-Region Replication, which continuously replicates entire server volumes and boots them as EC2 instances automatically during an outage.
CAWS Backup, which takes scheduled snapshots and requires restoring a full backup before any recovery instance can boot, typically taking hours.
DAmazon Data Lifecycle Manager, which is a snapshot-scheduling tool with no ability to launch recovery instances at all.
Correct answer: .
AWS Elastic Disaster Recovery (DRS) is purpose-built for this: it installs lightweight agents on source servers and continuously replicates their disks at the block level into a staging area in the target Region, using minimal, inexpensive compute and storage to hold that ongoing replica — full-sized recovery instances only get created at the moment of a drill or a real recovery, which is what keeps steady-state costs low. AWS Backup instead relies on periodic backup snapshots and a full restore process before anything can boot, which is a fundamentally slower path typically measured in hours. Amazon Data Lifecycle Manager only automates snapshot scheduling and retention; it has no mechanism to launch a runnable recovery instance at all. S3 Cross-Region Replication copies object storage, not EBS volumes or whole servers, and has no concept of converting replicated data into a bootable EC2 instance.
Source: AWS Elastic Disaster Recovery User Guide: What is Elastic Disaster Recovery?
AWS SAA-C03 · High Availability & DR · Card 008/020easy
A team's EBS backup strategy relies on regular incremental snapshots and assumes that same incremental chain automatically protects them against a full Regional outage. What is the flaw in that assumption?
AEBS snapshots are always full copies, never incremental, so cross-region protection already exists without any extra step.
BEBS snapshots cannot be copied to another Region under any circumstances, so Regional outages always require rebuilding volumes from scratch.
CAn EBS snapshot's incremental chain exists only within the Region it was taken in; protecting against a Regional outage requires explicitly copying snapshots to another Region, and the first copy into that Region is a full copy.
DCopying an EBS snapshot to another Region automatically deletes the original snapshot's incremental chain in the source Region.
Correct answer: .
Amazon EBS snapshots are incremental by nature, storing only the blocks that changed since the previous snapshot, but that incremental chain is tied to the Region where the snapshots were taken — it doesn't extend across Regions on its own. Protecting against a full Regional outage requires an explicit step: copying snapshots to another Region, and because that destination Region has no prior snapshot history for the volume, the first snapshot copied there is necessarily a full copy rather than an incremental one; later copies can then build their own incremental chain in that destination Region. Snapshots absolutely can be copied across Regions — that capability is a standard, supported operation, not something that's unavailable. And copying a snapshot elsewhere has no effect on the original snapshot or its incremental chain back in the source Region.
Source: Amazon EC2 User Guide: Copying an Amazon EBS snapshot
AWS SAA-C03 · High Availability & DR · Card 009/020easy
An application uses a DynamoDB global table replicated across three Regions so that if one Region becomes impaired, traffic can shift to another with minimal data loss. Which statement correctly describes how this works?
AOnly one Region's replica accepts writes at a time, and the others are strictly read-only until a manual promotion, similar to Aurora Global Database.
BGlobal tables are multi-active: any replica can serve both reads and writes, changes replicate asynchronously to the other Regions, and concurrent conflicting writes to the same item are resolved using a last-writer-wins rule.
CGlobal tables guarantee strongly consistent reads across every Region at all times, so no conflict-resolution mechanism is ever needed.
DEnabling global tables converts DynamoDB from a key-value store into a strictly single-Region service that merely mirrors backups to other Regions for cold storage.
Correct answer: .
DynamoDB global tables use a multi-active replication model, meaning every replica in every Region can accept both reads and writes directly, rather than restricting writes to a single elected primary the way some other multi-Region database features do — that rules out describing the setup as having one write-capable Region with the rest strictly read-only pending manual promotion. Because writes can land in more than one Region for the same item, DynamoDB needs a way to settle conflicting concurrent writes, and by default it does so with a last-writer-wins rule, which also rules out claiming no conflict-resolution mechanism is needed. Global tables replicate asynchronously between Regions rather than offering an always-on strong consistency guarantee across all Regions at all times. Global tables remain a fully functional, actively-written key-value store in every participating Region, not a passive backup mirror.
AWS SAA-C03 · High Availability & DR · Card 010/020hard
A workload's backups run automatically every 4 hours. A Regional outage begins at 14:00, 3 hours after the most recent backup completed at 11:00. The team restores service from that 11:00 backup and traffic resumes at 15:30. In terms of RTO and RPO, which statement is accurate?
ARPO is the 1.5-hour restoration window (14:00 to 15:30); RTO is the 3-hour gap since the last backup (11:00 to 14:00).
BRTO and RPO both refer to the same 1.5-hour window, since they are defined as the total time the workload was unavailable.
CRPO is the entire 4.5 hours from the last backup (11:00) to service resumption (15:30); RTO does not apply because the outage was Regional rather than a single-component failure.
DRPO is the up-to-3-hour data loss window between the last backup (11:00) and the outage (14:00), since restoring from that backup loses any changes made after it; RTO is the 1.5-hour restoration time between the outage (14:00) and service resuming (15:30).
Correct answer: .
Recovery Point Objective describes the maximum acceptable amount of data an organization could lose, measured as the time between the last good recovery point and the moment the interruption began — here, that's the gap between the 11:00 backup and the 14:00 outage, so up to three hours of changes made during that window are gone once the team restores from the 11:00 backup. Recovery Time Objective describes how long the workload was actually unavailable, measured between the interruption and service being restored — here, that's the interval between the 14:00 outage and the 15:30 recovery, so ninety minutes. Swapping those two definitions, treating them as the same window, or claiming RTO doesn't apply to Regional outages all describe the relationship incorrectly; RPO and RTO apply to any interruption regardless of its cause or scope, and they measure two different things: data loss versus downtime.
AWS SAA-C03 · High Availability & DR · Card 011/020easy
A team is setting up basic active-passive DNS failover for a public website using Route 53's failover routing policy, with exactly one primary record and one secondary record. Which statement correctly describes what health checking is required?
AOnly the primary record needs an associated health check (or, for an AWS alias target, Evaluate Target Health set to yes); Route 53 automatically returns the secondary record's answer whenever the primary is found unhealthy, without a health check needing to be attached to the secondary record.
BBoth the primary and secondary records must each have their own independently configured health check before Route 53 will accept either record as part of a failover routing policy.
CActive-passive behavior requires combining the failover routing policy with a weighted routing policy on the same records; failover alone only supports active-active configurations.
DRoute 53 decides which of the two records to return by comparing measured latency to each endpoint, falling back to the secondary only once the primary's latency exceeds a configured threshold.
Correct answer: .
With one primary and one secondary record under the failover routing policy, Route 53 only needs a way to judge the primary's health -- typically a health check, or Evaluate Target Health set to yes when the primary is an alias to an AWS resource -- to decide which record to return; while that primary is healthy, Route 53 answers queries with it, and only starts answering with the secondary once the primary is judged unhealthy, all without a health check ever being attached to the secondary record itself. The option requiring an independent health check on both records adds a mandatory step this simple one-primary/one-secondary setup doesn't need. The option describing weighted routing being required is backwards: failover routing policy alone is exactly what produces active-passive behavior, and active-active is what happens when other routing policies such as weighted are combined with health checks instead. The option describing latency-based comparison describes an entirely different routing policy (latency-based routing); failover routing policy makes its decision purely from health status, not measured response time.
Source: Amazon Route 53 Developer Guide: Active-active and active-passive failover
AWS SAA-C03 · High Availability & DR · Card 012/020medium
An Application Load Balancer has two registered targets in Availability Zone A and only one in Availability Zone B. A team wants to know how the ALB distributes requests across these unevenly distributed targets by default. Which statement is correct?
ACross-zone load balancing is off by default for an Application Load Balancer, so each Availability Zone's load balancer node only distributes requests among its own AZ's targets, giving the lone target in Availability Zone B a much larger share of that AZ's traffic.
BCross-zone load balancing can only be enabled by recreating the load balancer as a Network Load Balancer; Application Load Balancers have no such feature.
CCross-zone load balancing is on by default for an Application Load Balancer and cannot be turned off at the load balancer level; each load balancer node distributes requests evenly across all registered targets in every enabled Availability Zone, regardless of how many targets sit in each AZ. It can be turned off at the target group level if needed.
DCross-zone load balancing is on by default, but only within a single Availability Zone; requests never cross Availability Zone boundaries regardless of configuration.
Correct answer: .
For an Application Load Balancer, cross-zone load balancing is on by default at the load balancer level and that default can't be changed there, which is exactly why the two targets in Availability Zone A and the single target in Availability Zone B still each receive a roughly even share of total requests: every load balancer node spreads traffic across all registered, healthy targets in every enabled Availability Zone rather than confining itself to targets in its own zone. The option describing cross-zone load balancing as off by default confuses this with Network Load Balancers, where cross-zone load balancing is in fact off by default; Application Load Balancers behave differently. Recreating the load balancer as a Network Load Balancer isn't required to control this setting for an ALB: cross-zone load balancing for an Application Load Balancer can instead be turned off at the target group level, which is the one place this particular default can be overridden. The option claiming traffic never crosses Availability Zone boundaries describes the opposite of what cross-zone load balancing actually does.
AWS SAA-C03 · High Availability & DR · Card 013/020medium
An Auto Scaling group spans three Availability Zones. One Availability Zone becomes impaired, causing the group's instances to become unevenly distributed across the remaining healthy zones. Some time later, the impaired Availability Zone recovers. With the group's default settings, what happens next?
ANothing changes automatically; the uneven distribution persists until an operator manually terminates instances in the over-represented Availability Zones and launches new ones in the recovered zone.
BAmazon EC2 Auto Scaling automatically rebalances the group: it launches new instances in the Availability Zones with the fewest instances, including the recovered one, and only terminates the excess instances elsewhere after the new ones are up, so the rebalancing doesn't reduce capacity or availability while it happens.
CThe Auto Scaling group immediately terminates instances in the over-represented Availability Zones first, then launches replacement instances in the recovered zone, briefly reducing total running capacity during the transition.
DAutomatic Availability Zone rebalancing only happens if Capacity Rebalancing has been explicitly turned on for the group; without it, the group stays unevenly distributed indefinitely.
Correct answer: .
Amazon EC2 Auto Scaling continuously tries to keep an equivalent number of instances in each enabled Availability Zone, and when an impairment leaves the group unbalanced, it automatically corrects this once the zone recovers by launching new instances in whichever enabled Availability Zones currently have the fewest instances -- including the newly recovered one -- and only terminating the excess instances elsewhere afterward, which is what keeps this rebalancing from ever dropping the group's effective capacity partway through. No manual intervention is needed for this basic Availability Zone rebalancing, so treating it as a manual-only process is incorrect. Terminating the over-represented instances before the replacements are up would also get the order backwards from how Auto Scaling actually performs the swap, needlessly risking a capacity dip. Capacity Rebalancing is a separate, opt-in feature specifically for proactively replacing Spot Instances flagged as being at elevated risk of interruption; it has nothing to do with whether this ordinary Availability Zone rebalancing after an impairment happens automatically, which it does regardless of that setting.
Source: AWS documentation: Amazon EC2 Auto Scaling — Distribute instances across Availability Zones (rebalancing activities)
AWS SAA-C03 · High Availability & DR · Card 014/020easy
A team enables replication on an Amazon EFS file system to a destination file system in a second AWS Region for disaster recovery. Which statement about how this replication behaves is accurate?
AReplication is synchronous, so every write to the source file system is durably committed to the destination file system before the source write is acknowledged, giving a recovery point objective of zero.
BThe destination file system accepts both reads and writes from applications at all times, the same as the source, with no separate failover step required.
CAmazon EFS does not support replicating a file system to a different Region at all; only same-Region backups are available.
DReplication is asynchronous and, after the initial sync completes, Amazon EFS maintains a recovery point objective of about 15 minutes for most file systems; using the destination for reads and writes requires an explicit failover step to promote it, and returning to the original source afterward requires a separate failback.
Correct answer: .
Amazon EFS replication works asynchronously: after the one-time initial sync finishes, Amazon EFS keeps the destination file system's data caught up to within a recovery point objective of about 15 minutes for most file systems, which is measured from the last successful sync rather than guaranteed on every individual write, so describing it as synchronous with a zero recovery point objective overstates what the feature provides. The destination file system is not simultaneously read-write alongside the source; it exists specifically to be promoted during a disaster or planned game-day exercise, and an explicit failover step is what makes it usable for reads and writes, with a separate failback step needed to resume normal operation on the original source afterward. Amazon EFS does support cross-Region replication as a purpose-built feature, so claiming only same-Region backups exist is incorrect.
Source: AWS EFS User Guide: Replicating EFS file systems (replication performance and RPO)
AWS SAA-C03 · High Availability & DR · Card 015/020hard
A company runs its application across two AWS Regions and uses Amazon Route 53 Application Recovery Controller (Route 53 ARC) for its multi-Region disaster recovery plan. The team wants to continuously confirm ahead of time that the standby Region is scaled and configured to absorb full traffic if needed, and separately wants a safe, guardrailed way to actually redirect DNS traffic during a real Regional impairment. Which statement correctly matches ARC's components to these two needs?
AReadiness checks are for the ongoing, ahead-of-time confirmation that the standby Region's resources, quotas, and configuration could handle failover traffic, and are not meant to sit in the critical path during an actual event; routing control, which supports safety rules, is what should be used to actually redirect traffic during a real impairment.
BRouting control is for the ahead-of-time confirmation work, and readiness checks are what actually redirect DNS traffic during a live Regional impairment.
CRoute 53 ARC only supports shifting traffic away from a single impaired Availability Zone within one Region; redirecting traffic between two entire Regions requires a separate, non-ARC service.
DReadiness checks and routing control are two names for the same underlying capability, so either one can be used interchangeably for both the ahead-of-time confirmation and the live failover.
Correct answer: .
Route 53 ARC deliberately splits these two jobs into separate components: readiness check continually monitors resource quotas, capacity, and configuration in the standby Region so a team knows in advance whether a failover would actually succeed, and it's explicitly documented as not intended for the critical path during an actual failover event; routing control is the component built for that live moment, giving extremely reliable, guardrailed control over which Region's DNS traffic is active, with safety rules that prevent an operator from accidentally leaving an application with no healthy active replica. Swapping which component does which job gets the design backwards. Route 53 ARC's multi-Region recovery capability (routing control, alongside Region switch) handles failover between entire Regions, distinct from ARC's separate single-Availability-Zone zonal shift and zonal autoshift capabilities, so describing ARC as limited to single-AZ shifts within one Region ignores this multi-Region capability entirely. Readiness check and routing control are documented as distinct capabilities serving different purposes, not interchangeable names for one feature.
Source: AWS documentation: What is Amazon Application Recovery Controller (ARC) — Multi-Region recovery: routing control and readiness check
AWS SAA-C03 · High Availability & DR · Card 016/020hard
A company's disaster recovery runbook for a Regional failover calls for its automation to invoke EC2 and Auto Scaling APIs to launch and scale up a standby environment at the moment a disaster is declared. During an actual large-scale Regional outage, this automation runs far more slowly than expected, delaying recovery. Which Well-Architected Reliability Pillar principle does this design violate, and what is the recommended fix?
AThis is expected and unavoidable; every disaster recovery strategy, including Multi-AZ, depends on control-plane API calls during a failure, so no design change would help.
BThe design violates the shared responsibility model; the fix is to file a support case asking AWS to prioritize the account's API calls during the outage.
CThe design violates the static stability principle, which calls for a system to keep operating using resources that are already deployed and configured -- rather than depending on control-plane API calls made in the middle of the disruption -- precisely because control planes can be slower or degraded during the same large-scale event the design is trying to recover from; the fix is to pre-provision and pre-scale the standby capacity ahead of time so failover only requires redirecting traffic to already-running resources.
DThis is purely an RTO and RPO measurement problem; the fix is to redefine the workload's RTO and RPO targets to match whatever the automation currently achieves.
Correct answer: .
Static stability is the Well-Architected Reliability Pillar principle that a system should be able to keep running, or fail over, using capacity and configuration that already exist, rather than depending on control-plane API calls -- like launching new instances or scaling a group -- made in the middle of the very disruption it's trying to survive; those control-plane calls are exactly what can be slower or less reliable during a widespread event, since many other systems may be competing for the same API capacity at the same time. The recommended fix is to pre-provision and pre-scale the standby environment in advance, so a real failover only has to redirect traffic to resources that are already running rather than waiting on API calls to create them. Calling this unavoidable is wrong; static stability exists specifically to design around this dependency, and Multi-AZ standbys are themselves already provisioned in advance rather than launched reactively, which is part of why they achieve fast, predictable failover. This has nothing to do with the shared responsibility model or with AWS Support prioritizing API calls, and redefining RTO/RPO targets to match a flawed design's current performance addresses the symptom, not the architectural gap the question describes.
Source: AWS Well-Architected Framework, Reliability Pillar: Static stability using Availability Zones
AWS SAA-C03 · High Availability & DR · Card 017/020medium
An application's secondary cluster in an Aurora Global Database occasionally needs to run a write, and the team enables write forwarding on that secondary cluster instead of maintaining a separate connection to the primary Region for those rare writes. Which statement correctly describes what happens to a forwarded write?
AWrite forwarding lets the secondary cluster commit the write to its own local storage first, then asynchronously pushes that change to the primary cluster afterward.
BThe secondary cluster forwards the write's SQL statement to the primary cluster, which applies the change first and remains the source of truth; the change is then replicated back out to all secondary Regions as usual, so a forwarded write incurs the added cross-Region round-trip rather than committing locally first. Only data manipulation statements are forwarded this way -- most data definition language statements still must run directly against the primary cluster's writer instance.
CEnabling write forwarding turns every secondary cluster into an additional fully independent writer, eliminating the distinction between the primary cluster and secondary clusters entirely.
DWrite forwarding requires switching the entire global database from storage-level replication to engine-level logical replication.
Correct answer: .
With write forwarding enabled, a secondary cluster doesn't write to its own storage first -- it forwards the write's SQL statement, along with the necessary session and transactional context, to the primary cluster, which applies the change and remains the single source of truth before that change replicates back out to every secondary Region the normal way; that round trip to the primary and back is exactly why a forwarded write takes longer than a write issued directly against the primary. The option describing a local-commit-then-push order has this backwards -- the primary always changes first. Write forwarding is documented as supporting data manipulation statements such as inserts, updates, and deletes, while data definition language and certain other operations still need to run directly on the primary cluster's writer instance, so the feature doesn't universally forward every kind of write. It also doesn't turn secondary clusters into independent writers in their own right; the primary/secondary distinction and the underlying storage-level cross-Region replication mechanism both stay exactly as they were, which is also why the option describing a switch to engine-level logical replication is wrong -- write forwarding changes nothing about how the global database physically replicates data.
Source: Amazon Aurora User Guide: Using write forwarding in an Amazon Aurora global database
AWS SAA-C03 · High Availability & DR · Card 018/020easy
A compliance team needs certain S3 objects to be impossible for anyone to delete or overwrite before a fixed retention date -- including anyone with root access to the AWS account -- to satisfy a regulatory records-retention requirement. Which S3 Object Lock configuration meets this, and why don't the alternatives?
AGovernance mode, because by default no one, including users with special permissions, can ever bypass a governance-mode retention period.
BA legal hold by itself, because a legal hold enforces a fixed, predetermined expiration date that no user can shorten.
CEnabling S3 Versioning alone, without Object Lock, because keeping every version of an object already prevents any version from being permanently deleted.
DCompliance mode, because a compliance-mode retention period cannot be shortened or removed by any user, including the account's root user, before the configured retain-until date -- the only way to delete such an object earlier is to delete the entire AWS account -- whereas governance mode allows users with the specific bypass permission to override it, and versioning alone doesn't stop a locked object version from being deleted at all without Object Lock enabled.
Correct answer: .
Compliance mode is the Object Lock retention mode built for exactly this requirement: once an object version is locked under compliance mode, its retention mode can't be changed and its retention period can't be shortened by any user at all, including the AWS account's own root user, and the documented exception -- deleting the entire AWS account -- makes clear just how absolute that protection is meant to be. Governance mode looks similar on the surface but is deliberately less strict: it's designed so most users are blocked from deleting or altering a locked object, while users holding the dedicated bypass permission can still override the lock, which is the opposite of the 'no one, ever' requirement here. A legal hold has no expiration date at all -- it protects an object indefinitely until someone with the right permission explicitly removes it -- so it doesn't provide a fixed retention date and doesn't describe compliance mode's behavior. S3 Versioning by itself, without Object Lock enabled, only keeps prior versions recoverable after a new version or delete marker is added; it does not stop a version from being permanently deleted outright, which is precisely the gap Object Lock closes.
Source: Amazon S3 User Guide: Locking objects with S3 Object Lock (retention modes)
AWS SAA-C03 · High Availability & DR · Card 019/020easy
A team runs a production database on a Multi-AZ RDS DB instance deployment (single-standby). A mandatory operating system security patch becomes available and RDS applies it during the configured maintenance window. Which statement correctly describes how RDS applies this patch and the resulting impact?
ARDS applies the OS patch to the standby instance first, then promotes that newly patched standby to primary, and finally applies the patch to the old primary, which becomes the new standby -- so the deployment experiences only a brief Multi-AZ failover, typically well under a minute, rather than an extended outage.
BRDS patches the current primary instance first while it continues serving traffic, then patches the standby afterward, with no failover involved at either step.
CRDS patches both the primary and standby instances at exactly the same time, taking the entire Multi-AZ deployment offline for the full duration of the OS patch, identical to how a Single-AZ deployment would experience it.
DA Multi-AZ deployment defers all mandatory OS patches indefinitely until an operator manually initiates a failover first.
Correct answer: .
For a Multi-AZ DB instance deployment, Amazon RDS reduces the impact of a required OS patch by applying it to the standby instance first, then promoting that now-patched standby to be the new primary, and finally patching the old primary, which becomes the new standby -- the only visible impact to the application is the brief Multi-AZ failover this promotion causes, typically well under a minute, rather than downtime for the full patching duration. Patching the live primary first, while it's still serving traffic, isn't how this process works and would defeat the purpose of having a standby ready to absorb the failover. Patching both instances simultaneously describes what happens for a database engine version upgrade, a different maintenance operation, not routine OS security patching, and that scenario doesn't apply here regardless of Single-AZ versus Multi-AZ. Mandatory OS updates aren't deferred indefinitely either; if left unapplied, RDS automatically applies them at or after their specified mandatory apply date during a maintenance window, whether or not an operator manually intervenes first.
Source: AWS RDS User Guide: Maintaining a DB instance — Maintenance for Multi-AZ deployments
AWS SAA-C03 · High Availability & DR · Card 020/020easy
A team has defined an RTO and RPO for a production workload and wants an AWS service that will assess the workload's actual architecture against those specific targets and produce prioritized, concrete recommendations for closing any gaps -- rather than just monitoring metrics or executing backups. Which service fits, and how does it differ from Trusted Advisor and AWS Backup?
AAWS Backup, because its centralized backup policies are themselves the mechanism that evaluates whether an architecture meets a defined RTO and RPO.
BAmazon CloudWatch, because configuring alarms on the workload's resources is equivalent to assessing whether its architecture meets defined RTO and RPO targets.
CAWS Resilience Hub, because it evaluates an application's architecture against a resiliency policy built from the team's own defined RTO and RPO targets, identifies specific gaps, and produces prioritized recommendations to close them -- distinct from Trusted Advisor's general best-practice checks across cost, security, and service limits, and from AWS Backup's role in actually executing backup and restore operations.
DAWS Trusted Advisor, because its checks are specifically designed around a workload's individually defined RTO and RPO targets rather than general best practices.
Correct answer: .
AWS Resilience Hub is purpose-built for this: a team defines a resiliency policy expressing their RTO and RPO targets, and Resilience Hub assesses the application's actual architecture against that policy, surfacing specific gaps and prioritized, actionable recommendations for closing them, along with the ability to run tests that validate whether the recommendations actually improve resilience. Trusted Advisor instead runs general best-practice checks across categories like cost optimization, security, and service limits; it isn't built around a workload's own custom-defined RTO and RPO targets, which is why relying on it for this specific assessment is the wrong fit. AWS Backup is an operational service for centrally managing and executing backup and restore policies -- valuable for actually protecting data, but it doesn't itself assess an architecture's overall resiliency against defined recovery targets. Amazon CloudWatch provides monitoring and alarming on metrics; configuring alarms doesn't amount to an architectural gap analysis against RTO and RPO targets either.
Source: AWS Resilience Hub User Guide: What is AWS Resilience Hub