passdrill

High Availability & DR

20 cards · AWS SAA-C03 · answer each one, then read the explanation. Your score tallies below. Looking for AWS disaster recovery strategies: RTO/RPO worked example? Read the explainer.

0 / 20 answered · 0 correct

AWS SAA-C03 · High Availability & DR · Card 001/020 easy

A company runs a production database on a Multi-AZ RDS DB instance deployment (the single-standby kind, not a Multi-AZ DB cluster deployment). During a failover, which statement correctly describes what happens?

  1. The standby instance has been serving read queries all along, so the failover simply stops routing writes to the old primary.
  2. RDS automatically promotes the synchronously-replicated standby and repoints the same DB instance endpoint to it, without the application needing a new connection string.
  3. The application must be reconfigured with a new endpoint address, because the standby instance has a different DNS name from the primary.
  4. RDS asynchronously copies data to the standby, so some in-flight transactions are typically lost during the failover.
AWS SAA-C03 · High Availability & DR · Card 002/020 medium

How does a Multi-AZ DB cluster deployment for RDS differ from a traditional Multi-AZ DB instance deployment (single standby)?

  1. A DB cluster deployment has only one standby instance, identical to a DB instance deployment, but replicates asynchronously instead of synchronously.
  2. A DB cluster deployment eliminates standby instances entirely and instead relies on cross-Region read replicas for failover.
  3. A DB cluster deployment adds a second standby instance that also cannot serve read traffic, purely to add redundancy against a second simultaneous failure.
  4. A DB cluster deployment spans three Availability Zones with two reader-capable standby instances, so those standbys can serve read traffic in addition to providing failover support.
AWS SAA-C03 · High Availability & DR · Card 003/020 medium

A company wants to add a secondary AWS Region to an existing Aurora Global Database for disaster recovery, and later needs to relocate the primary Region there for planned maintenance with no data loss. Which statement is accurate?

  1. Aurora can support the primary cluster plus up to ten read-only secondary clusters in other Regions, replicating writes to secondaries with typical latency under a second, and a planned relocation of the primary should use the switchover (previously called managed planned failover) operation rather than an unplanned failover.
  2. Aurora Global Database limits a primary cluster to exactly one secondary Region, and any relocation of the primary — planned or unplanned — is called a failover.
  3. Secondary Aurora clusters in a global database are write-capable by default, so no special switchover operation is needed to relocate the primary.
  4. Cross-Region replication in an Aurora Global Database uses the database engine's binlog-based replication, the same mechanism used by standard read replicas, rather than dedicated storage-level replication.
AWS SAA-C03 · High Availability & DR · Card 004/020 hard

A company's compliance team sets a recovery objective allowing at most a few minutes of downtime and a few minutes of data loss after a full Regional outage, but cost is a secondary concern next to reliability. Per the AWS Well-Architected Reliability Pillar's disaster recovery strategies, which strategy best fits, and why do the cheaper options not fit?

  1. Backup and restore, because restoring from backups after any outage always meets a few-minutes recovery objective regardless of data size.
  2. Pilot light, because keeping only core data services running with everything else fully scaled down and de-provisioned reliably yields minute-level recovery for any workload.
  3. Warm standby (or multi-site active/active if the budget allows), because pilot light and backup-and-restore both require provisioning or scaling up infrastructure before traffic can be served, which typically pushes recovery time well past a few minutes.
  4. Any of the four strategies works equally well, since AWS Regions fail over to each other automatically regardless of which DR strategy a workload's architecture implements.
AWS SAA-C03 · High Availability & DR · Card 005/020 hard

A team configures a Route 53 health check on an application endpoint with the default consecutive-failure threshold, then observes that a single health checker briefly failing to reach the endpoint does not immediately mark it unhealthy in the Route 53 console. Which statement explains this correctly?

  1. This is a bug; a single consecutive-failure count from any one health checker location always overrides Route 53's overall assessment of the endpoint.
  2. Route 53 runs health checkers from many locations worldwide, and only marks an endpoint unhealthy once 18% or fewer of the reporting health checkers currently judge it healthy — a single checker's transient failure gets outvoted by the rest.
  3. Route 53 disables DNS failover entirely until the endpoint has failed every health checker location at least once, then never re-enables it automatically.
  4. The consecutive-failure threshold only applies to CloudWatch-alarm-based health checks, not to health checks that monitor an endpoint directly.
AWS SAA-C03 · High Availability & DR · Card 006/020 easy

A team enables S3 Cross-Region Replication on a bucket that already holds two years of existing objects, expecting all of that history to appear in the destination bucket automatically. What actually happens, and what should they do instead?

  1. Cross-Region Replication automatically backfills every existing object the moment replication is enabled, so no extra action is needed.
  2. Cross-Region Replication requires S3 Transfer Acceleration to be enabled first, and only then will pre-existing objects replicate automatically.
  3. Cross-Region Replication silently deletes objects created before replication was enabled, since it only tracks objects going forward.
  4. Cross-Region Replication (live replication) only replicates new and updated objects going forward; to replicate the pre-existing two years of objects, they need to run an S3 Batch Replication job.
AWS SAA-C03 · High Availability & DR · Card 007/020 easy

A company migrating its DR plan away from self-managed replication tooling wants continuous, near-real-time replication of its on-premises servers into AWS, with the ability to launch fully booted, native EC2 recovery instances within minutes of a declared disaster. Which AWS service matches this, and how does it keep steady-state costs down?

  1. AWS Elastic Disaster Recovery (DRS), which continuously replicates source servers at the block level into a low-cost staging area (minimal compute and storage) and only converts data into fully booted EC2 instances at the moment of a drill or actual recovery.
  2. Amazon S3 Cross-Region Replication, which continuously replicates entire server volumes and boots them as EC2 instances automatically during an outage.
  3. AWS Backup, which takes scheduled snapshots and requires restoring a full backup before any recovery instance can boot, typically taking hours.
  4. Amazon Data Lifecycle Manager, which is a snapshot-scheduling tool with no ability to launch recovery instances at all.
AWS SAA-C03 · High Availability & DR · Card 008/020 easy

A team's EBS backup strategy relies on regular incremental snapshots and assumes that same incremental chain automatically protects them against a full Regional outage. What is the flaw in that assumption?

  1. EBS snapshots are always full copies, never incremental, so cross-region protection already exists without any extra step.
  2. EBS snapshots cannot be copied to another Region under any circumstances, so Regional outages always require rebuilding volumes from scratch.
  3. An EBS snapshot's incremental chain exists only within the Region it was taken in; protecting against a Regional outage requires explicitly copying snapshots to another Region, and the first copy into that Region is a full copy.
  4. Copying an EBS snapshot to another Region automatically deletes the original snapshot's incremental chain in the source Region.
AWS SAA-C03 · High Availability & DR · Card 009/020 easy

An application uses a DynamoDB global table replicated across three Regions so that if one Region becomes impaired, traffic can shift to another with minimal data loss. Which statement correctly describes how this works?

  1. Only one Region's replica accepts writes at a time, and the others are strictly read-only until a manual promotion, similar to Aurora Global Database.
  2. Global tables are multi-active: any replica can serve both reads and writes, changes replicate asynchronously to the other Regions, and concurrent conflicting writes to the same item are resolved using a last-writer-wins rule.
  3. Global tables guarantee strongly consistent reads across every Region at all times, so no conflict-resolution mechanism is ever needed.
  4. Enabling global tables converts DynamoDB from a key-value store into a strictly single-Region service that merely mirrors backups to other Regions for cold storage.
AWS SAA-C03 · High Availability & DR · Card 010/020 hard

A workload's backups run automatically every 4 hours. A Regional outage begins at 14:00, 3 hours after the most recent backup completed at 11:00. The team restores service from that 11:00 backup and traffic resumes at 15:30. In terms of RTO and RPO, which statement is accurate?

  1. RPO is the 1.5-hour restoration window (14:00 to 15:30); RTO is the 3-hour gap since the last backup (11:00 to 14:00).
  2. RTO and RPO both refer to the same 1.5-hour window, since they are defined as the total time the workload was unavailable.
  3. RPO is the entire 4.5 hours from the last backup (11:00) to service resumption (15:30); RTO does not apply because the outage was Regional rather than a single-component failure.
  4. RPO is the up-to-3-hour data loss window between the last backup (11:00) and the outage (14:00), since restoring from that backup loses any changes made after it; RTO is the 1.5-hour restoration time between the outage (14:00) and service resuming (15:30).
AWS SAA-C03 · High Availability & DR · Card 011/020 easy

A team is setting up basic active-passive DNS failover for a public website using Route 53's failover routing policy, with exactly one primary record and one secondary record. Which statement correctly describes what health checking is required?

  1. Only the primary record needs an associated health check (or, for an AWS alias target, Evaluate Target Health set to yes); Route 53 automatically returns the secondary record's answer whenever the primary is found unhealthy, without a health check needing to be attached to the secondary record.
  2. Both the primary and secondary records must each have their own independently configured health check before Route 53 will accept either record as part of a failover routing policy.
  3. Active-passive behavior requires combining the failover routing policy with a weighted routing policy on the same records; failover alone only supports active-active configurations.
  4. Route 53 decides which of the two records to return by comparing measured latency to each endpoint, falling back to the secondary only once the primary's latency exceeds a configured threshold.
AWS SAA-C03 · High Availability & DR · Card 012/020 medium

An Application Load Balancer has two registered targets in Availability Zone A and only one in Availability Zone B. A team wants to know how the ALB distributes requests across these unevenly distributed targets by default. Which statement is correct?

  1. Cross-zone load balancing is off by default for an Application Load Balancer, so each Availability Zone's load balancer node only distributes requests among its own AZ's targets, giving the lone target in Availability Zone B a much larger share of that AZ's traffic.
  2. Cross-zone load balancing can only be enabled by recreating the load balancer as a Network Load Balancer; Application Load Balancers have no such feature.
  3. Cross-zone load balancing is on by default for an Application Load Balancer and cannot be turned off at the load balancer level; each load balancer node distributes requests evenly across all registered targets in every enabled Availability Zone, regardless of how many targets sit in each AZ. It can be turned off at the target group level if needed.
  4. Cross-zone load balancing is on by default, but only within a single Availability Zone; requests never cross Availability Zone boundaries regardless of configuration.
AWS SAA-C03 · High Availability & DR · Card 013/020 medium

An Auto Scaling group spans three Availability Zones. One Availability Zone becomes impaired, causing the group's instances to become unevenly distributed across the remaining healthy zones. Some time later, the impaired Availability Zone recovers. With the group's default settings, what happens next?

  1. Nothing changes automatically; the uneven distribution persists until an operator manually terminates instances in the over-represented Availability Zones and launches new ones in the recovered zone.
  2. Amazon EC2 Auto Scaling automatically rebalances the group: it launches new instances in the Availability Zones with the fewest instances, including the recovered one, and only terminates the excess instances elsewhere after the new ones are up, so the rebalancing doesn't reduce capacity or availability while it happens.
  3. The Auto Scaling group immediately terminates instances in the over-represented Availability Zones first, then launches replacement instances in the recovered zone, briefly reducing total running capacity during the transition.
  4. Automatic Availability Zone rebalancing only happens if Capacity Rebalancing has been explicitly turned on for the group; without it, the group stays unevenly distributed indefinitely.
AWS SAA-C03 · High Availability & DR · Card 014/020 easy

A team enables replication on an Amazon EFS file system to a destination file system in a second AWS Region for disaster recovery. Which statement about how this replication behaves is accurate?

  1. Replication is synchronous, so every write to the source file system is durably committed to the destination file system before the source write is acknowledged, giving a recovery point objective of zero.
  2. The destination file system accepts both reads and writes from applications at all times, the same as the source, with no separate failover step required.
  3. Amazon EFS does not support replicating a file system to a different Region at all; only same-Region backups are available.
  4. Replication is asynchronous and, after the initial sync completes, Amazon EFS maintains a recovery point objective of about 15 minutes for most file systems; using the destination for reads and writes requires an explicit failover step to promote it, and returning to the original source afterward requires a separate failback.
AWS SAA-C03 · High Availability & DR · Card 015/020 hard

A company runs its application across two AWS Regions and uses Amazon Route 53 Application Recovery Controller (Route 53 ARC) for its multi-Region disaster recovery plan. The team wants to continuously confirm ahead of time that the standby Region is scaled and configured to absorb full traffic if needed, and separately wants a safe, guardrailed way to actually redirect DNS traffic during a real Regional impairment. Which statement correctly matches ARC's components to these two needs?

  1. Readiness checks are for the ongoing, ahead-of-time confirmation that the standby Region's resources, quotas, and configuration could handle failover traffic, and are not meant to sit in the critical path during an actual event; routing control, which supports safety rules, is what should be used to actually redirect traffic during a real impairment.
  2. Routing control is for the ahead-of-time confirmation work, and readiness checks are what actually redirect DNS traffic during a live Regional impairment.
  3. Route 53 ARC only supports shifting traffic away from a single impaired Availability Zone within one Region; redirecting traffic between two entire Regions requires a separate, non-ARC service.
  4. Readiness checks and routing control are two names for the same underlying capability, so either one can be used interchangeably for both the ahead-of-time confirmation and the live failover.
AWS SAA-C03 · High Availability & DR · Card 016/020 hard

A company's disaster recovery runbook for a Regional failover calls for its automation to invoke EC2 and Auto Scaling APIs to launch and scale up a standby environment at the moment a disaster is declared. During an actual large-scale Regional outage, this automation runs far more slowly than expected, delaying recovery. Which Well-Architected Reliability Pillar principle does this design violate, and what is the recommended fix?

  1. This is expected and unavoidable; every disaster recovery strategy, including Multi-AZ, depends on control-plane API calls during a failure, so no design change would help.
  2. The design violates the shared responsibility model; the fix is to file a support case asking AWS to prioritize the account's API calls during the outage.
  3. The design violates the static stability principle, which calls for a system to keep operating using resources that are already deployed and configured -- rather than depending on control-plane API calls made in the middle of the disruption -- precisely because control planes can be slower or degraded during the same large-scale event the design is trying to recover from; the fix is to pre-provision and pre-scale the standby capacity ahead of time so failover only requires redirecting traffic to already-running resources.
  4. This is purely an RTO and RPO measurement problem; the fix is to redefine the workload's RTO and RPO targets to match whatever the automation currently achieves.
AWS SAA-C03 · High Availability & DR · Card 017/020 medium

An application's secondary cluster in an Aurora Global Database occasionally needs to run a write, and the team enables write forwarding on that secondary cluster instead of maintaining a separate connection to the primary Region for those rare writes. Which statement correctly describes what happens to a forwarded write?

  1. Write forwarding lets the secondary cluster commit the write to its own local storage first, then asynchronously pushes that change to the primary cluster afterward.
  2. The secondary cluster forwards the write's SQL statement to the primary cluster, which applies the change first and remains the source of truth; the change is then replicated back out to all secondary Regions as usual, so a forwarded write incurs the added cross-Region round-trip rather than committing locally first. Only data manipulation statements are forwarded this way -- most data definition language statements still must run directly against the primary cluster's writer instance.
  3. Enabling write forwarding turns every secondary cluster into an additional fully independent writer, eliminating the distinction between the primary cluster and secondary clusters entirely.
  4. Write forwarding requires switching the entire global database from storage-level replication to engine-level logical replication.
AWS SAA-C03 · High Availability & DR · Card 018/020 easy

A compliance team needs certain S3 objects to be impossible for anyone to delete or overwrite before a fixed retention date -- including anyone with root access to the AWS account -- to satisfy a regulatory records-retention requirement. Which S3 Object Lock configuration meets this, and why don't the alternatives?

  1. Governance mode, because by default no one, including users with special permissions, can ever bypass a governance-mode retention period.
  2. A legal hold by itself, because a legal hold enforces a fixed, predetermined expiration date that no user can shorten.
  3. Enabling S3 Versioning alone, without Object Lock, because keeping every version of an object already prevents any version from being permanently deleted.
  4. Compliance mode, because a compliance-mode retention period cannot be shortened or removed by any user, including the account's root user, before the configured retain-until date -- the only way to delete such an object earlier is to delete the entire AWS account -- whereas governance mode allows users with the specific bypass permission to override it, and versioning alone doesn't stop a locked object version from being deleted at all without Object Lock enabled.
AWS SAA-C03 · High Availability & DR · Card 019/020 easy

A team runs a production database on a Multi-AZ RDS DB instance deployment (single-standby). A mandatory operating system security patch becomes available and RDS applies it during the configured maintenance window. Which statement correctly describes how RDS applies this patch and the resulting impact?

  1. RDS applies the OS patch to the standby instance first, then promotes that newly patched standby to primary, and finally applies the patch to the old primary, which becomes the new standby -- so the deployment experiences only a brief Multi-AZ failover, typically well under a minute, rather than an extended outage.
  2. RDS patches the current primary instance first while it continues serving traffic, then patches the standby afterward, with no failover involved at either step.
  3. RDS patches both the primary and standby instances at exactly the same time, taking the entire Multi-AZ deployment offline for the full duration of the OS patch, identical to how a Single-AZ deployment would experience it.
  4. A Multi-AZ deployment defers all mandatory OS patches indefinitely until an operator manually initiates a failover first.
AWS SAA-C03 · High Availability & DR · Card 020/020 easy

A team has defined an RTO and RPO for a production workload and wants an AWS service that will assess the workload's actual architecture against those specific targets and produce prioritized, concrete recommendations for closing any gaps -- rather than just monitoring metrics or executing backups. Which service fits, and how does it differ from Trusted Advisor and AWS Backup?

  1. AWS Backup, because its centralized backup policies are themselves the mechanism that evaluates whether an architecture meets a defined RTO and RPO.
  2. Amazon CloudWatch, because configuring alarms on the workload's resources is equivalent to assessing whether its architecture meets defined RTO and RPO targets.
  3. AWS Resilience Hub, because it evaluates an application's architecture against a resiliency policy built from the team's own defined RTO and RPO targets, identifies specific gaps, and produces prioritized recommendations to close them -- distinct from Trusted Advisor's general best-practice checks across cost, security, and service limits, and from AWS Backup's role in actually executing backup and restore operations.
  4. AWS Trusted Advisor, because its checks are specifically designed around a workload's individually defined RTO and RPO targets rather than general best practices.