Exam Overview
Exam Details
- Duration: 180 minutes
- Questions: 65 + exam lab
- Passing Score: 720/1000
- Format: Multiple choice & multiple response
- Cost: ~$150-300 USD
- Validity: 3 years
Exam Domains
| Domain | Weight |
|---|---|
| Monitoring, Logging & Remediation | 20% |
| Reliability & Business Continuity | 16% |
| Deployment, Provisioning & Automation | 18% |
| Security & Compliance | 16% |
| Networking & Content Delivery | 18% |
| Cost & Performance Optimization | 12% |
CloudWatch Deep Dive
| Feature | Description | Key Details |
|---|---|---|
| Metrics | Numerical data over time | Default 1-min resolution; custom up to 1-second; retained 15 months |
| Alarms | Act on metric thresholds | States: OK, ALARM, INSUFFICIENT_DATA; notify SNS, trigger ASG |
| Logs | Centralized log storage | Log Groups → Streams; retention 1 day to indefinite |
| Log Insights | Query logs with SQL-like syntax | Query across multiple log groups |
| CloudWatch Agent | Collect OS-level metrics | Memory, disk usage (NOT default EC2 metrics) |
⚠ Memory & Disk are NOT default EC2 metrics!
CPU, Network, and EBS metrics ARE default. To get memory usage and disk space, install the CloudWatch Agent on the EC2 instance.
AWS Systems Manager (SSM)
| Feature | What It Does | Use Case |
|---|---|---|
| Session Manager | Browser-based SSH/RDP without port 22/3389 | Secure shell access; audit trail in CloudTrail |
| Run Command | Run scripts on many EC2 at once without SSH | Patch, configure, update many instances |
| Patch Manager | Automated OS patching | Schedule patching windows; compliance reporting |
| Parameter Store | Store config and secrets (strings, SecureString) | App config, non-critical secrets |
| Inventory | Collect software/hardware inventory | Compliance, auditing |
Backup & Disaster Recovery
Recovery Strategies (Cost vs RTO)
| Strategy | RTO | RPO | Cost | Description |
|---|---|---|---|---|
| Backup & Restore | Hours | Hours | $ | Cheapest. Backup to S3/Glacier; restore on disaster. |
| Pilot Light | Minutes-Hours | Minutes | $$ | Core systems in AWS; scale up on disaster. |
| Warm Standby | Minutes | Seconds-Minutes | $$$ | Scaled-down version of prod running; scale up fast. |
| Multi-Site Active/Active | Real-time | ~0 | $$$$ | Full duplicate; DNS routes to both. Zero downtime. |
CloudFormation
| Section | Purpose | Required? |
|---|---|---|
| AWSTemplateFormatVersion | Always "2010-09-09" | No |
| Parameters | Input values at deploy time | No |
| Mappings | Static key-value lookup (e.g., AMI by region) | No |
| Conditions | Create resources conditionally | No |
| Resources | AWS resources to create | YES |
| Outputs | Values to export or display | No |
- StackSets — deploy stacks to multiple accounts/regions
- Nested Stacks — reusable stacks as components
- Change Sets — preview changes before applying
- Drift Detection — find resources changed outside CloudFormation
📋 Study Checklist
Progress0%
- Install CloudWatch Agent for memory/disk metrics
- Create CloudWatch Alarms connected to Auto Scaling or SNS
- Use SSM Session Manager for SSH without port 22
- Use SSM Run Command for mass instance management
- Configure SSM Patch Manager with maintenance windows
- Know all 4 DR strategies and their RTO/RPO
- Configure AWS Backup with cross-region copy
- Understand IAM policy evaluation order (deny > allow > implicit deny)
- Explain SCPs vs Permission Boundaries vs Resource Policies
- Create cross-account IAM roles with sts:AssumeRole
- Write a CloudFormation template with Parameters, Resources, Outputs
- Use Change Sets and Drift Detection
- Understand CloudFormation StackSets
- Configure S3 lifecycle policies for cost optimization
- Know EC2 Health Check vs ELB Health Check differences
- Understand Auto Scaling lifecycle hooks
- Use AWS Config for compliance checking
- Configure CloudTrail for API audit logging
- Know Trusted Advisor check categories
- Understand AWS Cost Explorer and Compute Optimizer