Skip to content

17 — Cloud Architecture

Cloud architecture is the practice of designing systems that leverage cloud computing benefits — elasticity, pay-as-you-go, managed services, global infrastructure, and high availability — while avoiding vendor lock-in and runaway costs.

Analogy: On-premises IT is like owning a house — you buy everything, maintain everything, and pay whether you’re using it or not. Cloud architecture is like a hotel — you pay for exactly the room and services you need, only when you need them.


Poor cloud architecture leads to:

  • High costs — over-provisioned resources, idle capacity
  • Vendor lock-in — tightly coupled to one provider’s services
  • Performance issues — wrong region, wrong service, wrong tier
  • Security gaps — misconfigured S3 buckets, open security groups
  • Complex migration — lift-and-shift without refactoring

flowchart TB
OnPrem["On-Premises<br/>You manage everything"] --> IaaS["IaaS<br/>You manage: OS, apps, data<br/>Provider manages: servers, network"]
IaaS --> PaaS["PaaS<br/>You manage: apps, data<br/>Provider manages: runtime, OS, servers"]
PaaS --> SaaS["SaaS<br/>You manage: data<br/>Provider manages: everything else"]
SaaS --> FaaS["FaaS (Serverless)<br/>You manage: code only<br/>Provider manages: everything"]
style OnPrem fill:#3b82f6,color:#fff
style IaaS fill:#059669,color:#fff
style PaaS fill:#f59e0b,color:#fff
style SaaS fill:#ef4444,color:#fff
style FaaS fill:#7c3aed,color:#fff

Six pillars of a well-architected cloud system:

PillarWhat It MeansKey Practices
Operational ExcellenceRun and monitor systemsIaC, automation, observability
SecurityProtect data and systemsIAM, encryption, least privilege
ReliabilityRecover from failuresHA, DR, auto-scaling
Performance EfficiencyUse resources effectivelyRight-sizing, serverless, caching
Cost OptimizationMinimize costsReserved instances, lifecycle policies
SustainabilityMinimize environmental impactEfficient code, region selection

flowchart TB
Patterns["Cloud Architecture Patterns"] --> LiftShift["Lift & Shift<br/>Move existing app as-is<br/>Fastest migration"]
Patterns --> Refactor["Re-platform<br/>Minor changes for cloud<br/>Good performance gains"]
Patterns --> Re-architect["Re-architect<br/>Redesign for cloud-native<br/>Best long-term results"]
Patterns --> CloudNative["Cloud-Native<br/>Built from scratch for cloud<br/>Maximum benefits"]
LiftShift --> L_Details["Migration speed: ⭐⭐⭐<br/>Cloud benefits: ⭐"]
Refactor --> R_Details["Migration speed: ⭐⭐<br/>Cloud benefits: ⭐⭐"]
Re-architect --> RA_Details["Migration speed: ⭐<br/>Cloud benefits: ⭐⭐⭐"]
CloudNative --> CN_Details["Migration speed: N/A<br/>Cloud benefits: ⭐⭐⭐"]
style Patterns fill:#7c3aed,color:#fff
style LiftShift fill:#3b82f6,color:#fff
style Refactor fill:#059669,color:#fff
style Re-architect fill:#f59e0b,color:#fff
style CloudNative fill:#ef4444,color:#fff

flowchart LR
subgraph Cloud_A["Primary Cloud (AWS)"]
S3["S3 Storage"]
Lambda["Lambda Functions"]
Dynamo["DynamoDB"]
end
subgraph Cloud_B["Secondary Cloud (GCP)"]
PubSub["Pub/Sub"]
BigQuery["BigQuery Analytics"]
CF["Cloud Functions"]
end
subgraph OnPrem["On-Premises / Edge"]
Legacy["Legacy Systems"]
Cache["Edge Cache"]
end
Cloud_A <-->|"Cross-cloud communication"| Cloud_B
Cloud_A <-->|"Hybrid connectivity"| OnPrem
Cloud_B <-->|"Data sync"| OnPrem
style Cloud_A fill:#7c3aed,color:#fff
style Cloud_B fill:#3b82f6,color:#fff
style OnPrem fill:#f59e0b,color:#fff
Cloud StrategyProsCons
Single CloudDeep integration, simpler opsVendor lock-in
Multi-CloudBest-of-breed services, no lock-inHigher complexity, cross-cloud networking
Hybrid CloudKeep sensitive data on-prem, burst to cloudNetwork complexity, consistent management
Cloud-NativeFull benefits of cloudRequires redesign, new skills

flowchart TB
Cost["Cloud Cost Optimization"] --> RightSize["Right-Sizing<br/>Match instance type to workload"]
Cost --> Reserved["Reserved/Committed<br/>Steady workloads → reserves<br/>Up to 72% discount"]
Cost --> Spot["Spot/Preemptible<br/>Fault-tolerant batch work<br/>Up to 90% discount"]
Cost --> AutoScale["Auto-Scaling<br/>Scale down when not needed"]
Cost --> Lifecycle["Lifecycle Policies<br/>Move old data to cheaper tiers"]
Cost --> Serverless["Serverless<br/>Pay per execution,<br/>no idle cost"]
style Cost fill:#7c3aed,color:#fff
style RightSize fill:#3b82f6,color:#fff
style Reserved fill:#059669,color:#fff
style Spot fill:#f59e0b,color:#fff
style AutoScale fill:#6366f1,color:#fff
style Serverless fill:#ef4444,color:#fff

PrincipleDescription
Design for failureAssume everything fails — build redundancy
Build loosely coupledServices don’t depend on each other’s internals
Implement elasticityScale up AND down — don’t pay for idle resources
Think parallelDistribute work across many small resources
Automate everythingNo manual configuration — use IaC
Use managed servicesLet the cloud provider handle operations
Cache aggressivelyReduce load, improve latency
Secure by defaultLeast privilege, encryption everywhere

DecisionProsCons
Managed servicesLess ops workVendor lock-in, higher per-unit cost
ServerlessNo idle cost, auto-scaleCold starts, max execution time
Multi-cloudNo lock-in, best servicesCross-cloud latency, complexity
Reserved instancesMajor discountCommitment, less flexibility
All-in on one providerDeep integration, simplerHard to migrate away

StrategyCloud Implementation
Horizontal scalingAuto Scaling Groups, serverless (auto-scale by default)
Global distributionCDN (CloudFront), multi-region deployment
Elastic storageS3 (unlimited scale), DynamoDB (auto-scaling throughput)
Queue-based decouplingSQS / SNS / EventBridge between services
Infrastructure as CodeCloudFormation, Terraform, CDK

  1. What are the six pillars of the AWS Well-Architected Framework?
  2. When would you choose a single cloud vs multi-cloud strategy?
  3. How do you optimize cloud costs without sacrificing performance?
  4. What’s the difference between lift-and-shift and cloud-native?
  5. How do you choose between EC2, Lambda, and ECS for a workload?

SystemCloud Architecture
NetflixAWS-native (cloud-native), microservices, Chaos Monkey
AirbnbAWS with multi-region (US East + Europe), cloud-optimized
SpotifyGCP-native, migrated from on-prem to cloud
Capital OneFull AWS migration from on-prem — all-in on a single cloud

  • Cloud architecture = designing systems to maximize cloud benefits (elasticity, managed services, global reach)
  • Six pillars: Operational Excellence, Security, Reliability, Performance, Cost, Sustainability
  • Lift & shift is fastest but gets least cloud benefit — cloud-native is hardest but best
  • Multi-cloud = best tools, no lock-in — but adds complexity
  • Hybrid cloud = keep sensitive data on-prem, use cloud for burst
  • Cost optimization = right-size + reserved/spot + auto-scale + serverless + lifecycle policies
  • Design for failure, automate everything, use managed services — the three cloud mantras