Production AWS networking is less about memorizing service names and more about clear trust boundaries. This guide walks through battle-tested VPC patterns you can adapt for SaaS, internal platforms, and regulated workloads.
Why VPC design still matters
Even with serverless and managed container platforms, every packet still traverses a network you own. Poor subnet planning shows up later as:
- brittle NAT cost spikes
- security group sprawl
- impossible private connectivity for data stores
- multi-AZ failures that cascade into total outages
Info: Treat your first VPC as a product. Document CIDR plans, naming conventions, and ownership before the second team ships into it.
Managed runtimes still need subnet placement and egress policy. Front-loading those decisions is cheaper than a mid-quarter migration when a private database or partner network appears.
Reference topology
A durable baseline looks like this:
Account / Region
└── VPC 10.20.0.0/16
├── Public AZ-a 10.20.0.0/20 (ALB, NAT)
├── Public AZ-b 10.20.16.0/20
├── Private AZ-a 10.20.32.0/20 (ECS/EKS, app)
├── Private AZ-b 10.20.48.0/20
├── Data AZ-a 10.20.64.0/20 (RDS, caches)
└── Data AZ-b 10.20.80.0/20Keep data subnets isolated with no direct internet path. Application tiers reach the internet only through controlled NAT or VPC endpoints. Name subnets consistently (prod-app-private-a) so responders can map blast radius without reopening a diagram.
CIDR planning checklist
- Reserve space for future accounts and peered networks.
- Prefer
/16or/18VPCs for production unless you are tightly constrained. - Avoid overlapping CIDRs with corporate VPN or partner networks.
- Leave empty
/20blocks for shared services later.
# Example: allocate non-overlapping ranges per environment
# prod: 10.20.0.0/16
# staging: 10.30.0.0/16
# shared: 10.10.0.0/16Publish the allocation table as a living doc. New environments should check out a range from that table—not invent one in Slack that later collides with a partner CIDR.
Public vs private tier decisions
Public subnets
Use public subnets only for:
- Application Load Balancers
- NAT Gateways
- Bastion hosts (prefer Session Manager instead)
Private subnets
Run application compute here. Attach VPC endpoints for S3, ECR, CloudWatch, and STS to reduce NAT dependency. Gateway endpoints for S3 (and DynamoDB where applicable) usually cut NAT bytes first; add interface endpoints for ECR, Logs, and STS once you measure real egress.
Warning: Do not place RDS or ElastiCache in public subnets "temporarily." Temporary networking decisions become permanent blast radius.
Security groups as contracts
Model security groups as intent, not as IP spaghetti:
| Group | Ingress | Purpose |
|---|---|---|
sg-alb | 443 from Internet / CloudFront | Edge entry |
sg-app | 8080 from sg-alb | Application tier |
sg-db | 5432 from sg-app | Data tier |
Prefer security-group references over CIDR rules. When a service needs database access, attach it to sg-app (or sg-worker) rather than opening the data tier to a subnet CIDR. Keep NACLs mostly permissive unless compliance requires subnet denies—duplicating SG logic in NACLs creates asymmetric failures that are hard to debug.
Multi-AZ resilience pattern
Always deploy stateful and stateless tiers across at least two AZs. For NAT:
- High availability: one NAT Gateway per AZ
- Cost-sensitive non-prod: single NAT with accepted AZ risk
const privateRoute = {
destinationCidrBlock: "0.0.0.0/0",
natGatewayId: natGatewayPerAz[availabilityZone],
};Route each private subnet’s 0.0.0.0/0 to the NAT in the same AZ. Cross-AZ NAT looks cheaper until that AZ fails—or until cross-AZ data charges spike under load. Place database primary and standby in separate data subnets spanning the same AZs as the app tier.
Private connectivity without flattening the graph
Prefer narrow patterns over “peer everything”:
- VPC peering for simple one-to-one links with non-overlapping CIDRs
- Transit Gateway when many VPCs need hub-and-spoke routing or centralized inspection
- PrivateLink / interface endpoints when a consumer needs one service, not a whole VPC
A shared-services VPC with controlled attachments ages better than an ad-hoc peering mesh nobody can draw from memory.
Observability for the network layer
Track these signals from day one:
- NAT Gateway bytes and active connections
- Rejected security group packets (VPC Flow Logs)
- ALB 5xx correlated with target health
- Endpoint interface ENI saturation
Alert on NAT connection spikes and endpoint packet drops before users report timeouts. Correlate ALB target health with AZ-local NAT so egress failures are not misfiled as application bugs.
Common mistakes
- One giant public subnet for compute and data
- Tiny
/24VPCs that cannot grow - Single NAT for production (one AZ = full egress outage)
- Ticket-named security groups that never get cleaned up
- Skipping endpoints while image pulls dominate NAT bills
- Overlapping CIDRs with on-prem or partners, found only at VPN time
Rollout checklist
- CIDR plan reviewed against corporate, partner, and future ranges
- Public / private / data subnets in at least two AZs
- NAT Gateway per AZ for production (or documented risk acceptance)
- Endpoints for high-volume AWS APIs you actually call
- Security group matrix in IaC; Flow Logs and NAT/ALB alarms live
- No public IPs on data stores; Session Manager over bastions
- Documented AZ-loss survival assumptions
FAQ
Separate VPC per environment?
Usually yes for production. Staging may share an account with clear subnet and IAM boundaries, but shared blast radius with prod is rarely worth the savings.
When Transit Gateway?
When peering count or inspection needs outgrow pairwise links. A handful of VPCs can stay on peering; many spoke accounts usually cannot.
Are endpoints mandatory on day one?
No—but add them once ECR, S3, or logging traffic dominates NAT. Measure, then enable what moves cost and reliability.
Does this model fit Lambda?
Yes. Prefer private subnets when functions must reach private data stores, and keep AZ coverage in mind for ENI placement.
Related GyanTrails resources
- Read next: Designing Idempotent APIs
- Review IAM least privilege before expanding east-west traffic
Closing principle
Design for least privilege paths, not maximum convenience. A VPC that is slightly harder to bootstrap but easy to reason about will outlive every hot framework choice on top of it.
