Every AWS account comes with a network you didn’t build, and most people run production in it for a year before they look at what it actually is. Then something needs a private subnet, or a second VPC has to talk to the first, or the bill grows a line called “NAT Gateway” that costs more than the servers behind it, and suddenly the VPC matters.
This is the guide I wish I’d had at that point. What a VPC is made of, how the pieces fit, the two or three decisions that are hard to undo later, and where the money goes. Everything with an ID in it below was pulled from a real account with the AWS CLI (v2.32, read-only calls only; the resource IDs are real), so you can run the same commands against yours and compare.
What a VPC is, and what you already have
A Virtual Private Cloud is a private IP network inside an AWS region. You choose the address range, cut it into subnets, decide which subnets can reach the internet, and put firewalls on the instances. It behaves like a physical network you’d wire in a data centre, except every piece is an API call and nothing is shared with another customer.
Before you create one, look at the one AWS gave you:
$ aws ec2 describe-vpcs --query 'Vpcs[].{VpcId:VpcId,Cidr:CidrBlock,Default:IsDefault,State:State}' --output table
---------------------------------------------------------------
| DescribeVpcs |
+----------------+----------+------------+--------------------+
| Cidr | Default | State | VpcId |
+----------------+----------+------------+--------------------+
| 172.31.0.0/16 | True | available | vpc-0158b49ec514e60ac |
+----------------+----------+------------+--------------------+
$ aws ec2 describe-subnets --query 'Subnets[].{AZ:AvailabilityZone,Cidr:CidrBlock,Free:AvailableIpAddressCount,PublicIp:MapPublicIpOnLaunch}' --output table
------------------------------------------------------
| DescribeSubnets |
+-------------+-----------------+-------+------------+
| AZ | Cidr | Free | PublicIp |
+-------------+-----------------+-------+------------+
| us-east-1a | 172.31.0.0/20 | 4091 | True |
| us-east-1b | 172.31.80.0/20 | 4091 | True |
| us-east-1c | 172.31.16.0/20 | 4091 | True |
| us-east-1d | 172.31.32.0/20 | 4091 | True |
| us-east-1e | 172.31.48.0/20 | 4091 | True |
| us-east-1f | 172.31.64.0/20 | 4091 | True |
+-------------+-----------------+-------+------------+
That’s the default VPC: a /16, one /20 subnet per availability zone, every subnet public, every instance handed a public IP at launch. It’s fine for a weekend project. It’s the wrong shape for anything with a database in it, because there is no private side at all. Also notice Free: 4091 on a /20. A /20 has 4,096 addresses; AWS keeps five in every subnet (the network address, the VPC router, the DNS resolver, one spare, and broadcast). Plan capacity with that in mind: a /28 gives you 11 usable, not 14.
The CIDR block: the decision you can’t undo
A VPC gets one primary IPv4 range between /16 and /28, from the private space in RFC 1918. You can add secondary ranges later; you cannot shrink or change the primary. So think for five minutes before typing 10.0.0.0/16, because:
- It must not overlap any network this VPC will ever be peered with, VPN’d to, or reached through a Transit Gateway. Two VPCs on 10.0.0.0/16 can never be connected. Ever. Renumbering one is a rebuild.
- The default VPC took 172.31.0.0/16. Your office LAN is probably on 192.168.x. So 10.x is free, but carve it: 10.0.0.0/16 for prod in one region, 10.1.0.0/16 for staging, 10.10.0.0/16 for another region, and write the plan down.
- A /16 is 65,536 addresses. You will not run out. A /24 for a whole VPC, on the other hand, is a mistake people make when they confuse “VPC” with “subnet”.
If the arithmetic behind prefixes and subnet sizes is fuzzy, CIDR and subnetting explained with examples covers it with real ipcalc output, and it’s exactly the maths that VPC design runs on.
Subnets, and what “public” actually means
A subnet is a slice of the VPC range pinned to one availability zone. That AZ pinning is the point: put a subnet in us-east-1a and another in us-east-1b and you can survive one of them going dark. Two AZs minimum, three if you’re picky, and mirror the layout in each.
Here’s the part that trips people up. There is no “public” checkbox on a subnet. A subnet is public because its route table has a route to an internet gateway; it’s private because it doesn’t. That’s it. The MapPublicIpOnLaunch flag you saw above just decides whether instances get a public address handed to them, which is useless without the route and dangerous with it.
A layout that works for nearly everything:
| Subnet | AZ a | AZ b | Route to internet | What lives there |
|---|---|---|---|---|
| public | 10.0.0.0/24 | 10.0.1.0/24 | via internet gateway | load balancers, NAT gateway, bastion |
| private app | 10.0.10.0/24 | 10.0.11.0/24 | via NAT gateway (outbound only) | EC2, containers, Lambda in VPC |
| private data | 10.0.20.0/24 | 10.0.21.0/24 | none | RDS, ElastiCache, anything that never needs out |
Public subnets can be small; almost nothing should be in them. The Application Load Balancer sits there and forwards to app instances in the private tier, which is the whole reason the private tier can exist without public IPs.
Route tables
Every VPC has a main route table, and every subnet uses it unless you associate a different one. The default VPC’s main table is why everything in it is public:
$ aws ec2 describe-route-tables --query 'RouteTables[].{Rtb:RouteTableId,Main:Associations[0].Main,Routes:Routes[].[DestinationCidrBlock,GatewayId||NatGatewayId]}'
[
{
"Rtb": "rtb-039432768763e10b4",
"Main": true,
"Routes": [
[ "172.31.0.0/16", "local" ],
[ "0.0.0.0/0", "igw-0d9d5e63d13cb385b" ]
]
}
]
Two routes. The local one is created for you and can’t be removed: it’s how subnets in the same VPC reach each other, across AZs, without any gateway. The 0.0.0.0/0 to igw-... is the one that makes it public. Delete that route from a table and every subnet using it becomes private on the spot; no instance restart, no re-addressing.
Routing works by longest prefix match, same as on a Linux box. A packet to 172.31.5.9 matches both the /16 and the /0, the /16 wins, it stays local. That’s the mechanism behind every peering connection, VPN and endpoint below: each one is just a more specific route that beats the default.
The habit worth forming: leave the main table with only the local route, and create explicit tables for public and private subnets. Then a new subnet is private until you deliberately make it otherwise.
Internet Gateway
An IGW is a horizontally scaled, AWS-managed thing you attach to a VPC once. It does two jobs: gives the VPC a path to the internet, and performs the one-to-one NAT between an instance’s private address and its public or Elastic IP. There’s nothing to size or pay for.
$ aws ec2 describe-internet-gateways --query 'InternetGateways[].{Igw:InternetGatewayId,Attached:Attachments[0].VpcId,State:Attachments[0].State}' --output table
---------------------------------------------------------------
| DescribeInternetGateways |
+------------------------+-------------------------+----------+
| Attached | Igw | State |
+------------------------+-------------------------+----------+
| vpc-0158b49ec514e60ac | igw-0d9d5e63d13cb385b | available|
+------------------------+-------------------------+----------+
For an instance to be reachable from the internet it needs all four: a public or Elastic IP, a subnet whose route table points 0.0.0.0/0 at the IGW, a security group that allows the port, and a network ACL that doesn’t block it. When “I can’t SSH to my instance” comes up, it’s one of those four, in roughly that order of likelihood. And note that since early 2024 every public IPv4 address costs about $0.005 an hour whether it’s attached or not, which is $3.60 a month per address and adds up across a fleet.
NAT Gateway
Private instances still need to fetch packages, pull images and call APIs. A NAT gateway lives in a public subnet, and private route tables send 0.0.0.0/0 to it; it forwards the traffic out through the IGW using its own Elastic IP and only ever allows return traffic back. Nothing on the internet can initiate a connection to the private side.
It’s managed, it scales to 100 Gbps, and it is the single most common surprise on a VPC bill. In us-east-1 you pay around $0.045 per hour just for it to exist, about $32 a month, plus $0.045 for every gigabyte that passes through. Two things make that worse than it sounds:
- You want one per AZ, because a NAT gateway in us-east-1a is useless if 1a is down and your 1b instances route through it. So double the hourly cost, and make sure each private route table points at the NAT in its own AZ.
- Traffic to S3, ECR and DynamoDB from private subnets goes through the NAT by default and gets charged per gigabyte. A container fleet pulling images all day can rack up hundreds of dollars for traffic that never needed to leave AWS. That’s what endpoints are for.
The default VPC has no NAT gateway, because it has no private subnets:
$ aws ec2 describe-nat-gateways --query 'NatGateways[].NatGatewayId' --output text
(empty)
Security groups vs network ACLs
Two firewalls, and the difference is worth getting exactly right.
A security group is attached to a network interface, so effectively to an instance, a load balancer, an RDS instance. It’s stateful: allow inbound 443 and the replies go out automatically. Rules are allow-only; anything not listed is dropped. And a rule can reference another security group instead of a CIDR, which is the trick that makes tiered designs clean: “the database SG allows 5432 from the app SG”, and it keeps working as app instances come and go.
Here’s the default security group of the default VPC, which explains a lot of “why can my instances talk to each other but nothing else”:
$ aws ec2 describe-security-groups --group-names default --query 'SecurityGroups[0].{Id:GroupId,In:IpPermissions[].{proto:IpProtocol,src:UserIdGroupPairs[0].GroupId||IpRanges[0].CidrIp},Out:IpPermissionsEgress[].{proto:IpProtocol,dst:IpRanges[0].CidrIp}}'
{
"Id": "sg-0cfa3e16a6745085f",
"In": [ { "proto": "-1", "src": "sg-0cfa3e16a6745085f" } ],
"Out": [ { "proto": "-1", "dst": "0.0.0.0/0" } ]
}
Inbound: everything, but only from members of this same group. Outbound: everything, anywhere. That’s the default, and it’s why a fresh instance can’t be reached on port 22 until you add a rule.
A network ACL is attached to a subnet. It’s stateless, so you need an inbound rule and an outbound rule for the ephemeral return ports (1024 to 65535), it supports deny rules, and it’s evaluated in rule-number order with the first match winning. The default one allows everything, in both directions, and ends with a deny that never fires:
$ aws ec2 describe-network-acls --query 'NetworkAcls[0].Entries[].{rule:RuleNumber,egress:Egress,proto:Protocol,action:RuleAction,cidr:CidrBlock}' --output table
-----------------------------------------------------
| DescribeNetworkAcls |
+--------+-------------+---------+--------+---------+
| action | cidr | egress | proto | rule |
+--------+-------------+---------+--------+---------+
| allow | 0.0.0.0/0 | True | -1 | 100 |
| deny | 0.0.0.0/0 | True | -1 | 32767 |
| allow | 0.0.0.0/0 | False | -1 | 100 |
| deny | 0.0.0.0/0 | False | -1 | 32767 |
+--------+-------------+---------+--------+---------+
| Security group | Network ACL | |
|---|---|---|
| Attached to | network interface (instance, ALB, RDS…) | subnet |
| State | stateful | stateless |
| Rule types | allow only | allow and deny |
| Evaluation | all rules, any match allows | numbered, first match wins |
| Can reference | CIDR, prefix list, another SG | CIDR only |
| Use it for | nearly everything | blocking a CIDR at the subnet edge, compliance |
In practice: do your real access control with security groups, leave NACLs at the default, and reach for a NACL deny only when you need to drop a specific address range from a whole subnet. NACLs that try to be precise turn into a maintenance problem because of the stateless return-port rules.
VPC endpoints: talking to AWS without leaving AWS
By default, a private instance calling S3 goes out through the NAT gateway, across the public internet, and back into AWS. An endpoint gives it a private path instead. There are two kinds and they’re priced very differently.
Gateway endpoints exist for exactly two services, S3 and DynamoDB. They’re a route table entry pointing at a managed prefix list, they’re free, and they remove the NAT per-gigabyte charge for that traffic. There’s no reason not to create both in every VPC.
Interface endpoints (PrivateLink) work for nearly every other service: ECR, Secrets Manager, SSM, CloudWatch, KMS, and so on. Each one is an elastic network interface in your subnet with a private IP, and DNS inside the VPC resolves the service name to it. They cost roughly $0.01 per hour per AZ plus $0.01 per gigabyte, so around $7 a month each per AZ. Worth it for ECR and SSM in any private subnet; add the rest when the NAT bill says so.
$ aws ec2 describe-vpc-endpoints --query 'VpcEndpoints[].ServiceName' --output text
(empty)
Nothing in the default VPC, as you’d expect.
Connecting VPCs: peering vs Transit Gateway
Peering is a direct, private link between two VPCs, in the same account or different ones, same region or across regions. You accept the request, add routes on both sides, and the two ranges can reach each other. It’s cheap (data transfer in the same AZ is free; cross-AZ and cross-region are charged per gigabyte) and it’s non-transitive: if A peers with B and B peers with C, A cannot reach C. With three VPCs that’s three peerings. With ten, it’s forty-five, and someone has to maintain the route tables.
Transit Gateway is the hub. Attach every VPC and VPN to it once, and it routes between all of them, with route tables of its own for segmentation. It costs about $0.05 per hour per attachment plus $0.02 per gigabyte, which is why small estates peer and large ones use a TGW. The crossover is usually somewhere around four or five VPCs.
Both share the rule from the CIDR section: overlapping ranges cannot be connected. The account I pulled these examples from has neither yet:
$ aws ec2 describe-vpc-peering-connections --query 'VpcPeeringConnections[].VpcPeeringConnectionId' --output text
(empty)
DNS inside the VPC
Two VPC attributes decide whether private DNS works, and both need to be on for interface endpoints, RDS hostnames and most things to behave:
$ aws ec2 describe-vpc-attribute --vpc-id vpc-0158b49ec514e60ac --attribute enableDnsSupport --output text
vpc-0158b49ec514e60ac
ENABLEDNSSUPPORT True
$ aws ec2 describe-vpc-attribute --vpc-id vpc-0158b49ec514e60ac --attribute enableDnsHostnames --output text
vpc-0158b49ec514e60ac
ENABLEDNSHOSTNAMES True
The resolver lives at the VPC range plus two (10.0.0.2 in a 10.0.0.0/16), which is one of the five reserved addresses. On an instance it shows up as 127.0.0.53 via systemd-resolved on Ubuntu, but resolvectl status will show the .2 address behind it. If you run your own BIND or forward to on-prem, Route 53 Resolver endpoints are the supported way to do it rather than replacing that resolver.
VPC Flow Logs
Flow logs record every accepted and rejected connection at the interface, subnet or VPC level, to CloudWatch Logs or S3. They’re the answer to “did that packet get dropped by the security group or by the NACL”, and to “what is talking to what”. They’re not on by default. Turn them on at the VPC level, send them to S3 (cheaper than CloudWatch for volume), and you’ll thank yourself during the first incident. The format is a plain space-separated line per flow, which grep and the usual tools handle fine.
Building one properly
Here’s the two-AZ layout from the table above as CLI calls, in the order the dependencies require. Set your region first, and these create nothing that costs money until the NAT gateway step.
# the VPC, with DNS on
VPC=$(aws ec2 create-vpc --cidr-block 10.0.0.0/16 --query Vpc.VpcId --output text)
aws ec2 modify-vpc-attribute --vpc-id $VPC --enable-dns-hostnames
aws ec2 create-tags --resources $VPC --tags Key=Name,Value=prod
# subnets: public and private in two AZs
PUB_A=$(aws ec2 create-subnet --vpc-id $VPC --cidr-block 10.0.0.0/24 --availability-zone us-east-1a --query Subnet.SubnetId --output text)
PUB_B=$(aws ec2 create-subnet --vpc-id $VPC --cidr-block 10.0.1.0/24 --availability-zone us-east-1b --query Subnet.SubnetId --output text)
PRV_A=$(aws ec2 create-subnet --vpc-id $VPC --cidr-block 10.0.10.0/24 --availability-zone us-east-1a --query Subnet.SubnetId --output text)
PRV_B=$(aws ec2 create-subnet --vpc-id $VPC --cidr-block 10.0.11.0/24 --availability-zone us-east-1b --query Subnet.SubnetId --output text)
# internet gateway + public route table
IGW=$(aws ec2 create-internet-gateway --query InternetGateway.InternetGatewayId --output text)
aws ec2 attach-internet-gateway --vpc-id $VPC --internet-gateway-id $IGW
RT_PUB=$(aws ec2 create-route-table --vpc-id $VPC --query RouteTable.RouteTableId --output text)
aws ec2 create-route --route-table-id $RT_PUB --destination-cidr-block 0.0.0.0/0 --gateway-id $IGW
aws ec2 associate-route-table --route-table-id $RT_PUB --subnet-id $PUB_A
aws ec2 associate-route-table --route-table-id $RT_PUB --subnet-id $PUB_B
# NAT gateway in the public subnet (this is where the meter starts)
EIP=$(aws ec2 allocate-address --domain vpc --query AllocationId --output text)
NAT=$(aws ec2 create-nat-gateway --subnet-id $PUB_A --allocation-id $EIP --query NatGateway.NatGatewayId --output text)
aws ec2 wait nat-gateway-available --nat-gateway-ids $NAT
# private route table via the NAT, plus a free S3 gateway endpoint
RT_PRV=$(aws ec2 create-route-table --vpc-id $VPC --query RouteTable.RouteTableId --output text)
aws ec2 create-route --route-table-id $RT_PRV --destination-cidr-block 0.0.0.0/0 --nat-gateway-id $NAT
aws ec2 associate-route-table --route-table-id $RT_PRV --subnet-id $PRV_A
aws ec2 associate-route-table --route-table-id $RT_PRV --subnet-id $PRV_B
aws ec2 create-vpc-endpoint --vpc-id $VPC --service-name com.amazonaws.us-east-1.s3 --route-table-ids $RT_PRV
That’s a single NAT in AZ a shared by both private subnets, which is the budget version. For real production add a second NAT in $PUB_B, a second private route table pointing at it, and associate $PRV_B with that one instead.
If you’d rather not hold twelve IDs in shell variables, and you shouldn’t for anything that lives longer than an afternoon, the same thing in Terraform is a few resource blocks, and the Terraform guide covers the tooling. The public terraform-aws-modules/vpc module does all of the above, including one-NAT-per-AZ and the gateway endpoints, in about fifteen lines of input, and it’s what most teams actually use:
module "vpc" {
source = "terraform-aws-modules/vpc/aws"
version = "~> 5.0"
name = "prod"
cidr = "10.0.0.0/16"
azs = ["us-east-1a", "us-east-1b"]
public_subnets = ["10.0.0.0/24", "10.0.1.0/24"]
private_subnets = ["10.0.10.0/24", "10.0.11.0/24"]
enable_nat_gateway = true
one_nat_gateway_per_az = true
enable_dns_hostnames = true
}
What it costs
The VPC, subnets, route tables, internet gateway, security groups, NACLs, peering connections and gateway endpoints are all free to have. The things that aren’t, at us-east-1 list prices as I write this:
| Item | Hourly | Per GB | Roughly per month |
|---|---|---|---|
| NAT gateway | $0.045 | $0.045 processed | $32 + traffic, per AZ |
| Public IPv4 address | $0.005 | $3.60 each | |
| Interface endpoint | $0.01 per AZ | $0.01 | $7 per AZ each |
| Transit Gateway attachment | $0.05 | $0.02 | $36 per VPC + traffic |
| Cross-AZ data transfer | $0.01 each direction | depends entirely on you | |
| Data out to the internet | $0.09 after the first 100 GB | the big one |
Check the current pricing page before budgeting; these move. But the shape doesn’t: NAT per-gigabyte and cross-AZ transfer are the two that grow silently, and gateway endpoints plus AZ-aware placement are the two fixes.
Limits to know about
$ aws service-quotas get-service-quota --service-code vpc --quota-code L-F678F1CE --query 'Quota.{Name:QuotaName,Value:Value}' --output text
VPCs per Region 5.0
$ aws service-quotas get-service-quota --service-code vpc --quota-code L-FE5A380F --query 'Quota.{Name:QuotaName,Value:Value}' --output text
NAT gateways per Availability Zone 5.0
Five VPCs per region by default, five NAT gateways per AZ, 200 subnets per VPC, 60 inbound and 60 outbound rules per security group and 5 security groups per interface. All of these are soft limits that Service Quotas will raise on request, and the VPC one is the first you’ll hit if you do a VPC per environment per region.
Frequently asked questions
Is a VPC free?
The VPC and its plumbing (subnets, route tables, internet gateway, security groups) cost nothing. NAT gateways, public IPv4 addresses, interface endpoints, Transit Gateway and data transfer are what you pay for.
Should I use the default VPC?
For learning and throwaway tests, sure. For anything with a database, no: it has no private subnets, and every instance gets a public IP by default. Build a proper one and leave the default alone, or delete it if your security policy prefers.
What makes a subnet public or private?
Only its route table. A route for 0.0.0.0/0 to an internet gateway makes it public. No such route, and it’s private regardless of anything else.
Security group or NACL, which do I use?
Security groups for almost everything: they’re stateful and can reference other groups. NACLs for the rare case of blocking a range at the subnet edge.
Why can’t my two VPCs be peered?
Almost always overlapping CIDR blocks. Both on 10.0.0.0/16 is the classic. There’s no fix short of renumbering one of them, which is why the CIDR is the decision to get right on day one.
Can I change a VPC’s CIDR later?
You can add secondary CIDR blocks. You cannot change or shrink the primary one. Design for that.
Do I need a NAT gateway in every AZ?
For production, yes. One NAT is a single point of failure for every private subnet routing through it, and cross-AZ traffic to reach it is charged. For dev environments, one is fine and saves $32 a month per AZ you skip.


