Somewhere around the time one server stops being enough, or one server going down stops being an option, you end up needing a load balancer. Most people reach for a managed one first, an AWS ALB, a cloud provider’s LB, and that’s often the right call. But nginx does the job too, it’s what a lot of managed load balancers are quietly running under the hood anyway, and understanding it teaches you things a managed dashboard hides from you: what a health check actually does when a backend dies mid-request, why round robin isn’t always the fair choice it sounds like, why WebSockets break the first time you proxy them and what fixes it.
This post builds a real one. Three backend app servers, one nginx load balancer in front, and every example below is captured from that setup actually running: killing a backend and watching traffic reroute, timing round robin against least_conn against a slow server, capturing the exact headers nginx does or doesn’t forward for a WebSocket upgrade. Nothing here is copied from a diagram.
What you’re actually building
A load balancer sits in front of a pool of identical (or near-identical) backend servers and decides, per request, which one handles it. The client only ever talks to the load balancer’s address. Behind it, backends can come and go, get replaced, get patched, without the client noticing anything except maybe a slightly slower request if the timing is unlucky.
For this post, three backends is enough to see every algorithm behave differently. In a real deployment they’d be separate servers or containers; here they’re three tiny Python HTTP servers on different ports, which is honestly close to how you’d smoke-test this on a single box before rolling it out anywhere:
# three identical backends, ports 8081-8083, each just says who it is
$ curl -s http://127.0.0.1:8081/
Hello from app-8081
$ curl -s http://127.0.0.1:8082/
Hello from app-8082
$ curl -s http://127.0.0.1:8083/
Hello from app-8083
And the nginx config that ties them together, the shape every example in this post builds on:
upstream backend_pool {
server 127.0.0.1:8081;
server 127.0.0.1:8082;
server 127.0.0.1:8083;
}
server {
listen 8080;
server_name _;
location / {
proxy_pass http://backend_pool;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
}
}
upstream names the pool. proxy_pass points a location at it. Reload nginx and you have a load balancer. Everything from here is variations on this.
Installing nginx
Package manager, same as anything else:
# Ubuntu / Debian
$ sudo apt update && sudo apt install -y nginx
# RHEL / Rocky / Alma
$ sudo dnf install -y nginx
$ sudo systemctl enable --now nginx
Config lives at /etc/nginx/nginx.conf, which pulls in everything under /etc/nginx/conf.d/*.conf and, on Debian-based systems, /etc/nginx/sites-enabled/* as well. The load balancer config above goes in a site file (Ubuntu) or a conf.d file (RHEL), same content either way. sudo nginx -t checks syntax before you reload; get into the habit of running it before every reload, it catches typos before they take your site down instead of after.
Round robin, the default
You don’t ask for round robin, it’s what you get the moment you list more than one server with no other directive. Each new connection goes to the next server in the list:
$ for i in {1..6}; do curl -s http://127.0.0.1:8080/; done
Hello from app-8081
Hello from app-8082
Hello from app-8083
Hello from app-8081
Hello from app-8082
Hello from app-8083
Clean rotation, exactly as advertised. One thing worth knowing before you trust that in production: nginx’s round-robin counter is per worker process, not shared across all of them unless you’re using a paid feature. With worker_processes auto spawning one worker per CPU core, and enough concurrent traffic hitting different workers, the sequence you see from any single client won’t be this tidy, because different workers are independently doing their own round robin. It still balances out over volume, it just won’t look like a clean 1-2-3 pattern under load the way it does in a single-threaded curl loop like this one.
Weighted round robin
Add a weight when your backends aren’t equal, a bigger instance that should take more traffic, or a new one you’re easing into rotation:
upstream backend_pool {
server 127.0.0.1:8081 weight=3;
server 127.0.0.1:8082 weight=1;
server 127.0.0.1:8083 weight=1;
}
$ for i in {1..10}; do curl -s http://127.0.0.1:8080/; done
Hello from app-8081
Hello from app-8082
Hello from app-8081
Hello from app-8083
Hello from app-8081
Hello from app-8081
Hello from app-8082
Hello from app-8081
Hello from app-8083
Hello from app-8081
app-8081 took 6 of 10 requests against 2 each for the others, roughly the 3:1:1 the weights ask for. Not a perfect mathematical split in a run of 10, weighted round robin approximates the ratio over a longer run rather than guaranteeing it request by request, but the bias toward 8081 is obvious and consistent.
Least connections: the one that actually reacts to load
Round robin doesn’t know or care if a backend is struggling. It just keeps handing it every third request regardless. least_conn tracks how many connections are currently open to each backend and sends new requests to whichever has the fewest, which matters a lot the moment your backends aren’t uniformly fast.
To show the difference honestly, I made one backend slow, a 2-second delay before it responds, simulating a server under load or doing something heavier than the others, then fired 15 requests at 3 concurrent at a time against each algorithm:
# round robin, 15 requests, concurrency 3, app-8082 takes 2s to respond
$ time (seq 1 15 | xargs -I{} -P 3 curl -s http://127.0.0.1:8080/ > out-{}.txt; wait)
5 Hello from app-8081
5 Hello from app-8082 (slow)
5 Hello from app-8083
real 0m4.116s
# same 15 requests, least_conn enabled
upstream backend_pool {
least_conn;
server 127.0.0.1:8081;
server 127.0.0.1:8082;
server 127.0.0.1:8083;
}
$ time (seq 1 15 | xargs -I{} -P 3 curl -s http://127.0.0.1:8080/ > out-{}.txt; wait)
7 Hello from app-8081
1 Hello from app-8082 (slow)
7 Hello from app-8083
real 0m2.053s
Round robin sent a fair fifth of the traffic, 5 requests, into the slow backend no matter what, because fairness by count is all it knows how to do. Total time: just over 4 seconds, dragged out by however many requests landed there and had to wait in that queue. least_conn noticed the slow backend was still busy and routed almost everything else around it, 1 request instead of 5, and the whole batch finished in about half the time. This is the actual argument for least_conn over round robin: it’s not about even distribution, it’s about not punishing your fast servers by force-feeding a struggling one a fixed share of traffic regardless of whether it can handle it.
IP hash: sticky sessions without a session store
Round robin and least_conn both assume it doesn’t matter which backend handles a given client’s next request. That’s true for stateless APIs. It’s not true the moment a backend holds something in memory tied to that client, a shopping cart, a login session, a WebSocket connection, and you haven’t set up shared session storage (Redis, a database, whatever) to make every backend interchangeable.
ip_hash is the low-effort fix: hash the client’s IP, use that to consistently pick the same backend for that client every time:
upstream backend_pool {
ip_hash;
server 127.0.0.1:8081;
server 127.0.0.1:8082;
server 127.0.0.1:8083;
}
$ for i in {1..6}; do curl -s http://127.0.0.1:8080/; done
Hello from app-8083
Hello from app-8083
Hello from app-8083
Hello from app-8083
Hello from app-8083
Hello from app-8083
Same client, same backend, every single time, because it’s all coming from 127.0.0.1 here. In production with real client IPs, different visitors land on different backends the same way round robin would spread them, but any individual visitor keeps landing on whichever backend they got the first time, for as long as that backend stays up.
The catch, and it’s a real one: if that backend goes down, ip_hash’s whole pool of clients who were pinned to it gets redistributed to the survivors, which means their session state (if it lived only in that server’s memory) is gone. ip_hash gives you stickiness, not durability. It also falls apart behind anything that changes the apparent client IP between requests, a proxy or CDN that doesn’t preserve it, mobile carriers that rotate IPs mid-session. For real session persistence at any scale, shared session storage beats IP-based stickiness; ip_hash is the quick version for when that’s overkill.
Health checks: what happens when a backend actually dies
Everything above assumed all backends stay up. They don’t. Here’s the part that matters more than which algorithm you pick: what nginx does the moment one of them isn’t there anymore.
upstream backend_pool {
server 127.0.0.1:8081 max_fails=2 fail_timeout=10s;
server 127.0.0.1:8082 max_fails=2 fail_timeout=10s;
server 127.0.0.1:8083 max_fails=2 fail_timeout=10s;
}
server {
listen 8080;
location / {
proxy_pass http://backend_pool;
proxy_next_upstream error timeout http_502 http_503 http_504;
proxy_connect_timeout 1s;
}
}
max_fails=2 fail_timeout=10s means after 2 failed attempts, nginx marks that backend down for 10 seconds and stops sending it traffic, no health-check endpoint required, this is passive, it learns from real request failures. Baseline with everything healthy:
$ for i in {1..6}; do curl -s http://127.0.0.1:8080/; done
Hello from app-8081
Hello from app-8082
Hello from app-8083
Hello from app-8081
Hello from app-8082
Hello from app-8083
Then I killed app-8082 outright, the way a crash or an OOM kill would, mid-traffic, no graceful shutdown:
$ kill -9 $(pgrep -f app-8082.py)
$ for i in {1..8}; do curl -s http://127.0.0.1:8080/; done
Hello from app-8081
Hello from app-8083
Hello from app-8083
Hello from app-8081
Hello from app-8083
Hello from app-8081
Hello from app-8083
Hello from app-8081
Not one client-facing error. Every single request in that batch came back 200. And the error log confirms nginx genuinely hit the dead backend twice before writing it off:
$ sudo tail -2 /var/log/nginx/error.log
2026/02/25 07:41:14 [error] 7787#7787: *15 connect() failed (111: Connection refused)
while connecting to upstream, upstream: "http://127.0.0.1:8082/", host: "127.0.0.1:8080"
2026/02/25 07:41:14 [error] 7787#7787: *22 connect() failed (111: Connection refused)
while connecting to upstream, upstream: "http://127.0.0.1:8082/", host: "127.0.0.1:8080"
Two connection refusals logged, matching max_fails=2 exactly, then it stopped trying and every request after that went straight to the two survivors. This is the entire point of running a load balancer instead of pointing DNS at one server: a backend dying becomes a line in a log file instead of an outage. After 10 seconds (fail_timeout), nginx quietly retries app-8082 on the next request, and if it’s back, it’s back in rotation with no manual step.
nginx’s open-source build only does this kind of passive checking, reacting to real traffic failures. Active health checks, hitting a /health endpoint on a timer whether or not real traffic is flowing, are an NGINX Plus feature. For most setups passive checking is genuinely enough; the gap only matters if you need to catch a dying backend before it fails a real user’s request.
SSL termination: let nginx handle the certificate
Terminating TLS at the load balancer, decrypting there and talking plain HTTP to the backends, is normal and usually the right call. One certificate to manage instead of one per backend, and your backend servers get to stay simple.
server {
listen 8080;
return 301 https://$host:8443$request_uri;
}
server {
listen 8443 ssl;
server_name lb.example.internal;
ssl_certificate /etc/nginx/tls/lb.crt;
ssl_certificate_key /etc/nginx/tls/lb.key;
ssl_protocols TLSv1.2 TLSv1.3;
location / {
proxy_pass http://backend_pool;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-Proto https;
}
}
Plain HTTP redirects cleanly:
$ curl -sI http://127.0.0.1:8080/
HTTP/1.1 301 Moved Permanently
Location: https://127.0.0.1:8443/
And the HTTPS side does a real TLS 1.3 handshake and still hits the backend pool underneath:
$ curl -skv https://127.0.0.1:8443/ 2>&1 | grep -E "subject:|SSL connection"
* subject: CN=lb.example.internal
* SSL connection using TLSv1.3 / TLS_AES_256_GCM_SHA384 / X25519 / RSASSA-PSS
$ curl -sk https://127.0.0.1:8443/
Hello from app-8081
That X-Forwarded-Proto https line matters more than it looks. Without it, your backend application has no way to know the original request arrived over HTTPS, it only sees plain HTTP from nginx, which breaks anything that checks the scheme, redirect logic, secure cookie flags, “insecure content” warnings in a framework that thinks it’s serving HTTP. Pass it through and let your app trust that header from the load balancer.
For real certificates instead of the self-signed one used here, Let’s Encrypt with Certbot automates issuing and renewal against this exact kind of nginx config.
WebSockets: the header that has to be forwarded explicitly
This is the one that catches people who’ve proxied plain HTTP through nginx a hundred times and assume WebSockets work the same way. They don’t, and I proved it rather than just asserting it: set up a minimal WebSocket backend, proxied it through nginx two different ways, and captured exactly what headers arrived on the other side using a raw listener instead of trusting client-side success or failure.
First, the naive config, just proxy_pass and nothing WebSocket-specific:
location /ws {
proxy_pass http://ws_backend;
proxy_http_version 1.1;
}
Here’s what the backend actually received when a client sent an Upgrade: websocket request through that:
$ what the backend actually received:
Connection: close
No Upgrade header at all, and Connection: close instead of upgrade. nginx silently stripped the upgrade request into a normal HTTP connection on its way through. This is why WebSockets “don’t work” behind nginx for a lot of people on the first try, not an error, just a quiet protocol downgrade that looks like the backend is broken when it’s actually never seeing the upgrade attempt.
Add the two lines nginx’s own docs call out for exactly this:
location /ws {
proxy_pass http://ws_backend;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
}
$ what the backend actually received:
Upgrade: websocket
Connection: upgrade
Same request, same client, and now the backend gets exactly what it needs to complete the handshake. $http_upgrade forwards whatever the client actually sent rather than hardcoding the value, which matters because a normal HTTP request has no Upgrade header at all, and this same location block needs to keep serving those too without breaking them.
Which backend actually served that request
When something goes wrong three servers deep, “which one” is the first question, and the default nginx access log doesn’t answer it. A custom log format does:
log_format upstream_log '$remote_addr - [$time_local] "$request" '
'$status upstream=$upstream_addr '
'rt=$request_time urt=$upstream_response_time';
server {
listen 8080;
access_log /var/log/nginx/lb-demo.log upstream_log;
location / { proxy_pass http://backend_pool; }
}
$ cat /var/log/nginx/lb-demo.log
127.0.0.1 - [25/Feb/2026:07:43:11 +0500] "GET / HTTP/1.1" 200 upstream=127.0.0.1:8081 rt=0.000 urt=0.000
127.0.0.1 - [25/Feb/2026:07:43:11 +0500] "GET / HTTP/1.1" 200 upstream=127.0.0.1:8082 rt=0.001 urt=0.001
127.0.0.1 - [25/Feb/2026:07:43:11 +0500] "GET / HTTP/1.1" 200 upstream=127.0.0.1:8083 rt=0.001 urt=0.001
127.0.0.1 - [25/Feb/2026:07:43:11 +0500] "GET / HTTP/1.1" 200 upstream=127.0.0.1:8081 rt=0.000 urt=0.000
upstream_addr names the exact backend, rt is total request time as the client experienced it, urt is how long the backend itself took. When one backend is consistently slower, this is where you’d see it: urt climbing on one address while the others stay flat, without needing to log into each server separately to check.
Frequently asked questions
Do I need nginx Plus for any of this?
No. Everything in this post, round robin, weighted round robin, least_conn, ip_hash, passive health checks, SSL termination, WebSocket proxying, custom logging, is in open-source nginx. NGINX Plus adds active health checks (probing a health endpoint on a timer regardless of real traffic), session persistence by cookie instead of IP, and a live status dashboard. Worth it at real scale, not required to get a working load balancer.
Which algorithm should I actually use?
Round robin if your backends are identical and stateless, which covers more cases than people expect. least_conn the moment request times vary meaningfully, slow endpoints mixed with fast ones, background jobs, anything where “equal number of requests” doesn’t mean “equal load.” ip_hash only if you have session state stuck on individual backends and can’t move to shared storage yet, treat it as a stopgap, not the long-term answer.
Can nginx load balance TCP traffic that isn’t HTTP?
Yes, the stream module handles raw TCP and UDP, load balancing databases, mail servers, anything that isn’t speaking HTTP. Different top-level config block (stream { } instead of http { }) and its own upstream/server syntax, but the same round robin, least_conn and weighting concepts carry over.
Is nginx as a load balancer a single point of failure?
Yes, exactly as much as any single load balancer is. The usual fix is two nginx instances with keepalived managing a floating virtual IP between them, or putting a cloud load balancer in front of multiple nginx instances, load balancers behind a load balancer. Worth planning for once this is handling anything that actually matters; not something to bolt on as an afterthought.
How is this different from an AWS Application Load Balancer?
Functionally similar goals, different tradeoffs. An ALB is managed, scales automatically, and integrates with AWS health checks and auto scaling groups without you touching config files. nginx is self-managed, meaning full control over exactly what’s in this post, at the cost of you being the one who patches it, scales it and keeps it running. Plenty of real deployments run nginx as the load balancer for an EC2 fleet specifically to get that control; plenty of others take the ALB and never think about it again. Both are legitimate choices depending on how much you want to own.

