“The network is down” is never actually true. Something specific is broken: a cable, an ARP entry, a route, a firewall rule, a DNS answer, a certificate. The skill isn’t knowing every networking command that exists, it’s knowing which layer to check first so you’re not running traceroute against a problem that’s sitting in your DNS resolver the whole time.
This is the method I actually use: start at the bottom of the stack, confirm each layer works, and stop the moment you find the one that doesn’t. Everything below is real output from Ubuntu 24.04 and Red Hat Enterprise Linux 9.7, including a live demonstration of the one distinction that trips up almost everyone: the difference between a connection that’s refused and one that’s silently dropped by a firewall, because they look identical until you know what to check.
The method, in one paragraph
Work bottom-up: link, then address and routing, then the actual connection, then the application talking over it. Each layer depends entirely on the one below it, so there’s no point debugging an HTTPS certificate error if the TCP connection underneath it never completed, and no point staring at a routing table if the interface itself is down. Check the cheapest, fastest thing first. Most of the time that’s layer 1 or 2, and most guides skip straight past them to the interesting stuff.
Layer 1 and 2: is the link actually there
Before anything involving an IP address, confirm the interface is up and has a carrier signal, and that the machine can actually reach its neighbors on the local segment:
$ ip -br link show eth0
eth0 UP 00:15:5d:d3:55:84 <BROADCAST,MULTICAST,UP,LOWER_UP>
$ ip -br addr show eth0
eth0 UP 172.21.3.249/20 fe80::215:5dff:fed3:5584/64
UP twice is what you want: administratively up (someone or something brought the interface up) and LOWER_UP, which means there’s an actual physical or virtual link detected, a cable plugged in or a virtual NIC properly attached. An interface that’s UP but missing LOWER_UP has an IP address and thinks it’s ready, but there’s nothing on the other end, which is a completely different problem from anything above this layer and wastes real time if you don’t catch it here first.
Once the link is confirmed, the neighbor table tells you whether this machine can actually talk to the next hop, which for most LANs is the gateway:
$ ip route | grep default
default via 172.21.0.1 dev eth0 proto kernel
$ ip neigh show 172.21.0.1
172.21.0.1 dev eth0 lladdr 00:15:5d:51:37:68 REACHABLE
REACHABLE means ARP resolved and the kernel has recently confirmed it. STALE means it resolved before but hasn’t been re-confirmed, which is normal and not a problem by itself. FAILED is the one that matters: ARP was tried and nothing answered. If your gateway shows FAILED, nothing above layer 2 has any chance of working, and there’s no point checking DNS or firewall rules until that’s fixed. It’s almost always a cabling, VLAN, or switch port issue at that point, not a Linux configuration problem.
Layer 3: addressing and routing, briefly
If the link is up and the gateway resolves, the next question is whether the kernel picks the right route for where you’re trying to go, and whether the actual path beyond your gateway is intact. This is deep enough territory that it gets its own coverage: the CIDR and subnetting math behind why a given address does or doesn’t belong to a given network, and using ping, traceroute and mtr to find exactly which hop a path breaks at, including the extremely common false alarm where a hop shows nothing but asterisks and the path is actually fine. If layers 1 and 2 check out, that’s genuinely the next stop, not this article.
The thirty-second version worth repeating here: ip route get <destination> tells you exactly which route the kernel will use and which local address it’ll send from, without sending a single packet. It answers “is this even going where I think it’s going” before you reach for anything heavier.
Layer 4: the connection itself, and the distinction almost nobody explains
This is where most of the actual debugging time goes, and it’s the layer this article exists to cover properly. First, ss shows you the real state of every socket on the box, which beats guessing every time:
$ ss -tan
State Recv-Q Send-Q Local Address:Port Peer Address:Port
LISTEN 0 4096 127.0.0.53%lo:53 0.0.0.0:*
LISTEN 0 128 0.0.0.0:22 0.0.0.0:*
ESTAB 0 517 172.21.3.249:41052 185.151.30.186:443
TIME-WAIT 0 0 10.42.0.1:57892 10.42.0.57:8080
LISTEN is a process waiting for connections on that port, nothing wrong there. ESTAB is a live, working connection; the Send-Q column above, 517 bytes queued on a real outbound HTTPS connection I made while capturing this, shows data actually in flight. TIME-WAIT is completely normal cleanup after a connection closes, the kernel holding the socket briefly in case a stray packet still arrives; it’s not a leak and not a problem unless there are thousands of them piling up. What you’re actually hunting for when something’s broken is a port you expect to be in LISTEN that simply isn’t there at all, which tells you the service itself never started, not that the network is at fault. netstat shows much the same information on older systems that don’t have ss, and once a service is confirmed up locally, nmap is the tool for checking whether that same port is actually reachable from outside the box, which is a different question with a different answer.
Here’s the distinction that actually matters, demonstrated live rather than described. Three attempts to connect to the same port, three different real outcomes:
# 1. Nothing listening on this port at all
$ nc -zv 127.0.0.1 8081
nc: connect to 127.0.0.1 port 8081 (tcp) failed: Connection refused
# 2. Start a real listener, then connect
$ nc -l -p 8081 &
$ nc -zv 127.0.0.1 8081
Connection to 127.0.0.1 8081 port [tcp/*] succeeded!
# 3. Same listener running, but a firewall rule now drops the port
$ sudo nft add table inet demo
$ sudo nft add chain inet demo input { type filter hook input priority 0 \; }
$ sudo nft add rule inet demo input tcp dport 8081 drop
$ time nc -zv 127.0.0.1 8081
(hangs, no response at all)
real 0m3.003s
Refused is instant, and it means one specific thing: nothing is listening on that port, the kernel itself said no immediately with a TCP RST. Filtered looks completely different: no response at all, and the connection just hangs until whatever client you’re using gives up on its own timeout, in this case exactly the 3 seconds I capped it at. A firewall rule that drops a packet doesn’t send anything back, by design, so the connecting side has no way to distinguish “still trying” from “will never answer” except by waiting it out. If a connection fails instantly, look at whether the service is running. If it hangs and eventually times out, look at the firewall, a security group, or a network ACL sitting between you and the destination, because that’s what filtering looks like from the outside every time.
On Ubuntu, that live demo used nft directly; ufw is installed by default but inactive out of the box, so a fresh Ubuntu box has no filtering at all until someone turns it on. RHEL’s equivalent is firewalld, and the check and the rule look like this:
$ sudo firewall-cmd --state
running
$ sudo firewall-cmd --list-all
public (active)
target: default
ports: 22/tcp
$ sudo firewall-cmd --add-rich-rule='rule family=ipv4 port port=8081 protocol=tcp drop'
Same effect as the nftables rule above: a connection to a firewalld-blocked port produces the identical silent hang, not a refusal. The underlying mechanism differs, firewalld manages its own nftables tables rather than you touching them directly, but the symptom on the wire is exactly the same, which is the actual point: refused-versus-timeout tells you where to look regardless of which firewall tool is involved.
Layer 7: DNS and the application itself
If the connection layer is clean, “it doesn’t work” is usually either DNS resolving to the wrong thing, or the application on the other end doing something unexpected once a connection actually opens. Two different lookup tools check two different points in that chain:
$ dig +noall +answer +stats linuxpathfinder.com
linuxpathfinder.com. 3600 IN A 185.151.30.186
;; Query time: 335 msec
;; SERVER: 10.255.255.254#53(10.255.255.254)
$ getent hosts linuxpathfinder.com
185.151.30.186 linuxpathfinder.com
dig talks straight to a DNS server and shows you exactly what came back, bypassing every layer of local caching. getent hosts asks the C library the same question your actual applications ask, which goes through NSS and whatever order /etc/nsswitch.conf defines:
$ grep ^hosts /etc/nsswitch.conf
hosts: files dns
files before dns means /etc/hosts is checked first, every time, for every program on the box. This is the answer to the specific and common frustration of updating DNS, confirming the new record with dig, and having the application still connect to the old address: a stale entry in /etc/hosts wins before DNS is ever consulted, and dig alone will never reveal that, because dig doesn’t read /etc/hosts at all. Reach for getent, not dig, when the question is “what does this specific program actually see.”
On Ubuntu, resolution normally runs through systemd-resolved’s local stub at 127.0.0.53, with /etc/resolv.conf symlinked to a file that resolved lives, and resolvectl status shows the real upstream servers behind that stub. RHEL more commonly has NetworkManager write /etc/resolv.conf directly per active connection. Worth knowing which model a given box uses before troubleshooting DNS on it, since “edit resolv.conf directly” works fine on one and gets silently overwritten on the other the next time the network reconnects.
Past DNS, if a TCP connection opens fine but the actual response never comes or looks wrong, curl -v shows you precisely where in the request lifecycle things stand, which is invaluable for anything involving TLS:
$ curl -v https://linuxpathfinder.com/
* Host linuxpathfinder.com:443 was resolved.
* IPv4: 185.151.30.186
* Trying 185.151.30.186:443...
* Connected to linuxpathfinder.com (185.151.30.186) port 443
* TLSv1.3 (OUT), TLS handshake, Client hello (1):
* TLSv1.3 (IN), TLS handshake, Server hello (2):
* TLSv1.3 (IN), TLS handshake, Certificate (11):
Read this top to bottom and stop at the first line that doesn’t happen. Resolution failed: DNS. Never gets to “Connected”: layer 4, back to the refused-versus-timeout check above. Connects but the TLS handshake stalls or errors: certificate or TLS configuration, not a network problem at all by that point, which is exactly why chasing it with traceroute would have been the wrong move.
Putting the bottom-up order together
ip -br link showandip -br addr show: interface up, has a carrier, has the address you expect. Costs nothing, rules out an entire category instantly.ip neigh show <gateway>: local segment reachable.FAILEDhere means stop, this is a physical or switch-level problem, not a Linux one.ip route get <destination>, then ping/traceroute/mtr if the route looks wrong or the path needs mapping. Covered in depth in the linked articles above.ss -tanfor what’s actually listening, then test the specific port. Instant refusal points at the service; a hang that times out points at a firewall, security group, or ACL somewhere on the path.getent hostsbeforedig, so you’re checking what the application actually sees rather than what the DNS server says in isolation, thencurl -vif the connection opens but the response doesn’t look right.
Five steps, cheapest first, and the overwhelming majority of “the network is down” reports resolve inside the first two before anything resembling a deep investigation is needed.
Frequently asked questions
What’s the real difference between connection refused and a connection timing out?
Refused is immediate and specific: nothing is listening on that port, and the kernel says so right away with a TCP reset. A timeout means something between you and the destination silently dropped the packet, usually a firewall rule, a security group, or a network ACL, and your client is just waiting out its own patience before giving up.
Why does ip neigh show FAILED for one address but STALE for others?
STALE means ARP resolved successfully at some point and the kernel just hasn’t reconfirmed it recently, which is completely normal. FAILED means an ARP request went out and nothing on the local segment answered, which is a real problem worth chasing at layer 2 before anything else.
Should I check DNS before or after the network connection itself?
After. DNS resolution requires a working network connection to the DNS server in the first place, so if layer 4 (the actual connection) is broken, DNS troubleshooting is premature, even though DNS problems are often what people assume first.
Why does /etc/hosts sometimes override DNS changes I’ve already confirmed?
Because /etc/nsswitch.conf almost always lists files before dns, so /etc/hosts is checked first for every application on the box. A stale entry there wins regardless of what a fresh dig query shows, since dig talks to the DNS server directly and never reads /etc/hosts at all.
Is ufw or firewalld running by default?
Ubuntu ships ufw installed but inactive; a fresh box has no port filtering until someone explicitly enables it. RHEL and its derivatives ship firewalld active by default, usually in the public zone, which blocks everything except what’s explicitly allowed. Worth checking which state you’re actually in before assuming either way.
Do I need to memorize the full seven-layer OSI model for this?
No. In practice, Linux troubleshooting collapses cleanly into four practical checkpoints: link, address and routing, the connection itself, and the application on top. The extra OSI layers matter for exam questions more than for a terminal in front of you.

