Take “development”, take “operations”, jam them together. That’s the word. And honestly the word tells you most of the story, because it describes two teams who worked at the same company and spent about fifteen years quietly making each other’s lives harder.
This guide covers what DevOps actually is, where it came from, how the lifecycle works, the practices that make it real, and how teams measure whether it’s working. Then, because definitions only go so far, we’ll run a small example on a real Linux box that blocks a bad change before it ships, and look at how teams measure whether delivery is actually improving.
What Is DevOps?
DevOps is a way of building and running software where the people who write the code and the people who keep it running work as one team, share responsibility for production, and automate the path between a commit and a live release.
That’s the short version. The longer version is that it’s three things at once:
- A culture. Shared ownership. If you wrote it, you care whether it stays up.
- A set of practices. Continuous integration, continuous delivery, infrastructure as code, monitoring, small and frequent changes.
- A toolchain. Git, a CI server, containers, config management, observability tools. Useful, but the least important of the three.
Culture first, tools second. People get that order backwards constantly. I’ve seen shops running Jenkins, Ansible and Kubernetes who still had a ticket queue between dev and ops and a two-week wait to get anything deployed. The tools didn’t fix that. They were never going to.
A few things DevOps isn’t:
- A standard. There’s no spec and no governing body. It’s a collection of practices that keep proving they work.
- A product. Nobody can sell it to you, though plenty of vendors will try.
- Just a job title. It describes how an organisation works. There is a DevOps engineer role now, and we cover it in detail in what a DevOps engineer actually does, but one person can’t “do DevOps” for a company that hasn’t changed how it works.
The Problem DevOps Was Built to Fix
Think about how the incentives used to line up. Developers were judged on shipping features. Operations was judged on nothing breaking. So every release was a small fight, and both sides were behaving completely rationally.
Developers threw a build over the wall. Ops caught it, often at night, often with a runbook that was already out of date. When something broke, the first hour went on working out whose fault it was. Releases got bigger and rarer because they were painful, and bigger releases broke more things. A loop, and not a good one.
DevOps attacks the loop from both ends. Make changes small so they’re cheap to ship and easy to undo. Automate the steps that people get wrong. And put the same people on the hook for building it and running it, so nobody is optimising for half the problem.
A Short History of DevOps
DevOps grew out of Agile, and out of a problem Agile created.
Agile talked teams into shipping in small increments instead of disappearing for nine months. That worked. But it moved the bottleneck rather than removing it, because development could now produce changes faster than operations could safely absorb them. Same wall. More traffic hitting it.
- 2007 to 2008. Patrick Debois, a Belgian consultant working on a data centre migration, kept bouncing between the dev side and the ops side and got fed up with the gap. At Agile 2008 in Toronto he connected with Andrew Shafer, who had proposed a session on “Agile Infrastructure”.
- June 2009. John Allspaw and Paul Hammond gave their Velocity talk, “10+ Deploys per Day: Dev and Ops Cooperation at Flickr”. Ten deploys a day sounds ordinary now. Back then it landed on a room full of engineers who’d been told frequent deployment was reckless.
- October 2009. Debois ran the first DevOpsDays in Ghent. The hashtag got shortened to #DevOps, and the movement had a name.
- 2013 onwards. The Phoenix Project (2013), The DevOps Handbook (2016) and Accelerate (2018) turned the ideas into something managers could read and researchers could measure.
Everything since has been the industry catching up. Plus a staggering amount of tooling built to sell into it.
How DevOps Works: The Lifecycle
You’ll usually see the DevOps lifecycle drawn as an infinity loop. It’s a bit of a cliché, but the shape makes a real point: there’s no finish line. Every release feeds information back into the next one.
- Plan. Decide what to build, in small pieces. Tickets, user stories, whatever your team uses.
- Code. Everything lives in version control. Application code, yes, but also pipeline definitions, infrastructure and config.
- Build. Turn the code into an artifact. A binary, a package, a container image. Build it once and promote the same artifact everywhere.
- Test. Automated tests run on every change. Unit tests first because they’re fast, then integration tests.
- Release. The artifact is versioned and ready to go out.
- Deploy. Push it to an environment, ideally the same way every time, ideally with no manual steps.
- Operate. Run it. Scale it, patch it, keep it healthy.
- Monitor. Metrics, logs and traces tell you how it’s behaving and what users actually experience. That feeds straight back into Plan.
The left half of the loop is mostly development, the right half mostly operations. DevOps is the claim that they’re one process, not two departments.
Core DevOps Practices
Continuous Integration (CI)
Everyone merges small changes into the main branch often, at least daily, and every merge triggers an automated build and test run. The point is to find breakage within minutes of causing it. Long-lived branches are what CI exists to prevent. Merging three weeks of divergent work isn’t integration, it’s archaeology.
Continuous Delivery vs Continuous Deployment
These two get used interchangeably and shouldn’t be. Continuous delivery means every build that passes the pipeline could go to production with a button press. Continuous deployment means it does go, automatically, with no human in the loop. Most teams want the first and think they want the second. Continuous delivery is the useful baseline. Continuous deployment is a choice you make once your tests and monitoring are good enough to trust.
Infrastructure as Code (IaC)
Servers, networks, load balancers and permissions get defined in files you can review, version and roll back, instead of being clicked together in a console. Terraform is the common choice for provisioning. If you’re new to it, our guide on installing and using Terraform walks through the full workflow. If your infrastructure only exists in someone’s memory, you don’t really have infrastructure.
Configuration Management
Ansible, Puppet and Chef keep servers in a known state. The thing you’re fighting is drift: two servers that should be identical and quietly aren’t. Increasingly, teams skip patching altogether and rebuild from an image instead. Either way, structure matters, which is why a sane Ansible directory layout pays off early.
Containers and Orchestration
Containers package an app with its dependencies so “works on my machine” stops being a defence. Docker builds and runs them. Kubernetes schedules them across a cluster. Neither is required for DevOps, but they make the build-once, run-anywhere part much easier. Start with Docker images before worrying about clusters.
Monitoring and Observability
Knowing something is wrong before a customer tells you. Sounds obvious. Plenty of teams still find out from social media. Metrics show trends, logs explain events, traces follow one request across services. The tooling is the easy bit. Deciding what deserves to wake someone at 3am is the hard bit, and most alerting is far too noisy. Our guide to Linux system monitoring covers the host-level basics.
DevSecOps
Security checks run in the pipeline: dependency scanning, secret detection, image scanning, policy checks. Security that arrives as a review two days before launch isn’t security, it’s a formality. Move it left, to where fixing things is cheap. The server side still matters too, so a server hardening checklist belongs in the same conversation.
DevOps in Practice: A Simple Example
All of the above can sound abstract, so here’s the smallest version of the idea. A tiny Python function, one test, and a deploy script that only copies the file to “live” if the test passes.
#!/usr/bin/env bash
# deploy.sh: run the tests, only copy the file to "live" if they pass
set -uo pipefail
if python3 -m unittest -q test_greet.py; then
cp greet.py live-greet.py
echo "deployed: $(date +%T)"
else
echo "tests failed, nothing deployed"
fi
Run it when the code and the test agree, and it deploys:
./deploy.sh
deployed: 11:31:44
Now change the greeting text without updating the test, and run it again:
./deploy.sh
tests failed, nothing deployed
Nothing gets copied anywhere. Fix the test to match the change, and the next run deploys again. That’s the whole idea behind continuous integration, just without the extra moving parts: a gate that only lets tested code through.
What You’d Add For Real
A real pipeline builds more around this same gate, not a different idea. A Git push triggers it instead of someone running the script, a CI service such as GitHub Actions or GitLab CI runs it, the deploy target is a real server or container instead of a local copy, and a rollback step undoes the deploy if a health check fails afterwards. The gate itself, test first, deploy second, stays the same underneath all of that. Our Git branching guide, self-hosted GitHub Actions runner and zero-downtime deployment guides go further into each piece.
CALMS: A Way to Check Whether You’re Doing DevOps
CALMS is a handy checklist. Damon Edwards and John Willis started with CAMS, and Jez Humble added the L later.
- Culture. Shared ownership, blameless post-incident reviews, and people who are allowed to say “I broke it”.
- Automation. Builds, tests, deploys and infrastructure, all scripted. Anything you do by hand twice, and definitely anything you do at 2am, should be a script.
- Lean. Small batches, short queues, and less work sitting half-finished.
- Measurement. Decisions based on data about delivery and reliability, not gut feeling.
- Sharing. Knowledge, dashboards and on-call duties spread across the team, not stuck in one person’s head.
If a team scores well on automation and badly on everything else, they’ve bought tools. Not the same thing.
How to Measure DevOps: The DORA Metrics
The DORA research program (DevOps Research and Assessment, now part of Google Cloud) has studied thousands of teams and boiled software delivery performance down to four key metrics:
- Deployment frequency. How often you ship to production.
- Lead time for changes. How long a commit takes to reach production.
- Change failure rate. What share of deployments cause a failure that needs fixing.
- Failed deployment recovery time. How fast you recover when a deployment goes wrong. Older reports called this time to restore service.
The first two measure speed. The last two measure stability. The finding that made DORA famous is that the best teams are good at both. Speed and stability aren’t a trade-off, because small, frequent changes are easier to test and easier to undo.
You don’t need a dashboard product to start. Even the small example above could write one line to a log file on every successful deploy, and a line to a different log on every run that tests blocked. Count the first log over a week and you have deployment frequency. Count the second against the total and you have change failure rate. Add a timestamp to each line and you have lead time too. None of that needs new tooling, just the habit of logging it.
DevOps vs Agile vs SRE vs Platform Engineering
These overlap a lot, and the job market blurs them further. Here’s how I’d separate them:
| Approach | Main focus | What it adds |
|---|---|---|
| Agile | How software gets planned and built | Short iterations, customer feedback, adapting to change |
| DevOps | How software gets delivered and run | Shared ownership of production, CI/CD, automation |
| SRE | Reliability, measured | SLOs, error budgets, engineering time spent cutting toil |
| Platform engineering | Developer self-service | An internal platform with paved paths so teams can ship without filing tickets |
Agile and DevOps are complementary: Agile makes changes small, DevOps gets them to production. Google has described SRE as one concrete way to implement DevOps. Platform engineering is what many larger companies do once they realise every team shouldn’t have to build its own pipeline from scratch.
Benefits of DevOps
- Faster delivery. Small changes go out in minutes, not in quarterly release trains.
- Fewer and smaller failures. When a deploy contains one change instead of two hundred, finding the bad one is easy.
- Faster recovery. Automated rollbacks and good monitoring turn outages into blips instead of long incidents.
- Repeatable environments. Infrastructure as code means staging actually looks like production.
- Happier engineers. Fewer late-night release calls, less blame, more time on real work. Not a soft benefit, either. Burned-out teams make more mistakes.
Common DevOps Mistakes
- Renaming the ops team. Calling the same siloed team “DevOps” changes the org chart and nothing else.
- Tooling first. Buying Kubernetes before you have automated tests is putting a race engine in a car with no brakes.
- Skipping tests. A pipeline without meaningful tests just ships bugs faster.
- Noisy alerts. If the pager goes off for things nobody needs to act on, people learn to ignore it. Then they ignore the real one.
- Forgetting the fundamentals. Linux, networking and DNS still decide whether you can debug a failure. Our layer-by-layer network troubleshooting method and journalctl guide are good places to sharpen those.
How to Start With DevOps
You don’t need a transformation programme. Start small and let results do the convincing.
- Put everything in Git. Code, scripts, config, infrastructure. If it isn’t versioned, it isn’t real.
- Automate the build and the tests. One pipeline, running on every push. Fast feedback beats complete feedback.
- Script the deploy. Even a shell script is better than a wiki page of manual steps. You saw one above.
- Add a health check and a rollback. Every deploy should prove it worked, and undo itself if it didn’t.
- Measure. Start logging deploys and failures so you can track the DORA metrics over time.
- Share on-call. Developers who get paged for their own code write different code. Better code, usually.
If you want hands-on practice on a single Linux server, deploying a full-stack Node.js app on Ubuntu is a good next project, and you can automate it once it works by hand.
Frequently Asked Questions
What is DevOps in simple terms?
It’s developers and operations working as one team, sharing responsibility for production, and automating the path from code to a live release, so software ships faster and breaks less.
Is DevOps the same as CI/CD?
No. CI/CD is one of its core practices, and probably the most visible one. DevOps also covers culture, infrastructure as code, monitoring, security and how teams share ownership.
Is DevOps a tool or a methodology?
Neither, strictly. It’s a culture and a set of practices. Tools like Git, Jenkins, Docker and Terraform support it, but installing them doesn’t make a team DevOps.
Do I need Kubernetes to do DevOps?
No. Plenty of teams practise DevOps very well on plain virtual machines. Kubernetes helps once you’re running many services at scale. It’s not a starting requirement.
What’s the difference between DevOps and DevSecOps?
DevSecOps is DevOps with security built into every stage, such as dependency scans, secret detection and policy checks in the pipeline, instead of a security review at the very end.
Is DevOps hard to learn?
The ideas are simple. The breadth is what takes time: Linux, networking, scripting, Git, CI/CD, containers and cloud. Learn them in roughly that order and build small projects as you go. Our post on the DevOps engineer role lays out a realistic roadmap.


