What Are DORA Metrics? A Practical Guide for SRE Teams
Summary
What are DORA metrics? They are five measures of software delivery from the DORA research program: change lead time, deployment frequency, failed deployment recovery time, change fail rate, and deployment rework rate. The first three track throughput and the last two track instability, so you read them in pairs. Instrument them per service from four deployment timestamps, use medians, and treat them as a diagnosis of your delivery system, never as an individual scorecard.
What are DORA metrics? They are five measures of software delivery: change lead time, deployment frequency, failed deployment recovery time, change fail rate, and deployment rework rate. The first three describe throughput, the last two describe instability. Together they tell you how fast changes reach production and how often those changes hurt. Used well, they point at the bottleneck. Used badly, they become a report card that your team learns to game.
Where the numbers come from, and why the vocabulary stuck
A platform lead sits in a quarterly review. The CTO asks one question: are we shipping faster than last year, and are we breaking less? Without a shared vocabulary, the answer is a pile of anecdotes. DORA metrics exist to replace the anecdotes with four or five numbers you can pull from systems you already run.
The name comes from DevOps Research and Assessment, the research program that spent years surveying engineering teams and correlating their delivery practices with organizational outcomes. It is now part of Google Cloud, and the findings are published each year in the DORA research program. The book Accelerate popularized the original four. The framework has been revised since, and that revision matters more than most blog posts admit.
Here is the part that gets lost: the metrics were never meant as a leaderboard. They came out of a statistical finding. Teams that scored well on speed also scored well on stability. Speed and safety were not a trade-off, they moved together. That is a claim about correlation across many teams, not a target for yours.
The five metrics, defined the way an on-call engineer would define them
The current definitions on the official DORA metrics guide split into two groups. Read them with a specific service in mind, because every definition falls apart if you apply it to "the company".
Change lead time. The time from a commit to that commit running in production. Not from the ticket being opened, and not from the pull request being approved. From commit to prod. If your pipeline takes forty minutes and your release train leaves on Thursdays, your lead time is dominated by Thursday.
Deployment frequency. How often you ship to production. You can count deployments over a period or measure the gap between them. A service that deploys eleven times a day and a service that deploys once a month are different animals, and averaging them hides both.
Failed deployment recovery time. How long it takes to recover when a deployment causes a problem that needs intervention. This used to be called mean time to restore, and the rename is deliberate. It only counts incidents caused by your own change, not a cloud provider outage on a Tuesday.
Change fail rate. The share of deployments that need immediate intervention afterwards: a rollback, a hotfix, a forward fix pushed at speed. Ten deployments, two of them reverted, a change fail rate of twenty percent.
Deployment rework rate. The share of deployments that are unplanned, triggered by a production incident rather than by roadmap work. This is the newest addition, and it catches something change fail rate misses: the team that never rolls back but spends half its week shipping emergency patches.

Throughput versus instability: why you never read one metric alone
The first three metrics are throughput. The last two are instability. The grouping is the whole point.
If you only watch throughput, you will celebrate a team that deploys forty times a day and quietly reverts a quarter of them. If you only watch instability, you will reward a team that ships once a month with a spotless record, because nobody changes anything. Every single metric has a cheap way to improve it that makes the system worse.
Deployment frequency goes up when you split a deploy into five empty deploys. Change fail rate goes down when you stop counting hotfixes as failures. Lead time shrinks when you redefine the start of the clock. This is Goodhart's law doing what it always does: when a measure becomes a target, it stops being a good measure. The official guide lists it first among its warnings, and it is the one that bites hardest.
So the working rule is pairs. Read lead time next to change fail rate. Read deployment frequency next to rework rate. A move in one direction without a matching story in the other is a signal to look closer, not a win to announce.
Benchmarks circulate in every vendor deck, and they are a trap. The reported pattern from recent DORA reports is consistent in shape: the strongest cluster of teams deploys on demand, with a lead time under a day, and recovers from a failed deployment in under an hour. The slowest cluster sits in the weeks-to-months range on both.
Those figures are useful for one thing: calibrating your intuition about what is possible. They are poor as targets. A payments service with a mandatory audit gate will not match a marketing site, and it should not try. The official guidance is explicit that you should compare only similar applications or services, and that you should aim for improvement against your own baseline rather than competition between teams.
Start with your own median. Measure it for a month before you set any goal. The first number is almost always embarrassing, and that is fine. A baseline that embarrasses you is a baseline you actually trust.
How to instrument them without building a data platform
Most teams overbuild this. You do not need a warehouse project. You need four timestamps and one flag.
The timestamps: commit merged to the main branch, build finished, deployment started in production, deployment finished. The flag: whether the deployment was later reverted, hotfixed, or followed by an unplanned deploy within a defined window.
Here is a minimal sketch in TypeScript. It takes a list of deployment records and returns the numbers that matter.
type Deploy = {
service: string;
committedAt: Date;
deployedAt: Date;
failed: boolean; // reverted, hotfixed, or manual intervention
unplanned: boolean; // triggered by an incident, not by roadmap work
recoveredAt?: Date; // set when failed is true
};
const median = (xs: number[]) => {
const s = [...xs].sort((a, b) => a - b);
return s.length ? s[Math.floor(s.length / 2)] : 0;
};
export function doraSummary(deploys: Deploy[], days: number) {
const leadHours = deploys.map(
d => (d.deployedAt.getTime() - d.committedAt.getTime()) / 36e5,
);
const failures = deploys.filter(d => d.failed);
const recoveryMins = failures
.filter(d => d.recoveredAt)
.map(d => (d.recoveredAt!.getTime() - d.deployedAt.getTime()) / 6e4);
return {
leadTimeHoursP50: median(leadHours),
deploysPerDay: deploys.length / days,
changeFailRate: failures.length / Math.max(deploys.length, 1),
recoveryMinutesP50: median(recoveryMins),
reworkRate: deploys.filter(d => d.unplanned).length / Math.max(deploys.length, 1),
};
}Two decisions in that snippet carry the weight. First, it uses the median, not the mean. One bad week with a three-day recovery will wreck an average and tell you nothing about a typical Tuesday. Second, it computes everything per service. Aggregate to a team or a department only after you have looked at the services underneath.
The hard part is not the code. The hard part is the failed flag. Somebody has to decide what counts as a failure, write it down, and apply it the same way every time. If your incident tracker links incidents to the deployment that caused them, you get this nearly for free. If it does not, start there.

The tooling question comes after the data question. You already own most of the raw data. Your Git host has commit times. Your CI has build and deploy events. Your incident tool has the failures. The work is joining them on a shared deployment identifier.
Observability platforms are a natural place to overlay deployment markers on top of error rates and latency, so you can see which deploy preceded which spike.
If you would rather keep the stack open and self-hosted, a dashboard layer over your own database of deployment events does the same job at lower cost, with more plumbing on your side.
There is also a category of tooling that reads your repositories and delivery history and surfaces risk, such as hotspots where code churn and past defects cluster. It will not replace the four timestamps, but it explains why one service has a high change fail rate when its neighbours do not.
A caution on buying anything here. A dashboard that shows the numbers is not the same as a team that acts on them. If nobody owns the recovery time chart, buying a nicer chart changes nothing.
The levers that actually move each number
Metrics are a diagnosis. The treatment is somewhere else, and it is different for each one.
For lead time, look at waiting, not working. Pull a dozen recent changes and mark where each one sat idle: awaiting review, awaiting a build slot, awaiting a release window. In most teams the idle time outweighs the coding time several times over. Smaller pull requests and a review service-level expectation beat any pipeline tuning.
For deployment frequency, the lever is batch size. Smaller changes are easier to review, easier to reason about, and easier to revert. Trunk-based development and feature flags exist to let you merge unfinished work safely, which is how you decouple deploying from releasing.
For change fail rate, the lever is how much of the blast radius you can see before it is the whole fleet. A rollout that exposes one percent of traffic, watches an error rate against an SLO, and halts on a breach turns a would-be incident into a non-event. That is the gap between a failed deployment and a failed change that nobody outside the team noticed.
For recovery time, the lever is the revert. If the fastest fix is rolling back, then recovery time is how quickly you detect the problem plus how quickly you can press the button. Automating the revert on a threshold breach cuts the human out of the first ten minutes. At 3am that matters more than any runbook.
For rework rate, the lever is upstream. A high number means incidents are generating deployments. Look at which services produce the emergency patches, and read their post-mortems as a set rather than one at a time.
What changes when AI writes more of your code
The recent DORA research points at something uncomfortable for anyone selling a coding assistant. Assistants speed up low-level tasks, yet the gains did not clearly carry through to lead time or change fail rate. More code written per hour is not the same as more value delivered per week.
If anything, a larger volume of generated changes pushes pressure onto review and onto your delivery pipeline. Bigger batches hurt stability, and review capacity does not scale with typing speed. Watch your change fail rate and rework rate closely in the months after a team adopts an assistant. If lead time falls but rework climbs, you did not get faster, you moved the cost.

How to use them in a retro without breaking your team
Here is what works in practice. Put the trend lines on the screen for the last quarter and ask one question: what changed here, and why? Do not ask who. The goal is a hypothesis about the system, not a verdict about a person.
Three habits keep the numbers honest.
First, never put them in individual performance reviews. The moment a metric attaches to a name, people optimize the metric, and your data stops describing reality.
Second, pair every number with a story. A spike in recovery time is one incident with a name and a post-mortem. Read the post-mortem before you read the chart.
Third, change one thing at a time. If you adopt feature flags, shrink pull requests, and automate rollbacks in the same month, you will never learn which one moved the needle.
Skip the maturity-model slide that sorts your org into tiers and hands out a trophy. Worth the effort instead: a single page per service, updated monthly, listing the five numbers, the previous month's value, and one sentence about what changed.
Suppose you had ninety days of clean data on one service, and you saw the change fail rate creeping upward while deployment frequency stayed flat. Where would you look first: the size of the changes, the quality of the review, or the visibility you have into a rollout while it is still small? Your answer says more about your delivery system than any benchmark does.