Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- DORA Metrics: The Five Keys to Measure DevOps Performance
Overview
DORA metrics are five measures of software delivery performance defined by the DevOps Research and Assessment program: deployment frequency, change lead time, failed deployment recovery time, change fail rate, and deployment rework rate. The first three measure throughput and the last two measure instability. A team is read on both groups together rather than on any single number.
What Are the Five DORA Metrics?
- Deployment frequency: Deployment frequency measures how often a team releases to production over a period. It needs only production deployment events, which makes it the usual starting point for a rollout.
- Change lead time: Change lead time measures the interval from a commit landing in version control to that commit running in production. DORA renamed it from lead time for changes, and reports it as a median because stalled changes distort an average.
- Failed deployment recovery time: Failed deployment recovery time measures how long it takes to recover from a deployment that fails and needs immediate intervention. It replaced time to restore service in 2023, which had itself replaced MTTR.
- Change fail rate: Change fail rate is the share of deployments that require a rollback or hotfix. In the 2024 DORA report, elite performers record 5% and low performers 40%.
- Deployment rework rate: Deployment rework rate is the share of deployments that are unplanned and follow a production incident. DORA added it in 2024 as the genuine fifth metric.
Which Benchmark Table Is Current?
The 2024 Accelerate State of DevOps Report holds the most recent Elite, High, Medium and Low band table. The 2025 report dropped that scheme for seven team archetypes and dora.dev publishes no thresholds today, so quote the 2024 bands with their year attached.
Where Does Testing Fit?
Test execution time sits inside change lead time and pre-merge coverage sits ahead of change fail rate, so the test suite is one of the few levers an engineering team controls directly. Platforms such as TestMu AI shorten the verification leg, while your own deployment and incident records still produce the metric.
Most attempts to measure DevOps end up counting the wrong things. Tickets closed, story points burned down and raw deployment counts all describe activity without saying whether software reaches users faster or breaks less often when it does.
DORA metrics answer that narrower question with five measures drawn from data your pipeline already emits. This guide covers what each metric means today, the current benchmark bands and the year they come from, what DORA has renamed since 2021, and how to instrument the whole set. If you are still mapping out how the two disciplines relate, our breakdown of DevOps vs CI/CD is a useful starting point.
What Are DORA Metrics?
DORA stands for DevOps Research and Assessment, a research program that studies how organizations use DevOps to improve software delivery. It ran independently before Google Cloud acquired it in 2018, and it publishes the annual Accelerate State of DevOps Report along with the metric definitions at dora.dev.
The program is survey-based. Each year it clusters respondents by delivery performance and reports what separates the fastest, most stable teams from the rest. The metrics are the output of that clustering, which is why their definitions and thresholds move between reports rather than standing still.
DORA now splits the five metrics into two groups that pull against each other. Throughput covers how much change moves through the system; instability covers how well those deployments go.
| Throughput | Instability |
|---|---|
| Deployment frequency | Change fail rate |
| Change lead time | Deployment rework rate |
| Failed deployment recovery time | Not applicable |
Recovery time sitting under throughput surprises most teams, who expect it beside change fail rate. DORA groups it with throughput because it measures how quickly the system returns to moving changes, not how often changes go wrong.
Reading the two columns together stops the most common misuse of DORA. Shipping more often is trivial if you ignore the right-hand column, and a team can drive its change fail rate to zero by shipping nothing at all. The point of the pairing is that neither column improves honestly at the other's expense.
Note: Verification time is the part of change lead time you control directly. Shorten it with TestMu AI. Try TestMu AI Today!
The Five DORA Metrics
Each metric below carries the name DORA uses today, its definition, the calculation, and the lever that moves it. Where a name changed, the older one is noted, because most tooling still ships the old label.
1. Deployment Frequency
Deployment frequency counts how often a team releases to production over a given period, or equivalently the time between deployments. It is the cheapest metric to instrument, being a count of production deployment events with no commit join and no incident records, which is why most rollouts start here.
Calculation: number of production deployments divided by the time period.
A deployment here means a release that reaches production, which is why the difference between continuous delivery and continuous deployment changes what this number counts. To raise it, cut batch size so each change carries less risk, automate the validation gate, and remove manual approvals that add waiting rather than scrutiny.
2. Change Lead Time
Change lead time measures how long a change takes to go from committed in version control to running in production. DORA renamed it from lead time for changes, and you will still see the old label and the abbreviation MLT in dashboards.
Calculation: the median interval between a change's first commit and the deployment that carried it. Report the median rather than the mean, for the same reason DORA dropped Mean from the recovery metric: a handful of stalled changes drag an average away from the typical case.
The interval decomposes into code review wait, queue time, test execution, and deployment mechanics. Measure the four separately before optimizing, because teams routinely tune the pipeline when review latency was the real cost. Auditing the DevOps pipeline stage by stage is what makes that split visible.
3. Failed Deployment Recovery Time
Failed deployment recovery time measures how long it takes to recover from a deployment that fails and requires immediate intervention. It replaced time to restore service in 2023, which had itself replaced MTTR, and the scope narrowed with each rename.
Calculation: the median time from the start of a deployment-caused impairment to service returning to normal. Measure from user-facing impact, not from when someone opened the ticket.
Recovery speed depends on detection as much as repair, so you cannot time the start of an impact you never observed. That is where test observability and production monitoring earn their keep. Rollback automation, feature flags and rehearsed runbooks move the second half of the interval.
4. Change Fail Rate
Change fail rate is the share of deployments that require immediate intervention afterwards, typically a rollback or a hotfix. DORA shortened the name from change failure rate, and the definition is narrower than most teams assume: it counts deployments that needed rescuing, not bugs found later.
Calculation: failed deployments divided by total deployments, expressed as a percentage.
This is the metric most sensitive to definition drift, because widening what counts as a planned fix lowers it without changing anything real. Stronger review and continuous testing move it honestly, by catching regressions before the deployment rather than after.
5. Deployment Rework Rate
Deployment rework rate is the share of deployments that are unplanned and happen as a result of a production incident. DORA introduced it in 2024 to test a longstanding hypothesis that change fail rate was acting as a proxy for how much rework a team absorbs. The data confirmed the two are related, and DORA now reads them together as its measure of delivery stability, which is a factor change fail rate could not carry on its own. Rework rate was excluded from the performance clusters, which is why it has no band in the table above.
Calculation: unplanned incident-driven deployments divided by total deployments, expressed as a percentage.
Change fail rate counts deployments that went wrong; rework rate counts the extra deployments you had to make because something in production went wrong. A team can hold a respectable change fail rate while spending a large share of its release capacity on cleanup, and only the second metric shows it.
DORA Performance Benchmarks
The table below is from the 2024 Accelerate State of DevOps Report, page 13. Attach the year whenever you quote it, because this is a snapshot of one survey population rather than a fixed standard.
| Performance level | Change lead time | Deployment frequency | Change fail rate | Failed deployment recovery time |
|---|---|---|---|---|
| Elite | Less than one day | On demand (multiple deploys per day) | 5% | Less than one hour |
| High | Between one day and one week | Between once per day and once per week | 20% | Less than one day |
| Medium | Between one week and one month | Between once per week and once per month | 10% | Less than one day |
| Low | Between one month and six months | Between once per month and once every six months | 40% | Between one week and one month |
The change fail rate column looks like a transcription error and is not. The medium cluster records a lower change fail rate than the high cluster above it, and the 2024 Accelerate State of DevOps Report explains why on page 14: within every cluster throughput and stability are correlated, and the medium cluster is the one that deploys less often and fails less often. The same page records DORA's naming decision, calling the faster teams high performers and the slower, steadier teams medium performers, so the ordering reflects throughput rather than stability.
The band count is not fixed: the 2022 report found only three clusters and dropped Elite entirely, and the 2025 DORA report replaced performance levels altogether with seven team archetypes. No deployment rework rate bands have been published either, so the fifth metric has a definition but no thresholds.
The report's own framing is the one to keep: the best teams are those that achieve elite improvement, not necessarily elite performance. Your trend against your own baseline is more actionable than your distance from a survey cluster.
What Did DORA Rename, and Why?
Search for DORA metrics and you will find MTTR, time to restore service and failed deployment recovery time used as if they were the same thing. They are three successive names for one measurement, and each rename narrowed what it counts.
| Year | Change |
|---|---|
| 2021 | Availability is broadened to reliability and described as a fifth metric. |
| 2023 | MTTR and time to restore service are renamed and redefined as failed deployment recovery time. |
| 2024 | Deployment rework rate is added as a fifth software delivery metric. |
Dropping Mean. An average is a poor summary of incident data, because incident durations are not evenly distributed. Most resolve quickly and a rare few run for hours, and those outliers drag the mean somewhere no real incident lives. Nine incidents resolved in 10 minutes and one that took 20 hours produce a mean of roughly two hours, which describes none of the ten.
Dropping Recover. Recovery invites you to time the fix: the moment the server came back or the bad deploy was rolled back. Users do not experience your infrastructure, they experience your service, so the interval that matters ends when the user-facing failure is over. That can be earlier than the fix if a feature flag cleared the impact, and later if the box is healthy while a queue still drains.
Narrowing to failed deployments. The 2023 redefinition draws the line DORA cared about most. Earlier definitions did not separate a failure caused by a software change from one caused by something external such as a data center outage. Scoping the metric to impairments that follow a change to production makes it comparable with the other delivery metrics, which are all about changes you shipped.
Reliability was never the fifth metric. The 2021 report did present reliability that way, and DORA has since said the label was wrong: reliability measures operational performance rather than software delivery performance, and it sits outside the delivery set. Reliability remains a useful thing to track through your SLOs and error budgets. It is just not one of the five.
If your organization still calls the metric MTTR, there is no need to fight over the label. Agree on the two timestamps, report a median, and note which definition the dashboard is using.
What Counts as a Deployment, a Failure, and Production?
Most DORA rollouts stall here rather than on tooling. The formulas are arithmetic; the definitions underneath them are decisions, and two teams using identical dashboards will produce incomparable numbers until those decisions are written down.
- A deployment - Does a config change count? A feature flag flip? A database migration shipped separately from the code that uses it? Teams running multiple testing environments also have to name which ones count as production.
- A failure - The DORA definition is a deployment needing immediate intervention. A bug found three days later is not one, and neither is a known issue shipped deliberately behind a flag. Where the boundary sits determines both change fail rate and rework rate.
- Production - Straightforward for one service, ambiguous across a microservice estate. A change released to one of forty services is a production deployment, yet counting each service release equally inflates frequency against a team that ships a monolith once.
- The unit of measurement - Per service, per team, or per product? Aggregating across services hides the slow ones; splitting too finely produces samples too small to read.
Write the four answers down before the first dashboard goes up, keep them in version control beside the queries, and date any change to them. An undocumented change to a definition makes every historical trend meaningless, and it is the single most common reason DORA numbers get abandoned a quarter after launch.
How Your Test Suite Sets Lead Time and Change Fail Rate
Two of the five metrics are largely decided by the test suite, which is the part of the pipeline an engineering team can change without organizational permission.
Change lead time contains test execution as a direct term. When a suite takes four hours, every change carries four hours of unavoidable latency, and the usual response makes things worse: teams batch changes to amortize the wait, which raises batch size, which raises the blast radius of each release. Suite runtime and batch size pull change lead time and change fail rate in the same direction.
Change fail rate is set by what the suite catches before a merge rather than after. A gap in browser, device or locale coverage is a gap that surfaces in production, and flaky tests are worse than absent ones: a suite that fails at random trains a team to re-run until green, which is indistinguishable from disabling the gate.
This is where a cloud execution platform moves a DORA number rather than just reporting one. HyperExecute from TestMu AI runs an existing suite up to 70% faster than a traditional grid by placing the test script and its execution components on a single just-in-time virtual machine instead of routing every command across a hub and a node. Boomi cut its test cycle from 9.5 hours to under 2 hours, saving roughly 7 hours per cycle, while tripling test coverage. Adoption does not require a rewrite: one declarative configuration file points the CLI at the tests you already have, and the HyperExecute CI/CD integration docs cover wiring it into an existing pipeline.
Recovery time has a testing component too, though a narrower one. The diagnosis leg, working out why something failed, is often longer than the fix. Test Insights generates an agentic root cause analysis that correlates a failed test's network, console and framework logs into a single report, marking which step is the likely cause and which are downstream effects. TestMu AI is explicit that the output is a lead rather than a verdict, so an engineer confirms it before acting, and the build insights documentation shows the trend views it feeds.
None of this computes your DORA metrics. Test tooling moves inputs to them, and the metrics themselves still come from your own deployment and incident records.
How to Implement DORA Metrics?
Every DORA metric derives from two event streams you already produce: deployment events from CI/CD, and incident events from monitoring or on-call tooling. Implementation is mostly a matter of capturing both reliably and joining them.
1. Emit deployment events. Have the pipeline write one row per deployment to a deployments table: deployment_id, deployed_at, service, environment, commit_sha, first_commit_at for the timestamp of that change's first commit, and a boolean required_intervention set when the deployment needed a rollback or hotfix. Integrating your version control with the best CI/CD tools is what makes this automatic rather than a manual log.
2. Emit incident events. Write one row per incident to an incidents table: impact_started_at, restored_at, and deployment_id naming the deployment that triggered it. That last column is what separates failed deployment recovery time from generic incident duration, it is the one teams most often skip, and it is also what back-fills required_intervention on the deployment row.
3. Query both tables. With the two streams landing in a warehouse, the metrics are ordinary aggregations. The queries below are BigQuery standard SQL against the deployments and incidents tables exactly as described above.
-- Deployment frequency and median change lead time, by week
SELECT
DATE_TRUNC(DATE(deployed_at), WEEK) AS week,
COUNT(*) AS deployments,
APPROX_QUANTILES(
TIMESTAMP_DIFF(deployed_at, first_commit_at, MINUTE), 100
)[OFFSET(50)] AS median_lead_time_minutes
FROM deployments
WHERE environment = 'production'
GROUP BY week
ORDER BY week DESC;
-- Change fail rate over the trailing 90 days
SELECT
ROUND(SAFE_DIVIDE(COUNTIF(required_intervention), COUNT(*)) * 100, 1) AS change_fail_rate_pct
FROM deployments
WHERE environment = 'production'
AND deployed_at >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 90 DAY);
-- Median failed deployment recovery time, trailing 90 days
SELECT
APPROX_QUANTILES(
TIMESTAMP_DIFF(i.restored_at, i.impact_started_at, MINUTE), 100
)[OFFSET(50)] AS median_recovery_minutes
FROM incidents AS i
JOIN deployments AS d
ON d.deployment_id = i.deployment_id
WHERE d.environment = 'production'
AND i.impact_started_at >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 90 DAY);4. Set a baseline before changing anything. Measure for one month or one full sprint cycle first. Without a baseline you cannot separate an improvement from ordinary week-to-week variance, and DORA's own guidance weighs improvement against your starting point rather than against the bands.
5. Read the groups together and act on one. Pick the weakest metric, find the stage that owns it, and change one thing. If change lead time is long, decompose it into review wait, queue time and test execution before touching the pipeline. If recovery time is long, work on detection before repair.
Google's Four Keys Open-Source Pipeline
Google's DORA team published an open-source project called Four Keys that collects the four original metrics from existing tooling and renders them on a dashboard. It is worth reading before you buy a delivery analytics product or write your own collectors, because it answers the hardest question for you, which is where the numbers come from. Note upfront that the repository was archived in January 2024 and its README states it is no longer maintained.
The pipeline is an extract, transform and visualize chain on Google Cloud, and it will feel familiar if you have built a Google Cloud CI/CD pipeline before:
- Data extraction - Webhook events from version control and CI/CD systems, such as GitHub or GitLab pushes, deployments, and PagerDuty incidents, are sent to a Cloud Run endpoint and published to a Pub/Sub topic. Pub/Sub decouples ingestion from processing, so a burst of events or a downstream outage does not lose data.
- Data transformation - A per-source Cloud Run worker normalizes each payload on the way in and writes it to BigQuery, where the raw event body is kept as JSON. BigQuery views then derive the changes, deployments and incidents tables on read. Keeping the raw events means you can change a definition, such as what counts as an incident, and recompute history rather than starting the dataset over.
- The dashboard - A Grafana dashboard reads the BigQuery tables and colour-codes each metric against thresholds that roughly follow DORA's performance bands, so a team sees a category rather than a bare number.
- Deployment - Cloud Build builds the pipeline's container images and Terraform deploys the resources. Cloud Build is also one of the supported deployment event sources, alongside GitLab, CircleCI, Tekton and ArgoCD.
Set expectations accordingly. Four Keys is a reference implementation rather than a supported product, it is now read-only and unmaintained, and it bills against your own Google Cloud project. It also covers only the original four metrics, so deployment rework rate is not included. Even if you never run it, the source is a precise specification of how each metric can be derived from raw events.
Note: Flaky tests raise change lead time and change fail rate at the same time, and the cost compounds on every pipeline that retries instead of diagnosing. TestMu AI runs suites across 3,000+ browser and OS combinations on its test automation cloud.
DORA Metrics and Value Stream Management
DORA metrics start at the commit, which means they measure only the part of the process engineering already controls. Teams that automate hard often hit a wall here: the metrics look good and the business still waits months for features.
Value Stream Management maps the full path from idea to production value, including intake and prioritization wait, design and security approvals, and the handoffs between teams. Consider a team with an elite change lead time under a day whose features still take four months to reach customers. DORA cannot explain that, because the delay is not in the pipeline. A value stream map can: the idea sat in a backlog for eight weeks and waited three more on a quarterly review board.
Map the value stream first to find where time is actually lost, then use DORA to measure the delivery segment once you have confirmed that segment is worth optimizing. Where the map does point at delivery, automation testing in the CI/CD pipeline is usually where the time goes, and broader software delivery practices cover the upstream half.
DORA Rollout Pitfalls
The failure modes below account for most abandoned DORA programs, and several overlap with the CI/CD pipeline challenges teams already know.
- Measuring individuals - DORA is a team-level diagnostic. Attach it to a person's review and you get optimized numbers rather than improved delivery, which is the fastest way to lose the data's usefulness permanently.
- Gaming that hides in plain sight - Deployment frequency inflates through empty commits deployed to raise the count. Change fail rate falls when incidents get reclassified as planned work. Both are detectable: watch deployment frequency against change size, and audit how many incidents changed category after the fact.
- Tracking one column - Optimizing throughput alone produces a team that ships fast and breaks often; optimizing instability alone produces one that ships nothing. The pairing is the whole design.
- Scattered, inconsistent data - When deployment records live in one tool and incidents in another with no shared key, the join is guesswork. Integrate across development, testing and operations first, or accept that the numbers are directional.
- Expecting fast results - These are trend metrics. A month of data is a baseline, not a verdict, and structural improvements such as cutting suite runtime show up over quarters. Complementary QA metrics give faster feedback while the delivery trends accumulate.
Conclusion
Start by writing down what counts as a deployment, a failure, and production for your team, then instrument deployment frequency and change lead time from CI/CD data alone. Those two need no incident tooling and give you a baseline within a sprint.
If that baseline shows test execution dominating change lead time, the GitHub Actions integration guide walks through wiring a faster suite into an existing pipeline. Pair the metrics with solid DevOps best practices.
Citations
- DORA's software delivery performance metrics - current definitions of the five metrics and the throughput and instability grouping.
- A history of DORA's software delivery metrics - the dated record of every rename and addition since 2021.
- The DORA Quick Check - a five-question self-assessment against the research data.
- GitLab DORA metrics documentation - one vendor's implementation of the metric definitions.
- Four Keys, Google's open-source DORA pipeline (archived) - read-only since January 2024, still useful as a specification.
Author
Nazneen Ahmad is a freelance Technical Content SEO Writer with over 6 years of experience in crafting high ranking content on software testing, web development, and medical case studies. She has written 60+ technical blogs, including 50+ top-ranking articles focused on software testing and web development. Certified in Automation Basic and Advanced Training - XO 10, she blends subject knowledge with SEO strategies to create user focused, authoritative content. Over time, she has shifted from quick, keyword-heavy drafts to producing content that prioritizes user intent, readability, and topical authority to deliver lasting value.
Reviewer
Sandeep Yadav is a Senior Software Engineer at TestMu AI (formerly LambdaTest), where he builds the platform's test intelligence and AI-native engineering systems. He has architected autonomous GitHub Apps, vector-search code intelligence, and self-diagnosing QA workflows, and designed distributed platforms that process 2M+ daily test executions and 1B+ events, turning high-volume test, log, and code data into intelligent, self-optimizing systems. He works on embedding reasoning models into production infrastructure to power autonomous review, root-cause analysis, and analytics workflows. He brings over four years of engineering experience with deep expertise in the Elastic Stack, Apache Kafka, and Redis. Earlier he engineered a GDPR-compliant, end-to-end-encrypted secure web-chat application at Mithi. A Facebook Hackercup 2021 Round 2 qualifier and merit-scholarship recipient, Sandeep holds a B.Tech in Electrical Engineering from Delhi Technological University.
DORA Metrics FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests





