What Is Trunk Based Development? A Field Guide for SREs
Summary
Trunk based development is a branching model where developers merge small changes into one shared trunk at least daily and keep it releasable. Unfinished work stays dark behind feature flags or branch by abstraction. It needs fast CI, trustworthy tests and observability tied to each release. DORA ties it to better delivery performance. It fits poorly when builds are slow or versioned releases need long-lived maintenance branches.
It is 3am and the release branch will not merge. Forty commits, three weeks old, touching the same files as main. That pain is the case for trunk based development: what is trunk based development, exactly? It is a branching model where everyone merges small changes into one shared branch, called trunk or main, at least once a day, and keeps that branch releasable at all times.
This guide is for platform engineers who already know Git. It covers what the model demands, what it protects you from, and where it quietly breaks.
What is trunk based development, in operational terms?
Strip the definition to its constraints. All developers integrate into a single branch. Any other branch lives for hours, not weeks. The trunk builds and passes tests on every commit, so it can ship on demand.
DORA's capability page puts numbers on it: three or fewer active branches in the repository, merges to trunk at least once a day, no code freezes, and a build and test cycle that runs in a few minutes. Those numbers are the point. The model is a feedback-loop budget, not a branching style preference.

Small teams sometimes commit straight to trunk. Larger ones use short-lived branches and pull requests for review and build checks, but never to hold work back from integration. The trunkbaseddevelopment.com reference describes both modes and cites Google, which runs about 35,000 developers on a single trunk in a monorepo.
Why long-lived branches fail on a schedule you can predict
A feature branch is a loan. The interest is merge conflict, and it compounds with every commit that lands on main while you are away. The longer the branch lives, the larger the diff, and the larger the diff, the less anyone reads it carefully.
Here is what that does to your change failure rate. A 2,000-line merge gets a skim and an approve. A 60-line merge gets read. Reviewers are not lazy, they are rationing attention. Small batches are the only mechanism that scales review quality.
The second cost is invisible until the release. Two branches that each pass CI can still break each other when they meet. You find out at integration time, the worst possible moment, with a deadline attached.
What the trunk needs before you can trust it
Moving to trunk without the supporting practices is how teams end up reverting to feature branches within a quarter. Four things have to exist first.
A build and test run under about ten minutes. If CI takes 40 minutes, developers batch changes, and batching defeats the model.
Tests that fail for real reasons. A flaky suite trains people to retry and merge anyway.
A fast, honest review process. DORA lists heavyweight and asynchronous code review as a common obstacle, because it pushes developers to batch work.
A way to ship incomplete work safely. That is the next section.
Skip any of these and trunk becomes the place where breakage accumulates, rather than the place where it is caught.
How do you merge unfinished work without shipping it?
This is the question every skeptic asks, and it has a boring answer: you decouple deploy from release. Code reaches production dark. A flag decides who sees it.
// checkout.ts
import { flags } from "./flags";
export async function renderCheckout(user: User) {
if (await flags.isEnabled("new-payment-flow", { userId: user.id })) {
return renderNewPaymentFlow(user); // merged to trunk, dark by default
}
return renderLegacyCheckout(user);
}The new flow merges on day one, behind a flag that is off. You flip it for internal users, then 1 percent, then 10, with an SLO gate at each step. If the error rate crosses the budget, the flag goes off and the change is effectively reverted without a deploy. Teams that evaluate hosted flag services often start with LaunchDarkly, though the pattern works with any provider or an in-house config store.
The second technique is branch by abstraction, for large refactors. You introduce an interface, route callers through it, build the new implementation behind it, and swap when ready. No long-lived branch, no big-bang merge.

The honest cost: flags are debt with a half-life
Nobody on the conference circuit wants to talk about this. Every flag you add is a conditional in production that someone has to remove. We have seen teams of 80 engineers carry several hundred stale flags, each one a branch in the code that nobody can test exhaustively.
Treat flags as short-lived by contract. Give each one an owner and an expiry date at creation. Alert on flags older than their expiry. Delete the dead code path in the same pull request that removes the flag.
The combinatorial risk is real too. Ten independent boolean flags give you 1,024 possible configurations. You will never test all of them. Keep flag interactions few, and keep release flags separate from long-lived operational toggles such as kill switches.
Release strategies on top of trunk
There are two common ways to cut a release from trunk, and neither requires a long-lived branch.
Release directly from trunk. Every green commit is a candidate. Bugs are fixed forward, with a new commit, not by patching an old branch. This suits teams with strong automated tests and fast deploys.
Cut a release branch just in time. You branch from a known-good trunk commit, harden it, ship, then delete it. Fixes land on trunk first and are cherry-picked back. This suits teams with slower release gates or regulated sign-off.
Which one you pick depends on your detection time. If you can notice a bad deploy within five minutes and revert in two, fix forward. If your mean time to detect is an hour, a release branch buys you a place to stand.

Observability is the other half of the deal
Trunk based development raises your deploy frequency, which raises the number of moments something can go wrong. The model only works if you can see a regression within minutes of a merge. That means error rates, latency percentiles and saturation tied to a specific release, not a dashboard someone checks on Monday.
A useful test: after a merge, can you answer "did this change move the p99 or the error rate?" without opening five tabs? If not, you are not ready to merge ten times a day.
Whatever stack you run, the requirement is the same. Deploy markers on your graphs, SLO burn-rate alerts, and a revert path that a tired person can execute at 3am without thinking.
Trunk based development vs GitFlow vs GitHub Flow
The three get conflated, so separate them by one property: how long work stays off the shared branch.
GitFlow keeps a develop branch, feature branches, release branches and hotfix branches. Work can stay isolated for weeks. It was designed for versioned, scheduled releases, and it shows when you try to deploy ten times a day.
GitHub Flow is close to trunk. Short-lived branches, pull requests, merge to main, deploy. The difference the reference site draws is mostly about where releases come from.
Trunk based development is the strictest of the three about branch lifetime, and it assumes feature flags or branch by abstraction for anything that cannot ship in a single small change.
If your team already does GitHub Flow with branches that live under two days, you are closer to trunk than you think. The gap is usually flag discipline and review latency, not tooling.
Which DORA metrics move first
Instrument these before you start, or you will argue from feeling afterwards. Deployment frequency and lead time for changes respond first, because smaller batches flow through the pipeline faster. Change failure rate and time to restore follow, and they depend on your observability and revert path rather than on branching alone.
Do not turn the numbers into a report card for individuals. They describe a system. When lead time rises, look at the review queue and the CI duration before you look at the people.
When trunk based development is the wrong call
Skip it, or at least delay it, in a few cases. Be honest about which one you are in.
Your test suite takes an hour and you cannot parallelize it. You will batch, and batching breaks the model.
You ship versioned artifacts to customers who stay on old versions for years. You need maintenance branches. That is legitimate, and you can still integrate daily on trunk.
You have no flag system and no appetite to build one. Half-finished work will leak into releases.
Your team does not trust the build. Fix the tests first. Trunk will not.
Trunk based development is also not a guarantee of anything. DORA's finding is correlation from its 2016 and 2017 data, with teams that follow these practices showing better delivery and operational performance. It does not say that renaming your branches fixes your pipeline.
How to start without a big-bang migration
Do not announce a policy. Measure first, then shrink.
Record your current branch lifetime and the size of the median pull request. These are your baseline numbers.
Set a cap, such as no branch older than two days, and make it visible on a dashboard.
Fix the slowest part of CI. If the build takes 30 minutes, nothing else matters yet.
Introduce one flag for one real feature. Ship it dark. Retire it within a sprint.
Remove code freezes last, once your revert path has been exercised in anger.
Review again in a month. The numbers that move first are usually pull request size and time to merge. Change failure rate takes longer, and sometimes gets worse before it improves, because you are finally seeing breakage earlier.
What would your post-mortem say about your branching model?
What will the post-mortem say? That is the useful question before you ship. If your last incident traced back to a merge nobody could review, a branch that diverged for three weeks, or a release that bundled forty changes, then you already know where the loan comes due.
Look at your last three outages. How many involved a large, late integration? That count is your business case.