
For a while, our entire infrastructure was living out the familiar story of every growing platform: a single Terraform root module per environment, more than a dozen nested modules underneath it, and a terraform apply that got a little more nerve-wracking every quarter. Nothing was broken. It had simply stopped scaling with us as we grew.
This is the story of why we moved our entire Terraform infrastructure — across development, test, and production — onto Terragrunt, and what we actually got out of it once the migration was done.
The problem wasn't Terraform. It was how much a single change could trigger.
When your infrastructure is young, one shared state file feels efficient. You run plan, you see everything, you apply, you move on. But as the number of resources grows — databases, caching layers, search clusters, dozens of microservice pipelines, CDN and edge rules — that single state file becomes a liability in three concrete ways:
Plans get slower and noisier. A terraform plan that typically finished in fifteen seconds starts taking minutes; the change you actually care about gets buried under hundreds of unrelated lines that Terraform still insists on re-evaluating.
Blast radius becomes unpredictable. You want to change one service's autoscaling policy. You run plan. Terraform shows you that change — plus, say, a provider version bump touching forty unrelated resources, because everything shares the same root module. Now every change carries the risk of every other change.
It was hard to give a clear answer to "which code applies to which infrastructure." A new engineer couldn't look at the repo and know which module corresponds to which running infrastructure. A single shared state file was also a single lock: when two engineers wanted to work on two unrelated parts of the same environment at the same time, they ended up waiting on each other for the same lock.
None of this is a Terraform bug. It's what happens to any monolith — infrastructure or otherwise — once enough people and enough resources pile into it.
Why Terragrunt, specifically
Terragrunt doesn't replace Terraform — it adds a thin orchestration layer on top of it. You can actually build per-component state with plain Terraform too, using separate root modules; Terragrunt doesn't turn the impossible into the possible, it makes that pattern standard and maintainable without the rewrite overhead — and that's exactly what fit for us. We didn't want to rewrite our modules or switch IaC tools; we wanted to fix how those modules were organized and applied.

Per-component state without the duplication. In a single root file, root.hcl, that every unit includes, we wrote the backend configuration and the shared provider definitions once, and had them generated automatically for every unit — instead of copy-pasting the same boilerplate into every directory. The result: we could split one giant state into dozens of small, independent states — one per database engine, one per microservice, one per frontend app.

Independent apply surfaces. Once each component had its own state file, a change to one service's configuration could only ever be planned and applied against that service. The era of a one-line change producing a fifty-resource diff was over.
A repeatable structure across environments. Because every unit followed the same inheritance pattern, promoting infrastructure from development to test to production stopped being an exercise in remembering what was different between environments. The structure was largely the same; only the input values changed.
Explicit dependencies between units. When one unit needs another's outputs — a KMS key, a registry, a pipeline — it declares a dependency block, so cross-unit references and apply order are written down in code instead of living in someone's head.
How we actually migrated it — without downtime or state loss
This was the part that worried us most. You can't just "turn on" Terragrunt — every resource already lives in a real, live state file, and a single mistake in a state operation means orphaned resources, or worse, Terraform deciding to destroy and recreate something carrying production traffic.
We treated it as a state-surgery problem, not a refactor:
- Waves, not a big bang. We split the migration into phases by subsystem: shared platform components first (registries, build pipelines, secrets management), then application modules in batches — backend services, then frontend services — then the data engines (the databases, caches, and search clusters, which carried the scariest replace-in-place risk), with edge/CDN rules last.
- Zero-diff as the acceptance test. After every single wave, the rule was simple:
terraform planagainst the new structure had to show no changes at all. This gate wasn't a formality — it caught real deviations in a few waves (a misaddressed resource, a file that got missed), and we fixed those before moving on. state mv, never destroy/apply. Every resource moved from the old state into its new, smaller state using Terraform's state-move operations. The infrastructure itself never changed — only the bookkeeping of which file tracks which resource.- Development first, always. Every wave ran in our lowest environment first, got verified, then promoted to test, then production — because the biggest risk in any migration isn't the first environment, it's the drift that quietly accumulates by the third one. The batch count per phase varied by environment, and not every wave even needed a literal
state mv— a unit with no real resources behind it yet just got applied fresh — but the zero-diff check ran every time, in every environment.

As the application-migration phase wound down, we took the same approach one step further for one of our frontend services: we split its CI/CD pipeline definition out into its own independent Terragrunt unit, across dev, test, and production — the closing piece of that phase rather than a project of its own.
It was slower than a rewrite would have been. It also finished without any downtime, which was the actual goal.
What we actually got out of it
Once the migration was done, the gains are concrete:
- Faster, legible plans. A change to one component now produces a plan scoped to just that component — short enough for a reviewer to actually read.
- Safe parallel work. Two engineers can now change two different services' infrastructure at the same time without touching the same state lock — something the old, single shared lock used to serialize instead.
- Independent CI/CD per component. Because each unit is self-contained, a service's pipeline definition can live in its own unit instead of a shared one — we've already done this for one frontend service, and the same pattern is available to any other that needs it, without being hostage to a shared queue.
- Verifiable environment parity. Because development, test, and production now largely share the same structural skeleton, "does production match test?" became a question you can answer by comparing directory trees, not by relying on institutional memory.
- A genuinely smaller blast radius. The failure mode we migrated away from — one bad diff touching everything — is now structurally much harder to hit. A mistake in one unit's configuration mostly can't leak into another unit's plan.
The honest tradeoff
Terragrunt isn't free. There are now more files, more directories, and an extra configuration layer a new engineer has to understand before they can see the whole picture. The DRY inheritance pattern that keeps everything consistent also means a change to the root configuration quietly affects every unit beneath it — so that one file deserves the same review rigor as any shared piece of infrastructure, and it's still a point of friction that hasn't fully gone away.
Running everything at once needs discipline. Terragrunt can run a command across every unit in an environment with run-all, which is powerful and exactly the kind of power we don't want to reach for by default. In production, the default is a single unit; run-all is used only deliberately and announced, and it's never run from the repository root.
We chose this not because it was trendy, but because a single shared state file had become a major source of risk and slowness in how we shipped infrastructure changes.