September 14, 2026

The CI/CD Mistake That Was Costing Us 180 Redundant Docker Builds

Most backend languages read environment variables when a process starts. You set DATABASE_URL, launch the binary, and it picks up whatever's in the environment at that moment. The same binary works in staging, production, or your laptop, because configuration is a runtime concern.

Frontend frameworks often don't work that way. A React app (and plenty of Next.js setups too) resolves its public environment variables at build time. The bundler reads process.env while compiling, inlines the values directly into the JavaScript output, and from that point on they're baked into the artifact. There's no process.env lookup left at runtime to intercept: the string is just sitting in the bundle.

That distinction sounds academic until you look at what it does to a CI/CD pipeline running across real infrastructure. In one pipeline I inherited, this meant rebuilding the exact same application for every service, region, and environment combination it ran in: 30 services, 2 regions, 3 environments. That's 180 separate Docker builds, each one compiling identical application code, differing only in which API endpoint or feature flag got inlined.

Root cause

The pipeline wasn't rebuilding 180 times because the code changed 180 ways. It was rebuilding because the build step itself was the thing coupling code to configuration. Once you run docker build with NEXT_PUBLIC_API_URL=https://api.eu.staging.example.com baked in as a build argument, the resulting image is that one environment. It can't be deployed anywhere else without lying about where it's running.

So the CI matrix had no choice. Every environment/region pair needed its own build job, because the artifact itself encoded one specific combination. Diff any two of those images and the application code was byte-for-byte identical. The only thing that differed was a handful of strings that had no business being decided at compile time in the first place.

A useful test for any pipeline: does this variable change what the binary is, or what it does? A compiler flag or a dependency version changes what it is, and belongs at build time. An API URL, a region, or a feature flag changes what it does, and belongs at runtime. When the two get mixed together, your build matrix starts multiplying by things that were never architectural decisions.

This is a Twelve-Factor App violation

Once I traced the root cause, I recognized the shape of it immediately: it's a direct violation of two Twelve-Factor App principles at once. Factor III, config, says config belongs in the environment, strictly separated from code. Factor V, build/release/run, says those stages stay separate, with one build producing an artifact that many releases can reuse against different config.

A React app that resolves process.env.NEXT_PUBLIC_* at build time collapses both into a single step. Config gets welded into the artifact during build, so build and release become the same event, and "config lives in the environment" stops being true the moment compilation finishes.

That's not just a style violation. It has real costs:

  • No single build can be promoted through environments. The binary tested in staging is never the binary that ships to production; they're different artifacts by construction, so a passing staging test doesn't prove what production will run.
  • Rollback means rebuilding, not redeploying. Reverting a bad release requires recompiling an old commit for a specific environment, instead of pointing traffic back at a previous image.
  • Public bundles can leak more than intended. Anything inlined at build time ships to every browser that loads the app, so a build-time secret or internal endpoint that should never be public is one careless env var away from sitting in plain text in the JS bundle.
  • Config drift becomes invisible. Two environments can end up running builds from different commits under the same version label, because nothing guarantees "the build for prod" and "the build for staging" ever shared an artifact.

The fix

We redesigned the pipeline so environment variables were injected at deploy time instead of build time, restoring the separation Twelve-Factor asks for. Concretely, that meant building one image per service (not per service times environment times region), and having the container read its configuration from the runtime environment when it started, via an entrypoint script that wrote a small runtime config file before the app booted.

The application code changed too. Anywhere that used to read process.env.NEXT_PUBLIC_API_URL directly (letting the bundler inline it) now read from that runtime config instead. It's a small change in principle, but it moves the entire deployment matrix off the build pipeline and onto the deploy step, which is exactly where environment-specific decisions belong.

BEFORE: build coupled to config
  code + config(env, region) -> docker build -> image(env, region)
  30 services x 2 regions x 3 envs = 180 builds

AFTER: build decoupled from config
  code -> docker build -> image (built once per service)
                             |
                             +--> deploy to env A, region 1  (config injected at runtime)
                             +--> deploy to env A, region 2
                             +--> deploy to env B, region 1
                             +--> ... same image, every combination

The result

Build time dropped 60%. The 180 redundant builds collapsed down to 30, one artifact per service, because the same image could now be deployed unchanged to every environment and region it needed to run in. Configuration became a deploy-time input instead of a compile-time constant, which also meant a config typo no longer required a full rebuild to fix, just a redeploy.

There's a side benefit for teams running a BCDR (business continuity/disaster recovery) architecture too. Once an image is no longer tied to a region, you only have to build it once and can replicate it to other regions asynchronously in the background, using whatever cross-region replication your cloud provider's registry offers (ECR cross-region replication on AWS, Artifact Registry replication on GCP, ACR geo-replication on Azure). Getting a region's failover target ready becomes a registry replication job that happens off the critical path, instead of another entry multiplying your build matrix.

Cost and incident response

Two costs here go beyond "the pipeline felt slow."

The first is billed compute. On GitHub Actions, and on any runner billed by the minute, each of those 180 builds is metered time, not just wall-clock time engineers spent waiting. Cutting build time 60% didn't just make the pipeline feel faster, it cut billed Actions minutes by roughly the same proportion, because the fix removed redundant builds outright rather than making each one faster. Rebuilding identical code 6 times (2 regions times 3 environments) for every service was paying for 6x the minutes to produce zero additional code difference.

The second is time to ship an incident fix, and this one matters more than the CI bill. When configuration is coupled to the build, a one-line hotfix still has to flow through however many builds are needed to reach every affected environment and region before it can go out, even though the code change is identical across all of them. Every one of those builds sits directly on the critical path of your MTTR, stacked on top of the incident itself.

After the fix, an incident patch is a single build, deployed unchanged to every affected environment and region, with config supplied at deploy time. The minutes saved on a routine day become minutes shaved off an outage on a bad one.

Signs your pipeline has this problem

A few signals I'd watch for if you suspect your own pipeline has the same coupling:

  • Your build matrix grows every time you add an environment or region, not just when you add a service.
  • CI minutes climb steadily as the org adds environments, even though the application code barely changes.
  • Engineers aren't sure which build artifact is "the" one to deploy, because there isn't one: there are dozens of near-identical images and the right one depends on remembering which combination it was built for.

If any of that sounds familiar, it's usually a sign that configuration and code got coupled somewhere upstream of deployment, not that you need more CI capacity. Untangling build-time from runtime configuration is exactly the kind of CI/CD and Infrastructure as Code work I do in Cloud with Gus engagements.