In September I found a job called deploy-prod in backend-ci.yml. It ran on every push to main and called the backend deploy workflow with environment: prod. No approval step, nobody asked.

I had disabled the four deploy workflows precisely so that nothing would reach production on its own, and I had been describing that state as "deploys are off". The job had never fired for a simpler reason: the tests in front of it had been failing on main since the end of March.

This is for anyone who has turned off deploys by disabling workflows, or who is counting on a red CI run to keep a pipeline quiet. Both paths in this post come from the same mistake: I trusted a label, "disabled" or "the deploy workflow", instead of listing what could actually run. Disabled and unable to run are different claims, and a path that never gets selected looks exactly like a working one in the run history.

The job that deployed production on every push

deploy-prod:
  needs: build-image
  if: github.event_name == 'push' && github.ref == 'refs/heads/main'
  uses: ./.github/workflows/backend-deploy.yml
  with:
    environment: prod
    image_tag: ${{ github.sha }}
  secrets: inherit

deploy-prod depended on build-image, which depended on lint and tests. The file it called, backend-deploy.yml, was disabled in the repository settings. backend-ci.yml, the file it lived in, was not. If your production environment has no required reviewers, a call like this goes straight through.

What "disabled" covers

GitHub's page on disabling a workflow says it stops the workflow "from being triggered". It says nothing about a disabled workflow being called from another one with uses:. A community discussion from 2023 reports that such a call still runs, since a reusable workflow executes as part of its caller, and it has no official answer.

I never saw it happen in my repository, because the chain was always skipped, and I'm not going to find out by pointing it at a production environment. So I don't know for certain which answer is right. I decided I didn't need to. A job whose safety rests on undocumented behaviour is a job to delete.

A gate that only worked while it was broken

Every push run of backend-ci.yml on main since March 28 ended in failure. With lint and tests red, build-image never ran and deploy-prod was skipped each time. The tests were the only thing between a merge and a production deploy. That turns fixing them into a hazard: the day I repaired the test job, the next merge would have deployed the backend without anyone asking for it.

The timing could have been bad in a specific way. During a rollback of the production frontend, main was 26 commits ahead of what the production backend was running. A merge at that point would have pushed the backend forward 26 commits in the middle of a rollback. The mitigation at the time was to disable the CI workflow before merging to main and enable it again afterwards. That is a chore, and it fails silently the first time someone forgets.

Deleting it, and a test that pins the absence

The fix deleted the job instead of putting a flag in front of it. A production deploy is now started only by hand, with gh workflow run on the deploy workflow and the environment and image tag given explicitly. The dev deploy on pushes to develop stayed: develop to dev is the environment mapping I actually want, dev is recoverable, and it is where changes get checked.

GitHub's environment protection rules, such as required reviewers, are the native control for this, and they're worth setting on any production environment. They put a person in front of a path. Deleting the job removes the path.

A test now fails if a job like this comes back. Two details in it are worth copying. The first is about YAML:

def _push_branches(workflow: dict[str, Any]) -> list[str]:
    """``on`` is the YAML boolean ``True`` once parsed, because unquoted ``on:`` is
    a boolean key in YAML 1.1. Reading only the string key would silently find
    no triggers anywhere and make every assertion below vacuous.
    """
    triggers = workflow.get("on", workflow.get(True, {}))
    if not isinstance(triggers, dict):
        return []
    push = triggers.get("push")
    ...

PyYAML follows YAML 1.1, where on means true. Read the workflow with workflow.get("on") alone and you find no triggers in any file, so every assertion that follows passes on nothing. The test file opens by checking that it found the workflow files, that at least one push trigger parses, and that its detector recognises a real deploy call. Without those three checks, an empty result would read as a pass.

The second detail is that it's narrow on purpose. It flags a job that passes environment: prod to a deploy workflow and fires on a push to main. It doesn't flag environment: prod by itself, because another workflow uses that environment on pull requests to read production credentials for read-only checks. A test that reports a read-only job as a violation soon stops being read.

It also has a limit. It recognises one shape, a uses: call with an environment input, and a deploy written inline as raw CLI steps would get past it. That's acceptable here because deploys go through dedicated deploy workflows, but it narrows the risk rather than proving it away.

The second path: a fast deploy that went quiet

backend-deploy.yml had a shortcut of its own. If a change didn't touch infrastructure, it skipped CDK and ran aws lambda update-function-code against a list of functions typed into the YAML. When I looked at it, it couldn't have worked, for three separate reasons:

  • The list named 7 of the 22 functions the backend has. The other 15 were never updated by it, including one of a pair of functions that together run the same workload, so a fast deploy would have left the two halves on different commits.
  • It pushed one image to all seven, but the functions are split across two image variants, 19 and 3, so some of them would always get the wrong one.
  • The tag it pushed was the bare commit SHA, which the build no longer publishes. Each variant gets a suffixed tag instead.
A grid of 22 function chips in two rows. 7 are filled, labelled on the fast path's list; 15 are outlines, labelled never updated by it, and the last 3 of those have violet borders. Bands underneath mark image variant A with 19 functions and image variant B with 3. Caption: one image tag pushed to all 7, and that tag was no longer built.
The fast path's hand-written list against the roster it was supposed to deploy.

None of this ever showed up as a failure. The shortcut's last successful run was on April 30. In the ten deploys after that its condition never selected it, so it didn't appear at all. Nine of those ten never started at all, for reasons unrelated to the code, so they say nothing about it either. In a list of green and grey runs, a path nobody takes looks the same as a healthy one.

I deleted it rather than repair it. Which function takes which image variant is already recorded as data, in scripts/ecr/image-variants.json, and the CDK stack reads that file. The typed-out list in the YAML was a second copy, and the second copy is what drifted three ways. The guard is a test that fails if any workflow names a backend function, and it takes the names it searches for from that same JSON file, so a new function is covered without anyone editing the test.

Checking your own repository

  • Search every workflow for uses: ./.github/workflows/ calls into a deploy workflow, and check whether the caller is active. Disabling the callee is not the same as disabling the call.
  • For each environment, list every job that passes it and what triggers that job.
  • If you parse workflow YAML in a test, read the on key as a boolean as well as a string.
  • For every conditional path in a deploy workflow, find its last successful run. A path that hasn't been selected in months may not work, and nothing will tell you.
  • Prefer deleting a path to putting a switch in front of it. A switch is one more thing somebody has to remember.