I changed one function in a test file so that it returned an empty list, and ran the file again. Before: 177 passed. After: 9 skipped, 0 passed, 0 failed, exit code 0. 168 assertions had stopped running, and nothing in the output said so.
The file checks structural rules on every workflow spec in the product, and every test in it is parametrized over the list of specs. Given an empty list, pytest generates no cases, so each parametrized test is reported as skipped instead of failing. In every summary I actually read, a skip looks like a pass.
This is for anyone who owns a large test suite and treats green as proof. A check that still passes when its input is empty hasn't been shown to check anything. You find those checks by emptying the input on purpose and requiring the result to turn red.
A skip someone wrote on purpose
The file already had an escape hatch for exactly this situation:
@pytest.fixture(scope="module")
def spec_ids() -> list[str]:
if not _SPEC_IDS:
pytest.skip("No goal_lane.* specs found — registry unloaded.")
return _SPEC_IDS
It was written as a courtesy for the case where the spec registry fails to load. But a registry that fails to load is precisely when this file stops protecting anything, and the courtesy turned that moment into the one outcome nobody reads.
The fix removed the skip and added a separate test with a floor:
#: 20 exist today; 10 leaves room for a deliberate cull without leaving
#: room for "the registry returned nothing".
_MIN_SPECS = 10
def test_the_registry_actually_discovered_specs() -> None:
assert len(_SPEC_IDS) >= _MIN_SPECS, (...)
Two choices in it matter. It is a floor, not an exact count, so adding a spec doesn't turn the file red. And the guard itself isn't parametrized, because a guard that shares the failure mode it is guarding against protects nothing. The general rule I follow now: when a check depends on something in its environment being present, it fails when that thing is missing instead of skipping. Deliberate, declared skips are a different matter, such as a test that only runs when real cloud credentials are switched on. The kind to avoid is a skip that fires because something the test relies on quietly disappeared.
pytest can do part of this for you. The ini option empty_parameter_set_mark decides what an empty @parametrize does. The default is skip. I ran a three-line test with an empty list both ways:
$ pytest test_empty.py -rs
SKIPPED [1] test_empty.py:5: got empty parameter set for (spec)
1 skipped
$ pytest test_empty.py -o empty_parameter_set_mark=fail_at_collect
Empty parameter set in 'test_spec_has_a_name' at line 5
!!!!!!!!!!!!!!!!!!! Interrupted: 1 error during collection !!!!!!!!!!!!!!!!!!!!
1 error
The first exits with status 0 and the second with 2. Setting fail_at_collect in your pytest config is a one-line change and worth making. It only catches a list that is exactly empty, though. A registry that silently returns 3 specs instead of 20 still passes it, which is why this file also keeps the floor. And whatever you decide, -rs prints the reason for every skip in the summary, so at least you can see them.
Why an empty input makes an assertion true
Every one of these is true when the collection it reads is empty:
assert all(check(spec) for spec in specs) # specs == []
assert not (found - known) # coverage check; found is empty
assert forbidden not in scanned_names # the scan read nothing
assert not (a & b) # either side is empty
@pytest.mark.parametrize("spec", specs) # zero cases: reported as skipped
Coverage checks are the riskiest of the group, because their whole shape is "nothing I found is missing from the list". A scanner that stops finding anything produces an empty difference and a green test. The fourth line catches people out because it doesn't look like an absence check at all, yet it passes if either operand goes empty.
The same sweep found a second case. A test that asserted the development environment keeps no always-on capacity read a tuple of configuration knob names. Emptying the tuple turned one test into assert set() == set() and collapsed another to skipped: 7 passed, 1 skipped, no failures. A written cost rule pointed at that file as its proof, so an empty tuple would have left the rule claiming something that nothing checked.
The same failure, one level up
The larger versions don't involve a single empty list. The whole run measures something other than what it is named after, and the output still looks fine.
The wrong interpreter
The backend's virtual environment had ended up containing only pip. poetry run python used it, but poetry run pytest fell through PATH to a global pyenv shim whose packages were a snapshot from March. About 16,000 tests reported green against langgraph 0.6.11 while the product was built for 1.2.9. The 22 tests that did fail had been blamed on lock-file drift, and a plain poetry install fixed all of them. The guard is now a test that checks which interpreter is running and whether its installed versions satisfy pyproject.toml. Pointed at the stale interpreter, it reports 36 findings.
One import
A test file imported a function that another pull request had just moved. Pytest stops at collection when an import fails, so 21,689 collected tests never ran, and the summary was a single error line, which reads like one small problem. The two changes came from branches that had diverged: one side deleted the function, the other added a caller, and git merge-tree reported no conflict because no line overlapped.
A linter that checked nothing
When ruff can't initialise its cache directory, it prints nothing and exits 0, having examined zero files. On Windows that happens intermittently, so the same command is honest on one run and empty on the next. Seven real errors, three unused imports and four import-ordering violations, were committed after a clean-looking run. The gate now passes --no-cache.
The probe
For any assertion I want to trust, I list its operands, including @parametrize sources and the return value of any function that scans files. Then, one operand at a time, I make it empty, run the test, check that it goes red, and restore it. The last step is re-running to confirm the original result comes back.

I tried a shortcut first and it didn't hold up. A keyword search across one test directory flagged 15 files as having no guard against empty input. I probed the 4 with the highest stakes: 2 were already guarded, in forms a keyword search can't see, and 2 were genuinely vacuous. Half wrong on the files that mattered most. Emptying the operand takes longer, but it doesn't depend on guessing correctly.
Four shapes a real guard takes
| Shape | What it looks like | Example |
|---|---|---|
| A separate guard test | Another test fails on its own when the operand empties | The spec-count test above |
| Non-empty by construction | Look things up by name so that a miss raises instead of returning nothing | Importing each module by dotted path: rename the file and you get ImportError, not an empty list |
| A floor | len(x) >= N, with N set deliberately below today's count | _MIN_SPECS = 10 with 20 present |
| Presence direction | The assertion itself fails on an empty set | assert {"a", "b"} <= granted |
The second and third shapes are the ones a keyword search misses. A test that resolves modules by import name has no explicit emptiness check anywhere, and a floor looks like an ordinary assertion. Both are real guards.
On floors versus exact counts: an exact count goes red every time someone adds a case, and a check that goes red for good reasons soon gets its number bumped without anyone reading why. A floor well below today's count fails only on the thing it exists to catch.

