Several checks in this repo have turned out to pass whether or not the thing they name works. They were found one at a time during review, so this issue collects them and proposes a convention.
A test that asserts nothing is worse than a missing one, because it reads as coverage.
Instances, and who verified each
The code never executes. In test/auth.test.tsx, a test rendered <AuthCheck> while beforeEach had signed the user out, so the fallback rendered, UserDetails never mounted, and its two expect calls never ran. Armando re-confirmed on v5 by putting an assertion that cannot pass inside UserDetails: the suite still went green. Fixed in #782.
The test passes when the thing it names is broken. Three, all confirmed by mutation:
The harness cannot report a failure at all. The flake probe ran its test command under bash -e without a set +e guard, so the first failing iteration killed the step before the result was recorded. It could only ever produce a clean table, and the first run that genuinely reproduced the flake would have reported least. Fixed in #785.
The test catches a mutation for the wrong reason. On the startWithValue removal branch, a test appeared to catch a deliberate break but passed under that same mutation when run in isolation: the failure came from another test's warning in the shared suite. It would have surfaced as an order-dependent CI flake.
What would catch these
A convention rather than a framework: any test whose purpose is to guard a specific failure should be shown to fail against a deliberate break of that failure, and the PR should say so. That is what caught four of the five above, and it costs one run.
Two caveats worth stating. Mutating in the suite is not enough on its own, since coupling between tests can produce the failure for an unrelated reason, so run the mutation in isolation as well. And this only covers tests written deliberately as guards; it says nothing about coverage that quietly evaporates when a wrapper changes, which is what happened in #782.
If we want a tool rather than a convention, mutation testing is off-the-shelf for TS (Stryker), and that is worth pricing before writing anything bespoke.
Several checks in this repo have turned out to pass whether or not the thing they name works. They were found one at a time during review, so this issue collects them and proposes a convention.
A test that asserts nothing is worse than a missing one, because it reads as coverage.
Instances, and who verified each
The code never executes. In
test/auth.test.tsx, a test rendered<AuthCheck>whilebeforeEachhad signed the user out, so the fallback rendered,UserDetailsnever mounted, and its twoexpectcalls never ran. Armando re-confirmed onv5by putting an assertion that cannot pass insideUserDetails: the suite still went green. Fixed in #782.The test passes when the thing it names is broken. Three, all confirmed by mutation:
initialDatabranch ofgetServerSnapshotwas dead under test. Neuter it to always returnloadingand all 22 tests still passed. Found by Armando on fix(ssr): add getServerSnapshot to useObservable's useSyncExternalStore #779, which is still open.does not show a logged-out user after navigating awaysits indescribe('useUser')but stopped callinguseUserwhen feat(auth)!: remove the deprecated AuthCheck and ClaimsCheck components #782 replaced its wrapper. MakinguseUserthrow left it passing on that branch while the same mutation failed it onv5. Fixed in feat(auth)!: remove the deprecated AuthCheck and ClaimsCheck components #782.import(). Breaking therequire()half deliberately left every test green. Found on ci: add built-artifact release gate (#765) #766. That code has since been removed from the PR, so this one never landed.The harness cannot report a failure at all. The flake probe ran its test command under
bash -ewithout aset +eguard, so the first failing iteration killed the step before the result was recorded. It could only ever produce a clean table, and the first run that genuinely reproduced the flake would have reported least. Fixed in #785.The test catches a mutation for the wrong reason. On the
startWithValueremoval branch, a test appeared to catch a deliberate break but passed under that same mutation when run in isolation: the failure came from another test's warning in the shared suite. It would have surfaced as an order-dependent CI flake.What would catch these
A convention rather than a framework: any test whose purpose is to guard a specific failure should be shown to fail against a deliberate break of that failure, and the PR should say so. That is what caught four of the five above, and it costs one run.
Two caveats worth stating. Mutating in the suite is not enough on its own, since coupling between tests can produce the failure for an unrelated reason, so run the mutation in isolation as well. And this only covers tests written deliberately as guards; it says nothing about coverage that quietly evaporates when a wrapper changes, which is what happened in #782.
If we want a tool rather than a convention, mutation testing is off-the-shelf for TS (Stryker), and that is worth pricing before writing anything bespoke.