How I Deleted 8,000 Lines of Tests and Got Better Coverage
A story about refactoring the Vizzly CLI with an AI pair programmer, fighting with macOS security daemons, switching test frameworks mid-flight, and learning that sometimes you need to throw everything away and start over.
So I just shipped v0.20.1-beta.0 of the Vizzly CLI and honestly, there was a moment halfway through where I thought I’d completely broken everything.
The git history tells a story:
Dec 14: ♻️ Refactor TddService (1,769 lines → functional modules)
Dec 14: 🔥 Remove problematic tests (-7,877 lines)
Dec 14: ♻️ Migrate from Vitest to Node.js test runner (+17,215, -13,707)
Dec 15: 🔖 v0.20.1-beta.0
That’s 30,000+ lines changed in a weekend. Deleting nearly 8,000 lines of tests. Switching test frameworks halfway through. And at one point, macOS quarantined my Node binary and started scanning every file in node_modules, which meant my terminal was so wedged that even ls would just… die.
This is the story of that weekend.
How It Started
The Vizzly CLI had what looked like a reasonable architecture. Service classes with dependency injection. A nice container pattern. Separation of concerns.
I’ll admit, I didn’t really audit this architecture too heavily when I first implemented it. Looking back, I should have gone functional from the start. But hindsight is 20/20.
But every time I went to write a test, I’d end up with something like this:
vi.mock('node:fs/promises');
vi.mock('../../src/utils/git.js');
vi.mock('../../src/utils/output.js');
vi.mock('@vizzly-testing/honeydiff');
vi.mock('../../src/api/client.js');
vi.mock('../../src/auth/token-store.js');
vi.mock('../../src/config/core.js');
describe('TddService', () => {
beforeEach(() => {
// 30 lines of mock setup
});
it('does the thing', async () => {
// 20 more lines of mock setup
// 5 lines of actual test
// 10 lines of assertions on mock calls
});
});
Seven mocks just to get started. And if you got the import order wrong? The test would fail in mysterious ways and you’d spend 20 minutes debugging why the mock wasn’t working. I didn’t like these tests. They were overmocked and they weren’t catching bugs.
The Plan (Emphasis on “Plan”)
My AI pair programmer and I had been talking about functional decomposition. The idea was simple: break services into pure functions that don’t need mocking.
core.js → Pure functions (easy to test)
operations.js → I/O operations (inject dependencies)
index.js → Exports
We’d start with the worst offender: TddService. 1,769 lines handling signature generation, baseline management, screenshot comparison, hotspot detection, result aggregation, and probably making coffee for all I knew.
The refactor looked clean:
src/tdd/
core/
signature.js
hotspot-coverage.js
metadata/
baseline-metadata.js
hotspot-metadata.js
services/
baseline-manager.js
comparison-service.js
result-service.js
tdd-service.js
We wrote 126 tests with zero mocks—just pure functions with inputs and outputs. It felt amazing to test code without fighting the test framework.
The Mistake
Here’s where I got cocky. After TddService worked so well, we refactored everything else: ApiService, AuthService, ConfigService, ProjectService, BuildManager, ScreenshotServer, ServerManager, TestRunner, Uploader, ReportGenerator. Ten major refactorings, all following the same pattern.
The code was much clearer—clean separation, explicit dependencies, easy to reason about. But the test suite was a mess.
We had 910 tests, many of them tightly coupled to the old class-based implementation. They assumed internal details and used extensive mocking to work around the old architecture. And some integration tests were… problematic.
The Gatekeeper Incident
We had integration tests that spawned actual CLI processes using execSync. Great for testing the real thing, right?
Wrong. Something about the combination of Vitest + Volta + spawning processes triggered macOS’s security daemon (syspolicyd) to absolutely lose its mind.
I still don’t fully understand what happened. At first I thought it was because I’d recently tried switching to Fish shell and maybe broke something in my Volta setup. But that wasn’t it. What I do know is that syspolicyd started scanning my Node binary, then every file in node_modules, thousands of scans per second. The CPU maxed out, commands started hanging, then everything started hanging.
I tried to run ls—killed. grep—killed. git status—you guessed it. Exit code 137 (SIGKILL) on everything.
I had broken my terminal so badly that even basic Unix tools wouldn’t run.
The weird part? Once I switched to Node’s built-in test runner, syspolicyd never appeared again. Something about how Vitest spawns workers combined with Volta’s shim must have been triggering the security checks. But I honestly can’t tell you exactly why.
The Nuclear Option
Saturday afternoon, December 14th. I’m watching the Army-Navy game, feeling pretty good about the refactor and how it’s all coming together. My girlfriend’s coming over in 45 minutes. Everything’s under control.
Then it all falls apart. I’m staring at a test suite that’s half-broken, triggering security daemons, and coupled to code I’d just refactored away. And I’d already been frustrated with Vitest—the coverage wasn’t working for spawned processes, parallel execution was a nightmare, the configuration was getting complex.
I made a decision that felt insane at the time: delete all the tests and switch to Node’s built-in test runner.
git rm tests/integration/*.spec.js
git rm tests/commands/*.spec.js
git rm tests/services/*.spec.js
# ... keep going
The final tally was 44 files changed, 6,245 insertions, 7,877 deletions. Nearly 8,000 lines, gone. My commit message tried to sound confident: ”🔥 Remove problematic tests causing CI issues (#132)“—but really I was thinking I hope I know what I’m doing.
The Rewrite
This was the “fuck it” moment. I’d already been eyeing Node’s built-in test runner—no dependencies, no configuration, native V8 coverage that works with subprocesses. Just node --test. Plus, the main Vizzly codebase had been using Node’s test runner from day one and it’s been rock solid.
So I deleted all the Vitest tests and started rewriting them for the Node test runner and the new functional architecture at the same time.
This is what happens when you pair program with an AI at 2am on a Saturday—you make decisions that seem reasonable in the moment. The changeset was massive: 119 files changed, 17,215 insertions, 13,707 deletions. I renamed every test file from .spec.js to .test.js, changed every expect() to assert.strictEqual(), replaced every vi.fn() with simple functions, and changed every describe and it to test. I rewrote the entire test suite in one sitting.
What Actually Worked
Here’s the thing that still surprises me: it worked. The Node test runner was perfect for this because we’d designed the code to not need mocks—just inject fake functions:
test('saves screenshot and returns success', async () => {
let saveToFile = () => Promise.resolve('abc123');
let output = { debug: () => {} };
let result = await handleScreenshot({
data: { name: 'homepage', image: Buffer.from('...') },
deps: { saveToFile, output }
});
assert.strictEqual(result.status, 200);
});
No hoisting issues, no import order problems, no mysterious mock failures. The tests ran in 30 seconds—all 1,432 of them—with 87% code coverage. And because Node’s test runner doesn’t spawn workers the same way Vitest does, no more Gatekeeper meltdowns. Even ls worked again.
The Coverage Jump
The numbers are a lot better:
| Metric | Before | After |
|---|---|---|
| Line Coverage | ~30% | 87.07% |
| Function Coverage | ~25% | 90.79% |
| Branch Coverage | ~40% | 94.99% |
But here’s what matters: these aren’t “fake coverage” from mocking everything. These are real tests hitting real code paths. All CLI commands at 100% coverage, API client fully tested, auth flows fully tested, config system fully tested, error handling fully tested. And I’m not afraid to refactor anymore because the tests actually catch things.
That said, it’s not perfect. I might be over-testing some things and under-testing others. We could probably use more “big tests”—integration tests that exercise the whole system like a user would. But I’m way happier with this than what we had before.
The Stuff We Deleted
In the process of making things better, we deleted:
- 5 service wrapper classes
- 480 lines of legacy HTML generation code
- Vitest and its coverage plugin
- A bunch of test utilities we don’t need anymore
- My sanity (temporarily)
The service layer shrank by ~700 lines.
Commands now import directly from functional modules instead of going through a service container:
// Before
const services = createServices(config);
await services.apiService.request(...);
// After
const client = createApiClient(config);
const build = await getBuild(client, buildId);
Simpler. More direct. Easier to tree-shake.
What I Learned
Deleting code is scary but sometimes necessary. Those 7,877 lines of tests weren’t saving me time. They were costing me time. Every change required updating mocks. Every refactor broke tests in weird ways.
Starting fresh let me write tests the right way.
The Node test runner is actually really good. This is the real lesson. I’d been using Vitest because everyone uses Vitest and it has all the features. But Node’s built-in test runner? It’s genuinely great. No dependencies, no configuration, native V8 coverage that actually works with spawned processes, and it just runs with node --test. The lack of features is the feature. Sometimes the boring choice is the right choice, and honestly, more projects should give it a shot.
Sometimes weird things happen and you don’t fully understand why. The syspolicyd incident still confuses me. Something about Vitest + Volta + process spawning triggered it, but switching to Node’s test runner made it go away. I thought it was Fish shell, but it wasn’t. Maybe it was how Vitest spawns workers, maybe it was Volta’s shim, maybe it was both. I honestly don’t know, and I’m okay with that.
Functional decomposition works. Pure functions are easy to test. Dependency injection makes integration testing straightforward. The pattern scales.
Pair programming with an AI assistant at 2am leads to questionable decisions. But sometimes questionable decisions work out.
The Part Where It All Worked
Sunday morning, December 15th. I ran the build one more time:
$ npm run build
✓ Compiled successfully
✓ Tests pass (1432 passed)
✓ Coverage: 87.07%
I shipped v0.20.1-beta.0.
The CLI works, the tests work, the coverage is good, and I didn’t break backward compatibility (the old service classes are still there, just deprecated). 30,000+ lines changed in a weekend, and somehow it all works. I’m still not entirely sure how.
The full journey is in the git history if you want to see the carnage:
- TddService refactor (#125) - Where it started
- Remove problematic tests (#132) - The deletion
- Node.js test runner migration (#135) - The rewrite
- v0.20.1-beta.0 release - The result
The Vizzly CLI is open source. Feel free to judge my weekend decisions.
And if you’re curious what Vizzly actually does… we’re building visual testing tools that don’t make you want to delete everything and start over. Usually.