What We've Been Building: Vizzly Updates

I’ve been quiet on the blog for a bit. Here’s what we’ve been working on: CIEDE2000 color science, cluster-based noise filtering, BearDen-inspired local tokens, hotspot filtering in TDD, Swift SDK improvements, and more.

The hotspot filtering section describes a feature that is no longer part of Vizzly. The other release notes remain useful history.

I’ve been heads-down building instead of writing about building. Let me catch you up on what’s been happening with Vizzly over the past few months.

Better Color Science: CIEDE2000

Remember when I talked about Honeydiff’s speed? Turns out speed is only half the battle. What you’re actually measuring matters just as much.

Honeydiff was using YIQ color difference for threshold calculations. It’s fast, it’s simple, and it’s… not very accurate. YIQ treats all color differences equally, but that’s not how human perception works. A 10-unit shift in blue doesn’t look the same as a 10-unit shift in red. Your eyes are way more sensitive to some colors than others.

CIEDE2000 fixes this. It’s a perceptual color difference formula that models how humans actually see color. The math is gnarly (seriously, there are like 15 intermediate calculations), but the results speak for themselves: fewer false positives from colors that look identical to your eyes but differ slightly in RGB values, and better detection of changes that actually matter visually.

We shipped CIEDE2000 in Honeydiff v0.5.0, and it’s now the default across Vizzly. If you’ve been using custom thresholds, you might need to adjust them. CIEDE2000 values are scaled differently than YIQ, so we updated the defaults. Check your builds. They should feel more accurate now.

Cluster-Based Noise Filtering

This one came directly from a user. Kael from the Vuetify team opened an issue about scattered single-pixel differences causing false positives. You know the ones: anti-aliasing artifacts, font rendering variance, sub-pixel shifts that show up as noise across screenshots. They’re not real changes, but basic pixel diffing counts them anyway.

His workaround was clever: filter out small clusters. If the diff pixels aren’t grouped together, they’re probably noise.

So I built it into Honeydiff. The new minClusterSize option lets you set a threshold for what counts as a real change:

const result = await compare(baseline, current, {
  minClusterSize: 2  // Ignore isolated single pixels (default)
});

With minClusterSize: 2 (the default), scattered single pixels get filtered as noise. A solid 5x5 block of changes? That’s 25 connected pixels, definitely gets detected. But 5 random pixels scattered across the image? Filtered out.

You can tune this based on your tolerance:

  • minClusterSize: 1 - Exact matching, every pixel counts
  • minClusterSize: 2 - Default, filters isolated noise
  • minClusterSize: 3+ - More permissive, only detects larger clusters

This is the kind of collaboration I love. Someone hits a real problem, suggests a fix, and we ship it. Thanks Kael. This makes Honeydiff meaningfully better for everyone.

BearDen-Aligned UI Tokens

The site and product surfaces now share the same BearDen-inspired token language. That means the marketing pages, the docs, and the product-adjacent UI all point at one visual vocabulary instead of a bunch of stale one-off styles.

The old design-system label was useful during the migration, but that is a history note now. The current story is simpler: visual regression testing, shared local tokens, and a calmer visual system that doesn’t fight the content.

We’ve been aligning the public-facing pages and internal review surfaces piece by piece. The result is a UI that feels more consistent, is easier to scan, and gives us a cleaner foundation for the rest of the platform.

Hotspot Filtering in TDD Mode

Here’s a workflow improvement that I’m genuinely excited about. When you’re running vizzly tdd and iterating on UI changes, you want to see what you just broke. Not what’s been changing in every build for the past month.

Hotspots are regions that change frequently across builds: timestamps, session IDs, dynamic content, that kind of thing. Vizzly tracks these automatically and uses them for cloud-based auto-approval. But locally? They were just noise.

Now vizzly tdd filters out hotspots by default. You still see them in the UI if you want, but the diffs focus on actual changes you made, not the timestamp that shifts on every test run. Less noise, more signal, faster iteration.

You can toggle hotspot filtering on or off in the TDD dashboard. It’s on by default because that’s how most people want to work, but the raw diffs are always one click away.

Custom Baseline Signatures

This one’s for teams with complex testing setups. Vizzly matches screenshots to baselines using viewport dimensions, device info, OS version, that kind of thing. Works great for most cases.

But what if you’re testing the same screen with different data? Or running visual tests in different environments? Or testing A/B variations?

Custom baseline signature properties let you define your own matching rules. Pass custom metadata when you capture screenshots, configure which properties matter for baseline matching, and Vizzly creates separate baselines for each variation.

await vizzly.screenshot('dashboard', {
  metadata: {
    userRole: 'admin',
    featureFlag: 'new-nav-enabled'
  }
});

Now your admin dashboard screenshots won’t get compared against non-admin baselines. Your feature-flagged variants won’t clash with production screens. You get precise control over what gets compared to what.

This shipped a couple months ago and has been solid for teams with sophisticated testing needs.

Swift SDK Polish

We launched the Swift SDK in November for iOS and macOS visual testing. Since then, we’ve been refining the experience based on feedback.

Auto-detecting device info was the big one. The SDK now automatically captures device model, OS version, screen dimensions, and scale factor without you needing to configure anything. Just call app.vizzlyScreenshot(name: "screen") and it figures out the rest.

We also tightened up the TDD workflow integration. The Swift SDK now writes build metadata to a discovery file, so the TDD dashboard can show you exactly which build you’re working on. Small detail, but it makes local development feel more connected.

Platform Stability Improvements

Behind the scenes, we’ve been hardening the infrastructure.

Graceful shutdown: Vizzly now handles deployments with zero downtime. Servers finish processing in-flight requests before shutting down, so your builds don’t fail mid-upload when we deploy updates.

Better GitHub integration: Error handling for GitHub PR integration is way more robust now. Network issues, API rate limits, temporary GitHub outages. The integration handles all of it gracefully and retries intelligently.

Reliability fixes: Fixed edge cases in TDD run cleanup, screenshot reporting, JWT token handling, comparison ordering. The kind of bugs that only show up in production with real usage patterns.

Developer Experience

We migrated the entire codebase (both platform and CLI) from ESLint and Prettier to Biome. One tool instead of two, faster linting, better defaults, less configuration to maintain.

We also added hand-written TypeScript types to the CLI with automated tsd testing. The auto-generated types were technically correct but hard to read. Now the types are clean, documented, and actually helpful when you’re integrating Vizzly into your tests.

What’s Next

I mentioned accessibility testing in the Honeydiff benchmarks post. That’s coming. Honeydiff already has WCAG color contrast analysis, color blindness simulation, and gradient filtering built in. We’re integrating it into Vizzly’s UI so you can catch accessibility violations directly from your visual tests.

We’re also working on MS-SSIM (Multi-Scale Structural Similarity) and GMSD (Gradient Magnitude Similarity Deviation) as additional perceptual metrics for auto-approval. Different metrics catch different types of changes, and giving you multiple options means better automation for your specific use cases.

The broader UI story is still the same: keep tightening the review history, keep the local token language consistent, and keep making the workflow feel like one product instead of a pile of surfaces.

The Bottom Line

These past few months have been about refinement. Making Vizzly faster, smarter, more reliable, and easier to use. Not flashy feature announcements, just consistent improvement across the board.

CIEDE2000 makes diffs more accurate. Cluster filtering cuts the noise. Hotspot filtering makes TDD more focused. Custom signatures make complex testing setups possible. BearDen-style local tokens make the UI feel cohesive. The stability improvements make everything more reliable.

Visual testing should feel integrated into your development process, not bolted on afterward. Every one of these updates moves closer to that goal.

Back to building.


Want to try these updates yourself? Sign up for Vizzly and run vizzly tdd to experience the local visual development workflow. Read the docs to learn more.

Ready to improve your visual workflow?

Start using Vizzly today and bring visual regression testing into the workflow described in this article.