Visual Testing That Learns From You
Your visual testing tool should get smarter over time, not dumber. Here's how Vizzly learns from your approval patterns to surface what actually matters.
The learning and automatic approval features in this post are no longer part of Vizzly.
Every time you approve a screenshot in your visual testing tool, you’re making a decision. You’re saying “this change is fine, ship it.” That decision has value. It’s signal.
So why do most visual testing tools throw that signal away?
Your Approval History Is Training Data
Here’s something I think about a lot: you’ve been telling your testing tool what matters for months. Every approval is a data point. Every “yes, this timestamp change is expected” is a pattern you’ve already identified.
The tool should be learning from that. Not in some creepy, opaque ML way. Just… paying attention.
When you approve a screenshot where the only change is a timestamp in the header, that’s information. When you do it again the next day, and the next, that’s a pattern. And when you’ve done it twenty times in a row, your tool should probably stop asking.
That’s the core idea behind how Vizzly handles dynamic content. We watch what you approve. We track where changes happen. And over time, we start recognizing the patterns you’ve already identified.
The UI Side: Marking Dynamic Regions
Sometimes you don’t want to wait for patterns to emerge. You know that sidebar widget shows live data. You know that footer has a build timestamp. You’ve known it since day one.
In Vizzly, you can mark regions as dynamic directly in the comparison view. Click on the area, mark it as dynamic, and you’re done. The system incorporates that into its understanding immediately.
But here’s the key difference from ignore regions: we still show you the diff.
The region isn’t hidden. The change isn’t invisible. You can see exactly what happened. We just don’t make you manually approve it every single time when it’s behaving exactly as expected.
Why “Ignore Regions” Get It Wrong
I have strong opinions about ignore regions. Specifically: I think they’re a trap.
The pitch sounds reasonable. “Just tell us what to ignore, and we won’t bother you about it.” Simple. Clean. And completely wrong.
Here’s what actually happens. You set up an ignore region around your timestamp. Great, no more timestamp noise. But then six months later, a bug causes that entire header component to render incorrectly. The timestamp is still there, still changing, still ignored. But the layout around it broke. The spacing is wrong. A new element appeared where it shouldn’t be.
You never see it. The ignore region is doing exactly what you asked. It’s ignoring.
With Vizzly, we take a different approach. We still produce the diff. We still analyze what changed. We just don’t force you to manually approve it when the change matches a known pattern.
If the timestamp changes and nothing else does? Auto-approved. If the timestamp changes AND the layout shifts? That needs your eyes. The diff is there. We just got smarter about when to interrupt you.
Same Shape, Different Pixels
This is the mental model that made everything click for me: we’re looking for the same shape of change, not the same pixels.
A timestamp that says “Jan 22” today and “Jan 23” tomorrow is a different set of pixels. But it’s the same shape of change. Same location. Same size. Same single-line text replacement.
Our statistical analysis tracks this at the line level. We use honeydiff’s cluster analysis to group consecutive changed lines into regions. Then we look at historical patterns: does this region change frequently? Is it a single line of text (classic timestamp signature)? Does the change fall within a known hot spot?
When 80% or more of the current changes overlap with regions that historically change in approved builds, we auto-approve and tell you exactly why.
The Feedback Loop
Here’s where it gets interesting. The system improves as you use it.
Early on, when you only have a handful of builds, confidence is low. We might notice a pattern but not have enough data to act on it. So we ask you to review. But every approval adds signal. Every build adds data points. After ten or twenty builds, the patterns become clear. Confidence scores climb. Auto-approvals increase.
Not everything gets auto-approved, though. That’s intentional. If the change pattern is inconsistent, or it affects more than just known dynamic regions, it needs human review. The system is conservative by design - I’d rather show you something that could have been auto-approved than let something slip through.
You can see confidence levels in the UI: green for high, yellow for medium, orange for low. And it’s not a black box. You can see exactly why something was auto-approved:
- “85% hot spot coverage from last 15 builds”
- “Single-line text change in known dynamic region”
- “Confidence: 87 (high) based on approved build history”
No mystery. No “the AI decided.” Just statistics you can verify.
No ML Required (For Now)
One thing I want to be clear about: there’s no machine learning here. No neural networks. No models trained on other people’s data.
It’s statistical analysis of your own build history. Pattern matching on your own approval decisions. The “learning” is just paying attention to what you’ve already told us.
And it works today. No waiting for models to train. No inference latency. No API calls to external services. Just fast, local analysis that runs in milliseconds.
This matters for a few reasons:
Speed. Statistical analysis is instantaneous. Your CI isn’t waiting on model inference.
Transparency. You can see exactly why something was flagged or auto-approved. No black box.
Determinism. Same input, same output. Every time.
Privacy. Your screenshots stay yours. No external processing.
Statistical analysis gets you 95% of the way there for dynamic content detection. The remaining 5%? That’s where human judgment comes in today - though models will eventually get good enough to help there too.
Building the Foundation
Here’s what excites me about this approach: we’re not just solving today’s problem. We’re building a dataset.
Every comparison honeydiff processes captures rich metadata. Region boundaries. Change densities. Line-level clustering. Text-like signatures. And all of that gets stored alongside your approval decisions.
Right now, that data powers our statistical analysis. Pattern matching, confidence scoring, hot spot detection. It works. It ships value today.
But that same metadata becomes an even richer foundation for whatever comes next. Want to layer ML on top someday? The training data is already there - your own UI patterns, your own approval history, your own definition of “expected change.”
I didn’t want to wait until we had the perfect AI solution to ship something useful. Statistical analysis works now. It’s fast, it’s transparent, and it solves real problems. The metadata we’re capturing along the way just makes future improvements easier.
Giving Your Agent Eyes
We already plug this data into coding agents. Vizzly reads from the .vizzly/ folder - status, diff percentages, thresholds, hot spot coverage. Ask “what’s failing?” and your agent gets the rundown. Ask it to debug a regression and it can analyze the structured data, view the images, and suggest next steps.
The metadata does the heavy lifting. Diff percentage, region classification, confidence scores - that’s context the agent can reason about without burning vision tokens on every screenshot. The statistical analysis already classified the change. The agent just acts on it.
That’s the approach I like: ship what works today, collect the data that makes tomorrow better.
The Workflow Shift
The end result: you catch more bugs because you’re not fatigued from approving the same timestamp twenty times in a row. Your brain stays engaged because every screenshot that lands in front of you actually matters.
No configuration files full of ignore selectors. No maintenance overhead. Just a system that watches, learns, and gets out of your way when it should.
Try Vizzly and see what it’s like when your testing tool actually pays attention.