The full-scan version still needed a way to check its own work. The plan from that dev meeting was simple: run an existing accessibility tool for a baseline score, run the agent, run the tool again, see what moved. I picked WAVE.
Why WAVE
WAVE is built on the WebAIM Million, an annual study of accessibility across the top 1,000,000 home pages, and it scores a page out of 10. Its errors and warnings map closely to the study's own WCAG 2.2 failure data, so I didn't have to invent my own rubric for what "better" meant. If the agent's fixes moved a real WAVE score, that was a number I could trust instead of one I made up.
It sorts what it finds into a few buckets: errors (definite failures, like a form input with no label), contrast errors (text that doesn't meet the minimum ratio against its background), and alerts (likely problems that need a human look, like a redundant link or an empty heading). It also counts features, structural elements, and ARIA usage, the things a page is doing right. All of it rolls up into that single score out of 10.
Teaching the agent to read WAVE's output
I updated the agent to take WAVE's error and warning list as input, not just the raw HTML: instead of guessing at screen-reader issues from first principles, it now fixes the specific things WAVE flags. On the first system, the average page score went from 6.3 to 7+. Contrast errors were what held it back the longest. Those got resolved with a separate CSS-architect skill, built for exactly that kind of patch.
What's measured and what's still a projection
Two pieces of this are real elapsed time, not estimates: the core-file pass took about 2 hours, and the CSS contrast patches took about 1 hour. Together, roughly 3 hours closed about 85% of the score gap. The rest, a recursive pass through every remaining folder, now runs on its own, but I haven't clocked a full end-to-end run yet. My working estimate is about 10 hours total against a 50-to-70-hour manual audit. That's still a projection, same caveat as before, just a better-supported one.
What carries forward
The framework patches aren't specific to one system. Applying them to the next one should take less than the first 10 hours, since most of the core fixes and the contrast rules are already written. What I have now is a repeatable pattern: a skill, then an agent, then sub-agents, then an external benchmark to check the work against instead of my own judgment. That's the part I'd reuse first on the next thing I build, accessibility or not.