Back to Writing

Writeup

An ADA compliance agent that won't touch code without asking

I'd been using AI to help with an ADA compliance task I'd been assigned, because it was saving me time. It became worth building properly in a meeting, when a developer put a number on the work: about 50 hours per system. Resourcing was tight, and not every developer was cleared or available to take it on, so those hours were expensive in a way that wasn't only time.

Where the 50 hours go

Some of it is research: cross-referencing WCAG criteria to confirm what needs to change and why. Some is low-hanging fruit: every div or input acting like a button, missing ARIA labels, elements relying on visual cues a screen reader can't see. The rest is JavaScript that has to behave differently for assistive tech, not just markup with a tag swapped. None of it is hard in isolation. It's a lot of files, and few people free to work through them.

How it works

It's an AGENT.md file, GitHub Copilot's format for a persistent agent with its own instructions, running on Codex. I tested a few effort settings and landed on medium; it gave better code than low without the cost of going higher.

The first version reviewed whole folders and lost track of detail across that much code. Scoping it to one file at a time fixed that. It gives up the broader picture of the codebase and gets the close reading right, which matters more here. Within a file it looks for two things: non-semantic HTML, and JavaScript that touches the DOM in ways that affect screen readers, since a lot of the markup is generated rather than hand-written.

Why it proposes instead of applies

Every run produces two tables rather than a diff. One lists the WCAG rules it thinks apply. The other lists proposed changes, each scored for impact and feasibility by the agent. Semantic HTML fixes usually score high on both: foundational to screen-reader navigation, low-risk to implement.

The clearest thing it caught that's hard to spot manually was a focus-handling bug. When a client submitted a form, the success or error message wouldn't get read by the screen reader; sometimes the reader restarted from the top of the page, sometimes it announced nothing. The fix was a few lines of JavaScript to move focus to the message after submission. Small change, large difference for someone relying on a screen reader, and easy to miss unless you're testing with one.

The agent generates the WCAG citations from what it knows, without grounded access to the spec, so I built the table as a cross-checking reference. It hands you exactly what to go verify. From there the developer picks which proposals to implement. The output also reports a current ADA/WCAG compliance estimate and a projected one if the selected changes go in, so there's a number on the decision before any code changes.

The test run

Ten files, 50-plus individual changes. By hand, with the research and implementation both counted, roughly a 3-hour job. The agent got through it in about 30 minutes. A couple of proposed updates conflicted with existing code so the screen-reader fix didn't take effect as written; once flagged, the agent revised without much back-and-forth. That's what review is for.

I built the first working version alone and only showed it to one developer once I was confident enough to put my name on it. Before he could use it, we ran out of GitHub Copilot tokens. That's how I learned token efficiency in the output format is part of the design, not a detail.

What's next

The output format is efficient enough now that running out of tokens mid-test shouldn't happen again. The validation loop with another developer is the next thing to pick back up. The open question is the same one under every agent I build: how much stays something only I run, and how much becomes something the team can pick up.

Update: that validation loop happened. See the follow-up on teaching the agent to generate SVN patch files.