Back to Writing

Writeup

Letting the ADA agent scan an entire codebase on its own

The patch-file version worked, but I was still running it one file at a time: point the agent at a file, wait for the patch, move to the next. A senior dev pointed out that the per-file loop was still mine to run by hand, and didn't need to be.

From per-file requests to a multi-agent scan

I rebuilt it so the agent scans a whole codebase on its own, checking every file against the same screen-reader rules and generating a patch file with a diff for each one that needs a change. The first attempt hit token limits again, the same problem that slowed the original version. That's fixed now: I restructured it into a multi-agent system, a set of narrower agents and skills that each handle part of the job, so no single run has to hold the whole codebase in context.

The projected numbers

By hand the estimate is 50 to 60 hours a system. The full scan is projected to bring that to about 5 hours, including reviewing the diffs. That's still a target, not a measured result. The way we validate it, worked out in a dev meeting: run an existing ADA compliance tool against a system for a baseline score, run the agent, run the tool again, measure the difference.

What's still manual

The review step stays. Every diff gets read before it's applied, the same human-approval principle the first version was built around. What the full scan removes is the time spent waiting on me to point the agent at the next file, not the time spent deciding whether a fix is right.

Update: the validation loop from this article happened too. See the follow-up on scoring the agent's fixes against WAVE.