Is AI Deskilling Accessibility Work?

Summary

Ken Nakata and Thomas Logan talk about AI deskilling: what happens to accessibility skills when AI does the first pass, why accessibility experts need to agree on what counts as a failure before AI can judge it, and how to keep expert human review in AI remediation workflows.

Cartoon of Thomas Logan and Ken Nakata talking at an A11y Insights desk in front of a city skyline. In an inset, Knomo, the Equal Entry mascot, scratches its head at a computer whose developer console says 3 tests passed.

In 2025, a study in The Lancet Gastroenterology & Hepatology found that experienced doctors detected fewer precancerous growths on their own after their clinics started using AI during colonoscopies. Their detection rate without AI fell from 28.4% to 22.4%. Ken sent Thomas an article about that study, and it started this conversation.

An accessibility audit is also a judgment task. A tester has to know the WCAG success criteria and decide when something passes or fails. If AI does that work first, testers may lose the ability to question what it reports. Thomas is already seeing AI-generated code fail WCAG in ways he had not seen in 20 years of audits.

Ken and Thomas discuss using AI as a second reviewer instead of the first pass, why the field needs clear pass and fail examples that AI can learn from, and how remediation tools could lock in decisions an expert has already reviewed.

How AI Deskilling Shows Up in Accessibility Work

Thomas Logan: Hello, everyone. This is Thomas Logan from Equal Entry here with Ken Nakata of Converge Accessibility. In this episode of A11yInsights, we're back and talking to you about AI deskilling: what happens to accessibility skills when AI does the work first, and how to keep human judgment in the loop. Ken, how did we decide to start talking about this topic?

Ken Nakata: Well, Thomas, as I recall, it was a fairly funny discussion. I was talking not just about accessibility, but about how AI makes us all dumber because we don't have to use the same skills that we used before. I sent you an article from Medium that quoted an article from The Lancet about doctors losing their skills when they look over X-ray or colonoscopy results to find potential tumors. They're relying more on AI because it's reliable, but they're also losing the ability to visually discern those things by themselves.

Thomas Logan: When you sent me that article, I started reading it and thinking about all the ways it applies to my day-to-day work at Equal Entry. That includes training people and often auditing technology to determine if it meets the Web Content Accessibility Guidelines. That's a cognitive task: you have to remember all the success criteria and understand when things pass or fail. There hasn't been a study of it in accessibility, but I thought that research was showing evidence across a broad set of industries that this is happening. I would say the same thing is happening in accessibility, and it's a risk for our industry, just like it is in the healthcare situation you mentioned. If we become too reliant on this tool to identify that an accessibility issue exists, we may lose the ability to independently reason and challenge that assertion.

Ken Nakata: Yeah. When you said that, I was amused because I didn't think about it in terms of accessibility at all. I was just thinking about how we're losing our cognitive abilities. But when you mentioned it, I started thinking, "Well, yeah, it definitely has an impact on my writing," because I rely on it for writing client reports and proposals, the same things most of us use AI for. This conversation has also made me think about how it could be affecting accessibility professionals who do the actual testing, because AI can do a pretty good job of going through web pages and finding problems. Maybe that's going to affect humans' ability to do the same thing.

Thomas Logan: Right. It's both positive and negative. There are so many success criteria and so many points of view to consider that, on the human side, it was always difficult to weigh all of them when doing an audit. That's where I can see the AI superpower of getting a perspective like, "Hey, you missed this," or, "You missed that," which we didn't have before. On the flip side, from my direct experience with AI over the last couple of years, and even more in the last six months with people using it as part of their accessibility strategy, I've been seeing lots of new ways to fail the criteria when you check manually. The code may pass an automated check. AI is good at getting through those. But the mistakes it makes in the code, when you check it manually, are almost all new things that I haven't seen in my twenty years of doing accessibility audits. So we have new ways to find things and new ways to fail things. And it comes back to this: if we automate all the learning away, where will we be in the future? Will we even be able to have the opinion, like I do, that says, "Hey, that's actually a failure," when it's not something that used to be seen in accessible documents?

Ken Nakata: Yeah, exactly. If we rely too much on AI, I think it also blunts our sensitivity to the issues. For instance, if you have heading structures that don't follow 1.3.1, AI can catch all those things. And if I'm just going to rely on AI, then I won't even think about the screen reader implications of that failure over time. As far as I can tell, there are two ways you could use AI: at the front end or at the back end. If you use it at the front end, you're using AI to do the initial check of a web page, and then a human comes along and double-checks the AI's work. Clearly, deskilling is going to happen in that instance. But there's another way you can use AI, which is as an auditor, a follow-up to make sure that you actually caught everything. I think that's a more promising way to use AI because you're still required to use that human touch up front. I think it can also lead to better outcomes for our clients because it makes sure that we actually caught everything. There's somebody else looking over our shoulder. But what I found really interesting about that article is that deskilling happens no matter which way you do it. It probably happens more in the first instance, where you're relying primarily on AI for the first cut. But strangely enough, it also happens in the second instance, even when AI is acting as an auditor. I found that really strange, because I would have thought your skill set would improve because somebody's double-checking your work. Instead, humans still deskill even when they have a machine looking over their shoulder.

Why Accessibility Needs Agreement on What Fails

Thomas Logan: Right. It's very interesting that even when you have those skills, there's such a tendency to start relying on the tool and feeling great about going faster and being more efficient. It's a natural part of what this technology dangles out there: "Look, now you can do more in the same amount of time." Potentially what's happening in some of those cases is that the shortcut becomes ingrained: "Oh, wow, I'm able to process a higher number of things than I could with my manual processes." That's where things get left out and forgotten. Something that's come up quite a bit on our podcast is the complexity of all the different success criteria mandated under WCAG 2.1 AA, and hopefully we have a new law coming next year with state and local governments needing to meet that. We've discussed severity: when an issue needs to be raised as a significant issue versus a minor one. My feeling is that there's really not broad consensus among accessibility experts on a lot of those topics. That's another risk with AI tools. What today's LLMs give as the decision on those points is based on what's been posted on the web and what conversations have been made public, and I don't think we've really had those broad discussions as an industry. What do you think?

Ken Nakata: Yeah, I agree. If we're going to rely on AI, and it seems inevitable that we will, then I think we need some consensus around creating a model for AI to reliably use to know which things are failures and which are not. AI is really good at figuring out the difference between something that's good and something that's bad, but we have to be the ones to tell it those differences. When you get a bunch of accessibility experts in a room, we have enough trouble figuring out what good alt text is, or what the meaning or purpose of content really is. It's such a human-based decision that we can get lost in the weeds. But I think we need some really clear examples for AI to use where we can all agree and say, "Okay, this thing violates 1.3.1, or this thing violates 1.1.1. This other example does not, even though it looks very similar." If we can come to agreement on those basic examples, I think that's what will really help in guiding these AI agents.

Thomas Logan: And ideally, that consensus can be used to build that opinion across all the different models. There are so many companies making these models, and all of them have the potential to opine on what's accessible or not. So the benefit of having human consensus is that it could also bring consensus to competing models when they analyze these topics.

Keeping Human Oversight in AI Remediation

Thomas Logan: Lastly, I want to talk about the role of human oversight in, let's say, automated remediation. We touched on it in the first question, but I think it's worth taking as a given that AI is going to be used in a lot of accessibility work today and going forward. So it's also up to the expert practitioners in our industry to have a conversation about the best ways to build human oversight into that process. Do you have any thoughts on what works well right now in an accessibility workflow with AI and humans?

Ken Nakata: Well, I think a key thing is making sure that there are different types of entities in the loop. When I give something to AI, there's a natural tendency for it to hallucinate. One of the things I've found the most benefit from is using a different model as an auditor agent. If you have a different model, such as Claude Fable, as the auditor, and Claude Sonnet or Opus as the writer, and the auditor reviews what the writer produced in a clean session, the chances of hallucination are a lot lower. The same thing extends to humans. We're basically an agent too [laughs], whether we like to think of ourselves that way or not. We have AI producing content, and of course, we're the human auditor looking over that content. We don't share the same context or the same learning as an LLM, so we can provide that quality oversight. I guess the bottom line is that when we as a society first looked at LLMs and AI, we immediately distrusted what came out. I think that's largely because we were just letting one model do one thing. So use multiple models within the AI, but also human and AI as part of that workflow. This is one of those rare situations where the more chefs you have in the kitchen, the better.

Thomas Logan: Right. One of my observations here is about being able to annotate and mark when an expert human review was performed in this sequence. From what I've seen to date, remediation is often a start-over process. The human may have made some great remediations and, as we said, missed some other things, but it's a difficult workflow if the AI can overwrite the human effort. You get stuck in a loop where you always need to fully review all of the output. Instead of one review and a round of fixes, you're now going into two, three, four, because you're always wondering whether it hallucinated or changed something else in the content that you didn't check the fourth time. So when I publish it, even though I thought I'd nailed the image descriptions, for example, it actually changed one of them along the way. It would be nice to have a way to lock in certain decisions and have a change log, or a side-by-side view of the last iteration that says, "Hey, the only things that changed in this review were these one or two things." Right now, my experience is having to start back over and test everything again. That's not saving time.

Ken Nakata: Yeah, that's definitely not a time-saver, and it also loses quality. That seems to be a really common model with PDF remediation. I haven't seen it as much in web accessibility, but PDFs are such an easy place to say, "Oh, let's just start over," because PDFs are such a discrete thing.

Thomas Logan: We'd love to hear from you. Let's continue the conversation. Please add comments, and we'll be happy to respond. Thank you so much for your time, and we'll see you in our next episode.

This transcript has been edited to remove filler words and repeated phrases while keeping each speaker's meaning.

References

Related A11yInsights Episodes

Get Help Keeping Experts in the Loop

Using AI in your accessibility testing or remediation? Equal Entry's audits pair tools with expert human review, so every finding is checked by someone who knows why it fails. Contact us to talk about your next audit.