Improvements to the Web for AI Should Benefit All Users

Apple’s WebKit team recently opposed the WebMCP proposal, arguing that it creates a separate semantic layer for agents when we should instead improve the shared layers people, assistive technology, and agents already use. Jason Grigsby’s post led me to Apple’s comment, and I share the concerns raised in both. In fact, I’d made a similar case over email back in April, when I first learned about WebMCP.

Apple’s WebKit team put the heart of the issue plainly:

Although browser-integrated agents could struggle to act on interfaces built for humans, we do not think a parallel agent-facing tool layer is the right solution. When a site’s actions are hard for an agent to use, that is a gap in the page’s own semantics, and the fix, in our opinion, is to close it in the platform’s shared layers (HTML and ARIA), where the user, assistive technology, and agents all benefit.

That echoes what I proposed in April. Many of the declarative attributes being discussed for WebMCP duplicate info that’s already available in the markup. I suggested letting a developer opt a form in, then having the browser gather its labels, descriptions, and other details from what’s already there.

Browsers already derive a great deal of useful info from a well-built interface: labels, descriptions, relationships, states, and more. If we’ve done the work to make a form understandable and operable for people—including people using assistive technology—exposing that form to an agent should build on that work, not require us to encode the same info a second time for a “new” audience. What happened to the DRY principle?

What I’d like to see us do is…

  1. Write the labels, descriptions, and other semantics for the people using the interface.
  2. Let the browser draw from that existing info by default when exposing the UI to an agent.
  3. If the agent genuinely needs more specificity, provide a narrowly scoped, agent-specific override.

That last bit matters because a label or description written for a person might not give an agent enough info to interact reliably. An override should be an option of last resort, but there needs to be a path for supplying agent-specific instructions when the human-facing ones don’t suffice. It should be an enhancement, though, not a requirement.

Building on existing semantics would place less of a burden on authors and reinforce accessibility best practices. It would also reduce the likelihood that a control ends up with one description for people and another for agents—and that the two fall out of sync. The risk here is that a parallel semantic layer would lead some developers to lavish attention on agents while neglecting the people who also need to use their sites.

We should push incentives in the other direction: Writing accessible HTML should light up as much agentic functionality as possible. That would reinforce accessible authoring rather than undermine it.

There are broader opportunities here, too. I’m particularly interested in how agent capabilities could connect to web app technologies such as shortcuts and share targets. I’m also curious about the role Service Workers might play in this new era.

Those opportunities come with broader risks, of course. WebKit is right to call out the importance of consent, security, and human oversight when agents can invoke tools that act within authenticated sessions—especially in the background, with no open tab and no person present. These web platform technologies may prove useful for agents, but we’ll need to approach them carefully.

In his post, Jason connects WebKit’s position to the W3C’s Priority of Constituencies, which is the right way to frame this. His discussion is worth reading in full. Wherever agents eventually fit into that order, they shouldn’t leapfrog the people on whose behalf they’re meant to act.