Agents fail compliance checks — often. Our eval proves it.

AI agents miss ~60% of accessibility criteria. With our pack, 2%. Same agents, same pages, and every run is reproducible.

v1.0 · run 2026-08-02 · 100 pages · signed
Try Bellbriar for free →
What we measure

Same agents. Same pages. 58 points of difference.

Run August 4, 2026 on 100 AI-generated pages across 13 agents, scored against WCAG 2.2 AA. Harness v1.0.

WCAG check group
Agent alone
Agent + pack
Δ
Images & text alternatives
— · —
— · —
Color contrast
— · —
— · —
Keyboard & focus order
— · —
— · —
Forms & labels
— · —
— · —
ARIA & landmarks
— · —
— · —
Overall pass-rate
[pending]
[pending]

None of the evaluated pages were used to build the pack. Commit [short hash]. Feel free to clone it and reproduce these results.

How to check the numbers.

PAGES

Start with real pages from real agents. No "eval of the eval" pages or toy fixtures.

SCORING

Axe-core handles what's mechanically checkable. The rest is scored against reference answers written by hand.

THE RECEIPTS

In-depth checks for each page. Signed, dated, diffable — watch the numbers move release to release.

STILL SKEPTICAL?

Clone it and run it yourself. Our evals are free and open source. (We also offer audits for free, to save you the work.)

A missed check is not a bug report.

It arrives as a demand letter naming criteria you missed — the same criteria used above.

AI agents ship pages faster than anyone can review them, and every page inherits whatever the agent missed at whatever rate it misses. That rate is the first column of our table ("Agent Alone"). The pack moves the rate ("Agent + Pack"). The resulting report gives your auditor something signed and dated to look at.

You've seen our numbers. What are yours?

Send a URL. We'll run the same harness on your pages and send back a report on every criterion that fails. No charge.

Request a free audit How the eval works →