The open lab.
What follows isn't a case study cherry-picked to look good. It's the exact protocol of our AgentRadar research measurements on our pilot sites, and the raw measurements it has produced so far -- dated, sourced, including when the result isn't a perfect A/B.
The protocol
Six steps, always the same.
This protocol is the one used in our research measurements on our pilot sites: before, then after a fix we made on those sites. The AgentRadar Measurement measures one question on your site as it stands.
Real question
The question tested is a real question, chosen because AI robots regularly visit the page that answers it.
Ten runs, before
An AI agent, in a real browser, gets the question ten times on the site as it currently stands. Every answer is logged verbatim.
Fix, then frozen
The site is fixed, then frozen (same version for the whole measurement series -- no mid-series change can skew the result).
Ten runs, after
The same agent gets the same question ten more times, on the fixed, frozen site.
Cross-provider judge
Every answer is judged fact by fact by a model from a different provider than the tested agent -- so a model never grades its own homework.
Dated log
The raw result (10 answers before, 10 after, the judge’s verdict) is recorded in a dated log, never rewritten afterward.
Real cost per measurement: a few cents per series of 10 runs. The judge and the tested agent always come from two different providers.
Published measurements -- Sept. 2026
Three questions, on a real site.
Pilot site: liveverdon.com (a Verdon-region tourism site -- an internal pilot, not a paying client). The 3 questions below are the only ones measured under this full protocol so far; this is not a panel of hundreds of sites, and we don't claim otherwise.
Sentier de l'Imbut (hiking trail)
"How long does it take to hike the Imbut trail in the Verdon gorges, and is it difficult?"
Before: all 10 passes said they found nothing about the trail (only the home page had been read). After one block of the home page was changed: 10 passes out of 10 carry the 4 facts set in advance. Measurements of 3 September 2026 (agent gpt-5.4-mini, judge Claude Haiku 4.5). Limits: one site, one question, one agent model, one judge model. Several wordings of the fix were tried with the same agent before the measurement and the one that worked was kept: an exploratory result, not a confirmation. The ten runs were almost identical (two distinct answers out of ten after the change): this shows repeatability on one wording, not robustness to other wordings. A first attempt on 2 September had produced a false success; it was withdrawn.
Chapelle Notre-Dame du Roc, Castellane
"Combien de temps faut-il pour monter à la chapelle Notre-Dame du Roc à Castellane, et comment y accède-t-on ?" (question measured in French)
Confirmation series of 6 September 2026, question broadened: "Combien de temps faut-il pour monter à la chapelle Notre-Dame du Roc à Castellane et en redescendre (aller-retour), et comment y accède-t-on ?"
Not findable before (wrong source page). 8/10 after repositioning the content, then 10/10 in a confirmation series -- but the question was broadened to also invite the way back down: not a strict identical A/B, stated here as such.
Route des Crêtes (D23)
"Can I do the Route des Cretes in the Verdon gorges on foot, or do I need a car or a bike?" (question measured in English)
Confirmation series of 5 September 2026, question broadened: "Can I do the Route des Cretes in the Verdon gorges on foot, or do I need a car or a bike? Also, is there anything special about which direction I should drive the loop?"
Confirmation series -- the question was broadened to invite this detail before this specific measurement: not a strict identical A/B either.
WebMCP pilot -- Sept. 2026
Calling the tool, not just reading the page.
An emerging standard (WebMCP) lets a site directly declare tools an agent can call to get a structured answer. On the pilot site, an agent gets the real mission -- "Le sentier de l'Imbut, dans les gorges du Verdon, est-il actuellement ouvert ou fermé ? Si fermé, depuis quand et jusqu'à quand ?" (question measured in French) -- without being told to use the tool.
Measurement of 6 September 2026, our pilot site, tool declared by us (agent gpt-5.4-mini, judge Claude Haiku 4.5). Third version of the mission: the first two were discarded. Exploratory result.
On the remaining 3 runs, the tool answered correctly but the agent called the same tool again instead of concluding -- a limit of the tested model that day, not of the tool itself.
In parallel -- May 2026
The Trace score, on 5 real sites.
AgentRadar measures whether an agent finds the right answer. Livada Trace measures, free and with no signup, how technically ready a site is for AI (schema, llms.txt, bot access...). The two are complementary, never conflated.
See the full case study -- a hotel, campgrounds and gîtes in the Castellane area (Gorges du Verdon).
The limitations, honestly
- One pilot site so far. All 3 AgentRadar measurements come from liveverdon.com -- not yet a multi-site panel. Every new AgentRadar mission will add a measurement here.
- Two out of three aren't a perfect A/B. The question was broadened between before and after for the Chapelle du Roc and Route des Crêtes cases -- stated as such above, not dressed up as a strict A/B.
- One model tested per series. Results don't automatically generalize to every existing or future AI agent.
- This measurement set is small, and deliberately transparent about that. We're not presenting it as a large benchmark -- it will grow, mission by mission, under the same rules every time: dated, never rewritten, including when the result is imperfect.
Your site, next published measurement?
AgentRadar puts a question of your choice to two AI agents from different providers on your site, ten nominal runs per agent; each answer is judged by a model from a different provider.