Webclat / Truth
The product

AI audit services: the answer audit

Your analytics assistant answers from the data your site collected. We grade those answers against what actually happened on the site, then trace every wrong one to the tracking, taxonomy or consent fault that caused it. You get the graded set, the causes and the fixes.

What you get

One AI. One data source. Forty business questions, each graded, each with its evidence. For every wrong answer: the root cause and the fix. Three weeks from kickoff to delivery. You keep the question set and re-run it after every release.

The questions are yours. We take the ones your analysts are asked every week: revenue by channel last month, conversion by device, the drop in checkout on the day of the release, which campaign drove the return visits. Forty is the starting set. One hundred is the full set for an enterprise with several data views.

Each row has four parts: the question, the assistant's answer as it was given, the correct answer or the honest "not answerable from this data", and the evidence that decides between them. The evidence is a decoded beacon, a parsed rule, a measured session, a warehouse query. Never our opinion.

Which assistants we grade

Adobe's Data Insights Agent

In Customer Journey Analytics. It answers in natural language inside Analysis Workspace and builds the visualization for the answer. Its knowledge base is built from the component names of the data view, not from the data points. A data view that names a metric wrongly is therefore a wrong answer waiting to happen.

Grading the Data Insights Agent

Amplitude's Global Agent and specialized agents

The Global Agent runs on every page of Amplitude, queries the project, builds charts and runs investigations; every response links to the charts it used. It answers from the taxonomy your team instrumented.

Grading Amplitude's agents

Shopify Sidekick

It generates ShopifyQL queries, builds visualizations and exports reports from the store's own database. Its numbers and your Google numbers will differ, and Shopify documents why.

Grading Sidekick

Warehouse copilots

Any assistant that writes SQL over your event tables. We grade it the same way: the question, the answer, the truth, the evidence.

How the grading works

Six failure classes, from "correct and supported" to "answered when it should have refused". A confident answer to a question the data cannot support is a failure, even when the number looks plausible.

Every failed row is traced to one of three causes. Collection: the event never fired, fired twice, or fired in a frame the reading missed. Taxonomy: the name or the identifier differs between pages or between the plan and the site. Consent: the consent field and the network disagree, so the history is not what it claims to be.

The full method, with the rubric and the evidence standard, is on /how-we-grade.

Why the answers are wrong before the AI touches them

On one national retailer's site, closing the cookie notice leaves every consent category on, so the consent field in the history is a constant. 279 of 280 tag rules carry no gate. A checkout progress event fired six times in one pass. No assistant can see any of that from inside the data. It sees a complete table and answers.

Custom LLM evaluation on a benchmark tells you how the model does on someone else's data. The answer audit tells you how it does on yours, and why.

Three ways to engage

  • Answer audit. One AI, one data source, forty questions, one delivery. The starting engagement.
  • Re-audit. The same question set re-run after each tag container release, each taxonomy change, or each update the vendor ships to the assistant. Vendors ship fast: Adobe's Data Insights Agent shipped with one skill, data visualization, and announced summarization, root cause analysis and recommendation skills to follow. Each new skill is a new set of answers to grade.
  • Real-session collection. When the audit shows that the data cannot support the questions, we measure what actually fires on the site, per device and consent state, and specify the collection that would. The engineering fix is quoted per scope.

No prices on this page. Every engagement is quoted per scope after the intake on /cooperate.

What we do not do

We do not tune the model, write prompts for it, or sell it. We do not do SEO, GEO or content. We audit the data and the answers, and we fix the data.

The other audits, when the question set needs them first: website tracking audit, data quality audit, cookie audit. Worked examples across the three agents are on conversational analytics; the one case we can show is on /cases.

Questions

Do we need to give you access to the assistant?

Yes, a named, read-only seat for the audit window. We also need read access to the analytics tool and the tag container. Client data stays confidential and is never sent to an external service without your written yes.

Who writes the forty questions?

You do, with us. We start from the questions your analysts already answer and add the ones a leader would ask an assistant if they trusted it.

What if the AI refuses to answer?

A correct refusal is a pass. An answer to a question the data cannot support is a fail, whatever the number.

Can the audit be repeated?

Yes. You keep the question set. The re-audit tier runs it again after each release.

How long does it take?

Three weeks for the forty-question set, from kickoff to delivered rows.

One AI. One data source. Forty questions, graded.

Book an answer audit