An LLM summary on its own will always claim 100% confidence that it is correct.

The real test, and the one almost nobody asks about in a vendor demo, is simpler: can you check a specific claim in that summary against the source in under thirty seconds? If not, you don't have an insight. You have a confident guess with good formatting.

This is the part of the AI-in-analytics story a lot of people are skipping, and it's the real lesson in how Lucky Orange user Himedi is engaging with Discovery AI. Not blinding charging forward with its quick data summaries. Using the receipts behind those claims to drive teamwide alignment.

What this looks like when it goes wrong

A mid-market DTC team notices add-to-cart rates dip on mobile.

Someone pulls up an AI-generated weekly summary and reads: “mobile users are experiencing friction during checkout.”

It sounds specific enough to act on, so the team ships a checkout redesign.

Conversion doesn't move.

Three weeks later, someone finally opens the underlying sessions and finds what was really going on: a promo banner overlapping the add-to-cart button on one product template, hitting a narrow slice of paid traffic. Nothing to do with checkout.

Nobody caught it sooner because the summary read well enough that nobody felt the need to check it. That's the problem: a claim that sounds like an insight but was never built to be verified.

The industry keeps confusing a good summary with a trustworthy one

Rebecca Lee, Head of Product at Himedi, described what she originally wanted from an analytics layer, and on paper it sounds like every other AI-summarization pitch:

“We really wanted a high-level view of where frustration was building and where customers were dropping off. Not just the numbers, but a readable summary we could act on.”

Almost every analytics vendor will tell you they already do this. Feed an LLM your Google Analytics export or Shopify performance data, ask it to summarize, and get three paragraphs of plausible-sounding prose back. The output reads the same whether the model found a real pattern or invented one from noise, which is exactly what happened in the checkout example above. That's the default failure mode of asking an LLM to make sense of behavioral data without making it show its work.

A claim you can't check is a liability, not an insight

If a summary says “mobile checkout has a friction point” and there's no way to trace that back to a specific URL, segment, or session, you've got two options: trust it blindly, or go verify it manually. Neither works well. Blind trust means a bad guess quietly becomes the basis for a roadmap decision. Manual verification means you paid for the AI summary and then did the original analysis anyway.

There's a well-documented pattern in how people trust AI output: the more polished and fluent something sounds, the more confidence people extend to it, whether or not that confidence is earned. Better formatting doesn't lower the risk of an unverified claim. It hides it.

This is the same issue the AI research world has been dealing with under the name “hallucination” for a couple of years now, just showing up in analytics instead of chatbots. A model that's fluent and wrong is worse than a tool that's slow and honest, because fluency is what gets a claim into a deck unchallenged.

What actually changed at Himedi wasn't the summary. It was the citation.

Look closely at how Rebecca describes her actual workflow, because the important detail is easy to miss:

“It's already saving me time each week. I've made it my last step when pulling the prior week's analytics, and it's helped me move from raw numbers to a more educated, confident interpretation faster than before.”

Her confidence isn't coming from how well the summary reads.

It's coming from the fact that Discovery AI lists the URLs behind each point in its output. When she pulls the top three bullets to walk her team through in a weekly review, every claim is one click from the page, session, or edit that produced it.

That's not a writing feature. It's a way to check the work:

  • Claim — the summary names a specific pattern, not a general topic.

  • Source — that claim points to an exact URL, segment, or session.

  • Confirmation — someone checks it in seconds and either acts on it or throws it out, before it reaches a deck.

Take the citation away and you're left with the same fluent paragraph and none of the trust.

That's the whole difference between an AI summary worth acting on and one you're just hoping is right, the same difference that would have caught the mislabeled “friction” claim three weeks sooner.

This is the same standard now being forced onto public AI answers

Watch how Google's AI Overviews and other AI-based search results have changed over the past couple of years and you'll see this exact fight happen at a much bigger scale.

Early AI-generated answers got called out publicly for stating things confidently with no source behind them. The fix wasn't “be more careful.” It was structural: make every generated claim link back to something you can check, so the reader doesn't have to take the model's word for it.

Analytics tools are a couple of years behind that shift, the same way CRO tools were slow to add real synthesis. Most “AI insights” features in dashboards today are where AI search answers used to be: fluent, confident, and impossible to check. Discovery AI's URL-level detail does for analytics what citations did for AI search. The claim and the source show up together instead of the source being an afterthought.

The question nobody asks in a vendor demo

Every AI analytics pitch shows you a clean summary.

That's not the real test.

The real test is picking one sentence out of that summary at random and asking the vendor, on the spot, to show you the exact page, segment, or session behind it. Not run the analysis again. Just show the receipt. If that takes more than a couple of clicks, or someone has to go reconstruct the reasoning by hand, the tool is generating prose, not insight.

  • Does every claim point to a specific URL, segment, or session, not just a general topic?

  • Can someone who wasn't in the room check a claim without redoing the analysis themselves?

  • If the summary gets something wrong, is that obvious right away, or does it look just like a correct claim until someone checks?

Ask those three questions about any tool that claims to summarize your behavioral data, including this one. The answers matter more than the writing.

The Take-Back-to-Your-Desk

Open whatever AI-generated summary your team is using right now — dashboard insights, a weekly digest, anything with a "here's what happened" paragraph attached. Pick one sentence from it at random. Not the one that looks most important. A random one.

Now try to trace it back to source in under a minute. Which URL, which segment, which session is it actually describing? If you can find it fast, you're in good shape. If you're clicking through four tabs, guessing, or just re-reading the sentence more slowly hoping it'll explain itself — that's your answer. That claim has been sitting in a report, maybe informing a decision, with nobody able to check it.

Do this once a week, on a different tool each time. It takes two minutes and it'll tell you more about what you can actually trust than any vendor's accuracy claims will.

Sean McCarthy

Sean McCarthy

Director of Marketing