Skip to content
Build vs. buy · No. 02 Market pressure

Everyone's asking about your AI.The model was never the hard part.

Anthropic's own internal analytics agent answered just 21% of questions accurately until they fixed the data foundation under it. Same model, 95%+ after. Here's what that means for the quarter your board just attached to an AI commitment.

ACCURACY EVAL · ONE MODEL, TWO RUNS Analytics agent, internal eval suite Anthropic's own agent, scored before and after the data foundation work RUN 1 · WITHOUT THE DATA FOUNDATION 21% RUN 2 · WITH IT. SAME MODEL, SAME QUESTIONS 95%+ Nothing about the model changed between the two runs. Source: Anthropic, on its own internal analytics agent POWERED BY QUERRI What moved the number Your metric definitions Permissions on every path Evals on every change WHAT DIDN'T CHANGE The model. Same weights, same questions, different answers.

Two runs of the same eval, on the same model. Everything that moved sits underneath the model, not inside it.

What the numbers actually say
21%
of analytics questions answered accurately before the data foundation work, on Anthropic's own evals.
95%+
after that work, on the same model. Nothing about the model changed.
30 points
of accuracy lost in one month when documentation went stale against a changing data model.
46%
of software teams still plan to build analytics themselves, after 63% of them already tried. Infragistics, 478 respondents.
How this starts

Nobody agrees what it is. Everybody agrees it's due.

A competitor announces an assistant. A board member forwards the press release. A customer asks about your AI story on a renewal call. Suddenly it's on the roadmap with a quarter attached, and your name is on the commitment.

Competitor, launch postTuesday, 9:14am

"Excited to announce our new AI assistant."

Board member, forwardedWednesday, 6:40am

"What's our AI story? Two portfolio companies shipped this quarter."

Customer, on a renewal callThursday, 2:05pm

"Where does AI fit on your roadmap? We're being asked internally."

You, in the roadmap reviewFriday, 11:00am

"It's on the roadmap for Q3."

What shipping fast costsWire an LLM to the database, ship a chat box, announce it. The demo works, because demos always work. Week six is when a customer makes a pricing decision on a confident answer that turns out to be wrong. AI analytics is the one feature where shipping fast and shipping wrong are the same thing wearing different sprints.
What shipping accurate earnsThe product whose answers hold up is the one your customers make decisions inside. That's the renewal conversation, the expansion, and the bake-off you win on the second demo question instead of the first. Accuracy is the feature. The assistant is the wrapper around it.
The number, with its work shown

Same model, two answers.The difference was underneath.

The most useful public data on this comes from Anthropic, the company behind Claude, writing about its own internal analytics agent.

"Without skills, Claude's ability to answer analytics questions accurately didn't exceed 21% on our evals. Adding skills gets these numbers consistently above 95%."
Same model in both cases. The difference wasn't a smarter model, it was the context and data foundations underneath it. Read the post
Read that as a buyer instead of an engineer and it says something clarifying. The part of AI analytics your team is excited to build, the model integration, isn't the part that decides whether it works. The part that decides whether it works is the unglamorous data layer underneath, and that's a build measured in months, not sprints.
The part that never makes the estimate

The demo always works.Week six is the real test.

From the same Anthropic post, on what happened after their agent was live.

"We watched our offline accuracy drift from ~95% at launch to ~65% over a month before we treated this as an engineering problem."
One month, 30 points of accuracy, at a company whose entire business is this technology. The cause was documentation going stale as their data model changed underneath the agent. Read the post
HOW A CORRECT ANSWER GOES WRONG A sprint shipsa schema change The definitionsquietly go stale The answer staysconfident A customer acts ona wrong number ~95% ACCURATE AT LAUNCH ~65% ACCURATE A MONTH LATER One month. No model change. The same agent, answering worse every sprint.

Your data model changes every sprint too. Whoever builds your AI analytics layer signs up to re-verify accuracy continuously, forever. That's the line item that never makes the build estimate, and it's the difference between shipping a feature and staffing one.

Take this into your next demo

Four questions for any AI demo.Ours included.

Ask them of every vendor on your list, and ask them of your own team before you commit a quarter. The answers separate a product that stays accurate from a chat box that demos well.

1

What happens when our schema changes?

What a good answer sounds like

A named process that re-verifies customer-facing answers on every schema change, with someone accountable for running it. "We keep an eye on it" is the same answer Anthropic had before they lost 30 points.

2

Show me your eval numbers.

What a good answer sounds like

A standing regression suite that tests answers, not just code, run on every model upgrade and every schema change, with numbers they'll put on the screen. Spot-checking isn't an eval.

3

Who owns the definitions, and how often do they change?

What a good answer sounds like

A named owner, one place the definitions live, and a process for changing them. If the answer is that the model infers definitions from the schema, then two teams' versions of "active customer" are already in production.

4

What's your wrong-answer story?

What a good answer sounds like

A specific one. Everybody who has run this in production has a story, and how they found out, from their own evals or from a customer, tells you which of those two you'd be.

We'd rather you ask us these four than not ask anyone. A vendor who can't answer them is asking you to trust a demo.

Build vs. buy, side by side

Both paths, on the accuracy questions

The build column assumes the build goes well. The documented case is that accuracy is the thing that stops going well first.

 
Build in-house
Querri, white-labeled
Time to an AI story you can announce
6 to 12 months, and only if the accuracy holds
A live instance in about a week, customer-facing in two to four
Engineering lift
A second product on the roadmap
Two tickets: provision data access, drop in the embed SDK
Who re-verifies answers when the schema changes
Your team, every sprint, forever
Querri. It's the whole job
Eval suite for answers, not just code
You build it, staff it, and keep it running
Standing regression suite, run on every model upgrade
Accuracy six months in
95% to 65% in a month is the documented case
Maintained through every schema change
First-year cost
$150K to $340K
Less than the cost of one engineer
Security and compliance
You build and audit it
SOC 2 Type II, ISO 27001, HIPAA-ready
What your engineers ship instead
This
Your actual product

The build-column cost math is shown in full, line by line, on the engineering cost page. Short version: 1.5 to 2 engineers at a $240K to $280K fully-loaded cost, using Stack Overflow's 2024 salary data, for 6 to 12 months.

When building is the right call

Build your own AI analytics layer if:

AI analysis is your product

If what you sell is being the team that gets AI answers right, that competency belongs in-house. Own the evals, staff them, and treat drift as your engineering problem.

Your data model barely moves

Stable and simple enough that accuracy drift is a minor risk. When the schema stops changing, the maintenance argument mostly goes away.

You have evaluation capacity

Data engineers with standing time for continuous evaluation, and a preference for owning the stack over owning the roadmap time.

If the pressure is "we need an AI story this year" and the team is already busy, building is how the AI story becomes next year's story.

The buy path, specifically

The accuracy work is our job.The AI story is yours.

Querri is AI-native analytics, white-labeled inside your product. Your customers ask questions in plain language and get accurate answers under your brand. The data foundation work that separates 21% from 95%+ sits with Querri, not on your roadmap.

The chat is the easy eighth

To name the objection directly: this isn't a chat wrapper. Underneath it sits the data foundation that keeps answers accurate, your metric definitions applied on every question, your permission model enforced on every path the AI can take to the data, and a standing eval suite that catches wrong answers before your customers do. On top of it ship the things launch day actually needs: dashboards, embeddable chat, exports, and a tuning interface your product team drives without filing engineering tickets.

Your definitions

Defined once by you, applied on every question, so two teams can't get two answers.

Your rules

The AI inherits your access model and can never answer beyond what the user is allowed to see.

Accuracy, maintained

A standing regression suite on answers, run through model upgrades and schema changes, so drift gets caught here.

Your brand

Ships in your UI under your name. Your sales team demos it as yours, because it is.

SOC 2 Type II ISO 27001 HIPAA-ready Less than the cost of one engineer Viewing and sharing free
Find out if your data is AI-ready today, for free
Get your free assessment
What your customers are asking

These are the answers people act on

Notice what these have in common. Every one is a number somebody spends money against, which is why a confident wrong answer is worse for your customer than no feature at all.

Field serviceThey price service plans on it
Which techs and jobs are profitable · why first-time-fix dipped · where callbacks come from
Builders and GCsIt gets billed against
Job costs over budget · change orders still unbilled · what WIP looks like this month
Multi-location operatorsIt reaches a board deck
Occupancy or census across sites · funnel conversion by location · revenue per site against plan
Questions people ask before they decide

The short answers

How long does it take to go live?

Two gates. Your white-labeled instance is up with your data connected in about a week. Customer-facing launch typically lands two to four weeks in, after tuning, which your product team does without engineering time. The AI layer, the accuracy work, and the data foundation are Querri's side of the line. Compare either gate with the quarter your board just asked about.

What does it cost?

Less than the cost of one engineer, on consumption pricing with free viewing and sharing. The build alternative starts at $150K in year one before maintenance.

We can fine-tune our own model. Doesn't that solve accuracy?

Anthropic's numbers say no. The 21% and the 95%+ came from the same model. Accuracy lives in the data foundation and context layer, and that layer needs maintenance as your schema evolves, not a one-time training run.

What happens when our data model changes?

That's exactly the drift scenario. Keeping answers accurate through schema change is the core of what Querri maintains, and it's the part that turned Anthropic's 95% into 65% in a month when documentation went stale.

The model isn't the hard part.The data foundation is.

Start with one export, or challenge us with as many sources and connectors as you'd like. In 3 to 5 business days you get a branded 4 to 5 page report: what's clean, what's broken, and the questions your customers could already be asking your product. Free, and yours to keep.