← All posts
  • evaluation
  • experiments
  • AI

We Tested Whether Our Ranking Engine Can Actually Identify Companies Worth Selling To

A good AI sales demo is easy. Finding out whether the recommendations are actually any good is much harder, and the number we got means less than you'd think on its own.

Logan Etherton8 min read

There is a very easy way to build an impressive AI sales demo.

Pick a company. Give an LLM a bunch of information about it. Ask whether it looks like a good prospect. Put the answer on a nice-looking card.

Done.

The model will probably give you an answer that looks pretty good. It may even look great.

There is just one annoying question left.

Is it actually right?

"Looks good to me" is not a test

Some software is easy to test.

If you ask a system to extract a date from a document, you can check the date. If you ask it to add two numbers, you can check the arithmetic. If you ask whether a specific word appears in a job posting, you can search for the word.

But what is the correct answer to:

Is this company a genuinely good sales opportunity for this seller?

There isn't a database somewhere containing the objectively correct answer. There definitely isn't a column called actually_worth_calling = true.

Believe me, we checked.

The obvious solution is to ask people who know what they're doing.

Unfortunately, people disagree too.

So now we have two problems.

Welcome to AI evaluation.

The easy versionThe version we wanted
Pick examples.Define the test first.
Run the model.Hide the machine's answers from the grader.
Read the answers.Use sellers the system hasn't been tuned on.
Decide they look pretty good. Build a demo.Compare only after grading is complete.

So we treated it like an experiment

We didn't want to know whether someone could look at Kairo's recommendations and say, "Yeah, those seem pretty good."

That's useful feedback. It isn't enough to build a company on.

We wanted something that made it harder to accidentally fool ourselves. So we started running blind studies:

  1. New seller. Give Kairo a seller it has not already been tuned around.
  2. Machine ranking. Kairo identifies the companies it thinks are the strongest opportunities.
  3. Blind expert grading. Domain experts judge the candidates without seeing the machine's scores or choices.
  4. Compare afterward. Only after grading is complete do we compare the expert's judgments with the system's recommendations.

And, critically, decide how you're going to judge the result before you know what happened.

That last part matters. If you wait until the end to decide what counts as success, humans have an almost supernatural ability to discover that whatever happened was exactly what they were hoping for.

We wanted to make that harder.

Then we made the test harder

There is another wonderfully effective way to get a good result: test the stuff your system already knows how to handle.

We didn't want that either. So the studies used sellers the system had not been tuned on.

The question wasn't:

Can Kairo reproduce something we've already tuned it to do?

It was:

Can we give Kairo a new seller, let it understand what that company actually sells, and identify companies a knowledgeable person would also consider worth pursuing?

That is much closer to the real product. A new customer doesn't care whether we can do a good job for the sellers we used while building the system.

They care whether it works for them.

The number we got was 87%

Across three independent blind studies, domain experts judged 87% of Kairo's top recommendations to be good matches.

Measured agreement on top recommendations
Judged good matches by domain experts
87%
Independent blind studies
3
Graders saw the machine's output before judging?
No

We were excited about that.

We were also very careful about what it meant. Because "87%" by itself tells you almost nothing.

87% of what? Measured how? Against whom? On which companies? Did we pick the examples after seeing the results? Did the grader know which ones the machine liked? Did one seller do amazingly well while another completely fell apart?

Those questions matter more than the size of the number.

What 87% means, and what it does not
It does meanIt does not mean
Domain experts independently judged 87% of the engine's top recommendations in these studies to be good matches.87% of every company Kairo has ever considered is a good prospect.
The studies were blind, mostly used sellers the system had not been tuned on, and each study's pass/fail rules were written before its results existed.87% of recommendations will become meetings, opportunities or closed deals.
The machine is "87% accurate" at some universal definition of sales truth.
We get to stop testing.

One of the things we decided early was that a flattering overall average wasn't good enough if it hid a disaster for one seller.

If a system works beautifully for two customers and embarrasses the third, "the average looked good" is not a comforting answer to the third customer.

So we care about the individual seller results too.

Then something more interesting happened

The machine disagreed with people.

Of course it did.

Our first instinct in a disagreement was usually to ask what the model screwed up.

Sometimes the answer was: plenty.

But sometimes the human grader had made a mistake. Sometimes two smart people could reasonably disagree. And sometimes the disagreement exposed something worse:

We hadn't defined the question clearly enough.

Take something as simple as "good prospect." What does that mean?

  • A company that could theoretically buy the product?
  • A company that matches the seller's usual customer profile?
  • A company with a problem the seller solves?
  • A company showing evidence of that problem right now?
  • A company with enough money to buy?
  • A company where we can identify the person who probably owns the problem?

How many of those things have to be true? How much evidence is enough? What happens when the evidence points in opposite directions?

At some point, "just have the AI decide" becomes:

Fine. Decide what?

And that question turns out to be much harder than picking a model.

The human is not automatically the answer key either

This is an easy trap to fall into with AI evaluation: treat the human answer as obviously correct, then measure how often the model agrees with it.

Humans are useful.

Humans are also humans. We miss things. We misunderstand companies. We bring assumptions. We get tired. We occasionally look at something, feel very confident, and are simply wrong.

That doesn't mean the machine gets to overrule the person whenever they disagree. It means disagreements are valuable.

When a machine and a knowledgeable person disagree, there are several possibilities:

  • The machine is wrong.
  • The person is wrong.
  • Both answers are reasonable.
  • The evidence is too weak to support either answer confidently.
  • The question itself is underspecified.

Those are very different outcomes, and each one teaches you something different about the system.

Testing exposed product problems too

This is probably the most useful part.

Evaluation didn't just tell us whether the ranking was good. It repeatedly showed us where the product itself was making assumptions it couldn't defend.

A system might identify the right company for the wrong reason. It might surface a genuinely useful fact but connect it to the wrong seller problem. It might have enough evidence to say "interesting" but not enough to say "opportunity." It might rank a company correctly and then completely butcher the explanation shown to the salesperson.

That last one happened.

More than once.

So we now separate questions that looked, at first, like one problem:

  1. Did we find the right company?
  2. Did we find the right evidence?
  3. Did we correctly understand that evidence?
  4. Does it actually connect to what the seller does?
  5. Did we explain that connection honestly?

Those are different tests. Passing one doesn't mean you've passed the others.

And no, 87% was not permission to declare victory

It would have been very convenient if it were. We could have put the number on the website, congratulated ourselves, and moved on to easier work.

Instead, we kept looking for new ways to make the system fail.

We've changed ranking logic. We've changed what counts as useful evidence. We've changed how claims are extracted. We've changed how seller fit is checked. We've changed what the system is allowed to say. We've added cases where it refuses to produce a report at all.

And we keep grading the output.

That is not because the original studies were bad. It's because a good result answers one question:

Did this version work under these conditions?

It does not answer:

Are we finished forever?

Software doesn't work like that. AI definitely doesn't work like that.

Why bother doing all this?

Because eventually somebody has to act on the output.

A salesperson is going to trust the system enough to spend time on a company. They might research the people. They might ask for an introduction. They might write an email or prepare for a call. They might spend an hour trying to understand an account because we told them it was worth their attention.

That hour matters. "The AI thought so" is not a good enough reason to spend it.

The goal isn't to make the AI say "yes" more often. The goal is to make "yes" mean something.

That's why we test.

Not because blind studies make for a nice chart. Not because 87% looks good on a website. And not because we think a human expert is a magical source of perfect truth.

We test because the alternative is building something that sounds convincing and hoping we're right.

There is already plenty of that.

We're trying to build something worthy of people's time.


Next: The Hardest Part of Building an AI System Is Deciding What "Correct" Means