- AI
- classifiers
- evaluation
The Hardest Part of Building an AI System Is Deciding What "Correct" Means
"Just have the AI decide" sounds wonderfully simple until somebody asks you to define exactly what the AI is supposed to decide.
Logan Etherton8 min read
One of the first things you learn building an AI system is that the model can be wrong.
One of the more annoying things you learn later is that sometimes you haven't defined "right" well enough to know.
This becomes a problem.
At Kairo, we analyze a large amount of public information about companies. Job postings are one of those sources. Early on, we needed to answer what sounded like an embarrassingly easy question:
Is this job posting relevant?
Great.
Relevant to what?
"You know. Relevant."
Suppose we're looking for evidence about a company's engineering organization.
Software Engineer? Yes.
Engineering Manager? Sure.
Senior Site Reliability Engineer? Obviously.
Solutions Engineer? Well.
Developer Advocate? Depends.
A four-week engineering contractor? Probably not for what we were measuring, but now we need to explain why.
An internship? What if the intern is working on exactly the system we care about?
A job with "Engineer" in the title that is actually a sales role? A sales role whose buyer is engineering? A product role embedded in an engineering organization?
Congratulations.
# A simple filter, approximately 20 minutes before becoming a constitution
is_relevant(posting):
if engineering_role:
return YES
except solutions_engineering?
except internships?
except temporary_contracts?
except customer_implementation?
except misleading_titles?
except roles that expose engineering problems
without actually being engineering roles?
return PLEASE_DEFINE_RELEVANT
And this isn't an LLM problem yet.
This is a definition problem.
We actually tried the easy version
Before building something complicated, we did the reasonable thing. We filtered job postings using keywords and rules.
The pipeline worked. The scraper worked. The database worked. The filter worked exactly as written.
Then we checked what it was doing.
One of those rules was supposed to throw out hourly retail and food-service jobs, which tell you nothing about a software company's engineering. For the titles that were ambiguous, we checked every posting the keyword path had flagged: 9,158 of them.
51.3% of them were false positives. More than half of what it threw away should have been kept.
- 51.3%
- false positives: should have been kept
- The rest
- not false positives
| Flagged for exclusion | Where it was actually posted |
|---|---|
| "Manager, Server Onboarding" | A cloud hosting company |
| "Driver Operations" | A mapping and navigation software company |
| "Podcast Host" | A venture capital firm |
Nothing was broken except the idea.
I have a special affection for this category of engineering failure. The software faithfully does exactly what you told it to do, at scale, and thereby demonstrates that what you told it to do was stupid.
So we built a classifier.
That helped.
Then the real work started.
A classifier still needs rules
An LLM can understand that "Senior Platform Reliability Engineer" and "SRE" may belong to the same conceptual family even though the strings don't match.
Wonderful.
But understanding language does not eliminate the need for a decision boundary. It gives you a much more capable system for applying that boundary.
| Clearly relevant | The boundary | Clearly noise |
|---|---|---|
| A permanent software engineering role directly building or operating the systems we're analyzing. | Everything that makes you argue. | A role that happens to contain technical words but tells us nothing useful about the function we're studying. |
The easy examples don't teach you much. The boundary examples teach you everything.
Internships. Short contracts. Misleading titles. Hybrid roles. Managers who own the function but don't perform the work. People who sell to engineers but aren't engineers. People who work in engineering organizations but own something else.
Every disagreement forces another question:
What did we actually mean when we wrote the rule?
Then the same problem appeared everywhere
Job relevance was only the beginning. We moved on to questions that sounded equally reasonable:
- Does this statement describe a real company need?
- Is this evidence current?
- Does this evidence support the claim?
- Does this company fit what the seller actually sells?
- Is this a strong enough opportunity to put in front of a salesperson?
Every one of those sentences contains a word that seems obvious until you try to make thousands of consistent decisions with it.
| The word | The case that breaks it |
|---|---|
| "Need" | A company uses Kubernetes. Is that a need? No. It tells us they use Kubernetes. |
| "Current" | Our crawler found a page today. Does that mean the underlying event happened today? Absolutely not. |
| "Evidence" | A job asks for observability experience. Does that prove the company has an observability problem? Not by itself. |
| "Good prospect" | A company perfectly matches the seller's ideal customer profile. Does that mean there's a reason to contact them right now? Also no. |
Language quietly smuggles in assumptions
This is where things get dangerous.
Consider a job posting that says a new engineering leader will "improve reliability and reduce operational friction."
What can we safely say?
We can say the company publicly assigned someone responsibility for improving reliability and reducing operational friction.
Can we say their systems are unreliable? Maybe. But the posting didn't necessarily say that.
Can we say they urgently need an observability vendor? Definitely not.
Can we say an observability seller should pay attention? Potentially, if other evidence supports the connection.
Those can sound like tiny distinctions. They stop being tiny when software makes them thousands of times.
| Step | Statement | Status |
|---|---|---|
| Source says | "Improve reliability and reduce operational friction." | Quoted |
| Supported observation | Reliability improvement is an explicit responsibility for this role. | Supported |
| Plausible inference | The company may be investing in operational reliability. | Plausible |
| Unsupported leap | The company has severe reliability problems and urgently needs our product. | Not supported |
Every sentence in that chain sounds plausible. Only some of them are actually supported.
Humans are not magically consistent either
Once we started using human graders seriously, we discovered another fun problem.
People disagree.
Sometimes one grader notices something another misses. Sometimes somebody interprets a rule differently. Sometimes the model catches an edge case correctly and the human doesn't. Sometimes the human catches the model making an extremely confident mistake.
And sometimes everybody stares at the same example and realizes the rule itself is garbage.
This is why disagreements are useful. A disagreement is not automatically a point against the machine. It can reveal a missing definition, two rules that contradict each other, a distinction we thought was obvious but never actually wrote down, or an answer that depends on information the system doesn't have.
The machine cannot consistently follow a distinction that we have not consistently defined ourselves.
This is where a lot of the engineering actually lives
People understandably focus on models. Which model are you using? How big is it? What's the benchmark score? What's the context window?
Those things matter.
But for a system like ours, an enormous amount of the work happens somewhere less glamorous. Definitions. Rubrics. Boundary cases. Counterexamples. Human grading. Disagreement analysis. Tests. More tests.
Finding out that a rule which sounded perfectly sensible on Tuesday becomes ridiculous when confronted with Example 347 on Thursday.
Then rewriting it without breaking Examples 1 through 346.
It feels less like teaching a magical intelligence and more like writing law for an extremely fast junior analyst who has read the entire internet and will exploit every ambiguity in your employee handbook.
Eventually, "correct" becomes a stack
We stopped thinking about correctness as one question. There are several.
| Layer | What has to be correct |
|---|---|
| Collection | Did we get the right source material? |
| Interpretation | Did we understand what the source actually says? |
| Claim | Did we state only what the evidence supports? |
| Connection | Does that claim actually matter to what this seller does? |
| Recommendation | Is the combined evidence strong enough to justify somebody's time? |
| Communication | Did we explain all of that without overstating it? |
A system can be right at five of those and still produce a bad answer.
You can collect the perfect source and misunderstand it. You can understand it correctly and make a claim that's too strong. You can make a perfectly supported claim that has nothing to do with the seller. You can find a real opportunity and explain it so badly that the salesperson walks away with the wrong idea.
Those are different bugs. They need different fixes.
The goal is not to eliminate judgment
I don't think we're going to write one perfect rulebook and be finished.
Real companies are messy. Language is messy. Sales is messy. There will always be judgment at the boundaries.
The goal is to make that judgment inspectable.
If Kairo makes a decision, we should be able to ask why. If a human disagrees, we should be able to determine whether the system got it wrong, the person got it wrong, or the rule needs work. If we discover a new boundary case, we should be able to add it to the tests and make the system better.
That is a form of continual improvement I trust.
Not:
We changed the prompt and the outputs feel better now.
But:
We found a specific class of failure, defined the boundary, changed the system, and checked whether we fixed it without breaking something else.
The model was never the whole product
That may be the biggest lesson for me.
A powerful model is incredibly useful. But intelligence without a definition of success is just a very capable way to produce answers you still don't know how to judge.
Building a dependable system means deciding which answers you will accept. Which ones you will reject. What evidence is enough. Where inference becomes invention.
And what "correct" means before you ask the machine to be correct.
The hardest part of building an AI system is often not getting the AI to answer. It is earning the right to say whether the answer was good.
We're still working on that.
I suspect we always will be.
Next: If AI Makes Personalization Free, Personalization Becomes Worthless