← All posts
  • automation
  • engineering
  • AI

Automation Executes Your Assumptions

I have built a lot of systems that worked perfectly. They passed their tests. Then I introduced them to reality. Automation doesn't eliminate human assumptions. It executes them.

Logan Etherton10 min read

I have built a lot of systems that worked perfectly.

That is not the compliment it sounds like.

They passed their tests. They passed their backtests. They did what the architecture said they were supposed to do. The data flowed where it was supposed to flow. The numbers came out the other side.

Then I introduced them to reality.

And reality beat them senseless.

I've done this while building automated trading systems. I've done it while building network infrastructure and analytics platforms. And I did it again while building what eventually became Kairo.

The details changed.

The pattern didn't.

I would learn the domain. Build the system. Test the hell out of it. Measure everything I could think to measure. Eventually, I would have something beautiful. Then I would turn it loose on the world.

Crash and burn.

So I did what seemed obvious.

Surely the answer is more data

Maybe the system didn't understand reality because it hadn't seen enough of it.

Fine. Collect more.

I've spent months doing exactly that. More records. More sources. More history. More observations.

Run it again.

Crash and burn.

Okay. The problem must be data quality.

So now we clean it. Normalize it. Transform it. Enrich it. Reconcile it. Build ETL pipelines. Find better sources. Throw away questionable records. Fill gaps. Interpolate missing values. Create increasingly elaborate machinery to turn the messy world into something our system can understand.

Run it again.

Crash and burn.

And this is where things get psychologically weird.

Because I wasn't just failing.

I was succeeding.

The pipeline got faster. Coverage improved. Tests got better. The model became more accurate against the validation set. Infrastructure became more reliable. I fixed bugs. I reduced variance. I improved my metrics.

Sometimes I could look at the thing I had built and prove, mathematically, that it was performing extraordinarily well.

And then I would put it in front of actual reality.

It didn't work.

Which was confusing. It was also embarrassing.

Because I knew I wasn't an idiot. I had done the work. I understood the math. I could show you the tests. I could show you the results. And when those results weren't reflected in the real world, there was always another plausible explanation.

Source: This post's own words; the stage names and step 2's text are this figure's gloss.

And every time I turned one, the system got better at the things I had defined as success.

That last part took me a very long time to understand.

Surely the answer is more automation

Something isn't producing the result we want, but it takes a human being three hours to do it.

So: automate it.

Now it takes three minutes.

Great.

Except we haven't established that the three-hour activity was worth doing in the first place.

So we improve the automation. We add another data source. We make the research better. We personalize the output. We tune the prompt. We add another agent. We automate the follow-up. We automate the automation.

And pretty soon we have an absolutely magnificent machine for doing something that may not matter.

Worse, it can look like it's working.

That's important.

It can really, really look like it's working.

There are numbers everywhere. Messages sent. Accounts researched. Contacts enriched. Open rates. Read rates. Responses. Scores. Tasks completed. Pipeline stages moved.

The system is doing things. Your coworkers may tell you it's working for them. Your boss may tell you it's working for everyone else. The dashboard may be green.

And if the thing you actually care about still isn't happening, it is incredibly easy to conclude that you just haven't tuned it enough yet.

The finish line is right there. One more improvement. One more data source. One more workflow. One more model.

I know this feeling extremely well.

I did it again with Kairo

Early in the development of what eventually became Kairo, I thought I knew what I needed: more people, more companies, more technology data, more firmographics, and more relationships between all of them. If I could collect enough information and connect it properly, I could automate the whole thing.

And I had numbers telling me I was succeeding.

At one point, I showed the system to someone for the first time. I was excited. I showed him a person the system had identified and told him, with greater than 99% confidence, that he knew this person.

He didn't.

Okay, maybe not directly. Did he know somebody who knew the person? No. Somebody who knew somebody? No.

The system was spectacularly confident. Reality remained unmoved.

Then I showed him a company. The system gave me roughly 96% confidence that this company needed the services he was selling.

There was one slight problem.

They were a competitor.

Not a prospect with an unusual relationship to his business. Not a borderline case. A competitor.

I had built a machine capable of being wrong with two decimal places.

The system wasn't broken

The system wasn't broken.

The system was doing what I had asked it to do. That's why the tests passed. That's why the metrics looked good. That's why adding more data didn't solve the underlying problem.

I had defined relationships that I believed represented reality. I had defined characteristics that I believed represented a good prospect. I had defined proxies. I had defined thresholds. I had defined what counted as success.

Then I built increasingly sophisticated machinery to execute those definitions.

And it did. Beautifully.

The computer wasn't making my assumptions.

I was.

The computer was just applying them much faster than I ever could. I had gotten extremely good at automating my own assumptions.

Automation freezes human judgment

We tend to talk about automation as though it removes human judgment from a process.

It doesn't.

Very often, it freezes human judgment into the process.

Source: This post's own lists, in its own words.

Eventually nobody says:

"We hypothesized that this was a useful proxy."

They say:

"This is how the system works."

Automate a hypothesis long enough and eventually everyone forgets it was a hypothesis.

Then the system produces results consistent with the assumptions built into it. Of course it does. We measure those results using metrics derived from many of the same assumptions.

The numbers look good. The process appears to work. We have proven ourselves right.

Again. And again. And again.

This isn't because people are stupid.

This is damn near impossible to see from inside the system.

I know because I spent years doing it.

More data can make it worse

More data is useful when lack of data is actually your problem. Better data is useful when data quality is actually your problem. A better model is useful when model performance is actually your problem.

But none of them rescue a bad definition of reality.

Source: This post's own words; the title and stage names are this figure's gloss.

I know because I did all four.

More data didn't necessarily bring me closer to reality. Sometimes it just made my little version of reality look a hell of a lot more convincing.

This is not obvious

Eventually I started asking a different question.

Not: How do I automate this?

But: What do I actually know?

Not what do I strongly suspect. Not what usually happens. Not what an expert says should happen. Not what correlated with something useful in the particular data I happened to collect.

What can I actually observe?

Then: what does that observation allow me to conclude?

And just as importantly: what does it not allow me to conclude?

It is tempting to read that and think this was the embarrassingly obvious answer sitting in front of me the entire time.

It wasn't. I don't think it is obvious. It sounds obvious after somebody says it. That's different.

Learning to actually operate this way took me years.

Most good observations are small

This has been one of the harder adjustments.

I like answers. I especially like big answers.

Most observations don't give you one. They give you a small piece of reality. Then we have to decide what, if anything, we're entitled to infer from it.

That sounds limiting. I've increasingly come to see it as freeing.

The observation doesn't have to explain everything. It doesn't have to fit the theory. It doesn't have to support the conclusion I hoped to reach when I started collecting it. It can just be what it is.

And if it doesn't support the conclusion I wanted?

Fine. Reality wins.

Here's one that happened inside Kairo this week.

A company posted a job opening for a salesperson. That was the entire observation. A company was hiring someone to sell.

The system took that and wrote a fluent, confident paragraph: this company's sales team now needed to prioritize accounts based on evidence of a reason to buy. Which, if you squint, is a description of my product with their company name in front of it.

Then a second model, asked whether the conclusion actually followed from the evidence, agreed. The writing was excellent. The reasoning was coherent.

None of it was supported by the job posting.

What did the posting actually entitle me to conclude? That the company was hiring a salesperson.

That's it.

It took ten rounds of revisions to stop this. The problem wasn't that the reasoning was bad. The problem was that the reasoning was excellent!

Every version of the step that read the job posting could also see what I sold. So every version did its best to find a clever path from that observation all the way to a sale for my company.

We kept trying to make the reasoning better. Better prompts. Better instructions. Better checks. Another model judging the first model. It kept happening.

Eventually the fix was made, and it was so much simpler:

We stopped letting the step that reads the evidence know what I sell.

What the evidence says and what I sell are completely distinct. The problem was not that the model wasn't smart enough to reason across both. The problem was that it was.

Even the truth can answer the wrong question

Even when the observation is true and the result is real, there is another trap: you can still be answering the wrong question. That one gets its own post.

Automation executes as defined

That's the lesson I wish I had understood earlier.

Automation is extraordinarily powerful. But it doesn't relieve us of the responsibility of understanding what we've told it to do.

It amplifies that responsibility.

If you give a system evidence, hypotheses, experiments and reproducible measurements, automation can do extraordinary things with them.

That still doesn't mean the things it does will be useful.

Evidence doesn't guarantee we're right. Experiments don't guarantee we asked the right question. Reproducibility doesn't mean we've captured reality.

But they give reality somewhere to object.

Assumptions masquerading as facts don't.

And automation won't save us from them.

It will not complain. It will not become suspicious. It will not notice that your 99% prediction just introduced two strangers. It will happily inform you, with 96% confidence, that your customer should sell to their competitor.

And if you automate the whole process, it can make the same mistake ten thousand times before lunch.

Automation doesn't eliminate human assumptions.

It executes them.

Sometimes beautifully.

That's the problem.


Next: Even the Truth Can Answer the Wrong Question