- experiments
- engineering
Even the Truth Can Answer the Wrong Question
Sometimes the data is real. Sometimes the experiment works. Sometimes the result is reproducible. And sometimes none of that matters in the least, because you answered the wrong question.
Logan Etherton9 min read
After enough systems failed in contact with reality, I started taking evidence much more seriously.
That helped.
It did not solve the problem.
Because there is another failure mode that took me even longer to understand.
Sometimes the data is real. Sometimes the experiment works. Sometimes the result is reproducible. Sometimes the answer is correct.
And sometimes none of that matters in the least.
Because you answered the wrong question.
I spent six months teaching a machine how to trade
Years ago, I became convinced that reinforcement learning could be used to build an automated trading system.
This wasn't an entirely stupid idea.
Reinforcement learning is built around an agent interacting with an environment, taking actions, receiving rewards or penalties, and gradually learning a policy that maximizes its reward. Trading certainly looks like that kind of problem. Observe the market. Take an action. Make money or lose money. Learn. Repeat.
Easy.
Except reinforcement learning is one of the more difficult areas of machine learning, and I didn't know nearly enough about it.
So I learned.
I spent probably six months studying the foundations. Then I spent weeks getting a system to actually work against real market data. No synthetic prices. No conveniently cleaned-up version of history. As I remember it, I was working with roughly 30 days of one-minute OHLCV bars.
And eventually, I got it working.
It really worked.
If the market had agreed to repeat that particular month over and over again for the rest of human history, I would be writing this from a substantially larger house.
Unfortunately, the market had not agreed to those terms.
The conditions changed. The behavior changed. Relationships changed. The system stopped working.
I hadn't taught a machine how to trade.
I had taught it how to trade one particular month.
That was an important distinction.
It was not the most embarrassing distinction I learned.
Then I invented a market
At another point in this project, I became obsessed with stationarity.
Financial data is messy. Different assets behave differently. Trends appear and disappear. Relationships change. Different parts of the trading day behave differently. The statistical properties you observe over one period do not necessarily survive the next one.
This is extremely inconvenient when you would like to build models whose assumptions prefer a much better-behaved world.
So I decided I could make the data better behaved.
I had dozens of assets. I wanted to remove trends. I wanted to improve stationarity. I wanted one preprocessing system capable of turning all of this ugly, complicated market data into something my downstream models could work with.
And I worked my ass off on it.
I tried means. I tried random walks around means with standard deviations. I tried quadratic approaches. I tried increasingly sophisticated methods.
I transformed. I detrended. I removed. I filled. I tested.
And I got better at it.
That's the part I want to emphasize.
I got better.
The code got better. My understanding got better. The transformations got more sophisticated. The results got cleaner. The data became more stationary. The trends became less problematic. The models became happier.
I kept solving problems.
And with almost every problem I solved, I moved a little farther away from reality.
It looked incredible
Eventually I had something that worked.
Not "worked" in the sense that I had waved my hands over a chart and decided it looked promising.
It produced results. Beautiful results. The kind of results that make you start thinking maybe all of those months of frustration were worth it.
Someone else thought so too. Word of what I was building made its way to someone willing to invest in it.
I was offered $890,000 to finish building the system.
Real money. A real contract. Eight hundred ninety thousand dollars.
There was just one problem.
The market I had successfully modeled had never existed.
Not for one day.
Not for one minute.
It existed in my preprocessing pipeline.
I had taken reality, found the parts that interfered with what I was trying to accomplish, removed some of them, transformed others, and invented plausible values to bridge the resulting gaps. Then I had built increasingly clever machinery to operate on the result.
And the machinery worked.
Of course it worked.
I had created a reality in which the question had a beautiful answer.
I did not feel clever
I can laugh about this now.
I was not laughing then.
I was completely demoralized. I was frustrated. I was fed up.
And the especially cruel part was that I had been getting better the entire time.
I wasn't sitting there making the same stupid mistake over and over again. I was learning. I was writing better code. I was learning better mathematics. I was finding increasingly creative solutions to increasingly difficult problems. I could point to individual decisions and explain why I had made them. I could show you the improvement.
That made the eventual realization much harder to accept.
I had spent all that effort trying to make reality become what I needed it to be.
Reality had declined.
So I built my own.
Then I proved that my system worked there.
Real data doesn't save you either
It would be comforting if the lesson were simply:
Don't play with your data.
Fair enough. I learned that one.
But that wasn't enough either.
Remember the reinforcement learning system? That one used real market data. I didn't invent the month. Those prices happened. Those trades happened. That market existed. The system really did learn something useful about it.
The result was real.
It just wasn't nearly as general as I thought it was.
That's a much harder problem.
Because now I couldn't simply divide the world into fake data and real data. I had to ask a much more annoying question:
What exactly have I established?
I had established that a particular system could perform well under a particular set of historical conditions.
That is not nothing.
It is also not:
I have taught a machine how to trade.
The distance between those two statements cost me a lot of time.
I made the same mistake with Kairo
Years later, I found a new way to do it.
At one point, I pulled the technology stacks of roughly 52,000 companies. I analyzed them looking for interesting combinations.
One pattern stood out to me. Some companies appeared to be running painfully old frontend technology while simultaneously adopting newer things: agentic systems, OpenAI APIs, and other cutting-edge tools.
I looked at that and immediately saw a story.
Engineering leadership was struggling. Executives were pushing downward for modernization. The engineering organization wasn't being given enough resources. Teams were trying to bolt shiny new capabilities onto aging foundations while struggling to keep everything else afloat.
I still think that's a pretty plausible story.
Maybe it was even true.
I don't know.
That's the point.
What did I actually know?
I knew what technology I had observed. I knew that some companies appeared to be using old technology in one part of their stack and very new technology in another.
Everything after that was interpretation.
Reasonable interpretation, perhaps. Informed interpretation. Interpretation backed by decades of my own experience as an engineer.
Still interpretation.
And because the story made sense to me, it was incredibly easy to forget where the observation ended and I began.
Plausible is dangerous
Obviously ridiculous conclusions are relatively easy to catch.
The dangerous ones are the conclusions that make perfect sense.
Especially when you're an expert. You have seen this before. You understand the domain. You recognize the pattern. You can explain exactly why a company would end up in this situation.
Maybe you're right.
But "I can explain this" and "I have established this" are not the same sentence.
That distinction matters enormously when you're about to automate what happens next.
At one company, one experienced person making one reasonable inference might be fine.
At 52,000 companies, it becomes a system.
Now the inference can be scored. Ranked. Filtered. Fed into another model. Turned into a sales recommendation. Passed to a salesperson as though the conclusion came from the underlying evidence rather than from a chain of increasingly confident interpretation.
And somewhere along that chain, the words "maybe" disappear.
The proxy becomes the thing
I think this happens constantly in complex systems because the thing we actually care about is usually difficult to measure.
So we find something we can measure. We establish that it seems related to the thing we care about. Then we optimize it.
And slowly, almost imperceptibly, the proxy becomes the thing.
We stop saying:
"This measurement appears to tell us something useful about success."
We start saying:
"This is success."
That is a very different claim.
A metric can be perfectly accurate. A measurement can be completely real. A relationship can be statistically significant. An experiment can be reproducible.
And you can still be optimizing the wrong thing.
The truth doesn't rescue you from that.
The question comes first
This is one of the reasons I have become suspicious whenever a complicated problem produces an extremely simple question.
At first, some of the questions seem almost insultingly simple. Buy low. Sell high. Couldn't be easier. And in a sufficiently constrained model, you can produce all sorts of beautiful results.
Then you meet the market.
The market contains other human beings. Irrational human beings. Emotional human beings. Human beings with different information than you. Human beings deliberately trying to profit from what you're doing. Human beings who change their behavior because people like you are changing yours.
Suddenly the simple question wasn't simple.
More importantly, it may never have been the right question.
The same thing happens elsewhere.
Which companies use old technology? Which prospects opened the message? Which accounts are showing intent? Which model has the highest accuracy?
Those questions can have perfectly correct answers.
But complex systems contain feedback loops, changing conditions, hidden variables, other people, incentives, selection effects, and things you haven't thought to measure yet.
The hard part is often not answering the question.
The hard part is figuring out whether answering it gets you any closer to what you actually wanted to know.
We automate before we've defined the problem. We measure before we've established whether the thing we're measuring represents the thing we actually care about. Then we become extraordinarily good at answering a question we should never have asked.
I still get this wrong.
I expect I always will.
The difference now is that when I get a beautiful result, I try not to immediately ask: How do I automate this?
Sometimes the answer is disappointingly small.
Good.
I've learned to be much more afraid of the enormous answer that arrived too easily.
Evidence matters.
Hypotheses matter.
Experiments matter.
Reproducible results matter.
They are the best tools I know for forcing our ideas to collide with reality. But none of them guarantee that we asked the right question.
Automation can do extraordinary things with all of them.
Extraordinary things are not necessarily useful things.
Truths can be measured.
Results can be reproduced.
Conclusions can be correct.
And none of that matters very much if we were answering the wrong question.
I once spent months building a beautiful model of a market that had never existed.
I don't want to do that again.
But I also don't expect there to be some point where I finally become smart enough to stop making mistakes like these.
That turned out to be another assumption I had to get rid of.
There is no finish line.
Next: There Is No Finish Line