Cover of the Central Ohio Market Velocity Report, June 2026, showing the 15-county region and headline figures
The June 2026 issue. Every month it forecasts the next one, then grades itself when the data lands.

Every month I publish a Central Ohio housing report. It forecasts three things for each of 15 counties: list price, days on market, and active inventory. Then, the following month, it grades itself against what actually happened and writes the score to a file I do not get to edit after the fact.

Five months of that are now in: February through June 2026, 225 graded predictions.

I want to talk about the result, because it is not the one I wanted.

The scoreboard

Two models run side by side on the same data.

The first is a naive baseline. It is about as dumb as forecasting gets: assume next month looks like this month. There is no learning in it at all.

The second is an ensemble, which combines several models and is the kind of thing that sounds impressive in a sales conversation.

Naive baselineEnsemble
Median error4.5%12.7%
Forecasts within 5%54%12%
Got the direction right80%27%

The dumb one won. Not narrowly. It was roughly three times more accurate, and on calling the direction of a move it was better than a coin flip while the sophisticated model was considerably worse than one.

Why this happens

This is not a bug, and it is not unusual. It is what a small, slow-moving dataset does to a complex model.

Monthly county housing data gives you a handful of observations per county per year, and most months genuinely do look like the month before. A complex model has enough freedom to find patterns in that noise, and it commits to them. When the pattern was noise, and it usually is, the model is confidently wrong. The naive baseline has no freedom to be creative, so it cannot be creatively wrong.

The direction number is the one worth sitting with. A model that calls the direction of a move correctly 27% of the time is not merely useless. It is worse than guessing, which means acting on it would have cost you money on most months. If nobody had been scoring it, it would still be running, still producing confident output, and still wrong.

The part that actually matters

The scorecard page from the June 2026 report: last month's call shown in full against what actually happened, alongside the new call for July.
The scorecard page from the June 2026 report: last month's call shown in full against what actually happened, alongside the new call for July.

Nothing here would have been visible without the grading step.

Both models produce clean numbers. Both would look perfectly credible in a report. The only thing separating "this forecast is useful" from "this forecast is worse than guessing" is that something writes down the prediction, waits, and compares it to reality without being able to revise the story afterward.

That is the whole thesis of how I build. An AI system that cannot tell you when it is wrong is not a system, it is a confident stranger. Most AI tools sold to small businesses have no scoring step anywhere in them, which means nobody, including the vendor, knows whether they work.

So when someone asks what makes an AI deployment trustworthy, my answer is not the model. It is whether anyone is keeping score.

Here is what one county looks like when the scoring is working. Every number on this page is graded against a 505-county peer set, so a claim like "affordability drags the grade" has a percentile behind it rather than an adjective.

A county page from the report: Franklin County graded across affordability, taxes, jobs, schools, liquidity, and price momentum, each ranked against 505 peer counties.
A county page from the report: Franklin County graded across affordability, taxes, jobs, schools, liquidity, and price momentum, each ranked against 505 peer counties.

What I am doing about it

I am not deleting the ensemble. I am benching it until it can beat the baseline, and the baseline is what the published report leans on in the meantime. It has to earn its place against the dumbest possible competitor, which is the correct bar for any model.

That is also the standard I would want applied to anything I build for you. Ask what the simple version scores. If nobody can tell you, the sophisticated version has not been tested, it has only been shipped.

The takeaway if you are buying AI

Three questions worth asking any vendor, including me:

  1. What is the simple baseline, and what does it score? If there is no baseline, there is no evidence the clever thing is doing anything.
  2. How do you know when it is wrong? Real answers name a specific measurement and a schedule. Vague answers mean nobody is checking.
  3. What happens when it is wrong? A human should be in the loop for anything that touches a customer or a dollar.

Five months of scoring my own work produced a result that made my sophisticated model look bad, and I would rather publish that than a number I never checked.

Want this run for your business?

Book a free discovery call. We map where AI fits, show you the numbers, and point you to the right plan. No pressure, no jargon.

Book a Discovery Call →