Wrong twice, then right

How we test a claim before it goes near a plant: the bar written down first, data nobody has opened, the result published whatever it is. One question, two failures, and the answer.

Wrong twice, stopped once, then right.

How we test a claim before we let it near a plant: write the bar down first, test on data nobody has opened, publish the result whatever it is. This is the story of one question we got wrong two times, refused to fake a third time, and then answered.

Every building and every treatment plant makes thousands of small decisions a day. A thermostat feels a room is cold and opens a radiator valve. An oxygen sensor in a tank reads low and an air valve opens. Engineers call each of these a control loop, and a large plant has thousands. Which valve acts on which reading is written in drawings that are often decades old, or only in someone's head. Get one wrong and a room is heated for nobody while another goes cold, or a tank is aerated at full cost for nothing.

How we hold ourselves to account

Before each attempt we wrote down, and committed to a public record, exactly how we would decide and what score would count as a pass. Only then did we run it, on months of data we had never opened. It is how a new medicine is tested, and it is the only way to stop yourself from seeing what you hoped to see.

Attempt one: fooled by the weather

4 of 6 answers rightthe bar was 80%; verdict falsified

We matched each valve to the reading it moves with most. In spring, two radiator valves matched the outdoor temperature instead of their rooms. The heating is designed to open up when it is cold outside, so the valves followed the weather. The method found what the controller listens to, not what it controls.

Attempt two: fooled by the plumbing

4 of 7 answers rightno better than the first; verdict falsified

So we removed every reading a valve can never move, such as the outdoor air. In winter, three valves then matched the building's shared hot-water flow instead. Every radiator really does change that flow, and in the cold months it changes it more than any one room. The method found what each valve affects, which is not the same thing as what it is there to control.

Attempt three: stopped by us

A third idea reached 14 of 20 on the data we were allowed to look at, below its own 90% bar. We stopped it there, rather than spend fresh data on a test we already knew it would fail. In these buildings the settings never change, so the data simply does not contain the answer to a blind question.

The right question

7 of 7 loops confirmed correctly · 197 of 198 wrong pairings rejectedon months never opened before; verdict supported

An engineer bringing a plant into service does not guess its loops blind. They have a list, from the drawings or the names, and they check each loop against how the plant actually behaves. That is the question we turned to: given the list, can the data confirm the real loops and reject the wrong ones? On months nobody had opened, every loop it confirmed was right, it confirmed 88% of the loops that were active, and of 198 deliberately wrong pairings, the kind a mis-wired valve produces, it rejected 197. It passed every bar we had set, and it is now in the product.

Two plants are a small sample, and two of the confirmations went through air flows that are themselves calculated from the valves; the full report says so, with every number.

Why we show you the failures

Industry has heard “AI will optimise your plant” many times. It has rarely heard a company say “here is where we were wrong, with the numbers”. A system that will one day be trusted to act inside a plant has to be able to say “I do not know” and “I was wrong”. That is not a weakness to hide. It is the part of the future worth building.

Back to the front page