A plant should be able to explain itself.

In plain words: a plant already records thousands of readings, and almost none of them say what they are. We read the files it already exports, work out which machine each reading belongs to and which ones cannot be trusted, and do it on one machine inside the fence, without connecting to anything.

The problem, in one minute.

1

A plant is full of sensors.

A waterworks has thousands. A refinery has tens of thousands. Every few seconds, each one sends a number.

2

But a number on its own means nothing.

It arrives like this: 31TI0102.PV = 957. Is that a temperature? Of what? Is the sensor even working?

3

The answers live in a few people's heads.

And in old binders and drawings. Ask the right person on the right day and you get an answer. Ask anyone else, and nobody knows. When that person retires, the answer goes with them.

4

Think of a phone where every contact is saved only as a number.

You know they are all someone. You just do not know who. Now imagine running a waterworks, or a power station, like that.

+47 000 00 041?

+47 000 00 117?

+47 000 00 208?

5

Auge writes down what every number is.

It reads the numbers themselves, how they move and which ones move together, and whatever paperwork exists. Then it says what each one measures and which machine it belongs to, with the evidence. When it cannot tell, it asks instead of guessing.

6

Now anyone can understand the plant.

A new operator on their first night shift. An engineer planning maintenance. Or an AI that is asked to help run it. All on the plant's own computer, with nothing sent out.

Every instrument in a plant reports under a code. Auge works out what each one means, from the readings and the plant's own paperwork, and says so when it cannot.

Industry is about to run at the speed of AI. Most plants cannot yet read their own data.

AI has moved from writing text to running things. The plants that power it, cool it and supply it are the ones least ready for it: decades of readings under codes nobody can explain, and fewer people every year who still can.

AI runs on power plants

415 → 945 TWh

Data centres used about 415 terawatt hours of electricity in 2024, and are set to use around 945 by 2030, growing four times faster than everything else on the grid. Every one of those hours comes out of a power station, a substation and a grid. IEA, Energy and AI, 2025

And on water

560 → 1,200 bn litres

The data centre sector consumes over 560 billion litres of water a year, which could reach 1,200 billion by 2030. Cooling water comes from the same waterworks a town drinks from. IEA, cited in Water use in AI and Data Centres, 2025

The margin for error is shrinking

~30%

Norway loses about 30% of its drinking water to leaks, one of the highest shares in Europe, and has to bring it under a quarter by 2033. When demand rises and resources shrink, a reading nobody can name is a decision nobody can make. Norsk Vann, citing SSB

The people who know are leaving

Every year

What a plant's readings mean lives in the heads of the people who run it, and the sector is struggling to replace them. Each retirement takes a piece of the plant's dictionary with it, unless it was written down. Norsk Vann, on operator competence

The next generation of industry is not a bigger model. It is a plant whose every reading is understood, checked against physics and owned by the plant itself, so that whatever AI comes next can be trusted with it. That is the layer we build.

Your plant measures everything. It cannot tell anyone what the measurements are.

Every reading has a code for a name, like 31TT0102. That it is a temperature, on which machine, in which unit, and whether the instrument still works, is written in a drawing, a binder or the head of one person on shift. So every project that wants to use the data, a dashboard, a report, an AI, starts by building that map again by hand. It is the slowest and least loved step of every industrial data project, and it is the step we automate. What we cannot place, we list with the reason, instead of guessing.

1,469tags on the aluminium works we test against, each with a code for a name and nothing that says which machine it is on
245 hoursto name them all by hand at ten minutes a tag, which is a careful engineer with the drawings open
87%placed by the compiler with nobody involved, against 53% by matching the names the way an integrator would
33 hoursleft for a person at the same pace, and most of it is the list of tags we refused to guess at

Tested on benchmarks and real plants. Not yet at a customer. Looking for the first pilot.

The numbers here were measured on public benchmark data, on plants we simulated, and on two real Danish plants whose owners published their data, with the pass marks written down before each run. No customer runs this yet. We are looking for the first plant that will let us run on a few weeks of its exports, on its own PC, and tell us where we are wrong.

What a pilot asks of you

  • a few weeks of exports from your historian or control system, copied to a folder
  • one PC on your side of the fence
  • an afternoon of someone who knows the plant, to answer the refusal list

What you get back

  • a model of your plant you can argue with, and the evidence for every placement
  • the instruments that are dead or stuck, and the tags nobody can name that are changing
  • a written report with the same numbers we publish here: placed, refused, hours it took

One person, on purpose, until the first plant says where we are wrong.

The layer we are building is narrow enough for one person to be the best in the world at it, and the wrong size for a team until a pilot has been run. What the company needs next is written below, next to what it has.

AA

Armin Alaei

Founder, Auge Labs AS, Oslo

Working with data and machine learning since 2017. Built and cleaned the data foundation for more than 7,000 well paths on the Norwegian continental shelf, and has delivered complete AI systems into production, with technical responsibility from design to deployment and access control. Former Discipline Lead for Data and AI at Experis Norway, leading a team on deliveries across energy, health, finance, construction and public administration. Has spoken on how to evaluate AI systems at the MLOps Community and at VivaTech. Also CTO of Inveniq, a data and AI company.

The reason this company exists is the number of times that work started by asking the one person on site what the tags meant. Every project, from the beginning, again.

+47 947 99 883 · Auge Labs AS, org. 937 600 860

What the company has

  • a compiler that reads a plant from its exports and refuses what it cannot tell, with every number on this page measured against a pass mark fixed in advance
  • five experiments that failed, published, which is how you know the ones that passed are real
  • a method for using a small local model that cannot invent a number: it proposes, the code verifies, a person approves

What it still needs

  • a first plant, and the operator who will say where the model is wrong
  • an adviser who has run a water utility, because the founder has not
  • a security assessment somebody else signs

Start with one sector in one country, own the layer everyone leaves to a person, and let the plant keep what it learns.

The plan is the shape of every company that ended up owning a layer: a small market first, dominated, then the next one on the same compiler. What is measured is on the left in green. The next two years depend on a pilot and on funding. Everything after that is a direction, drawn dashed, and we will move it left only when it has been measured.

  1. 2026

    2026 · measured

    Where we are

    The compiler reads a plant from its own exports, and says what it could not read.

    • validated on two public benchmarks and two plants we simulated, pass marks fixed in advance
    • the console takes a real plant from a folder and follows the folder; the model proposes, the code verifies, a person approves
    • 7 preregistered experiments this autumn, published pass or fail: listed below the road
    • every plant in the console says on its face whether it is real, partly simulated or simulated, with its source; the real ones are two Danish sites, the Avedøre wastewater plant and an Aalborg office building, from the exports their owners published
    • runs on a plant's own Windows PC from one unzipped folder: no install, no admin rights, no internet
    • read only readers for OPC UA, PI and SQL historians, tested against servers we ran ourselves
    • one founder, no customers, grant applications being written, looking for the first pilot
  2. 2027

    2027 · planned, depends on a pilot and funding

    The first pilots

    Norwegian municipal water first, because the naming conventions there are concentrated in three or four integrators.

    • two or three pilots: municipal water, and one process plant
    • the first numbers from a working plant, published whether they pass or fail
    • the person hours it takes to understand a thousand tags, measured and published, which nobody has
    • an operations adviser who has run a water utility, and an independent security assessment
  3. 2028

    2028 · planned, depends on a pilot and funding

    A convention, not a plant

    One integrator's naming convention becomes a shipped connector, and every plant built on it is readable on the day you point at it.

    • the three Norwegian water integrators' conventions as connectors, approved rule by rule
    • the next compiler pass: structure, which station feeds which plant, scored like everything else
    • the overflow ledger and the February report composed from verified tags, with the evidence chain the regulator can read
    • ISO 27001 under way, because the tenders ask for it
  4. 2029

    2029 · ambition

    Nordic water

    Most Norwegian municipal plants readable without a connector written first; Sweden and Denmark through the same integrators.

    • the plant model linked to the asset register the sector already keeps
    • infiltration ranked per pump station, energy per cubic metre per pump, in daily use
    • a shared rule library, opt in: a rule carries which other sites approved it, never their data
  5. 2030

    2030 · ambition

    A hundred plants

    The sectors we already simulate, process and grid, on the same compiler; the first plants outside water.

    • agents acting under permission as the normal way a small plant runs its morning: the report, the alarm budget, the work order
    • the plant's own model as a file other systems import, so a plant that grows into a big platform takes its model with it
  6. 2032

    2032 · ambition

    The operating layer for the plants nobody builds platforms for

    Every plant that cannot justify a data platform and a data engineer gets a model that explains itself, on hardware it owns.

    • distribution through integrators and equipment makers, not a sales force
    • the local model keeps getting better; the verified evidence loop is what stays ours
  7. 2034

    2034 · ambition

    The plant's own dictionary is normal

    What the tags mean is a document the plant keeps, like a drawing, written by its own people and kept when they go.

    • the format published and read by vendors who compete with us, which is what makes it a format
    • regulators cite the evidence chain the way they cite a calibration record
  8. 2036

    2036 · ambition

    A plant that explains itself

    To whoever walked in this morning, in any plant in Europe, and all of it inside the fence.

    • thousands of plants, none of them sending anything anywhere
    • the question a new operator asks at seven in the morning answered by the plant, with the evidence, in their language

What we measured this autumn

Each one was written down before the data was opened, and is published whatever it came out as.

  1. Real plants, real data

    On the two real Danish plants, against the variable tables their owners published: Avedøre went from 2 of 24 tags placed to 24 of 24, the Aalborg building from 0 of 188 to 151 of 188, every quantity placed is right, and the 37 the papers list in units we have no word for (lux, seconds) are asked, not guessed; the water and building connectors were written for these plants, so a third plant is the real test; the building's heat meter obeys the first law, power = 3953 x flow x temperature difference, 5% from what water carries in l/s and W, which is the units its owners published, worked out from the readings alone; and its supply air flow is not measured at all: it is 308 x the square root of the fan's inlet pressure, to 0.06%, a calculation found from the readings, so the flow is only as good as that one pressure tap; its ventilation meter reads 1000 times the two fans' power added up, to 100.0% of its movement, which gives the units and the wiring from the readings alone

  2. Control loops

    Which valve controls which reading, read from the data alone and preregistered: FALSIFIED on held-out periods, 4 of 6 answers right against an 80% bar (refusing the unsure ones still doubled the precision of the plain method, 67% against 33%); weather-compensated heating made two radiator valves follow the outdoor air rather than their rooms; a second attempt that first drops readings an actuator cannot move was also FALSIFIED on new held-out months, 4 of 7: in winter the valves follow the shared district heating flow, which they do move, more than their rooms, so neither is in the console; what does work, and is in the console, is checking the loop list a plant's names imply: SUPPORTED on never-opened months, 7 of 7 confirmed loops right and 197 of 198 deliberately wrong pairings rejected

  3. Which placements are wrong, and proving how many

    Proving how many tags are wrong from a few checks, preregistered: PARTIAL. An engineer checks 50 random placements and gets an exact 95% ceiling on the plant's error rate; across 10 held-out plants it held every time and averaged 7.5%, tighter than the textbook bound's 8.4%, so that bound, not a model, is what is now in the console. A model that learns which placements are wrong found 72% of them in the riskiest tenth on simulated plants, against 22% for the compiler's own confidence, but on the one real plant it ranked worse than that confidence (0.62 against 1.00), so it stays out until it has learned from real plants; a bound that leans on its predictions held only 69% of the time, not 95%

  4. Attacks, and the relations a plant must keep

    Catching attacks on a real testbed (HAI) with the relations a plant must keep, a valve's position against its command, a meter against its twin, preregistered: PARTIAL. On 38 attacks never opened before, the relations alone caught 22, the strongest plain method 19, at a similar F1 (0.44 against 0.47) and with the broken relation named; the bar of 80% of attacks was missed, so we claim no detection rate; the relations now run in the console's morning report, each break named with the tags that parted

  5. Asking in your own words

    Asking the plant in your own words, Norwegian or English, preregistered: PARTIAL. A small model on the same machine reads what is asked and picks only a question type and a machine from the plant's own list; the answer is computed from the data and never written by the model. On 80 questions never seen before it read 81% right, against 20% for the keywords that shipped, but only 80% of the Norwegian ones against a bar of 85%, so we do not claim it understands Norwegian yet; it is in the console, and it shows how it read each question

  6. Laws found from data, and virtual sensors

    The laws a plant follows, found from its own data and written so a person can read them, and an estimate when a reading dies, preregistered: SUPPORTED. On a real testbed, held out, a law of at most 5 terms was only 13% less accurate than a regression on everything, and an estimate for a dead reading was 38% as wrong as the frozen last value a control system shows; afterwards we found some laws leaned on a second output of the same meter, and with those excluded one claim falls just short, which we publish; the laws and the estimates are in the console, every estimate marked

  7. Asking back when unsure

    Asking back when it is unsure what a question means, preregistered: PARTIAL. When two readers on the same machine disagree, the console offers both readings and the operator picks one, which is kept with the plant; on 60 fresh questions it asked 16 times and offered the right reading 14 times, but 6 of 11 wrong readings went by without a question, mostly where both readers made the same mistake, and we say so

The arithmetic behind the ambition, which is arithmetic and not a forecast: the published price is from NOK 450,000 per plant per year. A hundred plants is NOK 45 million a year; a thousand is NOK 450 million. Norway alone has 357 municipalities, three or four water integrators whose naming conventions cover most of them, and more process plants and substations than that. The market is every plant that cannot justify a data platform and a data engineer, which is most of them.