The Platform

Nothing that
goes wrong here
trips an alarm.

A monitoring platform I designed, built and run. It reads six solar installations belonging to several owners, on equipment from several manufacturers, every few minutes, around the clock, and writes down what it thinks is happening.

What it does

It watches. Then it remembers
what it saw.

A solar installation with batteries runs on a daily cycle. It generates through the day, stores what it does not immediately use, draws down overnight, and starts again. Almost everything that goes wrong with one goes wrong slowly.

A battery loses a little capacity. An array produces slightly less than it should. A load pattern quietly costs money every morning for years. None of these trip anything, because protection systems are built to catch faults, and none of these are faults yet. They are all below the threshold, and they stay there while the money goes.

The platform connects to each site through the monitoring system that site already has, and pulls a reading every few minutes. It separately pulls weather for that site's own coordinates, which is what lets it tell the difference between an installation that is underperforming and one that is doing exactly what it should on a cloudy afternoon. That single distinction is most of what the weather buys.

Every hour, an AI process reads everything collected in that hour, along with what the equipment actually is, what the owner was told the system would deliver, and its own record of what it has previously concluded about this site. It writes a short structured assessment: what is happening, why it matters, how serious it is, and how confident it is.

Those hourly assessments then feed a chain of longer summaries, each written at the close of a period the sun defines rather than the clock. Each level asks a question that only becomes answerable at that resolution. Whether the battery arrived at sunrise where it should have is a daily question. Whether it has been arriving progressively lower for three weeks is not visible at daily resolution at all.

What accumulates is a written, reasoned memory of the system. Somebody can then ask it a question in ordinary language and get an answer grounded in that memory rather than in whatever the instruments happen to be showing at that moment.

The periods are astronomic The four phases of the solar day are bounded by night midpoint, sunrise, solar noon and sunset, computed from each site's own position. They move through the year and they would move with latitude. A morning phase in Johannesburg opens around 06:30 in June and around 05:00 in December, and nothing has to be reconfigured for that to happen.
Solar noon is a boundary for a reason Generation is effectively absent overnight and in the evening, so the whole generation cycle lives inside the morning and afternoon phases. Splitting at solar noon gives one sample across the rising half and one across the falling half, which is the minimum needed to reconstruct that curve. Two days with identical total yield and identical total load can be completely different underneath, one having peaked early and faded, the other having built slowly and held. Only the four phase view separates them.
Why the app on your phone cannot see it

Looking twice a day
is not monitoring.

It is sampling, at the slowest rate that can represent anything at all, with no analysis applied to either look.

To represent something you have to observe it at least twice as often as the fastest thing you want to see in it. Below that you do not get a rougher picture. You get a wrong one that looks right, which is a different and much more expensive failure.

A solar system's fundamental period is one day. Somebody opening a monitoring app morning and evening is sampling at exactly the minimum with no margin at all. A monthly report is thirty times below the threshold. The problems are not absent from those systems. They are present at frequencies the observation rate cannot resolve, which is not the same thing and never has been.

The platform does not need human attention at the sampling rate. It needs human attention at the decision rate.

That is the difference, and it is worth being plain about what follows from it rather than listing features.

A dashboard has no memory

Close it and everything it showed is gone. Open it and it shows now. It has never shown you the pattern, it cannot show you the trend, and it does not know what it expected to see or whether what it saw matched. Memory is the whole of this platform's architecture: every period's reasoning is written down, compressed upward, and handed back to the next observation.

It knows what the system is

Not just what it is reporting. Inverters, charge controllers, strings with their orientation and tilt, individual battery units with their own commissioning dates, and metering points with a declared position in the electrical system. A reading is attached to the thing that produced it, which is what makes it possible to say which of two arrays is underperforming rather than that the total looks low.

It watches the slope, not the cliff

Protection systems are cliff edge instruments. They disconnect, alert or trip when something crosses a line. Everything on the way to that line is invisible to them. A battery loses a little capacity, so it arrives at sunrise lower, so the morning load hits it lower, so it either draws grid power or cycles near its limit, either of which accelerates the loss that started the loop. Nothing trips at any point. The system is within specification the whole way down.

It says how sure it is

Every observation carries an explicit confidence level, and low confidence is stated in the text rather than written around. A summary built from uncertain inputs carries that uncertainty upward instead of laundering it into a clean sentence. Asked something a particular site's data cannot support, it says so and explains why rather than producing a plausible number.

What it has caught

Four things nothing else
was going to find.

None of these breached a threshold. Every one of them was stable, plausible and producing no symptom at all.

Configuration

Paying for electricity it did not need to buy

A site drew a kilowatt from the grid to service a load of roughly the same size while its battery sat at eighty percent charge. Nothing here is broken. The battery is healthy, the inverter is healthy, the grid connection is working exactly as intended, and no protection system anywhere watches for this. It is a configuration deciding to buy power that was already sitting in the building, and left alone it would have gone on doing that indefinitely.

Physical model

Two arrays reported as one

A site has two arrays with materially different geometry: one on a roof facing close to north at a shallow tilt, and one on a boundary wall facing west at a very steep angle. Added together into a single figure, the second array's lower output is indistinguishable from a fault, and somebody eventually goes looking for one.

Read separately, each against its own declared orientation and the day's actual irradiance, it is exactly what that array should produce and can be reported as such. The difference is not better analysis. It is that the platform holds a model of what the installation physically is, rather than only consuming the numbers it emits.

Measurement

A reading taken at the wrong point in the system

A site has three separate metering points: one at the grid boundary, one at the inverter's supply, and one at its output. Loads exist on the property between the first two, and those loads never pass through the inverter at all. The platform had been taking grid figures from the inverter supply meter and reasoning about them as though they described the property's whole relationship with the grid. The owner experienced this as the platform talking nonsense.

What matters is that the resolution was not a bug fix. The platform was reasoning correctly from what it had. What was missing was any declaration of where in the electrical system each measurement is taken. Metering points are now modelled as components with a stated position and a parent relationship, so a reading is understood rather than assumed. The domestic version of the same condition is a water heater wired upstream of the inverter: it draws from the grid, the inverter never sees it, and the electricity bill and the monitoring system disagree with each other forever.

My own fault, found and stated

A number in the wrong unit, in the first card I read

An observation card stated that generation had reached 5.46 kilowatts one minute before a 06:45 sunrise. The true figure was 5.46 watts. The cause was that the readings reached the model as bare numbers while every other block in the same prompt stated its units, so it had a labelled correct answer and an unlabelled ambiguous one in front of it and took the unlabelled one. The corroboration was sitting in the field names: every field carrying its own unit in its name was reported correctly all morning, and the one that does not was the one reported wrongly.

The response was not to patch the card. It was to stop, audit six sites and every row of prompt content behind them, and fix the class of fault rather than the instance. That is the same audit that found a charge controller which had been silently overwriting its partner since the day that site was connected, understating generation on a shared bus for months without ever producing a figure that looked wrong. A wrong number that is stable and plausible is invisible. It stays invisible until somebody goes looking on purpose.

The boundary

It reads.
It never writes.

Nothing is installed at a site. No connection into a site is opened. No device is contacted directly.

In every case the platform authenticates to a service the site already uses and asks it for data. That is the entire extent of the contact. No setpoint is written, no configuration is changed, no control action is issued, and there is no code path through which any of that could happen.

This is architectural rather than a stage the platform happens to be at. It was a founding constraint and it holds without exception. It is also enforced one layer further out than the code, which is the part worth knowing: every AI process on the platform, from the hourly observer to each summary tier to the conversational layer, carries an explicit instruction never to suggest writing to, reconfiguring or remotely controlling an inverter or any connected device. The platform will not do it, and it will not recommend that anybody else's system does it either.

The reason is not caution for its own sake. An instrument that can change the thing it is measuring is no longer only an instrument, and everything it then reports has to be read with that in mind. Keeping the two apart is what makes the record worth anything.

Six sites, several manufacturers Five different manufacturer interfaces run across the six installations, each with its own authentication, its own vocabulary and its own idea of what a reading is, normalised into one internal structure. Adding a newly exposed field is a change made on a dashboard rather than a deployment. That work is most of what makes the platform brand agnostic rather than a tool for one make of equipment.
It records what it does not know Where a value is genuinely unavailable at a site, nothing reaches the reasoning layer at all rather than a gap being filled with an estimate. Where a summary is built from an incomplete period, it says so. The platform relies on reasoning honestly about incomplete data rather than on pretending the data is complete.
If you want to see it

It runs on live data
for real owners, so it is
not open to the public.

What I can do is show it to you. Not a rehearsed demonstration of a prepared account: the actual system, on real installations, including the parts that are still ugly.

Most people who ask are not asking about solar. They are asking because they have something in their own operation that is being decided on numbers nobody has checked, and they want to see what checking looks like when somebody has actually done it.

That conversation is the useful one and it costs nothing. Tell me roughly what you are looking at and I will tell you whether I think there is anything in it.

Ask For A Walkthrough