02:07 AM. A transformer bank erupts, 47 alarms in 90 seconds, and the operator on shift has been solo for four months. The only person who ever solved this exact cascade retired in 2019. Watch the always-on GridCORTEX agent correlate the flood to one root cause in seconds, surface the 2014 event record and the retired operator's own switching notes, and walk tonight's operator through the fix, step by approved step. Diagnosis: six hours, down to twenty-two minutes. Zero customers in the dark.
| Without | With GridCORTEX | Δ |
|---|
| Without | With GridCORTEX | Δ |
|---|
It is 2:07 AM at the fictional RIVERSIDE substation, where a large transformer steps power down from 138,000 volts to 12,470 volts for neighborhood delivery. Transformer BANK-2 starts failing, and 47 alarms fire in 90 seconds. The operator on shift has worked alone for four months, and the one person who ever solved this exact failure retired in 2019. The stakes: guess wrong and the transformer is shut off, putting 8,400 customers in the dark for six hours; ride it wrong and a multi-million dollar transformer is destroyed. The screens show the trouble building: the oil at the top of the tank is at 96 degrees Celsius and climbing, the tap changer (the mechanical part inside that adjusts voltage while power flows) is hunting back and forth between two positions, its operations counter reads 96,400 against a 100,000-operation design life, and a routine oil test, which works like a blood test for transformers, shows acetylene gas (a sign of internal electrical arcing) up from 12 to 120 parts per million in six weeks. One of the cooling fans failed three weeks ago, and its repair order, #48211, is still open.
At 2:08 the always-on GridCORTEX agent finishes its analysis in 11 seconds: 46 of the 47 alarms are side effects of one underlying problem. The alarm page regroups itself into a single root-cause line with the 46 others filed beneath it as evidence. Nothing is hidden; everything is explained. The agent then searches 22 years of the utility's recorded operating data and finds a 93% match to the night of August 14, 2014: same transformer, same alarm sequence. The operator who handled it, R. Delgado, retired in 2019, but his digitized logbook entry from that night appears in the feed: freeze the tap changer, cut the load to sixty percent, and do not shut the unit off on a suspicion about the bushings (the insulated posts where wires enter the tank). The first decision point asks tonight's operator to accept the diagnosis: worn contacts inside the tap changer, made worse by the broken cooling fan, and explicitly not a bushing failure. The gas pattern and the alarm order are shown as proof. Once the operator accepts, an acoustic sensor confirms electrical arcing at the tap changer, exactly where the 2014 record said it would be.
The second decision point proposes a four-step fix, and the operator approves each step individually: reroute part of the load to a neighboring transformer; reduce BANK-2 to 60% of its rated capacity; freeze the tap changer in one position so the worn contact stops grinding; and schedule the repair crew for 7:00 AM with the parts list from the 2014 repair already attached. As the steps land, the oil temperature falls from 96 toward 91 degrees, the hunting stops, and voltage steadies on all four distribution lines. By 2:29 AM the transformer is stable, only 2 routine alarms remain, zero customers lost power, and the shift report is already drafted. The simulation ends at 2:31, twenty-four minutes after the first alarm.
If the operator ignores the recommendations, the demo plays the other version of the night. The on-call engineer is paged and is 45 minutes away. Over the phone, the best guess is a bushing failure. At 2:21 AM the transformer is shut off on that wrong guess, and 8,400 customers go dark. A test crew is sent to examine the wrong component, and restoration is projected for 8:15 AM: six hours of outage that took twenty-two minutes on the other path. The closing scorecard puts the two nights side by side: correct diagnosis in 22 minutes versus more than 6 hours, 0 customers out versus 8,400, and an operator with 4 months of experience backed by 31 years of a retired expert's knowledge, because the agent carries that knowledge on every shift.
The alarm system does its job: it announces 47 problems. But announcing is not diagnosing, and nothing on site says they are all one problem. The answer exists, split across four places: the long-term data recorder, an archived logbook, the transformer's sensor records, and an open repair order. No system connects them at 2:07 AM, so the night runs on a phone call to an engineer 45 minutes away. The transformer is shut off at 2:21 on a wrong guess, 8,400 customers sit dark for about six hours, roughly 50 megawatt-hours of electricity go undelivered, an $18,000 test crew examines the wrong part, emergency callouts cost $14,000, and the night totals $96,000, plus a transformer switched off while it was arcing inside, which shortens its life.
The software does three things. It reads the alarm flood and finds the one root cause in 11 seconds. It searches decades of past events and retired operators' digitized notes, and surfaces the 93% match from 2014 with the fix that worked. And it checks the transformer's own health record: the gas trend, the wear counter, the broken fan, and how much heat the unit can still take. The operator stays in command, accepting the diagnosis and approving each of the four response steps one at a time. Result: correct diagnosis verified in 3 minutes, zero customers interrupted, the transformer nursed safely to a planned 7:00 AM repair, and a $21,000 night instead of $96,000.
| KPI | Without GridCORTEX | With GridCORTEX | Delta |
|---|---|---|---|
| Alarms to root causehow many separate alarms the operator must interpret before knowing the one real problem | 47 alarms, no synthesis | 1 cause in 11 seconds | 46 distractions removed |
| Time to correct diagnosishow long until someone knows what is actually wrong | 6+ hours (after misdiagnosis) | 3 minutes, verified | −98% |
| Customers interruptedhomes and businesses that lost power, and for how long | 8,400 · ~6 hrs | 0 | 3.0M customer-minutes of outage avoided |
| Event SAIDI contributionSAIDI is the industry's standard reliability score: outage minutes averaged across every customer the utility serves | 21 min | 0 min | 21 minutes avoided |
| Bank tripped under arcingwhether the transformer was shut off suddenly while arcing inside, which stresses and ages it | Yes, asset stressed | No, controlled ride-through | asset protected |
| Night calloutspeople woken up and sent out in the middle of the night | Duty engineer + 2 crews | None, morning schedule | everyone slept |
| Who carried the answerwhere the knowledge that solved the problem actually lived | An archived logbook | The always-on agent | knowledge on every shift |
| Emergency callout costpremium pay for staff called in overnight | $14K | $0 | −$14K |
| Wrong-component testingthe cost of a specialist crew examining a part that was never broken | $18K bushing crew | $0 | −$18K |
| Unserved energyelectricity customers wanted but could not get, measured in megawatt-hours | ~50 MWh | 0 MWh | −50 MWh |
| OLTC repairfixing the tap changer, the worn voltage-adjusting mechanism that caused the event | Emergency, parts chase | Scheduled 07:00, 2014 parts list | planned work, not a scramble |
| Asset life impactwhat the night did to the remaining life of a multi-million dollar transformer | Trip under arcing fault | Load reduced, tap frozen | damage avoided |
| Event O&M costoperations and maintenance: the total cost of working the event | $96K | $21K | $75K saved |
| Morning paperworkthe shift report and work orders the event generates | Starts at 08:00 | Drafted by the agent at 02:31 | done before dawn |
| Experience on shift (headline tile)the years of judgment actually available to tonight's operator | 4 months, solo | 4 mo + 31 yrs | a retired expert's years, on call |
The safety mechanism here is indirect but very concrete. A great deal of what a veteran knows is hazard knowledge: this vault gasses after rain, this switch has flashed before, this section is not configured the way the drawing says. Written down and searchable, that becomes a briefing before the crew goes in. Left in one person's head, it retires with them.
Counted in units you already track:
Search time comes back to every field employee who asks a question, and structured shadowing time comes back to both the veteran and the apprentice.
The numbers we need from you to run that formula:
| Cost driver | How it is calculated, from a rate you supply |
|---|---|
| Field search time | search hours avoided x your loaded field technician rate |
| Apprentice ramp | months of ramp to competency reduced x monthly loaded cost of an apprentice x apprentices per year, using your own definition of competent |
| Retiree callback | your current hourly or daily rate for calling a retiree back as a contractor x the engagements you would avoid |
| Repeat truck rolls | trips avoided x hours per trip x crew size x your loaded crew rate, plus miles x your fleet cost per mile |
| Curation cost, which is negative | interview hours plus curation hours x the loaded rates of the veteran and the training supervisor, subtracted honestly from the case rather than left out |
You pay for the GridCORTEX capture and retrieval service, for integration into your work order history and document store, and, most importantly, for your veterans' hours in the interview chair during the last months of their career, plus a training supervisor's continuing time as editor of record. Be clear eyed about this one: the veteran's time is the dominant cost and it is the scarcest time you have.
Payback is driven by apprentice ramp time and by field search time, both of which you can measure. The avoided cost of losing forty years of knowledge is real and unquantifiable, so state it as a risk position rather than trying to put a number on it.
The safety effect is indirect for the control room and direct for the crews it dispatches. The mechanism is that when the room knows the fault is one upstream lockout rather than fourteen separate problems, it stops sending crews to chase downstream symptoms, and it stops the fatigue driven errors that come from a person reading alarm text for ten hours straight.
Counted in units you already track:
Alarm reading hours come back to the operators on the desk, and the extra bodies called in to read alarms during storms largely stop being called in.
The numbers we need from you to run that formula:
| Cost driver | How it is calculated, from a rate you supply |
|---|---|
| Storm operator labor | alarm review hours avoided x your loaded operator rate at the overtime rate that actually applies during storm staffing |
| Call-in premium | call-ins avoided x your call-in minimum hours x your loaded rate x your call-in premium |
| Event reporting | reconstruction hours avoided x your loaded rate for the analyst or engineer who writes the event report |
| Avoided misdirected dispatch | downstream chase dispatches avoided x your fully loaded cost per truck roll during storm conditions, which you supply |
| Alarm system tuning | engineering hours you would otherwise spend on manual nuisance alarm review x your loaded engineer rate |
You pay for the scoped engagement that builds and runs this, for a read integration to your advanced distribution management system (ADMS) alarm stream and connectivity model, and for your own engineers' time. The engineering time is significant and honest work: correlation quality depends on your connectivity model being right and on somebody rationalizing the standing and chattering alarms first. Budget an operator validation period through at least one full storm season before you change staffing on the strength of it.
Payback is dominated by storm overtime hours and event report preparation, both of which you already track by event. Do not build the case on SAIDI, because storm days are often excluded from it and the argument will stall in the room.
A unit carrying a developing high energy discharge fault that stays in service is the failure mode that produces a tank rupture and an oil fire, in a substation where people work. Getting those units onto an inspection or de energization list before they let go is the exposure that comes off.
Counted in units you already track:
Sample by sample interpretation time comes back to the transformer engineering group, and the engineer reviews a short ranked list instead of the whole sample set.
The numbers we need from you to run that formula:
| Cost driver | How it is calculated, from a rate you supply |
|---|---|
| Engineering interpretation labor | interpretation hours avoided x your loaded transformer engineering rate |
| Sampling program | samples avoided on units the trend shows are stable x your all in cost per sample, including the truck roll to take it |
| Avoided failure | your all in cost of a transformer failure, including replacement, oil release cleanup, and load transfer, x the share you believe consistent fleet wide trending would have caught, a share you set |
| Emergency response | emergent response events avoided x your average callout and overtime cost per event |
| Truck roll | sampling trips avoided x your fully loaded cost per truck roll |
You pay for the scoped engagement that builds and runs this, for the integration that pulls lab results and unit records together, and for engineer time to review the first fleet ranking against units they already know. If your DGA history lives in PDFs and local spreadsheets, digitizing enough of that history to establish trends is real work and it is yours to do.
Payback is usually carried by interpretation labor and by right sizing the sampling program, not by avoided failure. Build the case on the first two and let avoided failure be the argument for expanding the program.
What is this, exactly? It is AI software: intelligent agents and models built and delivered by SoftServe, running on NVIDIA accelerated computing. It is not a hardware appliance and it does not replace the systems you run today. It deploys in your own cloud or on your premises, connects read-only to your existing systems, and recommends; your people approve every action, starting in shadow mode until it earns trust.
A real-time triage service that condenses alarm storms into a short list of correlated events, each with a plain-language narrative and suggested next steps. Operators see a ranked event queue beside the native alarm list. The demo above uses synthetic data; everything below describes what the real deployment needs from your organization.
| Your system | Typical products | How we connect |
|---|---|---|
| Advanced Distribution Management System (ADMS) | Schneider EcoStruxure ADMS, GE Vernova PowerOn, Hitachi Energy Network Manager | event stream (read-only) |
| SCADA historian | AVEVA PI System, AspenTech eDNA, GE Proficy | historian mirror (one-way feed) |
| Outage Management System (OMS) | GE PowerOn, Oracle NMS, ADMS outage module | read-only API |
| Geographic Information System (GIS) | Esri ArcGIS Utility Network, GE Smallworld | scheduled file export (CSV or CIM XML) |
| Document and knowledge stores | SharePoint, procedure libraries | document upload |
| Asset / work management (EAM/CMMS) | IBM Maximo, SAP PM, Oracle WAM | database replica refreshed nightly |
| Oil test laboratory results | Doble, SDMyers, or in-house lab reports and databases | scheduled file export (CSV or CIM XML) |
Runs in your cloud account on GPU instances or on an on-premises NVIDIA server, listening to a read-only alarm stream through your existing data zone, with no link to control systems and no control actions. It starts in shadow mode; operators keep working from the native alarm list.
The Approve button you just clicked in the demo above is the real workflow. This is what it looks like on the screen of the distribution system operator during a storm in the GridCORTEX console:
Accepting the grouping never suppresses or acknowledges alarms in your ADMS; the native alarm list is untouched. GridCORTEX posts the correlated event and narrative to the ADMS event queue as an annotation through its API, and operators act through their own ADMS controls.
Nothing to enter; the trigger is automatic from the live ADMS and SCADA alarm streams.
Correlates the live alarm stream as it arrives, seconds behind real time; each event card shows the as-of timestamp of the newest alarm it includes. Pilot replicas, if used, show their cadence on the card.
A ranked event queue in the GridCORTEX console beside the native ADMS alarm list; a mobile push when a correlated event exceeds 100 alarms. The console runs in a browser beside your existing screens on day one; embedding into your own systems is a roadmap step once the read-only phase has earned trust. Approve, Modify, and Decline are all captured in an audit trail your compliance team can pull, and GridCORTEX never blocks or overrides anything in the systems you run today.
The fair question: "We have an alarm system, a historian, DGA monitoring, and an EAM full of records, what's new here?" Here's the honest answer.
When someone asks "what did it actually calculate?", this is the list. In the simulation these factors drive the storyline; in a pilot they are computed from your SCADA, historian, DGA, EAM, and operator-log archives.
Presenter's one-liner: "The alarm system saw 47 problems. The agent saw one, because it remembered a night from 2014 that only a retired operator ever solved, read the asset's whole life story in a second, and walked a four-month operator through the fix step by approved step. That's what you just watched."