GridCORTEX Live · Scenario Demo #4  ·  ← All Demos
On this page What you are watching The business case Run this at your utility Where you see it and how you say yes

The 2 AM Alarm Synthetic Data · Simulation

02:07 AM. A transformer bank erupts, 47 alarms in 90 seconds, and the operator on shift has been solo for four months. The only person who ever solved this exact cascade retired in 2019. Watch the always-on GridCORTEX agent correlate the flood to one root cause in seconds, surface the 2014 event record and the retired operator's own switching notes, and walk tonight's operator through the fix, step by approved step. Diagnosis: six hours, down to twenty-two minutes. Zero customers in the dark.

02:07
ALARM CASCADE
⏳ DECISION POINT: TIME SLOWED
SYNTHETIC DATA
One-Line
Alarm Summary
ALARM SUMMARY, RIVERSIDE 138/12.47 kV
0 STANDING · 0 UNACK
TIME
PRI
POINT
DESCRIPTION
STATE
A raw SCADA alarm page, exactly what tonight's operator faces. Watch what GridCORTEX does to it at correlation.
Energized / closed Alarming Reduced load Isolated Back-fed via tie

Same alarms. Same shift. Different morning.

What thirty-one years of retired experience is worth at 2 AM, when the agent carries it
,
Time to correct diagnosis
,
Customers interrupted
,
Experience on shift
The Night: Operations
WithoutWith GridCORTEXΔ
Cost & Asset: This Event
WithoutWith GridCORTEXΔ
Illustrative simulation on synthetic data; asset history, costs, and timings are placeholders. In a GridCORTEX pilot, the agent is loaded with YOUR event history, YOUR operator logs, and YOUR asset twins, then backtested against real past events. See UC 3.7 "Demo and Proof Plan."
0
Active alarms
,
Root causes identified
96°C
BANK-2 top oil
0
Customers out
Intelligence Feed, always-on agent · human-in-the-loop
02:07
Diagnosis
Guided Fix
02:31
The Validated Use Cases Behind This Scenario
UC 3.7
Field Knowledge Capture
The retiring workforce's judgment (event fixes, switching notes, tribal knowledge) preserved as agent knowledge that never retires.
UC 1.2
Alarm Flood Triage
47 alarms correlated into one root-cause candidate in seconds, with an operator-facing narrative, not a blinking wall.
UC 4.2
Transformer DGA Interpreter
The asset twin behind tonight's call: gas trends, tap-changer wear, and cooling history read as one story.
187 UCs
One Framework
The 2 AM Alarm is one of 187 validated use cases across 10 solution areas and 23 utility domains.
Inside the Demo
What you are watching, and what it proves

It is 2:07 AM at the fictional RIVERSIDE substation, where a large transformer steps power down from 138,000 volts to 12,470 volts for neighborhood delivery. Transformer BANK-2 starts failing, and 47 alarms fire in 90 seconds. The operator on shift has worked alone for four months, and the one person who ever solved this exact failure retired in 2019. The stakes: guess wrong and the transformer is shut off, putting 8,400 customers in the dark for six hours; ride it wrong and a multi-million dollar transformer is destroyed. The screens show the trouble building: the oil at the top of the tank is at 96 degrees Celsius and climbing, the tap changer (the mechanical part inside that adjusts voltage while power flows) is hunting back and forth between two positions, its operations counter reads 96,400 against a 100,000-operation design life, and a routine oil test, which works like a blood test for transformers, shows acetylene gas (a sign of internal electrical arcing) up from 12 to 120 parts per million in six weeks. One of the cooling fans failed three weeks ago, and its repair order, #48211, is still open.

At 2:08 the always-on GridCORTEX agent finishes its analysis in 11 seconds: 46 of the 47 alarms are side effects of one underlying problem. The alarm page regroups itself into a single root-cause line with the 46 others filed beneath it as evidence. Nothing is hidden; everything is explained. The agent then searches 22 years of the utility's recorded operating data and finds a 93% match to the night of August 14, 2014: same transformer, same alarm sequence. The operator who handled it, R. Delgado, retired in 2019, but his digitized logbook entry from that night appears in the feed: freeze the tap changer, cut the load to sixty percent, and do not shut the unit off on a suspicion about the bushings (the insulated posts where wires enter the tank). The first decision point asks tonight's operator to accept the diagnosis: worn contacts inside the tap changer, made worse by the broken cooling fan, and explicitly not a bushing failure. The gas pattern and the alarm order are shown as proof. Once the operator accepts, an acoustic sensor confirms electrical arcing at the tap changer, exactly where the 2014 record said it would be.

The second decision point proposes a four-step fix, and the operator approves each step individually: reroute part of the load to a neighboring transformer; reduce BANK-2 to 60% of its rated capacity; freeze the tap changer in one position so the worn contact stops grinding; and schedule the repair crew for 7:00 AM with the parts list from the 2014 repair already attached. As the steps land, the oil temperature falls from 96 toward 91 degrees, the hunting stops, and voltage steadies on all four distribution lines. By 2:29 AM the transformer is stable, only 2 routine alarms remain, zero customers lost power, and the shift report is already drafted. The simulation ends at 2:31, twenty-four minutes after the first alarm.

If the operator ignores the recommendations, the demo plays the other version of the night. The on-call engineer is paged and is 45 minutes away. Over the phone, the best guess is a bushing failure. At 2:21 AM the transformer is shut off on that wrong guess, and 8,400 customers go dark. A test crew is sent to examine the wrong component, and restoration is projected for 8:15 AM: six hours of outage that took twenty-two minutes on the other path. The closing scorecard puts the two nights side by side: correct diagnosis in 22 minutes versus more than 6 hours, 0 customers out versus 8,400, and an operator with 4 months of experience backed by 31 years of a retired expert's knowledge, because the agent carries that knowledge on every shift.

Without GridCORTEX

The alarm system does its job: it announces 47 problems. But announcing is not diagnosing, and nothing on site says they are all one problem. The answer exists, split across four places: the long-term data recorder, an archived logbook, the transformer's sensor records, and an open repair order. No system connects them at 2:07 AM, so the night runs on a phone call to an engineer 45 minutes away. The transformer is shut off at 2:21 on a wrong guess, 8,400 customers sit dark for about six hours, roughly 50 megawatt-hours of electricity go undelivered, an $18,000 test crew examines the wrong part, emergency callouts cost $14,000, and the night totals $96,000, plus a transformer switched off while it was arcing inside, which shortens its life.

With GridCORTEX

The software does three things. It reads the alarm flood and finds the one root cause in 11 seconds. It searches decades of past events and retired operators' digitized notes, and surfaces the 93% match from 2014 with the fix that worked. And it checks the transformer's own health record: the gas trend, the wear counter, the broken fan, and how much heat the unit can still take. The operator stays in command, accepting the diagnosis and approving each of the four response steps one at a time. Result: correct diagnosis verified in 3 minutes, zero customers interrupted, the transformer nursed safely to a planned 7:00 AM repair, and a $21,000 night instead of $96,000.

The KPIs, side by side
KPIWithout GridCORTEXWith GridCORTEXDelta
Alarms to root causehow many separate alarms the operator must interpret before knowing the one real problem47 alarms, no synthesis1 cause in 11 seconds46 distractions removed
Time to correct diagnosishow long until someone knows what is actually wrong6+ hours (after misdiagnosis)3 minutes, verified−98%
Customers interruptedhomes and businesses that lost power, and for how long8,400 · ~6 hrs03.0M customer-minutes of outage avoided
Event SAIDI contributionSAIDI is the industry's standard reliability score: outage minutes averaged across every customer the utility serves21 min0 min21 minutes avoided
Bank tripped under arcingwhether the transformer was shut off suddenly while arcing inside, which stresses and ages itYes, asset stressedNo, controlled ride-throughasset protected
Night calloutspeople woken up and sent out in the middle of the nightDuty engineer + 2 crewsNone, morning scheduleeveryone slept
Who carried the answerwhere the knowledge that solved the problem actually livedAn archived logbookThe always-on agentknowledge on every shift
Emergency callout costpremium pay for staff called in overnight$14K$0−$14K
Wrong-component testingthe cost of a specialist crew examining a part that was never broken$18K bushing crew$0−$18K
Unserved energyelectricity customers wanted but could not get, measured in megawatt-hours~50 MWh0 MWh−50 MWh
OLTC repairfixing the tap changer, the worn voltage-adjusting mechanism that caused the eventEmergency, parts chaseScheduled 07:00, 2014 parts listplanned work, not a scramble
Asset life impactwhat the night did to the remaining life of a multi-million dollar transformerTrip under arcing faultLoad reduced, tap frozendamage avoided
Event O&M costoperations and maintenance: the total cost of working the event$96K$21K$75K saved
Morning paperworkthe shift report and work orders the event generatesStarts at 08:00Drafted by the agent at 02:31done before dawn
Experience on shift (headline tile)the years of judgment actually available to tonight's operator4 months, solo4 mo + 31 yrsa retired expert's years, on call
Live KPIs on the dashboard
Active alarmsHow many alarms are demanding attention right now. A healthy night reads 0 to 2; the 47-alarm flood is the wall the solo operator faces, and watching it fall back to 2 is the fix taking hold.
Root causes identifiedHow many underlying problems have been named. A dash means nobody knows yet, which is the dangerous state; 1 confirmed means the operator accepted the diagnosis and a sensor verified it.
BANK-2 top oilThe temperature of the oil at the top of the transformer tank. Around 90 degrees Celsius is manageable; 96 and climbing is a countdown, and falling back toward 91 means the load reduction is working.
Customers outHomes and businesses without power. Zero is the win. It stays at zero on the guided path and jumps to 8,400 if the transformer is shut off on the wrong guess; it is the single number that separates the two mornings.

The Business Case: Safety, Hours, and Cost

A utility does not buy a demo. It buys a safety exposure that goes away and a cost that goes down. Below is that case for every use case behind The 2 AM Alarm, written the way a plant manager, a safety lead, and a CFO each need to read it. Every hour and every dollar is a formula you run with your own rates and volumes. There are no vendor benchmarks in here and no invented percentages. If a number is not yours, it is not a number.
UC 3.7 Field Knowledge Capture for the Retiring Workforce

What happens today, without this

You know the retirement dates of your most knowledgeable field technicians. What they know about which vault takes water in a hard rain, which transformer hums before it fails, and which switching sequence works on the downtown network is not in the work order history and never was. Today the transfer plan is shadowing, which happens when schedules allow, plus an exit interview that is human resources paperwork. After they leave, the answer to a hard question is somebody calling a retiree at home, or a crew rediscovering the problem on site the hard way.

What it replaces or shrinks

  • The phone call to a retired employee to ask about a specific asset or location
  • Shadowing hours scheduled purely for information transfer rather than for supervised work
  • Rediscovery on site of a quirk that a veteran already knew about
  • The personal binder, the notebook, and the folder of notes that leave the building with the person
  • Shrinks the long search across work orders, drawings, and documents for an answer that a person could give in a sentence
  • The repeat trip caused by not knowing a location specific handling requirement in advance

Why it is safer

The safety mechanism here is indirect but very concrete. A great deal of what a veteran knows is hazard knowledge: this vault gasses after rain, this switch has flashed before, this section is not configured the way the drawing says. Written down and searchable, that becomes a briefing before the crew goes in. Left in one person's head, it retires with them.

Counted in units you already track:

  • Confined space entries into vaults and manholes with a known water or atmosphere history
  • Energized area entries at locations where the field configuration does not match the drawing
  • Switching operations on equipment with a known behavior that only a veteran could warn about
  • Road miles driven on second trips caused by not knowing a location specific requirement in advance

Man-hours it gives back

Search time comes back to every field employee who asks a question, and structured shadowing time comes back to both the veteran and the apprentice.

HOURS AVOIDED PER YEAR = field staff who ask questions x questions per person per week x minutes per question spent searching or calling around divided by 60 x 48 weeks, plus shadowing hours per apprentice scheduled purely for knowledge transfer x apprentices per year, minus the curation hours the training supervisor spends reviewing and publishing entries, which is a real and continuing cost.

The numbers we need from you to run that formula:

  • Field staff who would use the knowledge base and questions per person per week
  • Minutes currently spent finding an answer, including calls to retirees
  • Apprentices per year and shadowing hours per apprentice devoted to knowledge transfer
  • Technicians retiring in the next three years and hours each would sit for interviews
  • Loaded hourly rates for a field technician, an apprentice, and the training supervisor

Where the dollars come from

Cost driverHow it is calculated, from a rate you supply
Field search timesearch hours avoided x your loaded field technician rate
Apprentice rampmonths of ramp to competency reduced x monthly loaded cost of an apprentice x apprentices per year, using your own definition of competent
Retiree callbackyour current hourly or daily rate for calling a retiree back as a contractor x the engagements you would avoid
Repeat truck rollstrips avoided x hours per trip x crew size x your loaded crew rate, plus miles x your fleet cost per mile
Curation cost, which is negativeinterview hours plus curation hours x the loaded rates of the veteran and the training supervisor, subtracted honestly from the case rather than left out

Reliability and maintenance

Reliability
This touches reliability through avoided repeat trips and avoided misoperations on equipment with known quirks, which affects mean time to repair (MTTR) and, in the misoperation case, SAIFI (system average interruption frequency index). It is a modest and indirect effect and you should not let anyone size it as a large one.
Maintenance
Repeated observations captured from veterans, such as a class of transformer that hums before it fails or a manhole that always needs pumping before work, are exactly the inputs that let you tighten a preventive maintenance interval to condition rather than to calendar. Those patterns exist in your people's heads today and in no queryable system.

What else it moves

WorkforceThis is the primary case. It shortens the ramp for new technicians and it means a retirement is a staffing problem rather than a capability loss.
ComplianceTraining and qualification programs need documented, sourced content. Entries that cite a recorded clip and a work order are auditable in a way that oral tradition is not.
Insurance and riskHazard knowledge that exists only in one person's memory is an uninsurable single point of failure, and it is the kind of gap that gets named in an incident investigation.

What it costs you, stated honestly

You pay for the GridCORTEX capture and retrieval service, for integration into your work order history and document store, and, most importantly, for your veterans' hours in the interview chair during the last months of their career, plus a training supervisor's continuing time as editor of record. Be clear eyed about this one: the veteran's time is the dominant cost and it is the scarcest time you have.

How to build the payback case

Payback is driven by apprentice ramp time and by field search time, both of which you can measure. The avoided cost of losing forty years of knowledge is real and unquantifiable, so state it as a risk position rather than trying to put a number on it.

This is a planning model built from your headcount, your retirement schedule, and your rates, not a vendor claim. Measure question answering time and apprentice ramp before and after the first cohort of interviews, then re-run it.
UC 1.2 Alarm Flood Triage Assistant

What happens today, without this

In a storm or after a transmission fault, the distribution system operator watches an alarm list scroll faster than anyone can read it. They acknowledge alarms one at a time, mostly duplicates of the same underlying event, while trying to work out which upstream device actually operated. Supervisors call in a second or third operator whose whole job for the next several hours is reading alarms. Afterward, somebody spends days reconstructing the alarm timeline for the event report.

What it replaces or shrinks

  • Line by line reading and acknowledging of alarms that all trace to one upstream operation
  • Hand correlating a cluster of downstream alarms back to the device that actually operated
  • Calling in extra operators purely to have more eyes on the alarm list
  • Searching for the standard operating procedure that applies while the list keeps scrolling
  • Shrinks the post-event reconstruction of the alarm timeline for internal and regulatory event reports

Why it is safer

The safety effect is indirect for the control room and direct for the crews it dispatches. The mechanism is that when the room knows the fault is one upstream lockout rather than fourteen separate problems, it stops sending crews to chase downstream symptoms, and it stops the fatigue driven errors that come from a person reading alarm text for ten hours straight.

Counted in units you already track:

  • Road miles driven by crews dispatched to downstream symptoms of a single upstream fault
  • Night driving hours during storm response, which is when driving exposure is highest
  • Switching operations attempted before the true fault location is understood
  • Energized area entries by crews sent to a site that turns out not to be the problem

Man-hours it gives back

Alarm reading hours come back to the operators on the desk, and the extra bodies called in to read alarms during storms largely stop being called in.

HOURS AVOIDED PER YEAR = alarms per storm hour x storm hours per year x seconds per alarm the operator currently spends reviewing and acknowledging, plus extra operators called in per storm x hours per call-in x storms per year, plus event report reconstruction hours per event x events per year, minus the time operators still spend reviewing each correlated event group before acting.

The numbers we need from you to run that formula:

  • Alarms received per hour during a typical storm and total storm hours per year
  • Seconds an operator currently spends per alarm, from your own alarm system statistics
  • Extra operators called in per storm event and hours per call-in, with your overtime premium
  • Reportable events per year and hours spent reconstructing the alarm timeline for each
  • Loaded hourly rate for a distribution system operator, straight time and overtime

Where the dollars come from

Cost driverHow it is calculated, from a rate you supply
Storm operator laboralarm review hours avoided x your loaded operator rate at the overtime rate that actually applies during storm staffing
Call-in premiumcall-ins avoided x your call-in minimum hours x your loaded rate x your call-in premium
Event reportingreconstruction hours avoided x your loaded rate for the analyst or engineer who writes the event report
Avoided misdirected dispatchdownstream chase dispatches avoided x your fully loaded cost per truck roll during storm conditions, which you supply
Alarm system tuningengineering hours you would otherwise spend on manual nuisance alarm review x your loaded engineer rate

Reliability and maintenance

Reliability
This touches restoration time on multi-device events by getting the room to the true upstream device sooner, which shows up in CAIDI, the customer average interruption duration index, and in your mean time to repair. Note the honest limit: if your reported system average interruption duration index (SAIDI) excludes major event days under your own reporting practice, most of this benefit will sit outside your reported number even though the restoration hours are real and the crews are real.
Maintenance
The same points that generate chattering, duplicated, or standing alarms show up repeatedly in the correlation, which turns a vague complaint about a noisy alarm list into a named list of instrument, communication, and setting problems your engineers can schedule and fix.

What else it moves

WorkforceAlarm reading during a long storm is the least skilled and most exhausting thing an experienced operator does. Returning those hours to judgment work is also a fatigue and retention argument, not only a cost one.
CustomerGetting to the real fault sooner is what makes the first estimated restoration time you publish closer to the truth, which is most of what drives storm complaint volume.
ComplianceThe correlated event record, with each contributing alarm attached, is a cleaner starting point for regulatory event reporting than a raw alarm export.

What it costs you, stated honestly

You pay for the scoped engagement that builds and runs this, for a read integration to your advanced distribution management system (ADMS) alarm stream and connectivity model, and for your own engineers' time. The engineering time is significant and honest work: correlation quality depends on your connectivity model being right and on somebody rationalizing the standing and chattering alarms first. Budget an operator validation period through at least one full storm season before you change staffing on the strength of it.

How to build the payback case

Payback is dominated by storm overtime hours and event report preparation, both of which you already track by event. Do not build the case on SAIDI, because storm days are often excluded from it and the argument will stall in the room.

These are planning models driven by your alarm rates, your storm hours, and your overtime rates, not vendor claims. Re-run them after your first storm season with the actual counts, including how many correlations your operators overrode.
UC 4.2 Transformer Fleet DGA Interpreter

What happens today, without this

Oil samples come back from the lab as PDFs or spreadsheets, unit by unit. A transformer engineer opens them one at a time, applies Duval or Rogers ratios in a locally built spreadsheet, and compares the new numbers against whatever the previous sample said. With hundreds of units the review happens in batches whenever someone has time, and the units with a slow upward trend are the ones that fall through, because no single sample looks alarming on its own. When the engineer who does this interpretation is on vacation, nothing gets read.

What it replaces or shrinks

  • Opening and interpreting each dissolved gas analysis (DGA) result one spreadsheet at a time
  • Hand calculation of gas ratios and fault zone placement per sample
  • Manual comparison of a new sample against prior samples to spot a trend
  • The local spreadsheet holding fleet DGA history outside the asset system
  • Shrinks the periodic fleet review meeting where engineers rank which units need attention

Why it is safer

A unit carrying a developing high energy discharge fault that stays in service is the failure mode that produces a tank rupture and an oil fire, in a substation where people work. Getting those units onto an inspection or de energization list before they let go is the exposure that comes off.

Counted in units you already track:

  • Energized area entries near units carrying an undiagnosed internal fault
  • Hot work permits issued for emergency tank, bushing, and radiator repair after a failure
  • Switching operations performed as an emergency de energization rather than a planned one
  • Night driving hours for after hours response to a failed unit

Man-hours it gives back

Sample by sample interpretation time comes back to the transformer engineering group, and the engineer reviews a short ranked list instead of the whole sample set.

HOURS AVOIDED PER YEAR = samples per year x engineer minutes per sample for interpretation and trend comparison, plus fleet review meetings per year x attendees x meeting hours, minus the review minutes the engineer still spends confirming each flagged unit and setting its disposition.

The numbers we need from you to run that formula:

  • DGA samples taken per year across the fleet and how many units they cover
  • Engineer minutes currently spent interpreting and trending one sample
  • Number and length of fleet review meetings and who attends
  • Loaded hourly rate for a transformer engineer
  • Your current sampling interval by unit class, since the schedule itself is a candidate for change

Where the dollars come from

Cost driverHow it is calculated, from a rate you supply
Engineering interpretation laborinterpretation hours avoided x your loaded transformer engineering rate
Sampling programsamples avoided on units the trend shows are stable x your all in cost per sample, including the truck roll to take it
Avoided failureyour all in cost of a transformer failure, including replacement, oil release cleanup, and load transfer, x the share you believe consistent fleet wide trending would have caught, a share you set
Emergency responseemergent response events avoided x your average callout and overtime cost per event
Truck rollsampling trips avoided x your fully loaded cost per truck roll

Reliability and maintenance

Reliability
The failures this catches are the ones that take a unit out with no warning during a heat wave, which is exactly when you have the least room to back feed, so the SAIFI and CAIDI effect lands in the worst hours of your year. It only reaches the failure population that gasses before it fails, which is a large share of internal faults but not all of them and not the bushing and tap changer mechanical failures.
Maintenance
Consistent interpretation across the whole fleet is what lets you move sampling from a calendar interval to condition, sampling the drifting units more often and the stable ones less. A unit with a rising acetylene trend gets an internal inspection booked into a planned outage instead of an emergency pull.

What else it moves

ComplianceEach disposition carries the gas signature and trend it was based on, so the record of why a unit was left in service is defensible after the fact.
WorkforceDGA interpretation is a scarce skill that usually sits with one or two people. Applying it consistently to every sample means the fleet does not go blind when they retire.
EnvironmentA transformer that fails internally can release oil. Fewer in service failures means fewer reportable oil release events to remediate.

What it costs you, stated honestly

You pay for the scoped engagement that builds and runs this, for the integration that pulls lab results and unit records together, and for engineer time to review the first fleet ranking against units they already know. If your DGA history lives in PDFs and local spreadsheets, digitizing enough of that history to establish trends is real work and it is yours to do.

How to build the payback case

Payback is usually carried by interpretation labor and by right sizing the sampling program, not by avoided failure. Build the case on the first two and let avoided failure be the argument for expanding the program.

These are planning models built from your sample volumes, your rates, and your own failure history, not vendor claims. Re run them with a year of actuals before you change the sampling program.
Each of these opens in full on the use case page, alongside the integration plan, the data ask, the path to production, and the operator console. Open the use case library.
For Your Architects and Data Owners
Run this at your utility

What is this, exactly? It is AI software: intelligent agents and models built and delivered by SoftServe, running on NVIDIA accelerated computing. It is not a hardware appliance and it does not replace the systems you run today. It deploys in your own cloud or on your premises, connects read-only to your existing systems, and recommends; your people approve every action, starting in shadow mode until it earns trust.

A real-time triage service that condenses alarm storms into a short list of correlated events, each with a plain-language narrative and suggested next steps. Operators see a ranked event queue beside the native alarm list. The demo above uses synthetic data; everything below describes what the real deployment needs from your organization.

Systems it connects to

Your systemTypical productsHow we connect
Advanced Distribution Management System (ADMS)Schneider EcoStruxure ADMS, GE Vernova PowerOn, Hitachi Energy Network Managerevent stream (read-only)
SCADA historianAVEVA PI System, AspenTech eDNA, GE Proficyhistorian mirror (one-way feed)
Outage Management System (OMS)GE PowerOn, Oracle NMS, ADMS outage moduleread-only API
Geographic Information System (GIS)Esri ArcGIS Utility Network, GE Smallworldscheduled file export (CSV or CIM XML)
Document and knowledge storesSharePoint, procedure librariesdocument upload
Asset / work management (EAM/CMMS)IBM Maximo, SAP PM, Oracle WAMdatabase replica refreshed nightly
Oil test laboratory resultsDoble, SDMyers, or in-house lab reports and databasesscheduled file export (CSV or CIM XML)

Data it needs from you

How it runs on your systems

Runs in your cloud account on GPU instances or on an on-premises NVIDIA server, listening to a read-only alarm stream through your existing data zone, with no link to control systems and no control actions. It starts in shadow mode; operators keep working from the native alarm list.

Path to production

Weeks 1-4: data access and connections
The read-only alarm stream and history extracts are connected for the high-alarm-load region; live-feed approval sets the pace.
Weeks 5-16: shadow pilot through a peak season
The assistant runs through a summer or winter peak while operators grade its groupings against reality.
Weeks 17-18: evaluation
Operator feedback and measured restoration-time impact drive the go or no-go decision.
Months 5-7: production hardening
Security review, monitoring, and training finish and the triage view joins the control room displays.
Months 7-9: in production and scaling
Operators work the ranked event queue during every alarm surge, extending to more regions.

What we need from your team

Full integration, data, and timeline detail for each use case in this scenario: UC 1.2 · UC 3.7 · UC 4.2
For Your Operators and Dispatchers
Where you will see it and how you say yes

The Approve button you just clicked in the demo above is the real workflow. This is what it looks like on the screen of the distribution system operator during a storm in the GridCORTEX console:

GridCORTEX ConsoleSigned in: the distribution system operator during a storm
Notifications
Alarm storm: 412 alarms in 90 seconds correlated to 1 event on Feeder F-214; narrative and suggested actions ready
Daily model refresh complete; all connected feeds healthy
Recommendation
Treat 412 alarms as one fault event on Feeder F-214
  • 387 of 412 alarms trace to one upstream recloser lockout
  • Correlation cuts items needing operator review by roughly 60%
  • Matching procedure found: recloser lockout response, section 4.2
✓ Accept event groupingModifyDecline
After you approve: The correlated event and narrative post to the ADMS event queue as one working item with the procedure attached, and an audit entry records who approved it and why.
Computed from data as of 17:42:10 local; every card shows the timestamp of the data behind it.

What happens when you hit approve

Accepting the grouping never suppresses or acknowledges alarms in your ADMS; the native alarm list is untouched. GridCORTEX posts the correlated event and narrative to the ADMS event queue as an annotation through its API, and operators act through their own ADMS controls.

How you tell it what it cannot see

Nothing to enter; the trigger is automatic from the live ADMS and SCADA alarm streams.

Live data, not stale data

Correlates the live alarm stream as it arrives, seconds behind real time; each event card shows the as-of timestamp of the newest alarm it includes. Pilot replicas, if used, show their cadence on the card.

Where it lives day to day

A ranked event queue in the GridCORTEX console beside the native ADMS alarm list; a mobile push when a correlated event exceeds 100 alarms. The console runs in a browser beside your existing screens on day one; embedding into your own systems is a roadmap step once the read-only phase has earned trust. Approve, Modify, and Decline are all captured in an audit trail your compliance team can pull, and GridCORTEX never blocks or overrides anything in the systems you run today.

The Gap: Why Your Existing Systems Don't Already Do This

The fair question: "We have an alarm system, a historian, DGA monitoring, and an EAM full of records, what's new here?" Here's the honest answer.

What you own keeps doing its job

  • SCADA / alarm system, detected every one of the 47 alarms tonight, exactly as designed. Nothing changes.
  • Historian (PI), recorded every tag, including the 2014 event. It's all in there.
  • DGA monitors & asset sensors, the gas trend and oil temps were measured faithfully.
  • EAM / work management, the open cooling-stage work order existed. The 2014 repair record existed.
  • Your operators' judgment, stays exactly where it belongs: making the final call.

The gap GridCORTEX fills, above them, not instead of them

  • Detection isn't diagnosis. The alarm system announced 47 problems; nothing you own said "this is ONE problem, and here's which one." Correlation across SCADA + DGA + work orders + event history is the gap.
  • The answer existed, in four different systems and one retired head. The 2014 event was in the historian, the fix in a logbook, the wear data in the twin, the cooling failure in an open work order. No system connects them at 02:07 AM. The agent does, in seconds.
  • Experience doesn't transfer through documentation. Delgado's notes sat in an archive nobody searches during a cascade. Captured as agent knowledge, they surface exactly when the same signature returns.
  • Guidance beats a binder. Step-by-step validated switching guidance (each step approved by the human, checked against ratings) is not what a procedure PDF does at 2 AM.
  • Always-on matters. This agent was watching at 02:07 with the same attention as 02:07 PM. Duty engineers sleep; the runtime doesn't.
Accent, don't replace: GridCORTEX reads your alarm stream, historian, DGA, EAM, and digitized operator logs · correlates and recommends above them · and your operator executes through the systems you already run. The alarm wall keeps its job; it finally gets an interpreter.
Under the Hood: What GridCORTEX Took Into Account in This Scenario

When someone asks "what did it actually calculate?", this is the list. In the simulation these factors drive the storyline; in a pilot they are computed from your SCADA, historian, DGA, EAM, and operator-log archives.

🚨 Alarm Correlation

  • Temporal and topological clustering of the 47-alarm flood, which alarms are causes, which are symptoms, which are consequences of protection response
  • Alarm sequence fingerprinting: order and spacing of sudden-pressure, OLTC, cooling, and voltage alarms as a signature
  • Suppression of consequential alarms so the operator sees one story, not a wall

📼 Event-History Matching

  • Similarity search across decades of historian data; this cascade scored 93% against the Aug 14, 2014 event on the same bank
  • Retrieval of the full 2014 record: what was tried, what worked, what the misdiagnosis risk was
  • Digitized operator logbooks and shift notes, indexed and searchable, including R. Delgado's switching notes (synthetic persona)

🔬 The Asset Twin

  • DGA trend interpretation, acetylene signature consistent with selector arcing, not bushing degradation
  • OLTC operations counter vs. design life; contact wear model from operation count and load history
  • Cooling system state, stage-2 failure three weeks prior (open work order) as the enabling condition
  • Thermal model: how long the bank can safely carry reduced load at current temperature

🧭 Guided Response & Governance

  • Switching plan validated against ratings and topology before it's proposed, every step individually approved by the operator
  • Misdiagnosis protection: the bushing-fault hypothesis explicitly tested and rejected, with reasons shown
  • Always-on posture: the agent runs on the NVIDIA Agent Toolkit with NemoClaw, OpenShell-enforced approvals, every recommendation traced by NeMo Relay for the morning review
  • Morning handoff package: work order, parts list from the 2014 repair, and the full event narrative; written before the day shift arrives

Presenter's one-liner: "The alarm system saw 47 problems. The agent saw one, because it remembered a night from 2014 that only a retired operator ever solved, read the asset's whole life story in a second, and walked a four-month operator through the fix step by approved step. That's what you just watched."

GridCORTEX Live Scenario Demo · Synthetic data throughout, no utility's actual system is depicted; all personnel are fictional · GridCORTEX connects read-only to the systems you already run · SoftServe + NVIDIA · Created by Ronnie Mauldin, NVIDIA Solutions Director, Power & Utilities, SoftServe · JUL 2026