Field note

The Off-Shift Test: The Reason Plants Are Pulling Network Control

The fault rate does not care what day it is. The recovery does. Ask a paper mill why they moved control back to hardwired and they will not quote you a reliability number, they will tell you what happens at 2 a.m. on Thanksgiving.

Part 1 ended on a question MTBF cannot answer, and settled the reliability half honestly: on the hardware, for normal start and stop, hardwired and networked control are roughly a wash. The case for hardwiring control never rested on component reliability. It rests on supportability, and this is the part that shows why.

A Note on Where This Came From (Carried Over From Part 1)

Repeating this here on purpose, because it governs the numbers below just as much as the ones above, and Part 2 is where the numbers get specific.

This did not start with a dataset. It started with a paper mill, chasing the last two percent of downtime, telling me they had moved control back to hardwired to support it. The reason they gave was support: who could recover the system off-shift. No data was put in front of me. So what follows is not a report of someone’s spreadsheet. It is an attempt to test that field decision against what can be sourced and modeled, labeling as I go which numbers are published, which are my own engineering estimates, and which are first-hand testimony with no dataset behind them. The testimony is the hypothesis. The model is the test.

The First Question, and It Is Not a Number

Before the reliability math, before the switch datasheets, a decision on six drives starts with a question the spec sheets cannot answer: who supports this when it breaks on a night, a weekend, or a holiday? The most robust design an engineer can offer is not the one with the best numbers on paper. It is the one the customer can actually keep running with the people, the spares, and the knowledge they have on site.

So the real intake questions are about the plant, not the drive. Does the customer have PLC-savvy maintenance staff, or a crew of good electricians who have never opened a network configuration? If a drive dies at 2 a.m., is there a spare on the shelf, and can someone on shift put it in and bring it back? A hardwired design leans on skills every industrial electrician already has. A networked design quietly assumes a controls skill set that many plants only staff on day shift, if they staff it at all. Recommending fieldbus control to a plant without that bench is handing them a system they cannot recover without a phone call, and the engineer who specified it owns that outcome. That is why the support question comes first, and why the reliability comparison, the whole of Part 1, is the second question and not the first.

The Last Two Percent

The paper mill detail is not incidental. It locates the entire argument.

Every continuous operation eventually runs out of easy downtime to remove. The big, frequent failures get engineered out first, and availability climbs into the high nineties. Then the gains get elusive and expensive, because each additional fraction of a percent costs more than the last. The reliability tiers used in critical-facility work make the shape plain: moving from 99.98 percent availability to 99.99 percent is the difference between about 1.6 and 0.8 hours of downtime a year, and buying that last fraction can cost more than everything spent to reach it.

Here is what happens at that asymptote. Once the frequent failures are gone, the downtime that remains is not dominated by how often things fail. It is dominated by how long recovery takes when they do, and recovery is dominated by the worst shifts, the nights and weekends and holidays when the bench is thin. So a plant hunting the last two percent will find that the marginal lever is not component reliability. It is maintainability. And hardwiring the critical control is a maintainability lever, because it changes who can recover the system and how fast. That is why the mill made the call it made. The rest of this piece is the reconstruction of why that call is rational.

The Calendar Does Not Change the Fault Rate

Start with what does not change. A comm fault or a configuration fault is no more likely on a holiday than on a Tuesday. The wire ages the same, the connector loosens the same, the network drops a packet at the same rate at 3 a.m. as at 3 p.m. Whatever your control-path fault frequency is, it is flat across the calendar.

What is not flat is the recovery. The same fault that a controls engineer clears in twenty minutes on day shift can hold a line down for hours on a holiday, not because the fault got worse, but because the person who can read it is at home. That is the entire finding, and once you see it, the vendor reliability numbers look like they are measuring the wrong thing, because they are.

Two Kinds of Available

Reliability engineering has known this for decades, and it has two words for it.

Inherent availability uses only the active repair time: the interval when a qualified person is actually working the problem. It is what a datasheet implies, and by that measure it does not matter what shift you fault on. Operational availability adds the delays that come before the wrench turns: the time to detect the fault, the time to get the right skill and the right part to the panel, the restart. The classic name for that middle term is logistics delay, and it is where the off-shift problem lives.

This is not my framing. It is the standard maintainability distinction, and it shows up in the reliability literature in exactly these terms. A recent peer-reviewed availability study of dual-pump systems, the same duty-standby topology as the alternating pump panels I build, put it plainly: the main cause of repair delay is that unexpected failures require assembling a maintenance crew that may not be readily available and spare parts that may not be in stock, and once that unpredictability is removed, maintenance runs far more efficiently. Vendors quote inherent availability. Plants live operational availability. The difference between the two is a call tree.

Who Is the Required Skill

Now the piece that makes the two architectures diverge.

Break the down time into its parts: detection, then getting the required skill to the panel, then diagnosis, then the actual repair. For most faults the repair itself is quick once you know what is wrong. The time goes into the middle two terms, and the middle two terms depend on a single question: is the person already in the building the person who can fix this?

For a hardwired fault, the answer is yes. The maintenance electrician on shift can meter a wire, check a relay, read a discrete signal. The fault lives in their trade, so their down time barely moves whether it is noon on Wednesday or midnight on a holiday.

For a network or configuration fault, the answer off-shift is no. The electrician can confirm the drive is powered and the link light is green and still have no way to tell you the assembly instances are mismatched. The required skill is the controls engineer, and off-shift that skill is a phone call, a drive-in, or a remote login that has to be set up first. The fault did not get harder. The person who could close it went home, and now detection and diagnosis stretch to fill the gap until they are reached.

There is a quieter version of the same problem, and it is about seeing the fault at all. The managed switch’s advantage is that it can flag a degrading link before the process trips, but that flag usually travels by SNMP and switch traps, and those do not always land in the tools the automation team actually watches. More often they land in IT’s world, or nowhere. John Rinaldi of RT Automation, a specialist in industrial Ethernet, puts it plainly: the biggest weakness of the industrial switch is that its status reporting does not integrate into the tools control engineers use, and switch status has to be made visible to the automation team, not just to IT. So the visibility that could have shortened the diagnosis is stranded from the people at the panel, and off-shift the IT side is no more present than the controls engineer. A managed switch that no one is watching on the right screen is, for that shift, a blind one.

Even the Repair Is Not the Same

There is a second trap hiding in the word repair, and it is worse than the diagnosis problem.

Swap a hardwired drive and the job is bounded: land the power, land the control wires, set the parameters, run it. Any competent electrician can do it with the manual and a spare, and the replacement does not need to know anything about the drive it replaced.

Swap a networked drive and you are not finished when the wires are landed. You have to bring the replacement back onto the network exactly as the old one sat there, and that is not plug and play. The node address, the communication settings, the assembly mapping the PLC scanner expects, all of it has to match. And here is the part that catches people who did everything else right: the mapping does not transfer cleanly even between two drives of the same brand, because it is tied to the firmware revision. Run EtherNet/IP across a plant that has collected a dozen firmware versions over the years, drop in a spare off the shelf, and the PLC may not talk to it the way it talked to the one that failed. Now the drive swap, the part the electrician could have finished alone, is a controls task again, and you are waiting once more on the one person who can reconcile the EDS file, the firmware, and the scanner.

So the network does not only make the fault harder to find off-shift. It makes the fix harder to finish off-shift, in the exact moment you least want a surprise, with a line already down and a spare already in your hand. A hardwired swap ends when the wires are tight. A networked swap ends when the mapping agrees, and the mapping is a specialist’s problem.

What the Model Says

I built this as an availability model in the same framework I use for the ORCA reliability work: Monte Carlo for the distribution, a Markov chain as the analytical cross-check, and a sensitivity sweep on the parameter that carries the argument. To be fair to the question, I held the control-path fault rate constant across all three architectures, so every hour of difference you are about to see is recovery, not frequency. The time bands are transparent engineering assumptions, not field-logged data, and I will say so plainly. What matters is the shape, and the shape is not subtle.

Hardwired mean down time sits at about two hours and stays there across every shift, because the electrician who can fix it is always there. Networked mean down time climbs as coverage thins: for a managed-switch fault, from about three hours on day shift to six and a half on a holiday; for an unmanaged-switch fault, from about four hours to more than eight. Translate that into a year and hold the fault rate fixed, and the recovery penalty of networked control over hardwired grows from roughly thirteen extra down-hours per line per year on day shift to nearly forty on a holiday for the unmanaged case, with managed sitting partway between.

Three things keep this honest. The Monte Carlo and the Markov cross-check agree to within about a third of an hour a year across every cell, so the method is not carrying the result. The sensitivity sweep shows the conclusion survives even if I halve the assumed off-shift delay: the networked penalty is still large. And the numbers are conservative, because the model prices the extra time to detect and diagnose a network fault off-shift but does not yet price the recommissioning penalty on a drive swap. When a networked replacement will not map cleanly across firmware and a controls hand is needed just to finish the physical swap, the real off-shift gap is wider than the model shows, not narrower.

None of this depends on getting the one uncertain number exactly right, which is the only reason it is worth publishing when no one tracks that number publicly.

What It Costs

Down-hours only matter because of what they cost, and in continuous process the cost is brutal. Siemens’ True Cost of Downtime work puts automotive as high as two million dollars an hour and a typical large plant near a hundred and thirty million a year in downtime losses. Paper, metals, and chemical run their own six-figure-per-hour numbers because a continuous line does not stop cleanly: a paper machine web break means clean-up and re-thread, a furnace means a reheat, a chemical line means a full restart cycle.

Put even a modest continuous-process rate against the model. At fifty thousand dollars an hour, the holiday recovery penalty on an unmanaged networked control path runs into seven figures per line per year, and that is the avoidable part, the piece that exists only because the command rode a network the on-shift crew could not diagnose. That is not a rounding error in a maintenance budget. That is a capital decision hiding in a wiring choice.

And it explains the pattern. The plants pulling control back to hardwired are the paper mills, the metals operations, the chemical plants. They are continuous, so off-shift downtime is the most expensive downtime they have, and their off-shift skilled-trades bench is the thinnest. They are exactly where the recovery penalty is largest, so they are exactly where the math forces the decision first, and exactly where a mill chasing the last two percent goes looking.

Why This Never Shows Up in a Datasheet

Because it cannot. The recovery penalty is the operational-availability term, and it depends on your plant, your roster, your call tree, your parts crib, and how far the controls engineer lives from the gate. No vendor can put that on a spec sheet, so no vendor does, and the number that actually decides the architecture is the one number nobody publishes. That is not a gap in the analysis. It is the finding restated: the decision was never a reliability decision. It was a maintainability decision, and maintainability is measured in people, not in hours between failures.

Which is also why the paper mill had no data to hand me. There was none to hand. The number lived in their own roster and their own worst shift, and the only way to see it is to model it or to live it.

You Are Not Buying Reliability

Strip it all the way down and the vendor question and the plant question are different questions. The vendor sells you reliability, mean time between failures, a number that assumes the right person is always standing at the panel. The plant lives maintainability, mean time to recover, which is entirely about whether that person is there at all.

You are not buying reliability. You are buying maintainability under your actual roster.

Thanks.

Half power, no alarms (It’s a Navy thing…)

This is Part 2 of a two-part series. Part 1, Is Fieldbus Less Reliable Than Hardwired Control for VFDs?, makes the machine case, and the architecture argument underneath both isThe Bus Is for Data, Not for Control. The commissioning decisions that put control on a network are in theVFD Commissioning Guide, and the diagnostic method for comm faults is in theVFD Troubleshooting Guide. The full treatment, from control architecture through the loss-of-communication decision, is inBefore the First Fault, and theVFD training coursebuilds the diagnostic habit on live drives.

Author: Dr. Carl Lee Tolbert, PhD, CMRP, Wayward Leaders LLC, waywardleaders.com

Frequently Asked Questions

Does networked control fault more often on nights and weekends?

No. The fault rate is flat across the calendar. What changes is recovery time, because the skill required to diagnose a network or configuration fault is often not the skill on shift.

Why does a hardwired fault recover faster off-shift than a network fault?

Because the maintenance electrician who is already there can diagnose a wire, a relay, or a discrete signal, but usually cannot diagnose a mismatched assembly instance or an addressing conflict. The hardwired fault lives in the on-shift trade; the network fault waits for the controls engineer.

Isn't this just an argument for better on-call coverage?

Better coverage helps, and it is worth doing. But coverage is a recurring cost you pay forever, and it still leaves a mobilization delay on the worst shifts. Keeping the critical command hardwired removes the dependency instead of staffing around it.

How does the customer's own staffing change the recommendation?

It changes it more than any spec. A plant with PLC-savvy maintenance on every shift and firmware-matched spares can carry networked control that would strand a plant staffed only with electricians on days. The most robust design is the one matched to the bench that will actually support it, which is why the support question comes before the reliability question.

Why is swapping a networked drive harder than swapping a hardwired one?

A hardwired swap ends when the wires are landed and the parameters are set. A networked swap is not done until the replacement is mapped back onto the network the way the PLC expects, and that mapping is tied to firmware. Even a same-brand spare off the shelf may not come back cleanly across firmware versions, which turns a physical swap back into a controls task at the worst possible time.

How solid are the numbers?

The downtime costs come from published industry work. The recovery-time bands are engineering estimates, labeled as such. The direction is robust: the conclusion holds even at half the assumed off-shift delay, the two methods in the model agree, and the estimate is conservative because it does not yet price the recommissioning penalty. Treat the dollars as order of magnitude, not to the decimal.

Does this mean never network a drive?

No. Network the data, all of it: trending, runtime balancing, alarms, condition monitoring. Keep the start, the stop, and the last-resort safety on a path the on-shift crew can recover. That is the line, and it is the same line as the anchor piece in this series.

What method is behind these numbers?

The Universal Reliability Simulation Framework, part of the ORCA series of reliability tools, with its core model currently under peer review at a journal. Applied here to a system architecture rather than a single component, it runs Monte Carlo simulation cross-checked against a continuous-time Markov chain, holds to transparent parameter provenance so every input carries a published source or a labeled estimate, and puts a sensitivity sweep on the input that dominates the result. Comparing whole architectures this way, hardwired against networked, is the ORCA-Topology view: architecture as reliability optimization, the same framework that runs the bearing and motor cost-of-ownership comparisons. One point of standing: the framework is the peer-reviewed work; the fieldbus figures in this article are an applied illustration built on labeled estimates, not themselves a reviewed result.

Wayward Leaders® is a veteran-owned VFD training practice. We teach maintenance teams and plant electricians to install, commission, and troubleshoot variable frequency drives correctly, in person, at the plant that owns the equipment, anywhere in the United States.

Every class ends with a scored competency assessment. The plant gets a way to prove what each technician can do, not a roster of who sat in the room.

The instruction is field-trained. It is drawn from nearly 8,000 VFD, power, and motor documents and from more than 750 commissioned drives captured through personal field experience and data collection. Carl Lee Tolbert, PhD, CMRP, and ATD Master Trainer® candidate, leads it. He has delivered more than 5,000 hours of VFD instruction across 30 years and trained 8,000 industrial professionals, from International Paper to the U.S. Navy.

The premise is simple: VFDs do not fail. Installations fail them. The curriculum is built backward from that.

Installation, commissioning, troubleshooting, and fault diagnosis guides are published openly at waywardleaders.com.