Six drives, one decision: hardwire the control or run it over a fieldbus. The honest first question is not a reliability spec, it is who fixes this at 2 a.m. on a holiday. The argument almost always opens on the spec, so this is where we open too, and the spec does not settle it the way you would think.
A control engineer with six drives to bring online has one architecture decision to make before all the others: hardwire the control or run it over a fieldbus. The most robust answer is not automatic, and it is not the same for every plant, because a design is only as robust as the people who have to keep it running.
The first question a good design asks is not a number. It is who supports this on nights, weekends, and holidays, whether the customer has people who can swap a drive under pressure, and whether the spares and the know-how are actually on the shelf and in the building. Hold that question, because it is the one that should lead, and it is the whole of Part 2. Almost nobody leads with it. Faced with the choice, most of us reach for the reliability numbers first, so that is where this part starts. What it finds is that the numbers do not settle the thing the way you would expect.
A Note on Where This Came From
I want to be straight about the origin of this, because it changes how you should read the numbers that follow.
This did not start with a dataset. It started with a conversation. A paper mill, chasing the last two percent of downtime, the elusive and expensive tail that every continuous operation fights, told me they had moved control back to hardwired specifically to support it. The reason they gave was support: who could recover the system off-shift. No data was put in front of me. It was a credible practitioner telling me what they did and why.
So this series is not the report of someone’s spreadsheet. It is an attempt to take that field decision and test it against what can actually be sourced and modeled, and to be honest at every step about which is which. Where a number is published and sourced, I will say so. Where it is my own engineering estimate, I will label it an estimate. Where it is a first-hand account with no dataset behind it, I will call it what it is. The testimony is the hypothesis. The rest is the test.
The Switch Wins on Paper
Go looking for the number that proves you should keep control off the fieldbus, and the first one you find argues the other way.
Open the datasheet for a Moxa EDS-205, a five-port unmanaged switch, and it lists a mean time between failures of 3,915,945 hours, computed by the Telcordia method. That is about 447 years. Now open the datasheet for a Moxa EDS-G4008, a managed switch from the same vendor by the same method, and it lists 1,098,085 hours, or about 125 years, dropping to roughly 511,000 hours on the high-voltage model. Same brand, same standard, and the simple unmanaged box predicts a hardware life three to almost eight times longer than the managed one. If you were going to settle the control-architecture argument on box reliability, the dumb switch just won it.
The unmanaged switch has the higher predicted MTBF for a boring reason. It is a simpler device. Fewer components, no CPU running a management stack, no firmware doing spanning tree or IGMP snooping, so fewer things to fail. That is real, and it is honest to say so. Anyone who builds the case against networked control on “the switch is unreliable” is going to get handed these two datasheets in the first comment and lose the thread.
The managed switch does not earn its place by failing less often. It earns it by containing the failure modes that actually take drive networks down, the multicast floods and broadcast storms and slow topology recoveries, and by telling you a link is degrading before the process trips. That is a maintainability advantage, not a failure-rate advantage, and it does not show up in the MTBF line at all. Hold that thought, because it is the whole second half of this argument.
Why the Switch Is the Fair Proxy
It is worth saying why the switch stands in for the whole network here. Strip the two networked options down and the switch is the only real variable. The drives are the same drives. The communication option cards are the same cards. And if you pull premade, factory-tested Cat 6 instead of hand-terminating RJ45 in the field, the cabling reliability is close to a wash as well. What is left that separates a good network from a poor one is the switch: managed or unmanaged, diagnostic or blind. So when we weigh networked control against hardwired, the switch is the fair proxy for the network.
Two Comparisons, Not One
Here is the mistake I almost made, and it is worth watching for, because it is the most common unfair move in this whole argument.
There are two different things people call “control,” and they do not carry the same reliability data. Compare the wrong pair and you either overclaim or you argue past the person across the table.
The first is normal start and stop. That is ordinary control. On the hardwired side it is a PLC output driving an interposing ice-cube relay, then wire and terminals to the drive. It has no safety rating and no published probability of dangerous failure, because it is not a safety function. On the networked side it is the PLC scanner, the switch, a cable and connector, and the drive’s comm card.
The second is the last-resort stop, the E-stop into Safe Torque Off. That is a safety function, and it is required to carry a certified failure number because IEC 61508 and ISO 13849 say so.
You cannot use the safety number to win the normal-control argument. They are different functions with different data. So run the two comparisons separately and keep their numbers apart.
Comparison A: normal start and stop
This is the fair, symmetric one, because both sides have published component reliability.
On the hardwired side, the PLC digital output module carries a published MTBF (it lives in the ControlLogix technical data and SIL reference manuals, the 1756-TD002 and 1756-RM documents, rather than on the sales page), and the interposing relay carries a published life: mechanical life on the order of ten million operations, electrical life on the order of a hundred thousand at rated load. The relay is the wear item, but a start/stop that cycles a handful of times a day will take decades to reach a hundred thousand operations. On the networked side, the switch MTBF we already have, and the comm card carries its own.
Put honest component numbers on both and the result is uncomfortable for anyone hoping the hardware settles it: the two paths are roughly comparable, and as we just saw, the simple side can even come out ahead. Hardwire is not decisively more reliable than networked control for normal start and stop on the hardware alone. I am going to say that plainly rather than dance around it, because the case does not rest there, and pretending otherwise is exactly the overclaim a good controls engineer will catch.
Comparison B: the last-resort stop
This one is asymmetric, and the asymmetry is the finding.
The hardwired safety path carries a certified, published number because the standard forces it to. A Guardmaster safety relay lists a probability of dangerous failure per hour on the order of 1e-10 to 1e-9 at SIL 3 and PL e, with a twenty-year proof interval, on the manufacturer’s own functional safety data sheet. The drive’s Safe Torque Off input carries its own: Danfoss publishes the full table for its VLT drives, a PFH of 1e-10 per hour at SIL 2, PL d, with a mean time to dangerous failure of 14,000 years, and ABB and Yaskawa rate their integrated STO a notch higher at SIL 3 and PL e. Put the relay and the STO together and the hardwired stop has a certified dangerous-failure rate around one in a billion per hour, with a paper trail.
Now go find the equivalent for “stop this drive over EtherNet/IP.” It does not exist. Ordinary networked control is not held to a functional safety standard, so no one is required to quantify how often it fails dangerously, and no one does. The only way a networked command gets a certified number is if it runs as a safety protocol, PROFIsafe or CIP Safety, at which point that path must publish its own PFH.
That absence is the finding, and it belongs to Comparison B only. It is a statement about the safety function, not about normal start and stop. The safety PFHd is, arguably, all we have in the way of certified reliability data in this domain, and it is fair to use as a reference point as long as it stays labeled as the safety function and is not smuggled in to settle the normal-control question.
Count the Series Elements
There is one reliability argument that applies to both comparisons and needs no vendor to publish it.
A control command is only as reliable as the path it travels, and reliability in series multiplies down. Every element between the decision and the motor is another term in the product, another thing that has to be working.
Trace a hardwired stop. A relay contact, a length of field wire, a couple of terminals, the drive’s input. Four elements, give or take, all of them things a person can see and put a meter on.
Now trace the same command over the network. The PLC scanner, its uplink, the switch, the switch’s power supply, the switch’s firmware, a patch cable, a connector pair, the drive’s communication option card. Eight or nine elements, several of them invisible, at least one of them a software state. You did not make the command more reliable by moving it onto the bus. You added series terms, and every term you add can only lower the product. The hardwired path is shorter because it is dumber, and in a command path dumber is a virtue.
What the Fleet Data Actually Says
One more number, because it closes the door on the hardware-reliability argument from the other side.
ABB’s own field data, worked through their MTBF technical note, puts random drive hardware failure at about one drive every four years across a population of two hundred, roughly 0.0013 random failures per drive per year. Random drive hardware almost never fails. Which means the faults that actually interrupt a networked control system are not random hardware failures at all. They are communication and configuration faults, and as I have written before, those are installed, not manufactured. A wrong assembly instance, a mismatched baud rate, a missing termination resistor, a comm cable sharing a wireway with a motor lead. None of that is on any MTBF datasheet, because none of it is a hardware failure. It is an installation decision wearing a fault code.
So MTBF is the wrong question twice over. The hardware barely fails, and when the network does, the cause was not the hardware.
So Is Hardwire More Reliable, or Not
Here is the honest answer, and it is more useful than a slogan.
On component hardware, for normal start and stop, no. The two are close, and the datasheets will not hand you a win. If your whole case is “hardwire fails less often,” you have picked the one ground where you are weakest.
Where hardwire actually wins is not failure rate. It is three other things. It has fewer failure modes, because the configuration and communication faults that take networks down have no counterpart on a wire and never appear in an MTBF number. It is more supportable, which is the entire subject of Part 2. And its last-resort stop is certifiable, while the networked one is not. The thesis was never that hardwire is more reliable. It is that hardwire is more supportable and its critical function is certifiable, and for a plant chasing the last two percent of downtime, that is the lever that moves.
MTBF answers a question nobody was really asking. The real question is availability, and availability is not printed anywhere, because it depends on your roster. That is where this goes next.
The Number and the Question
You can win the datasheet and lose the plant. The unmanaged switch has the longest MTBF on the shelf, the drive hardly ever fails, the normal-control paths are a wash, and none of that tells you what you actually need to know, which is how long the line is down when the command does not arrive, and who is standing there when it happens.
MTBF answers a question nobody was really asking. The question is availability, and the answer to that one is not printed anywhere, because it depends on your roster. That is where this goes next.
Thanks.
Half power, no alarms (It’s a Navy thing…)
This is Part 1 of a two-part series.Part 2, The Off-Shift Test,makes the support case and reports the model, and the architecture argument underneath both isThe Bus Is for Data, Not for Control. The wiring and routing decisions behind a clean control path are in theVFD Installation Guide, the commissioning parameters that put control on a network are in theVFD Commissioning Guide, and the diagnostic method for comm faults is in theVFD Troubleshooting Guide. The full treatment is inBefore the First Fault, and theVFD training courseputs your team in front of live drives until this is habit rather than debate.
Author: Dr. Carl Lee Tolbert, PhD, CMRP, Wayward Leaders LLC, waywardleaders.com
Frequently Asked Questions
Is it fair to compare a networked start/stop against an E-stop safety rating?
No, and that is the most common unfair move in this debate. Normal start/stop is ordinary control with no safety rating. The E-stop into STO is a safety function with a mandated one. Compare normal control to normal control (PLC output plus interposing relay versus switch plus comm card), and compare the safety function separately. Do not let the certified safety number settle the normal-control question.
On the hardware, is hardwired control actually more reliable?
For normal start/stop, not decisively. Component MTBF for the two paths is roughly comparable, and the simple switch can even predict a longer life. The case for hardwiring control does not rest on component reliability. It rests on fewer failure modes, better supportability, and a certifiable last-resort stop.
If drive hardware rarely fails, why do networked drives drop offline?
Because the faults are communication and configuration faults, not hardware failures. Random drive hardware runs on the order of one failure every four years across a large fleet. Comm faults are installation-driven, and they are diagnosed, not replaced.
Why does the switch stand in for the whole network?
Because it is the only real variable between a good network and a poor one. The drives, the option cards, and premade tested Cat 6 are common to both networked options. Managed versus unmanaged is what is left, so the switch is the fair proxy.
What does the certified STO number actually buy me?
A published probability of dangerous failure, on the order of one in a billion per hour for a relay-plus-STO path, with a defined proof interval, that you can put in a risk assessment and defend. Ordinary networked control gives you no equivalent number for the same function.
How was this reliability comparison done?
With the Universal Reliability Simulation Framework, part of the ORCA series of reliability tools, whose core model is currently under peer review at a journal: transparent parameter provenance, where every figure carries a published source or a labeled estimate, Monte Carlo cross-checked against a Markov chain, and a sensitivity sweep on the input that dominates the result. Applied to competing architectures rather than a single part, it is the ORCA-Topology view, architecture as reliability optimization. Part 2 works the availability model in full.