[Prompt]
custom: In our episode about data centers, we looked at some of the amazing things done in the name of maxim | Hosts: corn, herman

[Response]
Corn: There are four hundred and fifteen terawatt hours of electricity going into data centres every year. Roughly one and a half percent of the world's electricity. And the thing Daniel noticed last time, the thing that actually stuck, was that the equipment pulling that power has two separate cords coming out of it.

Herman: Dual-corded. Two inputs, meant to pull from two independent networks.

Corn: Right. And Daniel's written back in with a whole set of follow-ups on the back of it. He wants to know, first: if the guiding principle is that every piece of equipment should have redundant power, wouldn't it make more sense to implement that redundancy centrally rather than run parallel connections out to every single device? Or has he misunderstood the principle entirely?

Herman: That's a good place to start, because I think he's half right.

Corn: Second: he assumes somebody has implemented this at the building level. Say a critical government facility with connections to two independent electricity networks plus a backup generator. Is there an equivalent of centrally managed failover for electrical distribution? And is there something like intelligent load balancing, the way network traffic gets balanced? His analogy, roughly: send power from this connection to these circuits, and from the other source to different circuits.

Herman: That's the fun one.

Corn: And third, he wants to know what these redundant power systems actually look like, where they're implemented, and what they cost to run. So, three questions, plus the "did I misunderstand" check. Where do we start?

Herman: Start with the misunderstanding, because it unlocks the other two.

Corn: Alright. Unlock it.

Herman: The principle isn't "every device has two power cords." That's the visible symptom. The principle is eliminating single points of failure across the entire electrical chain, top to bottom. And the trick is that "the chain" is the right word — it's not one system, it's a stack of layers, and each layer protects against a different class of failure.

Corn: Define the terms for people, because the episode lives or dies on these two.

Herman: N plus one means one extra component beyond what's needed. So if you need four generators to carry the load, you have five. 2N means two complete, fully independent systems, each capable of carrying the entire load on its own. A side and B side. And the whole point of A and B is that they share nothing. Separate utility feeds, separate UPSs, separate power distribution units, separate cable routes. If they share anything, they aren't really two.

Corn: And the vertical stack, top to bottom?

Herman: Utility feeds at the top, ideally two independent ones. Then generators, plus battery banks or UPS to bridge the gap while the generators spin up. Then transfer switches to move load between sources. Then PDUs, floor-mounted and rack-mounted. And at the very bottom, the dual-corded servers, two power supply units per server, one on A, one on B.

Corn: So the "parallel connections to every device" that Daniel's asking about isn't the whole architecture.

Herman: It's the bottom of the stack. It's the last meter. Which is exactly why it can't be replaced by anything upstream of it.

Corn: Alright, let's answer his first question properly then. Why do this at every layer instead of just centrally?

Herman: They do both, and they have to. Central redundancy — dual utility feeds, 2N UPS, N plus one generators — protects against upstream failures. Grid loss, substation failure, transformer failure. But the last few feet are still failure points. The PDU, the power cord, and the power supply inside the server itself. No amount of central redundancy protects against a failed power supply inside a chassis. Only the server's own second PSU can do that.

Corn: Because the failure is downstream of where the central system stops.

Herman: And you can see it as a table. Utility feeds protect against a grid or substation outage, that's 2N. Generators protect against extended utility loss, usually N plus one. UPS and batteries bridge the gap before the generator spins up, and they clean up power quality, N plus one or 2N. Transfer switches protect against source failure. PDUs protect against distribution failure, A and B pairs. Server PSUs protect against internal PSU failure, dual-corded. Each layer covers a failure class the layer above literally cannot.

Corn: So the answer to "did Daniel misunderstand" is no, not really. He just assumed redundancy was one decision, and it's actually six or seven decisions stacked.

Herman: And they're complements, not alternatives. That's the thing the question is really probing at. If you only did it centrally, you'd lose device-level protection. If you only did it at the device, you'd lose everything upstream. You'd have a beautifully dual-corded server connected to two cords that both run back to a single failed substation.

Corn: Which is the classic failure case.

Herman: Two power supplies plugged into the same socket. They fail together. And that's the load-bearing assumption in the whole design — the probability math assumes independence of failure events. If the two paths share any component, the redundancy is arithmetic on paper and nothing in reality. Two "independent" utility feeds that trace back to the same substation are not independent.

Corn: So the strictness of 2N — physically separate utility entrances, separate cable routing — that's not gold-plating. That's the only thing stopping you from building an expensive illusion.

Herman: It's the entire value proposition. What 2N contradicts. A facility can have two feeds, two UPSs, two PDUs, and still go down, and the post-mortem reads "common-cause failure." The design had a single point of failure wearing a costume.

Corn: And the cost of the device-level layer specifically?

Herman: Double the cabling. Double the PDU ports. Much more complex cable management. It's expensive, and it's still the only mechanism that survives a PSU or cord failure. There's no cheaper way to protect the last meter, because the last meter is inside the thing you're trying to protect.

Corn: Hold on. I want to push on something before we move up the stack. Redundancy is supposed to buy reliability. But the research keeps insisting that adding it can actually reduce availability. Walk me through why that isn't just contrarian noise.

Herman: It's a serious engineering argument, and it dates back to Charles Perrow's Normal Accidents. The claim is that redundancy increases complexity, and more complex systems are more error-prone. And there's a specific mechanism — redundancy can cause shirking of responsibility. If there are two of something, nobody owns it. The person who would have maintained the single critical unit now assumes the other one will cover.

Corn: So the backup becomes a reason not to maintain the primary.

Herman: And production pressure erodes safety margins. You have a spare, so you push the system harder, closer to its limits, because you can afford a failure. And then the failure happens while something else was already degraded.

Corn: So redundancy is a trade, not a free reliability multiplier.

Herman: It's a trade. And the honest version of the argument is: redundancy is worth it when it's done with independence and monitoring, and it's a liability when it's done as a checkbox. Marc Brooker's framing is that redundancy increases complexity, which reduces availability unless the components fail independently and the system can actually detect which ones are healthy. If your monitoring is bad, you have two systems and no idea which one is lying to you.

Corn: That lands as a general rule about infrastructure, honestly. The spare is only a spare if you know the state of both.

Herman: And you can see exactly why the data centre industry still pays for it. The alternative — one cord into one PDU off one feed — fails in a way you cannot design out. So you accept the complexity, and you spend the money on the monitoring so the complexity doesn't eat the benefit.

Corn: So Daniel's first question is answered: they do it centrally, they do it at the device, and the reason is that the layers cover different failures and neither replaces the other.

Herman: And the second half of his question — whether anyone's built building-level centrally managed failover — the answer is yes, and it's been standard for decades. The technology that does it is the transfer switch, and it's a much more interesting device than it sounds.

Corn: So the hypothetical government facility with two independent grid connections and a generator — that's real.

Herman: That's a Tuesday. That's how you build a hospital, an airport, a telecom exchange, a serious data centre. And the transfer switch is the thing that makes it automatic, which is where the "centrally managed" part comes from.

Corn: Walk the taxonomy.

Herman: Automatic Transfer Switch. It senses loss of the primary source, commands the generator to start, breaks the utility connection, connects the generator, and when utility comes back and stabilises, it transfers back and shuts the generator down after a cool-down. That's open transition. Break-before-make. There's a brief interruption — typically under a sixth of a second — and that interruption is a deliberate feature, because it prevents backfeeding the utility. Linemen working on a line that's supposed to be dead expect it to be dead.

Corn: So the tiny gap in power exists to keep a human being alive.

Herman: That's the trade. You can eliminate the gap with closed transition. Make-before-break. Zero-interruption transfer. But now you're momentarily paralleling your generator with the grid, so both sources have to be synchronised. Voltage difference under five percent, frequency difference under zero point two hertz, phase angle under five degrees, and the whole overlap under a hundred milliseconds. And you need utility approval, because you're feeding the grid for a moment whether you mean to or not.

Corn: And closed transition is the one that gets specified where?

Herman: Explicitly recommended for data processing and electronic loads. Anywhere a load interruption of even the shortest duration is objectionable. Which is the entire reason the dual-corded server exists in the first place.

Corn: So there's a feedback loop here. The more sensitive the load, the more you care about the transfer quality, the more you care about the two-source topology.

Herman: And there are two more variants. Soft-loading transfer switch. It synchronises and runs the on-site generation in parallel with utility power, not just for backup but for peak shaving. You generate your own power during expensive hours, and you're already running the generator when the grid goes down. And a static transfer switch, which uses power semiconductors instead of moving parts and transfers within a quarter-cycle. Used where even a few cycles of interruption is unacceptable.

Corn: So the "centrally managed failover" is not a metaphor. It's a physical device, and it comes with a decision tree about how much interruption you can tolerate.

Herman: And the best real-world example of the whole thing is Toronto Pearson Airport. Four redundant electrical lines, each capable of supplying the entire airport. They use a spot network substation with reverse-current relays that open breakers to failed lines while power keeps flowing. So a line dies, the relay trips, everything carries on, and nobody in the terminal notices.

Corn: Four lines, each capable of carrying everything. That's 4N, effectively.

Herman: It's beyond what most data centres do. And it's exactly Daniel's hypothetical, built at building and campus scale. Two independent sources plus a generator is the conservative version. Pearson just keeps adding lines.

Corn: Now his third question, the network analogy. Which is the one I want to get into, because I think it holds right up to the point where it doesn't.

Herman: It holds further than you'd expect. Intelligent load balancing is literally listed as a feature of larger floor-mounted PDUs, right alongside power filtering, remote monitoring, and SNMP control. And intelligent PDUs meter power at the inlet, at the outlet, and at branch-circuit level, and they let you switch individual outlets on and off remotely.

Corn: There's software in the loop at the rack level already.

Herman: At the grid level. Transmission networks are built with redundant pathways precisely to prevent a single point of failure. If a line fails, power reroutes across the remaining lines. Circuit breakers disconnect overloaded lines, and power is redistributed across what's left. Functionally, that's identical to network traffic routing around a failed link.

Corn: The spot network at Pearson is doing the same thing. Reverse-current relays isolating failed feeds while the others carry on.

Herman: Demand response, load shedding — that's traffic shaping. A utility or a large facility shifts or drops load dynamically to keep the system in balance.

Corn: Where does it break?

Herman: Three places, and they're all physics. First: electricity is synchronous and instantaneous. Packets can be buffered, queued, reordered. AC power has to be in phase and at matching frequency to be combined or switched between sources. You can't just route a watt from source A to circuit B without respecting the phase relationship. That's why closed transition has those absurdly tight sync parameters — under five percent voltage, under zero point two hertz, under five degrees phase angle. Those numbers exist because you physically cannot combine unsynchronised AC.

Corn: You can't buffer a packet of electricity at the switch.

Herman: Not cheaply. Network switches buffer in memory — that's trivial. Electrical switches generally can't. Storage is the buffer, and storage is batteries or pumped hydro, and it's expensive. So the network analogy assumes a buffering layer that the electrical system mostly doesn't have.

Corn: Second point?

Herman: The "load balancer" is largely mechanical and electrical, not software. Transformers, tap changers, protective relays, transfer switches. Software — SCADA, smart grid — is increasingly layered on top, but the foundation is physical. A network load balancer is a software function. An electrical load balancer is a copper-and-steel device with a computer bolted onto it.

Corn: The third, which I suspect is the real one.

Herman: The failure modes are physical and dangerous. Backfeeding, arc faults, phase mismatches — those can destroy equipment or kill people. Which is why the strict isolation requirements have no network equivalent. A dropped packet is an inconvenience. A phase mismatch at a closed-transition switch is an explosion.

Corn: The analogy is real at the level of "reroute around failure." It's false at the level of "just send it wherever you like."

Herman: That's the honest answer to Daniel's question. Yes, there's intelligent load balancing. No, it isn't the same thing as a network load balancer, because the medium is fundamentally different.

Corn: Which brings us to the smart grid, because that's the closest thing to closing the gap.

Herman: The smart grid is the closest thing to a true software-defined electrical load balancer. Two-way communication, distributed intelligent devices, smart meters, smart distribution boards and breakers integrated with demand response, and automatic surplus distribution through auto-smart switches.

Corn: And microgrids?

Herman: The modern framing. A local grid that can disconnect from the regional grid and island itself, running autonomously on its own resources. Generators, batteries, renewables. That's the architecture for a facility that wants to be able to operate independently of the grid entirely. Which is the logical endpoint of everything we've been describing.

Corn: Now the cost side, because Daniel asked what these systems cost to run, and I suspect the number is going to be bigger than people expect.

Herman: Power is the largest recurring cost to a data centre user. Electricity is over ten percent of total cost of ownership for high-power-density facilities. A high-availability one-megawatt data centre is estimated to consume around twenty million dollars in electricity over its lifetime, and cooling alone accounts for thirty-five to forty-five percent of that ownership cost.

Corn: Twenty million for one megawatt.

Herman: The way we measure the efficiency of it is PUE — Power Usage Effectiveness. Total facility power divided by IT equipment power. Average US data centre sits around two point zero. State of the art is around one point two. And the best-in-class, two-phase immersion cooling, gets down to about one point zero one.

Corn: Which means almost all the power you buy is doing compute work.

Herman: Almost all of it. And the reason PUE matters here is that the redundancy sits in that denominator, not the numerator. Every UPS, every transfer switch, every redundant PDU is facility power that isn't doing compute. So 2N roughly doubles the electrical infrastructure capital, and it runs at lower utilisation, because you have two systems each sized for the full load and each normally carrying half of it. That hurts efficiency.

Corn: Redundancy has a running cost, not just a build cost.

Herman: A real one. There's a line in the literature I like — in two years, the cost of powering and cooling a server can equal the cost of the server hardware itself. The metal and silicon are a one-time purchase. The electricity is forever.

Corn: Which explains why the industry keeps pushing efficiency even while it piles on redundancy.

Herman: It explains the scale problem. Data centres consumed roughly four hundred and fifteen terawatt hours globally in twenty twenty-four. About one and a half percent of world electricity. That's projected to roughly double to about nine hundred and forty-five terawatt hours by twenty thirty. The largest facilities now consume over three thousand megawatts. In the US, data centre power demand is projected to account for almost half the growth in electricity demand between now and twenty thirty.

Corn: Half the growth in national electricity demand. That's not an engineering footnote any more, that's a grid planning question.

Herman: There's a practitioner's caveat I want to bring in, because it cuts against the clean picture. A working engineer reported twice experiencing data centre hard outages caused by the power distribution system failing oddly — once switching between mains and UPS, and once between UPS and generator. The transfer mechanism itself was the failure.

Corn: The thing designed to prevent failure became the failure.

Herman: His conclusion was: the failover mechanism may have dependencies you didn't plan for, so you have to actually test it. Which means the architecture on paper and the architecture in the room are two different things.

Corn: Which, honestly, sounds like the part nobody writes about.

Hilbert: It's usually not the thing you designed for.

Hilbert: The failure. It's almost never the thing you designed for. I watched a transfer test once where the generator started, the switch threw, the load came back up, everybody ticked the box. And it passed. Only it hadn't. The UPS downstream of that switch had been sitting on bypass for months. Nobody had noticed. So the successful transfer had been riding on utility power the whole time. We'd been testing the switch against a live grid and calling it a generator test.

Herman: The test passed because the source it was supposed to move away from never actually went away.

Hilbert: The switch did exactly what it was told. It just wasn't carrying what we thought it was carrying. That's the whole thing in one line — the design is right, the as-built is right, and then somebody puts a unit on bypass during maintenance and nobody updates the drawing.

Corn: How long was it on bypass before anyone caught it?

Hilbert: Long enough that the maintenance window it was set for had been closed out twice. The paperwork said it was back in service. It wasn't. And here's the part I'd add to what you two were saying — the failure pattern is almost never the component. It's the breaker somebody racked out six weeks ago for a repair and never racked back in. It's the bypass nobody wrote down. Everything's correct in the schematic, and the room doesn't match the schematic.

Herman: You're describing a gap between the design and the as-maintained state.

Hilbert: I'm describing a room. The drawing's a drawing. I spent a lot of nights in rooms like that and the drawing was always older than the room. It's not that people are lazy. It's that the person who racked the breaker out was on shift and the person who was supposed to rack it back in wasn't, and between those two there's nobody whose actual job is the state of the thing.

Corn: The architecture isn't the hard part.

Hilbert: The architecture's the easy part. The architecture's what gets sold. What's hard is somebody at three in the morning knowing which of the two identical-looking panels is actually fed from the generator.

Corn: That reframes the whole episode, actually.

Hilbert: Anyway. That's what I'd have added.

Corn: You know what this does to Daniel's question? He's asking whether the redundancy should be central or distributed, and the answer might be "yes, and neither," because the failure lives in the maintenance state, not the topology.

Herman: That's exactly why the testing regime matters more than the design premium. You can spend double on 2N and have a better drawing and a worse outcome than an N plus one facility where somebody actually walks the room.

Corn: The misconception to kill here is the one everybody holds. That redundant power means every device has two cords.

Herman: Every device having two cords is the bottom layer of a layered architecture, and it's the layer most visible and least important on its own. Strip it out of context and it looks like waste. In context it's the only thing covering the failure classes nothing upstream can reach.

Corn: Second misconception, and this is the one Daniel's question is really circling: that two independent feeds are automatically independent.

Herman: If they trace back to the same substation, or share a transformer, or a breaker, or a cable tray, they are not independent, and the redundancy is arithmetic on paper. That's common-cause failure, and it's the reason 2N insists on physically separate entrances and routing.

Corn: The third: that adding redundancy automatically improves reliability.

Herman: Perrow's argument is that it can do the opposite. More complexity, more failure paths, shirked ownership, and production pressure that erodes the margin you built. And Hilbert's version of that is the one I'll remember — the redundancy was fine, and the system had quietly stopped being the system the drawing said it was.

Corn: If the failure pattern is almost always the gap between the design and the as-built state, what does that mean for the 2N versus N plus one trade-off as data centre demand roughly doubles by twenty thirty?

Herman: It changes the calculus, because 2N buys you redundancy on paper, and the maintenance regime is what decides whether that redundancy exists in the room. If you're planning to add nine hundred and forty-five terawatt hours of demand and data centre growth is going to be almost half the US electricity demand increase, then the cost of redundancy stops being an engineering footnote and becomes a grid planning question.

Corn: The network analogy Daniel reached for is real, but bounded. It holds at the level of rerouting around a failure. It breaks on synchronisation, on the cost of buffering, and on the fact that the failure pattern are physical and can kill somebody. The smart grid is the closest thing to closing that gap, and it is not there yet.

Herman: Not yet. Islanding microgrids, software-defined switching, intelligent PDUs with per-outlet metering — that's the direction. But the medium still does whatever the medium does.

Corn: If you want to pitch in on where this is going, leave us a review wherever you get your podcasts. Thanks as always to Hilbert Flumingtop, our producer, for keeping the desk running while we argue about switchgear.

Herman: He does it quietly.

Corn: That's been My Weird Prompts. You can find everything at my weird prompts dot com.

Herman: We'll be back soon.

Corn: See you then.