Power, cooling and redundancy: the physical layer that keeps a server running
The two chains that keep a server up: power in and heat out, and what redundancy covers.

Everything a server does depends on two things that contain no software at all: electricity arriving continuously, and heat leaving continuously. Both are chains, and both fail at their weakest link. That is why “redundant power supply” is not the same statement as “the server cannot lose power”, and why a rack that is cool at the front can still be cooking the equipment inside it.
The model: two chains
POWER HEAT
utility components dissipate
↓ ↓
UPS (battery, transfer) air moved front to back
↓ ↓
PDU ──┬──── PSU #1 ──┐ fans
└──── PSU #2 ──┴──→ mainboard ↓
RAM, disks, CPU room cooling (CRAC / in-row / contained)
Each chain has components that can be duplicated and components that cannot. Redundancy is the practice of making sure that no single failure cuts the chain — and only that.
Terms used here
- PSU — power supply unit: converts mains AC to the DC the server uses.
- Redundant PSU / N+1 — more capacity than the load needs, so one unit can fail or be replaced while the server keeps running.
- Hot swap — replacing a component without powering the system down.
- Hot spare (power supply) — a redundant PSU deliberately held in standby to reduce power overhead.
- Dual corded — the server has two power inputs, ideally fed from two independent distribution paths.
- PDU — power distribution unit: the strip or panel feeding the rack.
- UPS — uninterruptible power supply: battery plus the electronics that decide when the load runs from it; the designs differ in how the load is fed (see below).
- Heat load — the energy the equipment turns into heat and the cooling system must remove.
- Inlet temperature — the temperature of the air entering the equipment, which is what the equipment actually cares about.
- BMC — baseboard management controller (iDRAC, iLO, IPMI class): the out-of-band controller that also reports power and thermal sensors.
Power: from the socket to the board
A server with two PSUs is described as redundant, and the documentation is precise about what that means. Dell’s Installation and Service Manual for a current PowerEdge model documents two redundant PSUs (as opposed to a single non-redundant configuration) and a hot spare feature whose purpose is to reduce the power overhead of redundancy: when it is enabled, one of the redundant PSUs is held in standby while it is not needed; if the load on the active PSU rises above a threshold (the manual gives 50 percent of rated power) the standby unit is brought back, and if the load falls low again (below 20 percent) it returns to standby. The feature is configured through the BMC settings. In other words: redundancy is maintained, but one unit may be idling rather than contributing.
What the component-level statement does not say is anything about the rest of the path. Two PSUs plugged into two sockets of the same PDU share that PDU’s failure; two PDUs fed by the same UPS share the UPS; two UPS units fed from the same panel share the panel. Redundancy is a property of the entire chain, and the useful question is not “how many PSUs does the server have?” but “where does each path go?”.
UPS designs matter for the same reason. The APC white paper The Different Types of UPS Systems groups the common designs into topologies — standby, line-interactive and on-line — and describes how each feeds the load: a standby UPS is documented as having “bypass normal mode as their only normal mode”; a line-interactive design can “de-energize their automatic voltage regulation (AVR)” transformers to correct voltage without touching the battery; and an on-line UPS feeds the load through its conversion stage continuously, which is why the paper credits it with addressing voltage fluctuation events that a simpler topology passes through. A transfer switch is what moves the load between normal, bypass and battery sources.
Heat: what the cooling system has to remove
Two numbers frame cooling design. The first is how much heat there is. APC’s white paper Calculating Total Cooling Requirements for Data Centers works from the IT load and adds the rest of the facility’s load to it: in the example used by the paper, the total the cooling system must handle comes out at “approximately 50% more than the IT load”. The equipment’s own consumption is therefore the starting figure, not the target.
The second number is what the equipment tolerates. The ASHRAE Equipment Thermal Guidelines for Data Processing Environments reference card defines classes of air-cooled equipment and states that facilities should be designed and operated to target the recommended inlet range for the class — for the A1 to A4 classes the reference card gives an 18 °C to 27 °C range, with the allowable range wider still. Inlet temperature is the number to design for, because that is the air the components actually see.
Airflow is the third part: equipment is built to move air in one direction (front to back in a rack server), which is why cold aisle / hot aisle arrangements exist and why blocking the front of a server with a door, a cable bundle or another device is a thermal decision. Manufacturer manuals also document thermal restrictions for specific configurations, because a chassis filled with adapters and high-power CPUs is not the same thermal object as the base configuration.
Redundancy is a property of the path, not of the component
Putting the two chains together gives the rule that actually keeps systems up:
- duplicate the path, not only the component (two cords, two PDUs, two UPS units, two feeds);
- verify that the paths are genuinely independent — different breakers, different panels, different routes;
- remember that out-of-band management is a monitoring tool: the BMC reports the sensor values, but it does not create redundancy;
- and keep the physical layer documented, because “which PSU is on which feed” is exactly the information that is missing during an incident.
A common misconception: “redundant power means no power loss”
Redundant PSUs protect against a PSU failing, not against the power path upstream failing, and not against a cord being unplugged from the wrong strip. In the same way, a UPS is a battery with a transfer or a conversion stage rather than an energy source of its own: it bridges the load to the next source, so the design question is how the load is fed and for how long the battery can carry it.
What to remember
- Availability has a physical layer: power in, heat out, and both are chains.
- Redundant PSUs are one link of the chain; the path upstream decides whether they matter.
- The cooling system removes more than the IT load; the inlet temperature is what the equipment sees.
- Airflow direction and thermal restrictions are documented per chassis and per configuration.
- The BMC reports the state of the physical layer — it does not replace it.
Level and prerequisites
L1 — fundamentals: the components, the chains and the redundancy vocabulary. Prerequisites: none. Sizing a UPS, calculating heat load for a room, configuring PDU redundancy or reading BMC sensor thresholds is operational material (L2–L3), and power/cooling design for a facility belongs to L5.
Where to go next
- Server & Virtualization — the area this sheet belongs to.
References
- Dell, PowerEdge R750 Installation and Service Manual — redundant and non-redundant PSU configurations, the hot spare feature and its load thresholds, configuration through iDRAC, thermal restrictions.
- APC (Schneider Electric), White Paper: The Different Types of UPS Systems — standby and on-line (double-conversion) designs, the transfer switch, and how the load is fed in each design.
- APC (Schneider Electric), White Paper 1: Calculating Total Cooling Requirements for Data Centers — the method that adds facility load to the IT load, and the worked example in which the total is approximately 50% more than the IT load.
- ASHRAE, Equipment Thermal Guidelines for Data Processing Environments (reference card) — equipment classes A1–A4, the recommended inlet range for those classes, and the instruction to design and operate toward the recommended range.