Zeb Hering 1f7db6d4b0 Document matching Dell on CPU rather than out-bidding it
With the service stopped under sustained load, Dell held 5880-6120 RPM
with CPU2 at 70-74c. The CPU ramp calibrated from an idle reading asked
for 36-50% over that same range - up to 9720 RPM - while the drives ran
2c warmer. The extra airflow bought nothing.

Dell has the chassis thermal model and does CPU cooling well; what it
cannot see is drive temperature. Keeping the CPU ramp out of the way
until 75c and reaching full at 77c (iDRAC's own warning point) matches
stock RPM for 50 of 60 samples under load.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 23:05:58 -07:00

fan_speed

Fan control for Dell PowerEdge servers over IPMI, driven by CPU, GPU and disk temperature, with trend tracking and email alarms.

Upstream (tigerblue77/Dell_iDRAC_fan_controller_Docker) sets fan speed from CPU and GPU only. On a chassis with two dozen drives that misses the thing most likely to be quietly cooking, so this fork adds disks to the curve, records a rolling temperature history, and mails when something looks wrong.

For the MD1200 disk shelf, see the separate md1200-fan-control repo — that enclosure has its own controller and its own serial protocol.

How the fan speed is decided

Each heat source asks for a fan speed and the loudest request wins:

speed = max( interpolate(hottest of CPU/GPU, 45 -> 75),
             per-drive interpolation against each drive's own limit )

Sources are never compared as raw temperatures — an 85c CPU and a 45c disk are both "fine", and a single max() over the numbers would be meaningless.

The result is then held unless it moves by FAN_SPEED_DEADBAND (5%). This matters more than it sounds. The CPU curve is roughly 1% of fan per degree and idle CPU noise is ±2c, so without a deadband the setpoint changes on almost every pass and the fans never settle — measurably worse than stock. Measured on iz-pve0: stock iDRAC held one speed for 5 minutes straight through 57→59c jitter, while this script made 7 distinct changes in 16 minutes off the same signal. A request for full speed is never held back, and the held value is re-pushed every FAN_REASSERT_INTERVAL in case the BMC forgets it.

A deadband alone makes a machine permanently louder: it catches the fan on the way up and then nothing brings it down until the request falls a full deadband. Measured on iz-pve0 under load, 17 of 20 samples ran at 38% while the curve wanted 34-36%. So a re-assert - and only a re-assert - is allowed to settle the speed downward to what is actually wanted. Upward moves still need a full deadband crossing, which is what stops this reintroducing the drift it was built to prevent.

Note that iDRAC's own profile regulates to a target RPM (closed loop, rock steady), while manual control sets a PWM percentage (open loop, so measured RPM still wanders a percent or two at a fixed setpoint). That residual is inherent to manual control, not something the deadband can remove.

Why disks are measured against their own limits

Every drive reports its own maximum operating temperature (SATA: Min/Max Temperature Limit; SAS: Drive Trip Temperature). Each drive's ramp is derived from that number:

ramp starts at  limit - DISK_RAMP_LOW_OFFSET   (default 18)
full speed at   limit - DISK_RAMP_HIGH_OFFSET  (default 8)

So a Samsung SSD rated to 70c ramps 52→62, and a Toshiba spinner rated to 60c ramps 42→52. The 50c SSD asks for less airflow than the 45c HDD, which is correct — it is further from its own limit.

This matters because temperature does not tell you what you think it does. On these servers the rear-bay SSDs idle 10c hotter than the front-bay spinners: they sit in preheated exhaust air. A shared threshold would peg the fans for drives that are fine and ignore the ones that are not. Measuring each drive against its own envelope makes the curve independent of both drive technology and bay position.

A drive with a lower tolerance drives the fans even while a hotter drive does not. With iz-pve0's settings (floor 28%, ramp limit-8 to limit-5):

Drive Limit Ramp Asks at 45c at 50c at 55c
Samsung 860 EVO 70c 62-65c 28% 28% 28%
Toshiba / Seagate HDD 60c 52-55c 28% 28% 50%
a 55c-rated drive 55c 47-50c 28% 50% 50%

At 55c the SSD contributes nothing while the spinner asks for full speed, which is the whole point: the same temperature means different things to different drives.

Full cooling must arrive no later than the alarm. DISK_RAMP_HIGH_OFFSET must be >= DISK_ALARM_OFFSET, or the alarm mails you while there is still cooling left unused. check_ramp_ordering() clamps this at startup and logs when it does — it clamps rather than exits, because a fan controller that refuses to start leaves the fans wherever they happened to be.

Reported limits are clamped by class. They are not uniformly trustworthy — Samsung and Kioxia report a real operating maximum (70), Toshiba and Seagate report 60, and WD Reds report 85, which is the SCT critical limit and not somewhere you want a drive living. HDD_LIMIT_CAP (60) and SSD_LIMIT_CAP (70) bound whatever the drive claims, and supply the value when a drive reports nothing.

Drive class comes from /sys/block/<dev>/queue/rotational. That is derived from the device's RPM flag, which a few SAS drives behind HBAs report incorrectly; if you hit one, SMART's Rotation Rate field is the fallback.

Alarms

Email via mail to ALERT_EMAIL, which the host's postfix relays. Four conditions:

Alarm Fires when
disk_temp A drive is within DISK_ALARM_OFFSET (5c) of its own limit
trend Temperatures climbing and the fans cannot keep up — see below
smart smartctl -H reports anything other than PASSED/OK
cpu / gpu Threshold crossed; cpu also means fan control was handed back to Dell's profile
disk_count Fewer drives answered than were present at startup

The trend alarm is the one worth having, but a rising temperature is not by itself a fault — it is what a busy machine looks like, and a multi-hour zpool migration will do it all afternoon. What matters is whether the fans are answering. trend_verdict() returns one of three things:

Verdict Meaning Alarms
no_headroom Climbing while the fans are already at HIGH_FAN_SPEED yes — nothing left to give
not_converging Climbing, fans ramping with it, and the rise is not slowing down yes — cooling is losing ground
quiet Anything else, including a large climb the fans absorbed and that has plateaued no — that is just load

It tells a plateau from a runaway by comparing the first half of the window against the second: a normal load step decelerates once the fans catch up, a failing fan or a blocked intake does not. A dying fan still gets caught long before an absolute threshold, without mailing you every time the machine gets busy.

Every alarm is rate limited per key with a one hour cooldown, and sends a single recovery notice when it clears. On a 10s loop an un-throttled alarm sends 360 emails an hour, at which point the alarm is the outage.

smartctl -H across two dozen drives runs hourly, not per loop. Absolute and trend checks stay on the fast loop.

Ramp width controls how bouncy the fans are

Changing fan noise is more irritating than constant fan noise, and ramp width is the lever for it. The span sets the sensitivity: 22 points of fan over a 7c ramp is 3.14%/degC, so 1.6c of CPU drift is enough to cross a 5% deadband. The same 22 points over 11c is 2%/degC, needing 2.5c.

Replaying 180 recorded samples from iz-pve0 under load (CPU 57-69c) through candidate curves, counting how often the fan would actually move:

Ramp Deadband Fan changes Distinct speeds
62-69c 5 19 28,31,34,37,40,43,46,50
62-69c 8 6 28,37,46,50
62-73c 5 5 28,34,40
62-73c 8 1 28,36

Confirmed live against a multi-hour zpool migration: 19 fan changes became 2 over a comparable window, average fan speed dropped slightly (40% to 38%), and CPU temperature did not move (64-68c against 64-69c before). The extra airflow had been buying nothing.

Widen the ramp and raise the deadband together. The cost is lag - a deadband of 8 at 2%/degC is about 4c of slack - which is affordable when you are 8-14c below any threshold.

Calibrating a host

The defaults are not portable and the fans will be wrong on an uncalibrated host. Airflow, bay layout and the PWM→RPM relationship are all chassis specific. Calibration is a one-off, takes about five minutes, and is what keeps this from being louder than stock:

  1. Hand the fans back to Dell and let them settle, then record what stock actually does: ipmitool raw 0x30 0x30 0x01 0x01, wait a minute, then read ipmitool sdr type fan and ipmitool sdr type temperature.
  2. Take manual control (ipmitool raw 0x30 0x30 0x01 0x00) and sweep PWM to find the percentage that reproduces that RPM: ipmitool raw 0x30 0x30 0x02 0xff 0x1c (0x1c = 28%), waiting ~25s per step.
  3. Set LOW_FAN_SPEED to that percentage. This is the floor the host idles at.
  4. Set CPU_RAMP_LOW_OFFSET so the CPU ramp starts a few degrees above the host's normal idle, so idle sits on the floor rather than permanently part-way up a ramp.

Measured example — an R720xd (iz-pve0) whose CPUs idle at 60c and whose stock profile holds 6240 RPM. Here 28% PWM == 6240 RPM, so LOW_FAN_SPEED=28 and CPU_RAMP_LOW_OFFSET=15 (ramp starts at 62c). An R730xd (iz-pve1) idling at 48c is fine on the 18/24 defaults. Put the values in the unit as Environment= lines.

Put the disk ramp above the drives' normal working temperature

The single most common cause of a noisy host is a disk ramp that starts inside the range the drives normally live in. Then every load spike pushes the fan up, the fan drops back, and it cycles audibly forever.

Worked example, iz-pve0: sda/sdh are the rpool mirror in rear bays — busy root SSDs that idle at 53-55c and touch 60-61c under load, rated to 70c. With the default ramp starting at 52c they were permanently part-way up it, dragging the fans to 7920 RPM (+27% over stock) to cool drives that were never in trouble. Starting the ramp at 60c still cycled, because load spikes crossed it. Starting it at 62c, just above their working peak, holds 6240 RPM flat — and the drives settle at 53-54c anyway.

The rule: find the drives' normal peak, put DISK_RAMP_LOW_OFFSET a couple of degrees above it, and let the alarm (DISK_ALARM_OFFSET, 5c below the limit) catch anything real. Cooling a healthy drive from 57c to 53c is not worth 27% more fan.

Let Dell do the CPU; the disks are what it cannot see

Dell's profile has the chassis thermal model and is good at CPU cooling. What it does not do is look at drive temperature at all. So the CPU curve should approximate Dell's rather than out-bid it, and the disk curve is where this script earns its keep.

Measured on iz-pve0 under a sustained migration, with the service stopped so Dell had control: Dell held 5880-6120 RPM with CPU2 at 70-74c — 3c below its own 77c warning point. An earlier CPU ramp of 62-73c asked for 36-50% across that same range, i.e. up to 9720 RPM against Dell's 5880, while the drives ran 2c warmer. The extra 3800 RPM bought nothing.

The fix was to keep the CPU ramp out of the way until 75c and reach full speed exactly at 77c, where iDRAC itself starts warning. Result: 26% / 5880 RPM for 50 of 60 samples, matching stock, with brief excursions to 35% when CPU2 touched 76c.

Calibrating the CPU ramp against an idle reading is what caused this. Idle tells you where the floor should be; only a loaded CPU tells you where the ramp should start.

Expect to run above stock on some hosts, legitimately

Dell's profile does not look at drive temperature at all — only CPU, inlet and exhaust. On iz-pve0 two rear-bay SSDs sit in preheated exhaust air, and this was measured directly: raising DISK_RAMP_LOW_OFFSET so those drives stopped voting dropped the fans to 28% (exactly stock's 6240 RPM) and both SSDs climbed 54c → 61c within minutes. Restoring their vote brought them back to 53-55c at 34% / 7080 RPM.

So on that host ~13% more RPM than stock is the price of keeping two SSDs 6-8c cooler, and it is the whole reason this script exists. If you would rather have stock noise and hotter drives, raise DISK_RAMP_LOW_OFFSET until the drives stop asking — that is a preference, not a bug. What is not acceptable is being louder than stock for no reason, which is what an uncalibrated CPU curve does.

Trend history

Every pass appends to log/temps.csv:

epoch,cpu,gpu,hottest_disk,fan_speed

Trimmed to TREND_SAMPLES (90 = 15 minutes at a 10s interval). The trend check compares the newest sample to the oldest in that window, and reports zero until a full window has accumulated so a restart cannot alarm on a partial series. Flat file, tail to trim — no rrdtool, no database.

Usage

fan_speed.sh            # the service loop (default)
fan_speed.sh once       # a single pass, prints what it decided
fan_speed.sh disks      # every drive: temperature and the limit in use
fan_speed.sh selftest   # parsers, curve, trend detector, alert rate limiting

selftest touches no hardware and sends no mail — it runs the pure logic against canned smartctl output and a synthetic history file. Run it after any edit.

Install

install -m 755 fan_speed.sh functions.sh monitor.sh /root/fan_speed/
mkdir -p /root/fan_speed/log /root/fan_speed/state
install -m 644 fan_speed.service /etc/systemd/system/
systemctl daemon-reload
systemctl enable --now fan_speed.service

Tuning

Variable Default
CHECK_INTERVAL 10 Seconds between passes
LOW_FAN_SPEED / HIGH_FAN_SPEED 18 / 50 Percent
FAN_SPEED_DEADBAND 5 Hold the current speed until the request moves this far
FAN_REASSERT_INTERVAL 300 Re-push the held speed this often regardless
LOW_TEMPERATURE_THRESHOLD 45 CPU/GPU ramp start
CPU_TEMPERATURE_THRESHOLD 90 Above this, Dell's profile takes over
GPU_TEMPERATURE_THRESHOLD 75
DISK_RAMP_LOW_OFFSET 18 Ramp starts this far below each drive's limit
DISK_RAMP_HIGH_OFFSET 8 Full speed this far below it
DISK_ALARM_OFFSET 5 Alarm this far below it
HDD_LIMIT_CAP / SSD_LIMIT_CAP 60 / 70 Ceiling on what a drive may claim
TREND_SAMPLES 90 Window length, in passes
TREND_RISE_ALARM 8 Degrees of climb that alarms
DISK_EXCLUDE_PATTERN (empty) by-path substring to skip, e.g. an external shelf's HBA
ALERT_EMAIL Servers@ntfy1.izebra.xyz
ALERT_COOLDOWN 3600 Seconds between repeats of one alarm
SMART_CHECK_INTERVAL 3600 Seconds between SMART sweeps

DISK_RAMP_LOW_OFFSET is the knob to reach for first. At the default 18 an SSD rated to 70c starts ramping at 52c; raise the offset to react earlier and louder, lower it to stay quiet longer. It is set where it is because the hottest drive on iz-pve1 idles at 47c — close enough that a smaller offset would have the fans tracking normal daily drift, and you would lose the ability to tell "disks are warm" from "disks are fine".

Those defaults were sized against one chassis. Watch a day of log/temps.csv before trusting them anywhere else.

Notes

  • Kernel device names are not stable — a shelf rescan renamed sdaasdai to sdssdaa mid-session. Everything resolves through /dev/disk/by-path on every pass; never persist an sdX.
  • An external shelf must be excluded by hand. Set DISK_EXCLUDE_PATTERN to its HBA's PCI address in the systemd unit. This cannot be inferred: iz-pve1's MD1200 sits behind its own HBA at pci-0000:04:00.0, while iz-pve0's internal drives sit behind a SAS expander at pci-0000:02:00.0 — so "behind an expander" identifies an external enclosure on one host and the internal backplane on the other. iz-pve1 sets it, iz-pve0 has no shelf and leaves it empty.
  • Drive temperature limits are read once at startup. They do not change, and smartctl -x is far heavier than the -A used on the fast loop.
  • smartctl -n standby throughout, so a sleeping drive is skipped rather than spun up just to be measured. That is also why disk_count alarms on "fewer drives answered" rather than on a device disappearing.

Files

File
fan_speed.sh Main loop, fan speed decision, alarm conditions, selftest
monitor.sh Disk reading, trend history, alarm delivery — all local code
functions.sh Vendored upstream: IPMI, iDRAC, interpolation
fan_speed.service systemd unit
Description
No description provided
Readme 166 KiB
Languages
Shell 100%