Use case deep dive · L0 metal · data centre

Data centre: telemetry to gated action

The data centre is layer L0: the metal, the power envelope and the network the whole cognitive stack rides on. An FDE deployed here turns raw telemetry into deterministic actions — detect, explain, propose, gate, act, log — without ever letting a probabilistic model touch a power or isolation command directly.

$ python -m fde_toolkit deepdive --module datacentre

Fleet health

55%

weighted across 24 racks

Samples

170

synthetic L0 telemetry events

Racks needing attention

24

A1, A2, A3, A4, A5, A6, B1, B2, B3, B4, B5, B6, C1, C2, C3, C4, C5, C6, D1, D2, D3, D4, D5, D6

Control model

AI proposes · policy decides

Q-LOCK overrides any model output on thermal and isolation paths.

Deployment workflow

what the FDE does, in order

  1. 1

    Land and inventory FDE + site ops

    Walk the estate: racks, PDUs, top-of-rack switches, GPU pools, storage tiers, Kubernetes clusters. Record what exports metrics and what does not.

  2. 2

    Wire the telemetry spine FDE + platform

    Node exporter, IPMI/Redfish, switch counters, cluster events and traces into one canonical event schema (28 variables per sample).

  3. 3

    Baseline the normal FDE

    Fourteen days of quiet operation define per-rack bands for CPU, memory, network, latency, error rate, packet loss and inlet temperature.

  4. 4

    Model the signal FDE + data science

    Neural ODE for continuous drift, spectral (QFT) analysis for periodicity, isolation forest for point anomalies. Models score — they never act.

  5. 5

    Encode the runbook as policy FDE + SRE lead

    Every remediation becomes an OPA/Rego rule with an explicit blast radius, an approval requirement and a Q-LOCK deterministic override.

  6. 6

    Agent loop under a gate FDE

    ReAct agent gathers evidence via MCP tools (read-only), assembles a proposal, then the policy engine decides ALLOW / REVIEW / REJECT_ISOLATE.

  7. 7

    Act, log, learn FDE + ops on-call

    Approved actions execute through the change system; every hop lands in the append-only audit trail and feeds the weekly drift review.

Architecture

L0 → L9

Racks / PDUs / GPUs / ToR switches / storage        (L0 metal)
        |  IPMI · Redfish · SNMP · node-exporter · k8s events
        v
Canonical telemetry event (28 vars, 10s cadence)   (L3 pipeline)
        |
        +--> Vector store: incident postmortems      (L4 knowledge)
        +--> Graph: rack -> PDU -> circuit -> service
        v
Detectors: NeuralODE drift | QFT spectrum | anomaly (L5 AI fabric)
        v
Intent router -> ReAct agent -> read-only MCP tools (L6 runtime)
        |          (ipmi_read, k8s_describe, flow_stats)
        v
OPA/Rego runbook policy  +  Q-LOCK failsafe         (L7 control)
        |  ALLOW / REVIEW / REJECT_ISOLATE
        v
Human approval for anything with blast radius > rack (L8)
        v
Change execution + append-only JSONL audit          (L9)

Rack rollup simulator

reweight the health score and watch which racks fall out

RackNodesCPUMemNetP99 msErr %Loss %Inlet °CPressureHealth
A18566942983.300.8062.7Thermal56%
A275865531283.551.3055.9Thermal55%
A376062441052.540.7256.9Thermal57%
A475974461252.711.0058.5Thermal55%
A575774421053.231.0963.7Thermal54%
A67616846903.181.0754.4Thermal55%
B18627335802.010.7859.0Thermal56%
B276170381492.701.1353.8Thermal54%
B37567045812.830.7462.6Thermal57%
B476572441113.071.0958.1Thermal53%
B577669351123.171.2556.6Thermal51%
B677571331263.131.1954.3Thermal51%
C17656744942.480.9060.2Thermal55%
C27566353773.081.0454.4Thermal57%
C37646248893.340.7855.8Thermal56%
C476466521012.600.9557.1Thermal56%
C574976391023.401.1255.3Thermal56%
C676268361422.310.9059.4Thermal55%
D17657136822.180.9962.2Thermal55%
D276866361372.700.8060.0Thermal54%
D375467341442.771.1458.5Thermal56%
D475176511153.731.0962.8Thermal54%
D57636536672.520.7156.9Thermal57%
D67475736842.921.1962.0Thermal60%

Signal catalogue

what we watch, how, and what it costs when it breaks

SignalSourceNormal bandDetectorFailure modeGated action
CPU utilisationnode-exporter45–80%Rolling z-score + NeuralODE driftNoisy neighbour starves latency-critical podsPropose cgroup limit / reschedule — REVIEW
Memory pressurenode-exporter / cgroup50–85%Trend slope over 30 minOOM-kill cascade across a node poolDrain node, raise request floor — REVIEW
Network utilisationToR switch counters10–70%Spectral band energy (QFT)Microburst congestion, invisible in 1-min averagesShift replica set, enable ECN — ALLOW
P99 latencyservice mesh traces< 250 msLognormal tail breakQueueing collapse under retry stormCap retries, shed load — ALLOW
Error rateapp + ingress logs< 1.5%Beta-tail exceedanceSilent partial outage behind a healthy LB checkFail the readiness probe — ALLOW
Packet lossswitch / NIC counters< 0.4%Counter delta + link flap correlationOptic degradation mistaken for app latencyMark link suspect, request field swap — REVIEW
Inlet temperatureIPMI / Redfish18–27 °CThermal gradient across rack rowsCRAC failure → throttle → SLA breach in ~11 minThrottle GPU pool, page facilities — REJECT_ISOLATE if > 40 °C

Worked incidents

detection → agent trace → gate → remediation

REJECT_ISOLATEblast radius: Row (18 nodes, 2 GPU pools)

Cooling loop degradation, row C

Trigger: Inlet temperature rising 0.9 °C/min across 6 racks

Detection evidence

  • NeuralODE drift amplitude 0.447 → outside the 0.31 baseline band
  • Thermal gradient correlated across racks C3–C8, not a single sensor fault
  • GPU clock throttling already visible on 4 of 12 accelerators

Gate reasoning

Inlet > 33 °C with a single cooling path — deterministic failsafe overrides any AI proposal

MTTR before 46 minafter 7 min

ReAct trace (read-only MCP tools)

  1. ipmi_read(rack=C3..C8)

    → Inlet 34.2 °C mean, delta-T collapsing

  2. facilities_status(crac)

    → CRAC-2 fan speed 41%, alarm suppressed since 03:10

  3. k8s_describe(pool=gpu-c)

    → 12 pods, 3 flagged as SLA-critical inference

  4. graph_query(rack->circuit)

    → Row C on circuit A only — no thermal failover

Remediation

  1. Clamp GPU power limit to 60% on the affected pool (deterministic, no approval).
  2. Cordon row C; drain non-critical pods to row A.
  3. Page facilities with the CRAC-2 evidence bundle attached.
  4. Hold new scheduling on row C until inlet < 27 °C for 10 minutes.
$ nvidia-smi -pl 210 -i 0,1,2,3
$ kubectl cordon $(kubectl get no -l rack=C -o name)
$ kubectl drain -l rack=C,tier!=critical --ignore-daemonsets
$ python -m fde_toolkit seos --failsafe thermal

Outcome metrics

before / after the FDE engagement

MetricBeforeAfterHow
Mean time to detect18 min40 sContinuous drift scoring at 10 s cadence
Mean time to repair42 min8 minRunbook-as-policy with pre-approved reversible actions
False page rate34%9%Correlated multi-signal detection instead of single-threshold alerts
Change failure rate12%4%Every action gated and logged; blast radius declared up front
PUE drift1.621.44Thermal-aware scheduling and GPU power clamping
Unplanned capacity events6 / quarter1 / quarterForecast crossover alerts with 9-day lead time

Anti-patterns

how data centre AI deployments fail

  • Let the agent execute remediations directly

    A probabilistic planner will eventually cordon the wrong row at 03:00.

    Instead: Agent proposes, deterministic policy decides, change system executes.

  • Single-threshold alerting per signal

    Produces the 34% false page rate that trains humans to ignore the pager.

    Instead: Require correlated evidence across at least two independent signals.

  • Model on raw counters

    Counter resets and cadence gaps look exactly like incidents.

    Instead: Normalise into the canonical event schema first, with quality flags.

  • Skip the audit line for read-only calls

    You lose the evidence chain that explains why the action was proposed.

    Instead: Log every hop — routing, retrieval, tool call, gate verdict.