Use case deep dive · L0 metal · data centre
Data centre: telemetry to gated action
The data centre is layer L0: the metal, the power envelope and the network the whole cognitive stack rides on. An FDE deployed here turns raw telemetry into deterministic actions — detect, explain, propose, gate, act, log — without ever letting a probabilistic model touch a power or isolation command directly.
$ python -m fde_toolkit deepdive --module datacentre
Fleet health
55%
weighted across 24 racks
Samples
170
synthetic L0 telemetry events
Racks needing attention
24
A1, A2, A3, A4, A5, A6, B1, B2, B3, B4, B5, B6, C1, C2, C3, C4, C5, C6, D1, D2, D3, D4, D5, D6
Control model
AI proposes · policy decides
Q-LOCK overrides any model output on thermal and isolation paths.
Deployment workflow
what the FDE does, in order
- 1
Land and inventory FDE + site ops
Walk the estate: racks, PDUs, top-of-rack switches, GPU pools, storage tiers, Kubernetes clusters. Record what exports metrics and what does not.
- 2
Wire the telemetry spine FDE + platform
Node exporter, IPMI/Redfish, switch counters, cluster events and traces into one canonical event schema (28 variables per sample).
- 3
Baseline the normal FDE
Fourteen days of quiet operation define per-rack bands for CPU, memory, network, latency, error rate, packet loss and inlet temperature.
- 4
Model the signal FDE + data science
Neural ODE for continuous drift, spectral (QFT) analysis for periodicity, isolation forest for point anomalies. Models score — they never act.
- 5
Encode the runbook as policy FDE + SRE lead
Every remediation becomes an OPA/Rego rule with an explicit blast radius, an approval requirement and a Q-LOCK deterministic override.
- 6
Agent loop under a gate FDE
ReAct agent gathers evidence via MCP tools (read-only), assembles a proposal, then the policy engine decides ALLOW / REVIEW / REJECT_ISOLATE.
- 7
Act, log, learn FDE + ops on-call
Approved actions execute through the change system; every hop lands in the append-only audit trail and feeds the weekly drift review.
Architecture
L0 → L9
Racks / PDUs / GPUs / ToR switches / storage (L0 metal)
| IPMI · Redfish · SNMP · node-exporter · k8s events
v
Canonical telemetry event (28 vars, 10s cadence) (L3 pipeline)
|
+--> Vector store: incident postmortems (L4 knowledge)
+--> Graph: rack -> PDU -> circuit -> service
v
Detectors: NeuralODE drift | QFT spectrum | anomaly (L5 AI fabric)
v
Intent router -> ReAct agent -> read-only MCP tools (L6 runtime)
| (ipmi_read, k8s_describe, flow_stats)
v
OPA/Rego runbook policy + Q-LOCK failsafe (L7 control)
| ALLOW / REVIEW / REJECT_ISOLATE
v
Human approval for anything with blast radius > rack (L8)
v
Change execution + append-only JSONL audit (L9)Rack rollup simulator
reweight the health score and watch which racks fall out
| Rack | Nodes | CPU | Mem | Net | P99 ms | Err % | Loss % | Inlet °C | Pressure | Health |
|---|---|---|---|---|---|---|---|---|---|---|
| A1 | 8 | 56 | 69 | 42 | 98 | 3.30 | 0.80 | 62.7 | Thermal | 56% |
| A2 | 7 | 58 | 65 | 53 | 128 | 3.55 | 1.30 | 55.9 | Thermal | 55% |
| A3 | 7 | 60 | 62 | 44 | 105 | 2.54 | 0.72 | 56.9 | Thermal | 57% |
| A4 | 7 | 59 | 74 | 46 | 125 | 2.71 | 1.00 | 58.5 | Thermal | 55% |
| A5 | 7 | 57 | 74 | 42 | 105 | 3.23 | 1.09 | 63.7 | Thermal | 54% |
| A6 | 7 | 61 | 68 | 46 | 90 | 3.18 | 1.07 | 54.4 | Thermal | 55% |
| B1 | 8 | 62 | 73 | 35 | 80 | 2.01 | 0.78 | 59.0 | Thermal | 56% |
| B2 | 7 | 61 | 70 | 38 | 149 | 2.70 | 1.13 | 53.8 | Thermal | 54% |
| B3 | 7 | 56 | 70 | 45 | 81 | 2.83 | 0.74 | 62.6 | Thermal | 57% |
| B4 | 7 | 65 | 72 | 44 | 111 | 3.07 | 1.09 | 58.1 | Thermal | 53% |
| B5 | 7 | 76 | 69 | 35 | 112 | 3.17 | 1.25 | 56.6 | Thermal | 51% |
| B6 | 7 | 75 | 71 | 33 | 126 | 3.13 | 1.19 | 54.3 | Thermal | 51% |
| C1 | 7 | 65 | 67 | 44 | 94 | 2.48 | 0.90 | 60.2 | Thermal | 55% |
| C2 | 7 | 56 | 63 | 53 | 77 | 3.08 | 1.04 | 54.4 | Thermal | 57% |
| C3 | 7 | 64 | 62 | 48 | 89 | 3.34 | 0.78 | 55.8 | Thermal | 56% |
| C4 | 7 | 64 | 66 | 52 | 101 | 2.60 | 0.95 | 57.1 | Thermal | 56% |
| C5 | 7 | 49 | 76 | 39 | 102 | 3.40 | 1.12 | 55.3 | Thermal | 56% |
| C6 | 7 | 62 | 68 | 36 | 142 | 2.31 | 0.90 | 59.4 | Thermal | 55% |
| D1 | 7 | 65 | 71 | 36 | 82 | 2.18 | 0.99 | 62.2 | Thermal | 55% |
| D2 | 7 | 68 | 66 | 36 | 137 | 2.70 | 0.80 | 60.0 | Thermal | 54% |
| D3 | 7 | 54 | 67 | 34 | 144 | 2.77 | 1.14 | 58.5 | Thermal | 56% |
| D4 | 7 | 51 | 76 | 51 | 115 | 3.73 | 1.09 | 62.8 | Thermal | 54% |
| D5 | 7 | 63 | 65 | 36 | 67 | 2.52 | 0.71 | 56.9 | Thermal | 57% |
| D6 | 7 | 47 | 57 | 36 | 84 | 2.92 | 1.19 | 62.0 | Thermal | 60% |
Signal catalogue
what we watch, how, and what it costs when it breaks
| Signal | Source | Normal band | Detector | Failure mode | Gated action |
|---|---|---|---|---|---|
| CPU utilisation | node-exporter | 45–80% | Rolling z-score + NeuralODE drift | Noisy neighbour starves latency-critical pods | Propose cgroup limit / reschedule — REVIEW |
| Memory pressure | node-exporter / cgroup | 50–85% | Trend slope over 30 min | OOM-kill cascade across a node pool | Drain node, raise request floor — REVIEW |
| Network utilisation | ToR switch counters | 10–70% | Spectral band energy (QFT) | Microburst congestion, invisible in 1-min averages | Shift replica set, enable ECN — ALLOW |
| P99 latency | service mesh traces | < 250 ms | Lognormal tail break | Queueing collapse under retry storm | Cap retries, shed load — ALLOW |
| Error rate | app + ingress logs | < 1.5% | Beta-tail exceedance | Silent partial outage behind a healthy LB check | Fail the readiness probe — ALLOW |
| Packet loss | switch / NIC counters | < 0.4% | Counter delta + link flap correlation | Optic degradation mistaken for app latency | Mark link suspect, request field swap — REVIEW |
| Inlet temperature | IPMI / Redfish | 18–27 °C | Thermal gradient across rack rows | CRAC failure → throttle → SLA breach in ~11 min | Throttle GPU pool, page facilities — REJECT_ISOLATE if > 40 °C |
Worked incidents
detection → agent trace → gate → remediation
Cooling loop degradation, row C
Trigger: Inlet temperature rising 0.9 °C/min across 6 racks
Detection evidence
- NeuralODE drift amplitude 0.447 → outside the 0.31 baseline band
- Thermal gradient correlated across racks C3–C8, not a single sensor fault
- GPU clock throttling already visible on 4 of 12 accelerators
Gate reasoning
Inlet > 33 °C with a single cooling path — deterministic failsafe overrides any AI proposal
ReAct trace (read-only MCP tools)
ipmi_read(rack=C3..C8)→ Inlet 34.2 °C mean, delta-T collapsing
facilities_status(crac)→ CRAC-2 fan speed 41%, alarm suppressed since 03:10
k8s_describe(pool=gpu-c)→ 12 pods, 3 flagged as SLA-critical inference
graph_query(rack->circuit)→ Row C on circuit A only — no thermal failover
Remediation
- Clamp GPU power limit to 60% on the affected pool (deterministic, no approval).
- Cordon row C; drain non-critical pods to row A.
- Page facilities with the CRAC-2 evidence bundle attached.
- Hold new scheduling on row C until inlet < 27 °C for 10 minutes.
$ nvidia-smi -pl 210 -i 0,1,2,3 $ kubectl cordon $(kubectl get no -l rack=C -o name) $ kubectl drain -l rack=C,tier!=critical --ignore-daemonsets $ python -m fde_toolkit seos --failsafe thermal
Outcome metrics
before / after the FDE engagement
| Metric | Before | After | How |
|---|---|---|---|
| Mean time to detect | 18 min | 40 s | Continuous drift scoring at 10 s cadence |
| Mean time to repair | 42 min | 8 min | Runbook-as-policy with pre-approved reversible actions |
| False page rate | 34% | 9% | Correlated multi-signal detection instead of single-threshold alerts |
| Change failure rate | 12% | 4% | Every action gated and logged; blast radius declared up front |
| PUE drift | 1.62 | 1.44 | Thermal-aware scheduling and GPU power clamping |
| Unplanned capacity events | 6 / quarter | 1 / quarter | Forecast crossover alerts with 9-day lead time |
Anti-patterns
how data centre AI deployments fail
Let the agent execute remediations directly
A probabilistic planner will eventually cordon the wrong row at 03:00.
Instead: Agent proposes, deterministic policy decides, change system executes.
Single-threshold alerting per signal
Produces the 34% false page rate that trains humans to ignore the pager.
Instead: Require correlated evidence across at least two independent signals.
Model on raw counters
Counter resets and cadence gaps look exactly like incidents.
Instead: Normalise into the canonical event schema first, with quality flags.
Skip the audit line for read-only calls
You lose the evidence chain that explains why the action was proposed.
Instead: Log every hop — routing, retrieval, tool call, gate verdict.
