M03 · Data-centre telemetry · infra health scoring

Environment diagnostics

Establish whether the environment can actually run the system before anything else.

$ python -m fde_toolkit diagnose

Workflow

the order an FDE actually runs it in

  1. 1Collect the normalised infrastructure vector: CPU, memory, network, latency, error, loss, temperature, GPU, IO, pods.
  2. 2Score health and rank the dominant degradation term.
  3. 3Correlate paired signals (memory pressure with pod restarts, latency with packet loss).
  4. 4Emit the recovery workaround with a timebox.
  5. 5Fall back to the hosted environment when local recovery exceeds the box.

Function output

deterministic trace

CPU 72%  MEM 81%  NET 63%  LAT 410ms  LOSS 2.1%  PODS restarts 7  TEMP 71C
AI infrastructure health = 0.61
nudge: memory pressure + pod restart correlation — inspect eviction and container limits.