M14 · Evaluation · drift · champion / challenger

Neural feedback loop

Turn every override, false positive and incident into a governed model or policy change.

$ python -m fde_toolkit retro

Workflow

the order an FDE actually runs it in

  1. 1Capture outcomes, overrides and friction at the point of delivery.
  2. 2Label them into the evaluation set with provenance.
  3. 3Measure drift and online quality against the current champion.
  4. 4Train or tune a challenger and compare on the frozen eval set.
  5. 5Canary, monitor, then promote — or roll back with the evidence.

Function output

deterministic trace

false_positive(ALERT-8821) -> labelled, added to eval set v7
challenger v7: Δprecision +0.04  Δrecall +0.06  drift 0.12  PROMOTE (canary 5%)
issue recurrence 3 -> 1 per cohort, median time-to-resolution 42m -> 17m