M14 · Evaluation · drift · champion / challenger
Neural feedback loop
Turn every override, false positive and incident into a governed model or policy change.
$ python -m fde_toolkit retro
Workflow
the order an FDE actually runs it in
- 1Capture outcomes, overrides and friction at the point of delivery.
- 2Label them into the evaluation set with provenance.
- 3Measure drift and online quality against the current champion.
- 4Train or tune a challenger and compare on the frozen eval set.
- 5Canary, monitor, then promote — or roll back with the evidence.
Function output
deterministic trace
false_positive(ALERT-8821) -> labelled, added to eval set v7 challenger v7: Δprecision +0.04 Δrecall +0.06 drift 0.12 PROMOTE (canary 5%) issue recurrence 3 -> 1 per cohort, median time-to-resolution 42m -> 17m
