Module 12 · cognitive feature
Active intent routing
Not every query deserves a frontier model. A low-latency classifier reads the prompt and hands it to the specialist that owns that lane — cutting inference cost, cutting latency, and cutting hallucinations by keeping each agent inside its own tool set. When the classifier is not confident, it escalates rather than guesses.
$ python -m fde_toolkit router --query "..."
Route a query
workhorse-8bconfidence 1.00Matched: exposure, total
tools in lane: query_warehouse, describe_entity
- SQL2
- Compliance0
- General0
- Investigation0
- Ops0
Specialist roster
- SQL / Data Agent
workhorse-8bReads the semantic layer and answers quantitative questions.
$0.00080/1k · 320ms · hallucination 6.1%
- Compliance Agent
reasoner-midguardrailedInterprets policy, regulation and internal control documents.
$0.00900/1k · 950ms · hallucination 2.8%
- Investigation Agent
frontier-xlguardrailedMulti-step alert triage that must weigh conflicting evidence.
$0.04500/1k · 2400ms · hallucination 1.7%
- Ops / Integration Agent
workhorse-8bRuns the plumbing: job status, extracts, connector health.
$0.00080/1k · 320ms · hallucination 6.1%
- General Assistant
workhorse-8bFallback for chit-chat and open questions with no enterprise scope.
$0.00080/1k · 320ms · hallucination 6.1%
This query
600 output tokens
- Routed
- $0.00060
- Frontier for everything
- $0.02700
- Difference
- −97.8%
- Latency
- 365ms vs 2400ms
Batch of 5 typical queries
What is our total exposure to NMC Health this quarter?
SQL / Data Agent$0.00060 vs $0.02700
Is this transfer to an Iranian beneficiary permitted under our AML policy?
Compliance Agent$0.00552 vs $0.02700
Investigate alert AML-4471 and tell me whether to escalate it.
Investigation Agent$0.02712 vs $0.02700
The nightly KYC extract job failed again — what is the status?
Ops / Integration Agent$0.00060 vs $0.02700
Can you look at this?
General Assistant$0.00060 vs $0.02700