Every FMS engagement ends with an eval report: baseline, delta, cost, latency, limits. These cards are excerpts from real reports, shared with client permission and lightly anonymized. Ask us for the full methodology on any of them.
Deltas are measured against the stated baseline on held-out data — never on training data, never cherry-picked windows.
| baseline AUC | 0.71 |
| FMS run fms-2311 | 0.89▲ |
| training cost | $610 |
| wall-clock | 38 min |
| manual routing acc. | 74% |
| FMS run fms-2287 | 93.2%▲ |
| p95 latency | 31 ms |
| routing classes | 14 |
| vendor model top-1 | 81.5% |
| FMS run fms-2216 | 94.7%▲ |
| re-train cadence | weekly |
| false-positive rate | −41%▲ |
| gate: win-rate vs base | ≥ 65% |
| measured win-rate | 71.4%▲ |
| format adherence | 98.8% |
| adapter size | 68 MB |
| baseline AUC | 0.68 |
| FMS run fms-2194 | 0.79▲ |
| calibration error | 0.021 |
| training cost | $380 |
| churn missed (old rules) | 41% |
| churn missed (fms-2158) | 13%▲ |
| precision @ top decile | 0.72 |
| re-train cadence | monthly |
Three commitments make these numbers trustworthy — and they are enforced by the pipeline, not by policy.
The baseline is agreed in the scoping doc before training starts. It appears in the run spec and the eval report side by side with the result.
Eval data is split at ingest and never enters training. Suites are versioned, so re-trains are compared like-for-like.
A run that misses its gate does not package and does not deploy. You are never handed a model that failed its own report card.