Evidence, not testimonials

Report cards,
not case studies.

Every FMS engagement ends with an eval report: baseline, delta, cost, latency, limits. These cards are excerpts from real reports, shared with client permission and lightly anonymized. Ask us for the full methodology on any of them.

01 /

Recent results

Deltas are measured against the stated baseline on held-out data — never on training data, never cherry-picked windows.

Logistics · 2.1M rows
+0.18

Late-delivery risk model

baseline AUC0.71
FMS run fms-23110.89
training cost$610
wall-clock38 min
Fintech · 480k documents
+19.2

Support-ticket router

manual routing acc.74%
FMS run fms-228793.2%
p95 latency31 ms
routing classes14
E-commerce · 3.4M images
+13.2

Listing defect detector

vendor model top-181.5%
FMS run fms-221694.7%
re-train cadenceweekly
false-positive rate−41%
Retail · 88k documents
71.4%

Brand-voice fine-tune (LoRA)

gate: win-rate vs base≥ 65%
measured win-rate71.4%
format adherence98.8%
adapter size68 MB
Healthcare admin · 1.4M rows
+0.11

No-show predictor

baseline AUC0.68
FMS run fms-21940.79
calibration error0.021
training cost$380
SaaS · 620k events
−28%

Churn-risk scorer

churn missed (old rules)41%
churn missed (fms-2158)13%
precision @ top decile0.72
re-train cadencemonthly
02 /

How we measure

Three commitments make these numbers trustworthy — and they are enforced by the pipeline, not by policy.

Stated baselines

The baseline is agreed in the scoping doc before training starts. It appears in the run spec and the eval report side by side with the result.

Held-out suites

Eval data is split at ingest and never enters training. Suites are versioned, so re-trains are compared like-for-like.

Gates, not goals

A run that misses its gate does not package and does not deploy. You are never handed a model that failed its own report card.