Blutrain
  1. Home
  2. Capabilities

What we build

Ten capabilities, one delivery method.

The problems differ. The discipline does not: establish a baseline, build a thin slice end to end, harden it, and hand it over with the evaluation suite that proves it works.

01 / Method

Why every one of these is delivered the same way.

A recommendation engine and a defect-detection system share almost no domain knowledge and almost all of their engineering risk. Both fail for the same reasons: no baseline, untested data assumptions, no monitoring, and nobody owning the thing after launch.

  1. 01
    Diagnose before proposing
    A time-boxed paid review of your data, your current process and your constraints. Output is a written assessment — feasibility, approach, cost range, and an explicit list of what would make us recommend against proceeding.
  2. 02
    Measure the status quo
    Whatever the process does today becomes the number to beat. If nobody has measured it, measuring it is the first deliverable, and it frequently changes the brief.
  3. 03
    Thin slice, end to end
    One narrow path through ingestion, model, serving and interface, on realistic data. Deliberately narrow so that architectural mistakes surface while they are still cheap.
  4. 04
    Harden against reality
    Retries, timeouts, rate limits, cost ceilings, graceful degradation, an evaluation suite wired into CI, and an explicit answer for what the system does when it is wrong.
  5. 05
    Transfer ownership
    Repository, documentation, runbook, evaluation sets and a working session with the team who will operate it. Ongoing support is available and never assumed.

02 / Error economics

We agree what a mistake costs before we tune a threshold.

Every classifier trades one kind of error for another. Which trade is correct is a business decision, not a technical default — and it changes by sector. A missed fraud case and a wrongly-blocked customer land on different budgets and different people.

We put the four outcomes in front of the people who absorb them and get an explicit decision. It takes an afternoon and it prevents the most common category of post-launch argument.

Caughttrue positiveFalse alarmcosts review timeMissedcosts moneyCorrectly ignoredtrue negativemodel says yesmodel says no
We agree the cost of each error type before we tune a threshold

Next step

Not sure which of these you need?

That is a normal position to be in and a good reason to start with a diagnostic. Describe the process that is slow, expensive or error-prone, and we will tell you which capability applies — or whether none of them do.