- Home
- Capabilities
- Risk & Fraud
Capability
Scoring that survives an auditor and an adversary.
Fraud and credit models differ from most machine learning in two uncomfortable ways: someone is actively trying to defeat them, and someone else will demand an explanation for every decision.
01 / The problem
Where this usually goes wrong.
Fraud data is adversarial and non-stationary. A model trained on last year's patterns is being evaluated by people whose full-time occupation is finding what it does not catch. Performance decays for reasons that have nothing to do with code quality.
At the same time, the model operates under scrutiny. A decision that adversely affects a customer may need to be explained to that customer, to an internal reviewer, and to a regulator — years later, using the model version that was live at the time.
02 / What we build
The parts of the system.
- Features engineered for evasion
- Velocity, network and behavioural features rather than static attributes, because static attributes are the easiest thing in the system for an adversary to change.
- Explainable-by-construction scoring
- Model families and feature designs chosen so that every score decomposes into reasons a human reviewer can read, evaluate and, where necessary, defend.
- A full evidence trail
- Every decision recorded with the model version, feature values, threshold and outcome, retained per your regulatory obligations and reconstructable on request.
- A rapid response path
- New fraud patterns handled with rules deployable in hours while retraining proceeds on its own timescale. Waiting for a model retrain during an active attack is not a plan.
03 / How it is measured
What we agree to be judged on.
Set before the build starts, against a measured baseline, and reported honestly afterwards — including where the numbers are disappointing.
- Detection at a fixed review capacity
- Your team can investigate a finite number of cases per day. The meaningful question is how much fraud is caught within that budget.
- Cost-weighted, not count-weighted
- Catching ten small cases and missing one large one is a loss. Evaluation is denominated in money.
- Decay rate
- How quickly performance degrades after deployment, measured deliberately. This sets the retraining cadence instead of a guess.
- Fairness across segments
- Decision rates and error rates broken out by protected and proxy characteristics, examined before launch and monitored after.
04 / Honest limits
What this will not do.
No scoring system catches everything, and one tuned to try will block enough legitimate customers to cost more than the fraud. The operating point is a commercial decision, taken with the people who own both sides of that cost.
These models also require ongoing investment by their nature. A risk model is not a project that finishes; treating it as one is the most reliable way to have an expensive incident eighteen months later.
05 / When it applies
You probably need this if:
- Fraud losses are rising or manual review cannot keep up
- Rules have accumulated for years and nobody knows what they collectively do
- Credit decisions are inconsistent between reviewers
- A regulator or auditor has asked how a decision was reached
Recognising two or more of these is a reasonable trigger for a diagnostic. Recognising none of them is a reasonable trigger for not spending money here yet.
06 / Related
Often scoped together.
These capabilities share data, infrastructure or evaluation approach with risk & fraud, and are frequently part of the same engagement.
Next step
Is risk & fraud the right instrument for your problem?
The diagnostic exists to answer exactly that, including the possibility that the answer is no. It is time-boxed, fixed-fee, and the written assessment is yours.