Blutrain
  1. Home
  2. Capabilities
  3. Voice Infrastructure

Capability

Speech systems built for how India actually speaks.

Contact centre audio here is code-switched, accented, noisy and compressed. Models tuned on clean US English transcribe it badly, and the error rate that matters is on the words that carry the meaning.

01 / The problem

Where this usually goes wrong.

Vendor accuracy figures for speech recognition are quoted on clean read speech. Your reality is an 8kHz telephone channel, a customer on a scooter, two people talking over each other, and a sentence that starts in Hindi, names a product in English and ends in Punjabi.

Aggregate word error rate also hides the important failure. Getting 'the' wrong costs nothing. Getting the account number, the amount or the word 'not' wrong changes the meaning of the whole interaction, and those are exactly the tokens generic models handle worst.

02 / What we build

The parts of the system.

Transcription tuned to your audio
Recognition adapted on your own recordings, your product vocabulary and your channel conditions, with code-switching treated as the normal case rather than an edge case.
Diarisation and structure
Reliable speaker separation, turn segmentation and call structure, which is what makes downstream analysis possible at all.
Analysis that reflects the business
Intent, outcome, compliance-phrase checking and escalation detection, defined with your quality team against a labelled sample rather than a generic taxonomy.
Real-time assistance where it earns its place
Live agent prompting and automated after-call summaries, built within a latency budget that keeps the conversation natural.

03 / How it is measured

What we agree to be judged on.

Set before the build starts, against a measured baseline, and reported honestly afterwards — including where the numbers are disappointing.

Error rate on entities that matter
Measured specifically on numbers, names, amounts, product terms and negations — not as an undifferentiated average across all tokens.
Across accent and channel
Broken out by region, language mix, handset and line quality, because that is where the variance genuinely lives.
Against human QA
Your quality team currently reviews a sample by hand. Agreement with them on that sample is the benchmark for automated scoring.
Handling time and after-call work
The operational payoff, measured against a proper baseline rather than asserted.

04 / Honest limits

What this will not do.

Some audio is not recoverable. Heavy background noise, severe line degradation and crosstalk have a floor that no model clears, and the right answer is often to fix the capture path rather than to keep tuning the model.

Automated quality scoring should also inform human review, not replace it. Systems that assess staff performance need a review route and a human decision-maker, and we will design one in rather than ship a scoring system that grades people unsupervised.

05 / When it applies

You probably need this if:

  • Only a small percentage of calls are reviewed for quality
  • Agents spend significant time on after-call notes
  • Compliance checking is manual and sampled
  • An off-the-shelf speech vendor performed poorly on your recordings

Recognising two or more of these is a reasonable trigger for a diagnostic. Recognising none of them is a reasonable trigger for not spending money here yet.

Next step

Is voice infrastructure the right instrument for your problem?

The diagnostic exists to answer exactly that, including the possibility that the answer is no. It is time-boxed, fixed-fee, and the written assessment is yours.