- Home
- Capabilities
- Voice Infrastructure
Capability
Speech systems built for how India actually speaks.
Contact centre audio here is code-switched, accented, noisy and compressed. Models tuned on clean US English transcribe it badly, and the error rate that matters is on the words that carry the meaning.
01 / The problem
Where this usually goes wrong.
Vendor accuracy figures for speech recognition are quoted on clean read speech. Your reality is an 8kHz telephone channel, a customer on a scooter, two people talking over each other, and a sentence that starts in Hindi, names a product in English and ends in Punjabi.
Aggregate word error rate also hides the important failure. Getting 'the' wrong costs nothing. Getting the account number, the amount or the word 'not' wrong changes the meaning of the whole interaction, and those are exactly the tokens generic models handle worst.
02 / What we build
The parts of the system.
- Transcription tuned to your audio
- Recognition adapted on your own recordings, your product vocabulary and your channel conditions, with code-switching treated as the normal case rather than an edge case.
- Diarisation and structure
- Reliable speaker separation, turn segmentation and call structure, which is what makes downstream analysis possible at all.
- Analysis that reflects the business
- Intent, outcome, compliance-phrase checking and escalation detection, defined with your quality team against a labelled sample rather than a generic taxonomy.
- Real-time assistance where it earns its place
- Live agent prompting and automated after-call summaries, built within a latency budget that keeps the conversation natural.
03 / How it is measured
What we agree to be judged on.
Set before the build starts, against a measured baseline, and reported honestly afterwards — including where the numbers are disappointing.
- Error rate on entities that matter
- Measured specifically on numbers, names, amounts, product terms and negations — not as an undifferentiated average across all tokens.
- Across accent and channel
- Broken out by region, language mix, handset and line quality, because that is where the variance genuinely lives.
- Against human QA
- Your quality team currently reviews a sample by hand. Agreement with them on that sample is the benchmark for automated scoring.
- Handling time and after-call work
- The operational payoff, measured against a proper baseline rather than asserted.
04 / Honest limits
What this will not do.
Some audio is not recoverable. Heavy background noise, severe line degradation and crosstalk have a floor that no model clears, and the right answer is often to fix the capture path rather than to keep tuning the model.
Automated quality scoring should also inform human review, not replace it. Systems that assess staff performance need a review route and a human decision-maker, and we will design one in rather than ship a scoring system that grades people unsupervised.
05 / When it applies
You probably need this if:
- Only a small percentage of calls are reviewed for quality
- Agents spend significant time on after-call notes
- Compliance checking is manual and sampled
- An off-the-shelf speech vendor performed poorly on your recordings
Recognising two or more of these is a reasonable trigger for a diagnostic. Recognising none of them is a reasonable trigger for not spending money here yet.
06 / Related
Often scoped together.
These capabilities share data, infrastructure or evaluation approach with voice infrastructure, and are frequently part of the same engagement.
Next step
Is voice infrastructure the right instrument for your problem?
The diagnostic exists to answer exactly that, including the possibility that the answer is no. It is time-boxed, fixed-fee, and the written assessment is yours.