Precise recognition · Training data · Any language

Get the words that actually matter exactly right.

LLM-based speech systems are fluent, and they guess. Ask one for a part number, a customer's surname or a serial code and it returns something plausible. Inferret recognises your vocabulary: the fixed lists, proper nouns and alphanumeric codes your business actually runs on. We complement the large models rather than compete with them, and we build the training data behind both.

Building speech & language since 2007 Human and synthetic data pipelines All major languages
2007Founded
1M+Word vocabulary, on-device
2–3 moTo add a new language
100%On-device option, no audio leaves

Trusted by

LINE YahooLINE Yahoo Hitachi SystemsHitachi Systems UsedCar.comUsedCar.com JIEMJIEM Best Path ResearchBest Path Research
01 · Precise word recognition

The complement to LLM-based speech.

Large models transcribe conversational speech beautifully. They are far less reliable the moment a word has to be exactly right rather than merely plausible, because they predict the likeliest next token instead of matching against a known set. Feed SpeechFoundry your vocabulary and it recognises against that list, so the output is constrained to values your systems can actually accept.

Your list, not a guess

Supply the words that matter: part numbers, SKUs, serial ranges, customer surnames, drug names, station names, model codes. Recognition is constrained to that set, so results land on a real record instead of a convincing near-miss.

Alphanumeric codes

Strings like "K7-942A" or a fifteen-digit serial defeat general transcription, which renders them phonetically and inconsistently. We handle spoken letters, digits and separators as structured values, and validate them against the format you define.

Fuzzy and partial match

A user says a fragment of a long proper noun and still resolves to the right entry. Useful wherever names are long, formal or awkward to say in full: phonebooks, track titles, restaurant and business names.

Deterministic and auditable

Because the vocabulary is defined rather than inferred, behaviour is testable and repeatable. You can prove what the system will and will not accept, which matters in regulated and safety-critical work where an LLM's confident invention is unacceptable.

In practice the two approaches sit side by side: a large model handles open conversation and reasoning, while SpeechFoundry handles the fields that must be exact. We routinely build systems that route between them.

Products

Modular pieces, assembled into your product.

Take the whole platform or the one component you're missing. Every piece is customised to your domain. That is the part that decides whether voice actually works.

Voice agents & voice control

Give a device or an application a spoken interface. Users say what they mean, like "it's too hot in here", and the system resolves intent, holds multi-turn context and acts. Language-model reasoning where it helps; deterministic grammars where you need guarantees.

Voice search & assistants

Spoken queries against large catalogues and proper-noun databases: places, products, songs, people. Fuzzy and partial matching means a user can say part of a long name and still land on the right record.

Voice transcription

High-accuracy transcription of large volumes of audio and video, in real time or batch, with optional human verification. Search the transcript, jump straight to that moment in the audio, and get alerted the instant a keyword or phrase is spoken.

Language models, applied

Reasoning, summarisation, extraction and dialog on top of recognition, using open models you can host yourself where privacy or cost demands it. We fine-tune on data built for your domain, keep the deterministic parts deterministic, and evaluate against a benchmark rather than a demo.

Training data & annotation

Human-collected and synthetically generated datasets, annotated to spec, for speech and language models, ours or yours. See how we build it →

Acoustic quality monitoring

Problems have voices too. The same acoustic modelling that recognises speech can learn what a healthy machine sounds like and flag the moment it stops: a wind-turbine rotor changing pitch, a loose part on an assembly line, an unusual noise outside a building.

02 · Training data & annotation

Spoken language models are only as good as the data behind them.

Every model we have built in nearly two decades ran into the same bottleneck. Not architecture: data. So we build data as a product in its own right, collected from real people, generated synthetically at scale, or blended, then annotated and quality controlled to a standard you can actually train on.

Human

Real data, collected and labelled by people

Native speakers and domain specialists producing the ground truth a model is measured against. Speech collection in the acoustic conditions your product actually meets, transcription to a defined style guide, intent and entity labelling, preference and ranking judgements for model alignment, and red-team prompts written by people trying to break things.

  • Recruited to your demographic, accent and domain requirements
  • Written guidelines, calibration rounds, inter-annotator agreement
  • Multi-pass QA with a named reviewer, not a crowd average
  • All major languages in-house, more sourced on request
Synthetic

Generated data, at volumes people can't reach

Language models generating corpora under constraint, speech synthesised across voices and accents, and acoustic augmentation that puts clean audio into noisy rooms, moving cars and telephone channels. The point isn't volume for its own sake. It is deliberate coverage of the long tail and the edge cases your real data will never contain enough of.

  • Corpus generation steered by schema, ontology and domain vocabulary
  • TTS-driven speech data across speaker, accent and prosody variation
  • Noise, reverberation and channel simulation for robustness training
  • Rare-event and adversarial cases produced to order
01Specify

Schema, label taxonomy and quality bar agreed before anyone annotates anything.

02Produce

Human collection, synthetic generation, or a deliberate mix of the two.

03Annotate

Labelled against the guidelines, with agreement measured rather than assumed.

04Validate

Held-out evaluation sets and benchmarks, so you can prove the data moved the metric.

05Deliver

Your format, your infrastructure, with provenance and licensing documented.

Fine-tuning & alignment sets

Instruction, preference and evaluation data for adapting a language model to a domain, including the awkward cases where a general-purpose model quietly gets things wrong.

Evaluation & benchmarks

Held-out test sets built to expose real failure modes rather than flatter a model. If you can't measure it honestly, you can't improve it.

Data your licence allows

Provenance tracked and consent documented, so what you train on is defensible, and increasingly the part that decides whether a model can ship at all.

Capabilities

The details that decide whether people keep using it.

Accuracy in real noise

Designed for low signal-to-noise conditions such as moving vehicles, restaurants and station concourses, not quiet demo rooms.

Natural language understanding

People speak the way they'd speak to a person. No memorised command list, no unnatural phrasing.

Large vocabulary

Language models of a million distinct words and beyond, in both embedded and connected deployments.

Wake word

Always-listening activation on a keyword of your choosing, removing the press-to-talk button entirely.

Barge-in

Users interrupt a prompt mid-sentence and are heard immediately, the way a real conversation works.

Narrowband & wideband

Models tuned for telephony as well as PC and mobile audio, so call-centre channels are not an afterthought.

Dialog & context

Multi-step conversations that retain what was said earlier, so follow-up questions don't start from zero.

Speaker adaptation

Models tune to a primary user's voice and pronunciation over time, and accuracy climbs with use.

Speaker identification

Voices are as distinctive as fingerprints, so you can tell your users apart and personalise accordingly.

Footprint tuning

Accuracy, speed and memory traded off deliberately for your target hardware rather than left to chance.

Technology

Research-grade models, engineered to fit.

Voice and language are what make us human. We build the stack that lets a machine take part in that properly: hearing accurately, grasping meaning, and answering in kind.

01 / RECOGNITION

Speech recognition (ASR)

Deep neural network acoustic models combined with our proprietary, patented weighted finite-state transducer compression. That is the reason we get high accuracy and fast response inside a memory budget that would normally rule both out.

02 / UNDERSTANDING

Natural language understanding

Recognition turns sound into text. It does not tell a machine what the words mean. Our understanding layer closes that gap: users say things as they come to mind, and the system resolves the intent, the entities and what to do next, without forcing anyone to memorise a command list.

03 / REASONING

Language models in the loop

Where a task calls for open-ended reasoning or generation, we bring language models into the pipeline, including open models hosted on your own infrastructure, while keeping recognition and intent handling deterministic and testable. Fine-tuned on data built for your domain, and measured against a benchmark we build alongside it.

04 / DOMAIN

Custom language & semantic models

Every component is built for your project. Generic "cover-everything" APIs cannot be tuned to your vocabulary; a model built around your domain terms, your product names and your users' phrasing can, and that is where the accuracy gap opens up.

05 / RESPONSE

Speech synthesis & dialog

Text-to-speech, dialog management, context management and named entity recognition over very large proper-noun lists: the components needed to close the loop and let a device hold its side of a conversation.

06 / PRIVACY

Sovereign by construction

Because the full engine runs on-device, privacy isn't a policy promise. It is an architecture. Regulated, safety-critical and offline products get the same capability as connected ones, with audio that never leaves the box.

03 · Languages

Speak your users' language.

All major world languages are supported in production, with ready-made generic language models and large vocabularies. Every model comes in wideband for PC and mobile, and narrowband for telephony.

Including

US English UK English Japanese Mandarin Chinese Korean German French Italian NA Spanish + Your language

Need one we don't already run? We have a clear, transparent process for building a production-quality acoustic model for a new language in two to three months, depending on the availability of suitable training data, plus mechanisms to bootstrap a prototype far sooner for testing. Where the data doesn't exist, we can create it: see training data and annotation above. Tell us which languages you need and when, and we'll come back with a plan.

04 · Deployment

Embedded or cloud. Same engine either way.

Most voice stacks force a choice: a small, limited recogniser on the device, or a large capable one in someone else's cloud. SpeechFoundry runs identically in both, with the same core, the same acoustic models and the same language models. You are not locked into where your audio is processed, and you can change your mind late in the product cycle without rewriting anything.

SpeechFoundry Core

DNN acoustic models · patented WFST compression · constrained vocabulary decoding · natural language understanding · dialog & context

Embedded on-device

Ships as an SDK for your platform. Talks to your software over a local API and returns plain JSON. Nothing goes to a network.

  • Runs on constrained hardware, down to single-board class devices
  • Linux, iOS and Android client SDKs
  • Deterministic latency, with no round trip and no outage
  • Audio never leaves the product

Cloud yours or ours

The same engine as a service, on dedicated servers in your infrastructure or ours. Audio and JSON results over standard protocols.

  • Deploy inside your own server estate for data residency
  • High-volume batch transcription and analytics
  • Scale vocabulary and model size beyond device limits
  • Identical results to the embedded build
Industries

Deployed where the acoustics are hard.

Understanding voice makes work easier almost everywhere. Our team has put voice into production across a wide span of environments, most of them noisy, constrained, or both.

Automotive

Embedded voice with wide domain coverage and high accuracy inside the car, where other vendors fall back to a cloud connection they can't guarantee.

Public transport

Noise-robust, multilingual guidance on connections and nearby points of interest for travellers in busy metro stations.

Television

Real-time, high-accuracy subtitles for news and live programming.

Contact centres

Automated agents that handle routine calls end to end, plus intent analysis across very large volumes of recorded conversation.

Mobile apps

Voice replacing fiddly touch interaction in iOS and Android products.

Journalism

Sessions, speeches and panels transcribed live into searchable, formatted text, with instant alerts when a chosen keyword is spoken.

Medical

Language models customised to a clinical domain's terminology so reports can be dictated and searched automatically.

Insurance & claims

Damage descriptions and assessments captured by voice dialog on-site, so an assessor never has to look away to fill in a form.

Language learning

Pronunciation recognised and scored, so a learner gets feedback on how they actually sound.

IoT & smart home

Lights, climate and appliances, all controlled by speech, embedded or connected, with the same engine either way.

Industrial monitoring

Machinery, rotors and processes monitored acoustically, with alerts the moment the sound signature drifts.

Compliance

High-volume batch recognition over recorded interactions, processed inside your own infrastructure.

Services

We're good at listening.

Not every customer needs the same thing from voice, and nobody gets there by having a technology thrown at them. Customisation and support are part of every engagement, whether you have in-house speech experts or this is your first time near the subject.

01

Voice solution customisation

You have a product driven by screens and buttons and you want to add voice, or replace that interface with it. We gather the requirements, suggest what voice should and shouldn't control, and build a prototype that shows you the difference. If it works for you, we take it from prototype to a production-grade voice interface.

02

Custom language & semantic models

The single biggest lever on real-world accuracy. We build language and semantic models around your specific domain rather than handing you a generic model and hoping. This is exactly what the free voice APIs cannot do.

03

Build it yourself

Create, customise and test voice solutions in our browser-based development environment, free, with nothing to install. Get an account, do some initial setup with our team, then start building voice commands and testing them immediately with your own microphone. When you're ready, we help you optimise and port what you built to your embedded or production environment.

04

Transcription as a service

Upload audio or video and get transcripts back. Search them, jump to the relevant moment in the recording, or run the whole pipeline on your own hardware so that nothing ever leaves your building. That last option matters for legal, clinical and HR recordings, where sending audio to a third party is often simply not permitted. We can fold your domain vocabulary into the dictation models and set up keyword alerts that flag or email you the moment a phrase is spoken.

05

Dataset construction

A dataset built to your specification: human collection, synthetic generation, annotation, QA and a held-out evaluation set, whether the model it feeds is ours or entirely your own. This is the fastest-growing part of what we do, and often the highest-leverage thing a team can outsource.

06

Third-party voice consulting

We also improve voice built on someone else's stack. Our team has customised and deployed solutions across a range of third-party voice platforms and assistant frameworks. If your product already has voice and the user feedback is poor, that is a solvable problem and often a fast one.

Company

An engineering company, not a wrapper.

Inferret is a small, agile, international group of researchers and engineers who have been building speech and language technology since 2007, through the statistical era, the deep learning shift, and now the language model era. We own our engine: the acoustic models, the decoder, the compression method, the understanding layer.

That matters more than it used to. When the interesting part of a product is voice or language, depending on an API you cannot inspect, tune, or run on your own hardware is a structural weakness. We build the parts you need to control, and we make them yours.

Having trained models across many languages, we know where they break, and it is almost always the data. That is why dataset construction and annotation now sit alongside the engine as a business in their own right, for teams training their own models as much as for teams using ours.

Work with us
Company
Inferret Limited
Founded
19 July 2007
CEO
Dr. Edward Whittaker
Head office
Booth Rise, Northampton,
NN3 6HP, United Kingdom
Japan branch
Urbane Mitsui Building 7F,
Shintomi 2-4-6, Chuo-ku,
Tokyo, Japan
Registration
England & Wales #6317852
VAT
GB 927810806
法人番号
4700150006061
Share capital
£121,070
Contact

Tell us what you want to build.

Whether you know exactly which component you need or you're still working out whether voice belongs in your product at all, start here. We'll answer with something useful, not a brochure.

United Kingdom · HQ
Inferret Limited
Booth Rise
Northampton NN3 6HP
United Kingdom
Japan · 支店
Inferret Limited
Urbane Mitsui Building 7F
Shintomi 2-4-6, Chuo-ku
Tokyo 104-0041, Japan
03-6809-6789
Urbane Mitsui Building, 7F 〒104-0041 東京都中央区新富2-4-6 Nearest station: Shintomicho (Yurakucho and Hibiya lines)
Open in Google Maps ↗
We reply to every genuine inquiry. Your details are used to answer you and nothing else.