Precise recognition · Training data · Any language
Get the words that actually matter exactly right.
LLM-based speech systems are fluent, and they guess. Ask one for a part number, a
customer's surname or a serial code and it returns something plausible. Inferret
recognises your vocabulary: the fixed lists, proper nouns and alphanumeric
codes your business actually runs on. We complement the large models rather than
compete with them, and we build the training data behind both.
Building speech & language since 2007Human and synthetic data pipelinesAll major languages
SPEECHFOUNDRY / STREAMINGon-device
ASR
NLU
2007Founded
1M+Word vocabulary, on-device
2–3 moTo add a new language
100%On-device option, no audio leaves
Trusted by
LINE YahooHitachi SystemsUsedCar.comJIEMBest Path Research
01 · Precise word recognition
The complement to LLM-based speech.
Large models transcribe conversational speech beautifully. They are far less reliable
the moment a word has to be exactly right rather than merely plausible, because they
predict the likeliest next token instead of matching against a known set. Feed
SpeechFoundry your vocabulary and it recognises against that list, so the output is
constrained to values your systems can actually accept.
Your list, not a guess
Supply the words that matter: part numbers, SKUs, serial ranges, customer surnames, drug names, station names, model codes. Recognition is constrained to that set, so results land on a real record instead of a convincing near-miss.
Alphanumeric codes
Strings like "K7-942A" or a fifteen-digit serial defeat general transcription, which renders them phonetically and inconsistently. We handle spoken letters, digits and separators as structured values, and validate them against the format you define.
Fuzzy and partial match
A user says a fragment of a long proper noun and still resolves to the right entry. Useful wherever names are long, formal or awkward to say in full: phonebooks, track titles, restaurant and business names.
Deterministic and auditable
Because the vocabulary is defined rather than inferred, behaviour is testable and repeatable. You can prove what the system will and will not accept, which matters in regulated and safety-critical work where an LLM's confident invention is unacceptable.
In practice the two approaches sit side by side: a large model handles open conversation
and reasoning, while SpeechFoundry handles the fields that must be exact. We routinely
build systems that route between them.
Products
Modular pieces, assembled into your product.
Take the whole platform or the one component you're missing. Every piece is customised
to your domain. That is the part that decides whether voice actually works.
Voice agents & voice control
Give a device or an application a spoken interface. Users say what they mean, like "it's too hot in here", and the system resolves intent, holds multi-turn context and acts. Language-model reasoning where it helps; deterministic grammars where you need guarantees.
Voice search & assistants
Spoken queries against large catalogues and proper-noun databases: places, products, songs, people. Fuzzy and partial matching means a user can say part of a long name and still land on the right record.
Voice transcription
High-accuracy transcription of large volumes of audio and video, in real time or batch, with optional human verification. Search the transcript, jump straight to that moment in the audio, and get alerted the instant a keyword or phrase is spoken.
Language models, applied
Reasoning, summarisation, extraction and dialog on top of recognition, using open models you can host yourself where privacy or cost demands it. We fine-tune on data built for your domain, keep the deterministic parts deterministic, and evaluate against a benchmark rather than a demo.
Training data & annotation
Human-collected and synthetically generated datasets, annotated to spec, for speech and language models, ours or yours. See how we build it →
Acoustic quality monitoring
Problems have voices too. The same acoustic modelling that recognises speech can learn what a healthy machine sounds like and flag the moment it stops: a wind-turbine rotor changing pitch, a loose part on an assembly line, an unusual noise outside a building.
02 · Training data & annotation
Spoken language models are only as good as the data behind them.
Every model we have built in nearly two decades ran into the same bottleneck. Not
architecture: data. So we build data as a product in its own right, collected from real
people, generated synthetically at scale, or blended, then annotated and quality
controlled to a standard you can actually train on.
Human
Real data, collected and labelled by people
Native speakers and domain specialists producing the ground truth a model is measured against. Speech collection in the acoustic conditions your product actually meets, transcription to a defined style guide, intent and entity labelling, preference and ranking judgements for model alignment, and red-team prompts written by people trying to break things.
Recruited to your demographic, accent and domain requirements
Written guidelines, calibration rounds, inter-annotator agreement
Multi-pass QA with a named reviewer, not a crowd average
All major languages in-house, more sourced on request
Synthetic
Generated data, at volumes people can't reach
Language models generating corpora under constraint, speech synthesised across voices and accents, and acoustic augmentation that puts clean audio into noisy rooms, moving cars and telephone channels. The point isn't volume for its own sake. It is deliberate coverage of the long tail and the edge cases your real data will never contain enough of.
Corpus generation steered by schema, ontology and domain vocabulary
TTS-driven speech data across speaker, accent and prosody variation
Noise, reverberation and channel simulation for robustness training
Rare-event and adversarial cases produced to order
01Specify
Schema, label taxonomy and quality bar agreed before anyone annotates anything.
02Produce
Human collection, synthetic generation, or a deliberate mix of the two.
03Annotate
Labelled against the guidelines, with agreement measured rather than assumed.
04Validate
Held-out evaluation sets and benchmarks, so you can prove the data moved the metric.
05Deliver
Your format, your infrastructure, with provenance and licensing documented.
Fine-tuning & alignment sets
Instruction, preference and evaluation data for adapting a language model to a domain, including the awkward cases where a general-purpose model quietly gets things wrong.
Evaluation & benchmarks
Held-out test sets built to expose real failure modes rather than flatter a model. If you can't measure it honestly, you can't improve it.
Data your licence allows
Provenance tracked and consent documented, so what you train on is defensible, and increasingly the part that decides whether a model can ship at all.
Capabilities
The details that decide whether people keep using it.
Accuracy in real noise
Designed for low signal-to-noise conditions such as moving vehicles, restaurants and station concourses, not quiet demo rooms.
Natural language understanding
People speak the way they'd speak to a person. No memorised command list, no unnatural phrasing.
Large vocabulary
Language models of a million distinct words and beyond, in both embedded and connected deployments.
Wake word
Always-listening activation on a keyword of your choosing, removing the press-to-talk button entirely.
Barge-in
Users interrupt a prompt mid-sentence and are heard immediately, the way a real conversation works.
Narrowband & wideband
Models tuned for telephony as well as PC and mobile audio, so call-centre channels are not an afterthought.
Dialog & context
Multi-step conversations that retain what was said earlier, so follow-up questions don't start from zero.
Speaker adaptation
Models tune to a primary user's voice and pronunciation over time, and accuracy climbs with use.
Speaker identification
Voices are as distinctive as fingerprints, so you can tell your users apart and personalise accordingly.
Footprint tuning
Accuracy, speed and memory traded off deliberately for your target hardware rather than left to chance.
Technology
Research-grade models, engineered to fit.
Voice and language are what make us human. We build the stack that lets a machine take
part in that properly: hearing accurately, grasping meaning, and answering in kind.
01 / RECOGNITION
Speech recognition (ASR)
Deep neural network acoustic models combined with our proprietary, patented weighted finite-state transducer compression. That is the reason we get high accuracy and fast response inside a memory budget that would normally rule both out.
02 / UNDERSTANDING
Natural language understanding
Recognition turns sound into text. It does not tell a machine what the words mean. Our understanding layer closes that gap: users say things as they come to mind, and the system resolves the intent, the entities and what to do next, without forcing anyone to memorise a command list.
03 / REASONING
Language models in the loop
Where a task calls for open-ended reasoning or generation, we bring language models into the pipeline, including open models hosted on your own infrastructure, while keeping recognition and intent handling deterministic and testable. Fine-tuned on data built for your domain, and measured against a benchmark we build alongside it.
04 / DOMAIN
Custom language & semantic models
Every component is built for your project. Generic "cover-everything" APIs cannot be tuned to your vocabulary; a model built around your domain terms, your product names and your users' phrasing can, and that is where the accuracy gap opens up.
05 / RESPONSE
Speech synthesis & dialog
Text-to-speech, dialog management, context management and named entity recognition over very large proper-noun lists: the components needed to close the loop and let a device hold its side of a conversation.
06 / PRIVACY
Sovereign by construction
Because the full engine runs on-device, privacy isn't a policy promise. It is an architecture. Regulated, safety-critical and offline products get the same capability as connected ones, with audio that never leaves the box.
03 · Languages
Speak your users' language.
All major world languages are supported in production, with ready-made generic language
models and large vocabularies. Every model comes in wideband for PC and mobile, and
narrowband for telephony.
Including
US EnglishUK EnglishJapaneseMandarin ChineseKoreanGermanFrenchItalianNA Spanish+ Your language
Need one we don't already run? We have a clear, transparent process for building a
production-quality acoustic model for a new language in two to three months,
depending on the availability of suitable training data, plus mechanisms to bootstrap a
prototype far sooner for testing. Where the data doesn't exist, we can create it: see
training data and annotation above.
Tell us which languages you need and when, and we'll come back with a plan.
04 · Deployment
Embedded or cloud. Same engine either way.
Most voice stacks force a choice: a small, limited recogniser on the device, or a large
capable one in someone else's cloud. SpeechFoundry runs identically in both,
with the same core, the same acoustic models and the same language models. You are not
locked into where your audio is processed, and you can change your mind late in the
product cycle without rewriting anything.
Ships as an SDK for your platform. Talks to your software over a local API and returns plain JSON. Nothing goes to a network.
Runs on constrained hardware, down to single-board class devices
Linux, iOS and Android client SDKs
Deterministic latency, with no round trip and no outage
Audio never leaves the product
Cloud yours or ours
The same engine as a service, on dedicated servers in your infrastructure or ours. Audio and JSON results over standard protocols.
Deploy inside your own server estate for data residency
High-volume batch transcription and analytics
Scale vocabulary and model size beyond device limits
Identical results to the embedded build
Industries
Deployed where the acoustics are hard.
Understanding voice makes work easier almost everywhere. Our team has put voice into
production across a wide span of environments, most of them noisy, constrained, or both.
Automotive
Embedded voice with wide domain coverage and high accuracy inside the car, where other vendors fall back to a cloud connection they can't guarantee.
Public transport
Noise-robust, multilingual guidance on connections and nearby points of interest for travellers in busy metro stations.
Television
Real-time, high-accuracy subtitles for news and live programming.
Contact centres
Automated agents that handle routine calls end to end, plus intent analysis across very large volumes of recorded conversation.
Mobile apps
Voice replacing fiddly touch interaction in iOS and Android products.
Journalism
Sessions, speeches and panels transcribed live into searchable, formatted text, with instant alerts when a chosen keyword is spoken.
Medical
Language models customised to a clinical domain's terminology so reports can be dictated and searched automatically.
Insurance & claims
Damage descriptions and assessments captured by voice dialog on-site, so an assessor never has to look away to fill in a form.
Language learning
Pronunciation recognised and scored, so a learner gets feedback on how they actually sound.
IoT & smart home
Lights, climate and appliances, all controlled by speech, embedded or connected, with the same engine either way.
Industrial monitoring
Machinery, rotors and processes monitored acoustically, with alerts the moment the sound signature drifts.
Compliance
High-volume batch recognition over recorded interactions, processed inside your own infrastructure.
Services
We're good at listening.
Not every customer needs the same thing from voice, and nobody gets there by having a
technology thrown at them. Customisation and support are part of every engagement,
whether you have in-house speech experts or this is your first time near the subject.
01
Voice solution customisation
You have a product driven by screens and buttons and you want to add voice, or replace that interface with it. We gather the requirements, suggest what voice should and shouldn't control, and build a prototype that shows you the difference. If it works for you, we take it from prototype to a production-grade voice interface.
02
Custom language & semantic models
The single biggest lever on real-world accuracy. We build language and semantic models around your specific domain rather than handing you a generic model and hoping. This is exactly what the free voice APIs cannot do.
03
Build it yourself
Create, customise and test voice solutions in our browser-based development environment, free, with nothing to install. Get an account, do some initial setup with our team, then start building voice commands and testing them immediately with your own microphone. When you're ready, we help you optimise and port what you built to your embedded or production environment.
04
Transcription as a service
Upload audio or video and get transcripts back. Search them, jump to the relevant moment in the recording, or run the whole pipeline on your own hardware so that nothing ever leaves your building. That last option matters for legal, clinical and HR recordings, where sending audio to a third party is often simply not permitted. We can fold your domain vocabulary into the dictation models and set up keyword alerts that flag or email you the moment a phrase is spoken.
05
Dataset construction
A dataset built to your specification: human collection, synthetic generation, annotation, QA and a held-out evaluation set, whether the model it feeds is ours or entirely your own. This is the fastest-growing part of what we do, and often the highest-leverage thing a team can outsource.
06
Third-party voice consulting
We also improve voice built on someone else's stack. Our team has customised and deployed solutions across a range of third-party voice platforms and assistant frameworks. If your product already has voice and the user feedback is poor, that is a solvable problem and often a fast one.
Company
An engineering company, not a wrapper.
Inferret is a small, agile, international group of researchers and engineers who have
been building speech and language technology since 2007, through the statistical era,
the deep learning shift, and now the language model era. We own our engine: the
acoustic models, the decoder, the compression method, the understanding layer.
That matters more than it used to. When the interesting part of a product is voice or
language, depending on an API you cannot inspect, tune, or run on your own hardware is a
structural weakness. We build the parts you need to control, and we make them yours.
Having trained models across many languages, we know where they break, and it is almost
always the data. That is why dataset construction and annotation now sit alongside the
engine as a business in their own right, for teams training their own models as much as
for teams using ours.
Whether you know exactly which component you need or you're still working out whether
voice belongs in your product at all, start here. We'll answer with something useful,
not a brochure.