AI & ML interests
Speech foundation models for transcription, speaker diarization, vocal emotion, and delivery. We help AI understand what people say, how they say it, and what they mean—directly from the audio.
Recent Activity
Articles
Organization Card
We train speech models to help AI understand what people say, how they say it, and what they mean—directly from the audio.
Try the demo · Explore our models · API docs · Research
Speech understanding, from the audio
Words and speakers. Transcribe speech and follow who said what, with time-coded segments.
Emotion and delivery. Add context from tone, rhythm, emphasis, and vocal expression.
One API. Bring transcripts and acoustic context into the products you build.
Open research
- Orukeet — our open multilingual speech recognition model, with code, weights, evaluation results, and a technical report.
- oruk-bench — published speech-emotion evaluation results, with methods and limitations documented alongside the scores.
- Research highlights — experiments, model releases, and interactive explanations from the team.
Built by researchers from Stanford, Berkeley, and Cambridge. Meet the team →
Website · GitHub · Brand assets · Contact
