Rime — Machine Learning Scientist
job posting·professional role·active
Official active full-time posting for Machine Learning Scientist at Rime, located in United States.
Recorded facts
| Official site | https://jobs.ashbyhq.com/rime/b76243da-3ce0-463e-9f29-2b6d1b78666e ↗ |
|---|---|
| Geography | United States |
| job board posting id | b76243da-3ce0-463e-9f29-2b6d1b78666e |
| team | Modeling |
| location | United States |
| seniority | unspecified |
| department | Modeling |
| exact title | Machine Learning Scientist |
| role family | machine_learning_and_research |
| closing date | unknown |
| compensation | disclosed: false · parse state: not_publicly_disclosed |
| posting date | 2026-05-15 |
| employer name | Rime |
| posting status | active |
| workplace type | remote |
| employment type | full-time |
| required skills | Deep familiarity with the speech synthesis literature, contemporary and historical — Tacotron, FastSpeech, VITS, VALL-E, the codec-LM lineage. Opinions on what worked and why.; Hands-on training with neural codecs (EnCodec, DAC, Mimi, etc.) and multiple representation choices.; Experience with full- or half-duplex multi-modal modeling (Moshi, LLaMA-Omni, streaming S2S).; Strong attention to detail on data quality. You notice when an annotation pipeline is silently degrading or when an eval set has leakage.; Willing to roll up your sleeves on unglamorous data and training work — paired with the agency to build pipelines so the team isn't stuck doing it by hand.; Working knowledge of TTS frontend (G2P, normalization, prosody) and experience working with linguists.; Strong PyTorch fundamentals. Comfortable with training loops, distributed training, model internals.; PhD or equivalent research experience in speech, audio, ML, or computational linguistics or a track record that makes the credential irrelevant. |
| selection basis | official_current_board_freshness_or_domain_critical_role |
| employer subtype | speech_audio_ai_company |
| preferred skills | Multilingual TTS experience.; Background in prosody or paralinguistics.; Published work in speech, audio, or core ML venues.; Experience taking research models to production: quantization, distillation, streaming inference.; WHY JOIN RIME; Category-defining voice AI infrastructure, not incremental research deltas.; Direct collaboration with founders, including a CEO with a Stanford computational linguistics PhD.; Real impact on company trajectory.; Meaningful equity upside.; High ownership, high standards, low bureaucracy. |
| responsibilities | Design, train, and evaluate speech synthesis models, autoregressive and non-autoregressive.; Drive research on full-duplex and half-duplex multi-modal architectures, including unified S2S systems.; Choose and iterate on speech representations: neural codecs, semantic tokens, mel features, continuous latents.; Build rigorous evaluation, objective and perceptual. Hold the bar on quality and prosodic control.; Collaborate with our linguists on TTS frontend behavior so modeling and frontend choices reinforce each other. |
| work authorization | unknown |
| education requirements | PhD or equivalent research experience in speech, audio, ML, or computational linguistics or a track record that makes the credential irrelevant. |
Current
| posted by | Rime 2026-05-15 — nowSource 1 ↗ VERIFIED high confidence |
|---|
Sources & changes
Checked 10d ago · highhow verification works
Field-level evidence
Public change history
- Status unknown → active
- Official URL unknown → jobs.ashbyhq.com/rime/b76243da-3ce0-463e-9f29-2b6d1b78666e
- Record maintenance · 12 fields updated
Is this yours? Claim this record →·See something wrong? Report a correction →
