Rime — Machine Learning Scientist

job posting·professional role·active

Official active full-time posting for Machine Learning Scientist at Rime, located in United States.

Recorded facts

Official sitehttps://jobs.ashbyhq.com/rime/b76243da-3ce0-463e-9f29-2b6d1b78666e
GeographyUnited States
job board posting idb76243da-3ce0-463e-9f29-2b6d1b78666e
teamModeling
locationUnited States
seniorityunspecified
departmentModeling
exact titleMachine Learning Scientist
role familymachine_learning_and_research
closing dateunknown
compensationdisclosed: false · parse state: not_publicly_disclosed
posting date2026-05-15
employer nameRime
posting statusactive
workplace typeremote
employment typefull-time
required skillsDeep familiarity with the speech synthesis literature, contemporary and historical — Tacotron, FastSpeech, VITS, VALL-E, the codec-LM lineage. Opinions on what worked and why.; Hands-on training with neural codecs (EnCodec, DAC, Mimi, etc.) and multiple representation choices.; Experience with full- or half-duplex multi-modal modeling (Moshi, LLaMA-Omni, streaming S2S).; Strong attention to detail on data quality. You notice when an annotation pipeline is silently degrading or when an eval set has leakage.; Willing to roll up your sleeves on unglamorous data and training work — paired with the agency to build pipelines so the team isn't stuck doing it by hand.; Working knowledge of TTS frontend (G2P, normalization, prosody) and experience working with linguists.; Strong PyTorch fundamentals. Comfortable with training loops, distributed training, model internals.; PhD or equivalent research experience in speech, audio, ML, or computational linguistics or a track record that makes the credential irrelevant.
selection basisofficial_current_board_freshness_or_domain_critical_role
employer subtypespeech_audio_ai_company
preferred skillsMultilingual TTS experience.; Background in prosody or paralinguistics.; Published work in speech, audio, or core ML venues.; Experience taking research models to production: quantization, distillation, streaming inference.; WHY JOIN RIME; Category-defining voice AI infrastructure, not incremental research deltas.; Direct collaboration with founders, including a CEO with a Stanford computational linguistics PhD.; Real impact on company trajectory.; Meaningful equity upside.; High ownership, high standards, low bureaucracy.
responsibilitiesDesign, train, and evaluate speech synthesis models, autoregressive and non-autoregressive.; Drive research on full-duplex and half-duplex multi-modal architectures, including unified S2S systems.; Choose and iterate on speech representations: neural codecs, semantic tokens, mel features, continuous latents.; Build rigorous evaluation, objective and perceptual. Hold the bar on quality and prosodic control.; Collaborate with our linguists on TTS frontend behavior so modeling and frontend choices reinforce each other.
work authorizationunknown
education requirementsPhD or equivalent research experience in speech, audio, ML, or computational linguistics or a track record that makes the credential irrelevant.

Current

posted byRime 2026-05-15nowSource 1 VERIFIED high confidence

Sources & changes

Checked 10d ago · highhow verification works
  • Status unknown → active
  • Official URL unknown → jobs.ashbyhq.com/rime/b76243da-3ce0-463e-9f29-2b6d1b78666e
  • Record maintenance · 12 fields updated

Is this yours? Claim this record →·See something wrong? Report a correction →