Rime — Software Engineer, ML Serving
job posting·professional role·active
Official active full-time posting for Software Engineer, ML Serving at Rime, located in United States.
Recorded facts
| Official site | https://jobs.ashbyhq.com/rime/3ce06fc8-1896-4fa3-99d4-54b615480f59 ↗ |
|---|---|
| Geography | United States |
| job board posting id | 3ce06fc8-1896-4fa3-99d4-54b615480f59 |
| team | Infra and Engineering |
| location | United States |
| seniority | unspecified |
| department | Infra and Engineering |
| exact title | Software Engineer, ML Serving |
| role family | machine_learning_and_research |
| closing date | unknown |
| compensation | disclosed: false · parse state: not_publicly_disclosed |
| posting date | 2026-06-18 |
| employer name | Rime |
| posting status | active |
| workplace type | hybrid |
| employment type | full-time |
| required skills | Hands-on experience with real-time multinode ML serving infrastructure — ML serving framework experience: NVIDIA Dynamo/Triton, vLLM, SGLang, or equivalent.; Experience with distributed or disaggregated model serving (Tensor Parallel, Pipeline Parallel, or equivalent).; Strong cloud infrastructure fundamentals: Linux internals, networking, containerization (Docker, Kubernetes).; IaC experience — Terraform, Packer, or comparable. You should have opinions about how to do this right.; On-call is part of the job. You treat production reliability as a shared responsibility. |
| selection basis | official_current_board_freshness_or_domain_critical_role |
| employer subtype | speech_audio_ai_company |
| preferred skills | Experience with multinode training (DDP, FSDP, etc.).; Experience with gRPC or other bidirectional binary streaming protocols.; Experience with audio streaming and related technologies (WebRTC, WebSockets, etc.).; Experience with a multilingual monorepo where you pick the best language out of merit more than personal experience.; Experience with multi-cloud infrastructures (AWS, GCP, OCI, etc.).; Comfort with configuration management tooling (Ansible, Chef, Puppet, or similar).; SRE, DevOps, or platform engineering background at a startup.; Experience at an early-stage company.; WHY JOIN RIME; Build the serving infrastructure behind a category-defining voice AI company from the ground up.; You will bring in experience that no one else currently has at the company: you can help us set the vision.; Direct collaboration with the inference, platform, and ML teams — no handoff culture.; The systems you build determine what experiences our customers can deploy at scale.; Meaningful equity upside at an early stage.; High ownership, high standards, low bureaucracy.; SF / Bay Area. |
| responsibilities | Architecture and implementation of Rime's TTS serving infrastructure, from GPU-backed inference engines to the API surface.; Model optimization from a single-node to disaggregated fleet serving.; Compatibility with different NVIDIA hardwares from Hopper to Blackwell and beyond for on-prem and cloud deployments.; Continuous integration and deployment workflows for the model serving pipeline.; Site reliability: on-call rotation, monitoring, alerting, and observability across the serving stack.; Resource provision, cost management across our GPU fleet. |
| work authorization | unknown |
| education requirements | unknown |
Current
| posted by | Rime 2026-06-18 — nowSource 1 ↗ VERIFIED high confidence |
|---|
Sources & changes
Checked 10d ago · highhow verification works
Field-level evidence
Public change history
- Status unknown → active
- Official URL unknown → jobs.ashbyhq.com/rime/3ce06fc8-1896-4fa3-99d4-54b615480f59
- Record maintenance · 12 fields updated
Is this yours? Claim this record →·See something wrong? Report a correction →
