Senior MLOps Engineer
About the role
About the Institute of Foundation Models (IFM)
The Institute of Foundation Models is a dedicated research lab for building, understanding, deploying, and risk-managing large-scale AI systems. We drive innovation in foundation models and their operationalization, empowering research, education, and industry adoption through scalable infrastructure and real-world applications.
As part of our engineering team, you will operate at the intersection of machine learning and systems design - building the cloud, orchestration, and deployment layers that power the next generation of intelligent applications at MBZUAI. You'll work alongside world-class AI researchers and engineers to productionize LLMs, voice models, and multimodal systems at scale.
The Role
As a Senior MLOps Engineer, you will design, build, and maintain robust ML(Machine Learning) infrastructure across training, inference, and deployment pipelines. You will take ownership of the model lifecycle - from data ingestion to real-time serving - and ensure our LLM and speech models are deployed efficiently, securely, and reproducibly in Kubernetes-based environments.
This position requires deep hands-on experience with Kubernetes (EKS), Helm, AWS cloud infrastructure, and modern MLOps toolchains (e.g., vLLM, SGLang, OpenWebUI, Weights & Biases, MLflow). Familiarity with speech/voice AI frameworks like ElevenLabs, Whisper, and RVC is also valuable.
Responsibilities
Key Responsibilities
- Design and manage scalable ML infrastructure onAWSusingEKS,EC2,RDS,S3, andIAM-based access control.
- Build and maintainKubernetes deploymentsfor LLM and TTS inference usingHelm,ArgoCD, andPrometheus/Grafanamonitoring.
- Implement and optimizemodel serving pipelinesusingvLLM,SGLang,TensorRT, or similar frameworks for high-throughput inference.
- Develop CI/CD and MLOps automation for data versioning, model validation, and deployment (GitHub Actions, Jenkins, or AWS CodePipeline).
- IntegrateOpenWebUI,Gradio, or similar UIs for user-facing model demos and internal evaluation tools.
- Collaborate with ML researchers to productize models - including TTS (e.g., ElevenLabs API), ASR (Whisper), and LLM-based chat systems.
- Ensureobservability, cost optimization, and reliabilityof cloud resources across multiple environments.
- Contribute to internal tools fordataset curation, model monitoring, and retraining pipelines.
- Maintain infrastructure-as-code usingTerraformandHelm chartsfor reproducibility and governance.
- Support real-time multimodal workloads (voice, text, vision) across inference clusters.
Professional Experience - Preferred
- Extensive Experience withvLLM, K8s, Elevenlabs,Whisper,Gradio/OpenWebUI, or custom TTS/ASR model hosting.
- Familiarity withmulti-GPU scheduling,NCCL optimization, andHPC cluster integration.
- Knowledge ofsecurity,cost management, andnetwork policyin multi-tenant Kubernetes clusters andcloudflaresystems.
- Prior work inLLM deployment,fine-tuning pipelines, orfoundation model research.
- Exposure todata governanceandresponsible AI operationsin research or enterprise settings.
Requirements
Academic
Qualifications
- 4+ years of experience inMLOps,DevOps, orCloud Infrastructure Engineeringfor ML systems.
- Strong proficiency inKubernetes,Helm, andcontainer orchestration.
- Experience deploying ML models viavLLM,SGLang,TensorRT, orRay Serve.
- Proficiency withAWS services(EKS, EC2, S3, RDS, CloudWatch, IAM).
- Solid experience withPython,Docker,Git, andCI/CD pipelines.
- Strong understanding ofmodel lifecycle management,data pipelines, andobservability tools(Grafana, Prometheus, Loki).
- Excellent collaboration skills with ML researchers and software engineers.