Senior MLOps Engineer
We are seeking an MLOps Engineer to build and own model-serving infrastructure, GPU compute pipelines, and observability foundations for an AI research environment.
In this role, you will lead the transition from research experiments to production-grade deployments. You will work alongside AI Principals, Data Engineers, and Data Scientists to establish robust LLM serving engines, manage multi-cloud GPU compute, and implement modern LLM tracing and automated evaluations.
This is a high-ownership role for an engineer who wants to build modern generative AI infrastructure from the ground up with significant autonomy.
&
• Deploy and optimize high-throughput inference endpoints for open-weight LLMs using serving frameworks such as vLLM, SGLang, and Triton.
• Manage heterogeneous GPU compute allocation across AWS and RunPod, balancing latency, throughput, and hardware efficiency.
• Package model artifacts, PyTorch execution environments, and Hugging Face weights into production-ready containerized deployments.
, &
• Build and maintain an LLM gateway using LiteLLM for API proxying, load balancing, fallback routing, and cost controls.
• Implement tracing, prompt logging, and automated evaluation runs using Langfuse.
• Maintain MLflow pipelines for experiment tracking, model registry management, and artifact versioning.
• Build automated CI/CD deployment pipelines using Docker and GitHub Actions for fast, repeatable model releases.
&
• Partner with Data Engineers to connect ML pipelines and feature sets with a Databricks data foundation.
• Establish FinOps practices and monitoring guardrails to track, optimize, and control GPU and cloud compute spending.
• Deliver reliable, OpenAI-compatible APIs for downstream backend integration.
• Significant experience in software engineering, platform engineering, or MLOps, including direct experience with production ML and AI infrastructure.
• Hands-on experience serving open-weight LLMs using modern inference engines such as vLLM, SGLang, or Triton.
• Experience orchestrating GPU compute on AWS and bare-metal or ephemeral cloud providers such as RunPod.
• Production experience with modern LLM gateway, tracing, or registry tools such as LiteLLM, Langfuse, MLflow, or similar platforms.
• High proficiency in Python, containerization with Docker, and environment management for PyTorch and Hugging Face models.
• Proven ability to independently manage production releases without dedicated DevOps support.
• Experience integrating data pipelines with Databricks or Snowflake.
• Familiarity with Infrastructure as Code and deployment tools such as Terraform and CloudFormation.
• Exposure to vector search engines, retrieval pipelines, and retrieval-augmented generation.
• Production LLM endpoints maintain high throughput, low time to first token, and strong availability.
• Fast, frictionless transitions from model checkpoints to production environments.
• Transparent and predictable cloud compute spending with minimal idle GPU time.
• Robust tracing and evaluation across all production model endpoints.
• Competitive salary.
• Great work culture with a high level of ownership and autonomy.
• Opportunities to work with modern generative AI and MLOps technologies.
• Professional development and continuous learning opportunities.
• Collaborative working environment with experienced AI and data professionals.