Skip to main content
Posted 24 July, 2026

Senior MLOps Engineer

Unbeatable
Auckland,Auckland,New Zealand,1010 Full Time
Reference: 424_600728_db_39b0fa0731f4c9c9d56b3c0a6a41b942__1093

We are seeking an MLOps Engineer to build and own model-serving infrastructure, GPU compute pipelines, and observability foundations for an AI research environment.

In this role, you will lead the transition from research experiments to production-grade deployments. You will work alongside AI Principals, Data Engineers, and Data Scientists to establish robust LLM serving engines, manage multi-cloud GPU compute, and implement modern LLM tracing and automated evaluations.

This is a high-ownership role for an engineer who wants to build modern generative AI infrastructure from the ground up with significant autonomy.

&

• Deploy and optimize high-throughput inference endpoints for open-weight LLMs using serving frameworks such as vLLM, SGLang, and Triton.

• Manage heterogeneous GPU compute allocation across AWS and RunPod, balancing latency, throughput, and hardware efficiency.

• Package model artifacts, PyTorch execution environments, and Hugging Face weights into production-ready containerized deployments.

, &

• Build and maintain an LLM gateway using LiteLLM for API proxying, load balancing, fallback routing, and cost controls.

• Implement tracing, prompt logging, and automated evaluation runs using Langfuse.

• Maintain MLflow pipelines for experiment tracking, model registry management, and artifact versioning.

• Build automated CI/CD deployment pipelines using Docker and GitHub Actions for fast, repeatable model releases.

&

• Partner with Data Engineers to connect ML pipelines and feature sets with a Databricks data foundation.

• Establish FinOps practices and monitoring guardrails to track, optimize, and control GPU and cloud compute spending.

• Deliver reliable, OpenAI-compatible APIs for downstream backend integration.

• Significant experience in software engineering, platform engineering, or MLOps, including direct experience with production ML and AI infrastructure.

• Hands-on experience serving open-weight LLMs using modern inference engines such as vLLM, SGLang, or Triton.

• Experience orchestrating GPU compute on AWS and bare-metal or ephemeral cloud providers such as RunPod.

• Production experience with modern LLM gateway, tracing, or registry tools such as LiteLLM, Langfuse, MLflow, or similar platforms.

• High proficiency in Python, containerization with Docker, and environment management for PyTorch and Hugging Face models.

• Proven ability to independently manage production releases without dedicated DevOps support.

• Experience integrating data pipelines with Databricks or Snowflake.

• Familiarity with Infrastructure as Code and deployment tools such as Terraform and CloudFormation.

• Exposure to vector search engines, retrieval pipelines, and retrieval-augmented generation.

• Production LLM endpoints maintain high throughput, low time to first token, and strong availability.

• Fast, frictionless transitions from model checkpoints to production environments.

• Transparent and predictable cloud compute spending with minimal idle GPU time.

• Robust tracing and evaluation across all production model endpoints.

• Competitive salary.

• Great work culture with a high level of ownership and autonomy.

• Opportunities to work with modern generative AI and MLOps technologies.

• Professional development and continuous learning opportunities.

• Collaborative working environment with experienced AI and data professionals.

Employment Type: FULL_TIME

Sign up for Job Alerts