I will deploy and manage mlops, llm and gpu infrastructure on aws and kubernetes

Einige Informationen werden in englischer Sprache angezeigt.

Indien

Ich spreche Hindi, Englisch

19 Aufträge abgeschlossen

Senior AI DevOps Engineer: AWS, Kubernetes, ComfyUI APIs, DevSecOps

AI DevSecOps and Cloud Infrastructure Engineer, 4 years building production systems. I ship what most AI projects are missing: the infrastructure underneath. I work with AWS and EKS, Terraform, Docke...
Über diesen Service

I put AI models into production and keep them running. A training script is not a product. A monitored, versioned, autoscaling service is.

WHAT I BUILD
MLOps pipelines: MLflow or Weights and Biases tracking, model registry, automated retraining and safe rollout
LLM serving: vLLM, Ollama, TGI and Hugging Face models behind an OpenAI compatible API
GPU infrastructure: AWS EC2 g5 and g6, EKS with GPU node groups, RunPod and Vast.ai, spot instances and autoscaling to cut your bill
Containers: Docker, Kubernetes, Helm, GPU scheduling, health checks and zero downtime rollout
CI/CD for models and code: GitHub Actions, Jenkins, ArgoCD and GitOps
Observability: Prometheus, Grafana, request tracing, latency and cost per token dashboards
Security: IAM least privilege, private VPC endpoints, secrets management and SSL
WHAT YOU GET
Working infrastructure, the Terraform or Helm source that you own, runbook documentation, and a handover call so your team can operate it without me.
Four years in DevOps and AI infrastructure. Tell me your model, expected traffic and cloud, and I will size it and give you a realistic monthly cost before you order. I work US hours.

Tools:

Kubernetes

Docker

Amazon EKS

Frameworks:

Terraform

Ansible

Cloud-Provider:

Amazon Web Services

Google Cloud Platform

Programmiersprache:

Bash

Go

Python

Expertise:

Installation

Migration

Debuggen

Meine weiteren Dienstleistungen im Bereich DevOps-Engineering