Description de l’emploi
About the Role We're looking for a DevOps Engineer to build the infrastructure foundation that keeps our AI-powered product fast, reliable, secure, and cost-efficient. You'll set up the cloud, deployment, and monitoring systems that let a small team ship confidently and scale smoothly through the pilot and beyond.
What You'll Do
• Design, provision, and manage cloud infrastructure (AWS, GCP, or Azure)
• Build and maintain CI/CD pipelines for fast, safe, automated deployments
• Implement infrastructure-as-code (Terraform, Pulumi, or similar) for reproducible environments
• Containerize and orchestrate services (Docker, Kubernetes) as the system grows
• Set up observability: monitoring, logging, alerting, and tracing across the stack
• Manage and optimize the cost, latency, and reliability of LLM and AI workloads
• Own security, secrets management, and access controls across environments
• Establish backup, disaster-recovery, and incident-response practices for the pilot
• 4+ years of DevOps, SRE, or infrastructure engineering experience
• Strong hands-on experience with at least one major cloud provider (AWS/GCP/Azure)
• Proficiency with infrastructure-as-code tools (Terraform, Pulumi, CloudFormation)
• Experience with containerization and orchestration (Docker, Kubernetes)
• Solid CI/CD experience (GitHub Actions, GitLab CI, CircleCI, or similar)
• Strong scripting skills (Bash, Python, or Go)
• Experience implementing monitoring and observability tooling (Prometheus, Grafana, Datadog, etc.)
• Security-first mindset and experience with secrets and access management
• Experience managing infrastructure for AI/ML or LLM-heavy workloads
• Familiarity with cost optimization for high-throughput API usage
• Experience with GPU infrastructure or inference serving
• Early-stage experience standing up infrastructure from scratch
Originally posted on Himalayas