Senior DevOps Infrastructure Engineer (MLOps / GPU Platforms)
We are looking for a Senior to Staff Infrastructure Engineer to lead the design and evolution of large-scale, multi-GPU compute infrastructure used to train next-generation robotics and AI models. This role sits at the intersection of DevOps, MLOps, and distributed systems — owning architecture, reliability, and performance at scale in a fast-moving, cutting-edge environment.
Essential functions
Lead architecture and long-term technical direction of multi-GPU, cross-cloud training platforms
Build and evolve infrastructure-as-code for provisioning, orchestration, and lifecycle management
Architect and improve CI/CD systems for infrastructure and ML training workflows
Optimize distributed training workloads — scheduling, resource utilization, observability
Partner with ML engineers and researchers to enable efficient experimentation and productionization
Mentor engineers and drive operational excellence across the org
Document architecture, systems, and key technical decisions
Qualifications
Production-grade Kubernetes experience (CKA preferred)
Hands-on Terraform (infrastructure-as-code)
Kubernetes packaging and release management via Helm
AWS cloud operations experience
CI/CD pipeline experience including self-hosted runners (GitHub Actions)
Prometheus/Grafana monitoring and alerting
Linux administration, containerization, scripting (Python & Bash)
Availability for on-call rotation
Would be a plus
GPU-accelerated Kubernetes clusters (NVIDIA)
Cluster autoscaling (Karpenter)
Workflow orchestration (Prefect)
Gang scheduling, fair-share resource allocation
High-performance storage (FSx for Lustre, EFS)
We offer
- Opportunity to work on bleeding-edge projects
- Work with a highly motivated and dedicated team
- Competitive salary
- Flexible schedule
- Benefits package - medical insurance, sports
- Corporate social events
- Professional development opportunities
- Well-equipped office
About us
Grid Dynamics (NASDAQ: GDYN) is a leading provider of technology consulting, platform and product engineering, AI, and advanced analytics services. Fusing technical vision with business acumen, we solve the most pressing technical challenges and enable positive business outcomes for enterprise companies undergoing business transformation. A key differentiator for Grid Dynamics is our 8 years of experience and leadership in enterprise AI, supported by profound expertise and ongoing investment in data, analytics, cloud & DevOps, application modernization and customer experience. Founded in 2006, Grid Dynamics is headquartered in Silicon Valley with offices across the Americas, Europe, and India.Apply to the position
Thank you!
You applied for the position Senior DevOps Infrastructure Engineer (MLOps / GPU Platforms) successfully. We will get back to you soon. Have a great day!
Something went wrong...
There are possible difficulties with connection or other issues. Please try to use another browser (it's recommended to use the latest version of Google Chrome browser). If the problem still persists, please send your application to cv@griddynamics.com
Something went wrong...
Please double-check the information filled in the form, and make sure to provide valid data.
Don’t see the right opportunity?
Contact us anyway and let’s talk! To apply, send your resume and cover letter to jobs@griddynamics.com
Grid Dynamics is an equal opportunity employer. We are committed to creating an inclusive environment for all employees during their employment and for all candidates during the application process.
All qualified applicants will receive consideration for employment without regard to, and will not be discriminated against based on, age, race, gender, color, religion, national origin, sexual orientation, gender identity, veteran status, disability or any other protected category. All employment is decided on the basis of qualifications, merit, and business need.
