ROLE: SRE Solutions Architect – AI, HPC & GPU
Location: Remote | Preferred: Santa Clara, CA, Washington, Oregon, Texas
Duration: 12+ Months Contract
Travel: Occasional
Visa: USC/GC/H1B/EAD GC/H4EAD/TN Visa/L2)
Experience: 12 Years
We are looking for an experienced SRE Solutions Architect with strong expertise in AI/GPU infrastructure, HPC, cloud operations, and production reliability.
Key Skills:
8+ years in SRE, Cloud Engineering, Infrastructure, HPC, or Solutions Architecture
5+ years specialized experience with large-scale GPU/AI infrastructure preferred
GPU infrastructure, AI platforms, HPC & distributed systems
Kubernetes, Slurm, Linux, Python, Bash
InfiniBand, NCCL, UFM, High-Speed Ethernet
Prometheus, Grafana, OpenTelemetry
Terraform, Ansible, Argo CD, GitOps
DCGM, BMC, Redfish, firmware/driver lifecycle
Lustre, GPFS, WEKA, VAST or other HPC storage
Strong troubleshooting, observability, incident response and capacity planning
Experience with production GPU environments and customer-facing technical leadership
Experience with GB200/GB300 NVL72, Spectrum-X, GPU Operators, automated diagnostics, and AI-assisted operations is highly preferred.
Ideal Background: CoreWeave, Lambda, Crusoe, Vast.ai, AWS, Azure, Google Cloud, Oracle Cloud, HPC, SRE, AI Infrastructure, or GPU Cloud environments.
Thanks & Regards,
Pavan Ronanki
Delivery Manager, Siatech Solutions.
33 Wood Avenue South, Suite # 600 | Iselin, NJ 08830

[We are an E-Verify and Equal Employment Opportunity Employer with adherence to EEO policy]
