Infra SRE consultant/engineer required in CA(12+yrs exp. needed) - Onsite

0 views
Skip to first unread message

Radha Venkatraman

unread,
Oct 22, 2025, 3:26:47 PM10/22/25
to c2c requirements, staffing partners, urgent reks, urgentreqz, us staff augmentation, us staffing india, Alex, Anil, Arun1, Bala, Charan, David, Feroz, Koti, Mady, Maruthi, Phani, Pradeep, Praveen, Rani, Ratnakar, Rik, Simha, Siva, Sridhar, Srikanth, Srini, Tarun, Umesh, Venkatesh

 

Dear friends,

 Please respond with suitable profiles with contact details, visa copy, DL, Passport number and linkedin for the below urgent requirement.

 Position: Infra SRE consultant/ engineer (12+yrs exp. Needed)

Location: Santa Clara, CA - Onsite

Type: long term contract

Rate: open

Position summary:

Top Skills:

· Strong focus on observability tools like Prometheus, Grafana and practices

· Excellent problem-solving and troubleshooting skills, capable of handling escalations from L1 up to L4.

· Able to identify feature level issues and work with development teams to resolve them.

· Kubernetes expertise , should be able to manage deployments, perform deep-level debugging, and handle Kubernetes administration tasks.

· Strong Linux/Unix fundamentals, including system-level operations / engineering skills.

 

Job Description/Responsibilities:

● On-prem infrastructure management

Manage on-prem infrastructure. Maintain uptime, reliability and readiness of on-prem engineering cloud spread across multiple data centers.

● Guard SLAs Guard service level agreements (SLAs) for critical engineering services. Implement monitoring, alerting, and incident response procedures to ensure adherence to defined performance targets. Perform root cause analysis and post-mortems of incidents for any threshold breaches.

● Observability

Set up and manage monitoring and logging tools such as Prometheus, Grafana, or the ELK Stack to oversee system health and performance. Maintain KPI pipelines using Jenkins, Python and ELK.

Improve monitoring systems by adding custom alerts based on business needs.

● Automation & Optimization

Help in capacity planning, optimization and better utilization efforts.

● Day-to-Day Support

Support user reported issues & issues. Monitor alerts and take necessary action.

Actively participate in WAR room for critical issues

● Collaboration & Documentation

Create and maintain documentation for operational procedures, configurations, and troubleshooting guides.

Tech stack

– Baremetal data center machine management tools like IPMI, Redfish, KVM etc.

– Automation using Jenkins, Python, Go, Bash.

– Infrastructure tools like Kubernetes, MySQL, Prometheus, Grafana and ELK.

– Any familiarity with Nvidia hardware like GPU & Tegras is a plus

Years of Experience:     12.00 Years of Experience

 

 

Regards,

Radha Venkatraman

Recruitment Lead
ReqRoute,Inc
Desk: 408-600-2008; Fax: 888-400-2698
Email: ra...@reqroute.com

 

 

Reply all
Reply to author
Forward
0 new messages