Hi,
Hope you are doing Good !!!
This is Pavan working with VRDM Technologies.
we are looking for the below Resource if you are interested in this position,
please share me updated resume along with the details below to pa...@vrdmtech.com.
Title : Site Reliability Engineer
Location : Phoenix AZ(Hybrid)[Locals Only]
Duration : Long Term Contract
Visa : USC, GC, GC EAD,
L2 EAD, H4 EAD, Canadian-TN
Job Description
We are seeking a highly experienced Journey-Centric Lead Site
Reliability Engineer (SRE) with Banking and Financial Services (BFS) domain
experience to drive end-to-end reliability, observability, automation, and
operational excellence across critical customer and business journeys. This
role will lead the design and implementation of modern SRE practices, unified
observability platforms, self-healing capabilities, AI-driven operations, and
workflow automation to ensure highly resilient, scalable, and intelligent
digital services.
The ideal candidate combines deep SRE expertise with strong
platform engineering, cloud operations, networking, observability, and AI Ops
experience, with a particular focus on AWS AgentCore-powered operational
intelligence and autonomous operations.
Key Responsibilities
SRE & Reliability Engineering
- Define
and implement enterprise-scale SRE best practices across critical
applications and digital journeys.
- Establish
reliability frameworks, operational standards, and governance models.
- Drive
proactive reliability engineering initiatives to improve system
availability, resilience, and performance.
- Lead
incident management, postmortem analysis, root cause investigations, and
reliability reviews.
Unified Observability & Monitoring
- Design
and implement a unified observability strategy encompassing metrics, logs,
traces, events, and user experience telemetry.
- Build
comprehensive observability dashboards for business and technology
stakeholders.
- Implement
distributed tracing and end-to-end monitoring across complex microservices
ecosystems.
- Define
observability standards and instrumentation frameworks across engineering
teams.
Log Analytics & Trace Correlation
- Enable
unified logs, metrics, and trace correlation capabilities for rapid issue
detection and troubleshooting.
- Deploy
intelligent correlation engines for root cause analysis.
- Improve
Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through
observability-driven insights.
- Establish
service dependency mapping and journey-centric operational visibility.
Self-Healing & Autonomous Operations
- Design
and implement self-healing capabilities using event-driven automation and
AI-assisted remediation.
- Develop
automated recovery processes for common failure scenarios.
- Create
autonomous operational workflows that minimize manual intervention.
- Integrate
predictive alerting and automated response mechanisms.
Automation & Workflow Engineering
- Build
scalable operational automation frameworks.
- Develop
infrastructure, application, observability, and operational workflows
using Infrastructure as Code (IaC), Monitoring as Code (MaC), and
Observability as Code (OaC).
- Automate
deployments, monitoring, remediation, and operational runbooks.
- Reduce
operational toil through intelligent engineering solutions.
SLO, SLA & Error Budget Management
- Define
and govern measurable Service Level Objectives (SLOs), Service Level
Agreements (SLAs), and Error Budgets.
- Partner
with engineering and business teams to align reliability targets with
customer expectations.
- Establish
service maturity metrics and reliability scorecards.
- Drive
data-driven operational decision-making through reliability KPIs.
Network & Platform Reliability
- Apply
deep understanding of:
- Understanding of network layer
to troubleshoot critical bandwidth/latency issues
- No need to pass all these:
- TCP/IP
- DNS
- Load Balancing
- CDN
- API Gateway architectures
- Service Mesh technologies
- VPC and cloud networking
- Troubleshoot
complex network performance and availability issues.
- Ensure
end-to-end reliability across cloud and hybrid environments.
AI Ops & AWS AgentCore
- Implement
and operationalize AI Ops platforms and autonomous operations
capabilities.
- Leverage
AWS AgentCore to build intelligent operational agents for functions like:
- Incident response
- Root cause analysis
- Predictive remediation
- Capacity forecasting
- Automated operational
workflows
- Drive
adoption of GenAI-powered operational intelligence across the enterprise.
- Integrate
AI-assisted observability, automation, and service management
solutions.
Kindly share with us the following details to proceed further
Word / Pdf format resume.