Job
Title: Sr_SRE
Location:
Atlanta GA – ONLY Local for In-person Interview
Client:
Delta Air Lines
Backfill
role
Work
Mode for delta roles:
50%
Hybrid – Work From Office Schedule
The
role follows a 50% hybrid work model. Employees are required to work from the
office for 5 consecutive business days, from Wednesday through the following
Tuesday, on alternate weeks.
Job
Description:
Enterprise
Site Reliability Engineer
Position
Summary
We
are seeking an experienced Enterprise Site Reliability Engineer to advance
reliability, resilience, observability, automation, and operational excellence
across the organization.
This
is a highly visible, hands-on engineering role that will work across SRE,
application, architecture, cloud, platform, infrastructure, security, data, and
other IT domains. The successful candidate will combine deep technical
expertise with curiosity, thorough analysis, and strong engineering judgment to
identify risks others may overlook, connect technical findings to business
impact, and develop scalable enterprise solutions.
The
role requires someone who can move effectively between detailed technical
analysis, software development, application architecture reviews, complex
incident leadership, enterprise standards, and executive-level communication.
The candidate must be able to influence SRE and engineering practices across
organizational boundaries without relying on direct authority.
Key
Responsibilities
- Define and advance enterprise SRE standards for SLIs,
SLOs, error budgets, production readiness, incident management, and
operational excellence.
- Establish practical observability standards for
OpenTelemetry, traces, logs, metrics, events, telemetry correlation,
tagging, data quality, and service ownership.
- Conduct thorough application maturity assessments
across observability, reliability, resilience, operability, automation,
and incident readiness.
- Analyze application architecture and production
operations to uncover dependencies, failure modes, capacity constraints,
performance bottlenecks, and risks that may not be immediately visible.
- Connect technical and operational findings to customer
experience, business impact, service risk, and investment priorities.
- Translate assessments and operational data into clear,
prioritized, and measurable improvement roadmaps.
- Influence architecture and engineering decisions by
recommending reliability patterns such as fault isolation, graceful
degradation, circuit breakers, retries, rate limiting, load shedding, high
availability, and disaster recovery.
- Design end-to-end observability across distributed
services, APIs, business transactions, customer journeys, cloud platforms,
and cross-domain dependencies.
- Develop production-grade automation, applications,
APIs, and integrations that reduce toil, improve consistency, and scale
reliability practices across the enterprise.
- Automate service onboarding, telemetry validation, SLO
reporting, production readiness checks, incident enrichment, and
remediation workflows.
- Integrate observability, cloud, CI/CD, IT service
management, incident management, and configuration platforms using APIs,
SDKs, webhooks, and event-driven patterns.
- Lead complex incident investigations and post-incident
reviews, challenge assumptions, identify contributing factors, and drive
corrective actions through completion.
- Analyze operational and observability data to detect
patterns, quantify risks, identify systemic gaps, and generate actionable
insights for engineers and leaders.
- Proactively identify opportunities to improve
application operability, resilience, telemetry, support readiness, and
operational processes before issues affect customers.
- Define observability for AI agents and AI-enabled
applications, including workflows, model and tool interactions,
dependencies, latency, failures, quality, token consumption, cost, and
reliability.
- Use AI-assisted engineering tools such as Kiro,
GitHub Copilot, and AI agents to accelerate development, incident
analysis, correlation, pattern detection, and operational insights.
- Develop reusable reference architectures, engineering
patterns, assessment frameworks, maturity models, scorecards, and
implementation guidance.
- Work with SREs and engineering teams across the
organization to resolve cross-domain reliability issues and promote
consistent practices.
- Mentor engineers, facilitate difficult technical
decisions, constructively challenge existing approaches, and build
alignment across teams.
- Communicate complex technical risks, business impact,
recommendations, and progress clearly to engineering teams and senior
leadership.
Required
Qualifications
- 8+ years of experience in Site Reliability Engineering,
Production Engineering, Platform Engineering, or software engineering for
large-scale production systems.
- Deep knowledge of SRE principles, operational
excellence, distributed systems, microservices, APIs, and cloud-native
architecture.
- Deep hands-on experience architecting, operating, and
troubleshooting large-scale AWS environments across compute, containers,
serverless, networking, databases, storage, identity, and cloud
observability, including EC2, EKS, Lambda, VPC, Elastic Load Balancing,
Route 53, RDS/Aurora, DynamoDB, S3, IAM, and CloudWatch.
- Strong hands-on experience with Kubernetes and Red
Hat OpenShift Service on AWS (ROSA) architecture, operations,
performance, and troubleshooting.
- Strong experience with OpenTelemetry, distributed
tracing, logs, metrics, events, and Dynatrace or a comparable enterprise
observability platform.
- Experience defining and operationalizing SLIs, SLOs,
error budgets, production readiness criteria, and reliability scorecards.
- Demonstrated ability to assess application architecture
and operational maturity from observability, reliability, resilience, and
operability perspectives.
- Strong software development skills in Python, Java, Go,
JavaScript/TypeScript, or similar languages.
- Experience developing production-grade APIs,
integrations, automation services, and internal engineering tools.
- Proficiency with REST APIs, SDKs, Git, automated
testing, CI/CD, Infrastructure as Code, secure development, and software
lifecycle practices.
- Experience leading complex incidents, technical
investigations, root cause analysis, post-incident reviews, and
corrective-action programs.
- Strong analytical and investigative skills, with the
curiosity and technical judgment to look beyond immediate symptoms and
uncover systemic problems.
- Demonstrated ability to convert technical findings and
operational data into clear, actionable insights and enterprise
recommendations.
- Ability to connect reliability risks and technical
decisions to customer experience, business impact, and organizational
priorities.
- Excellent written, verbal, technical, and executive
communication skills.
- Proven ability to mentor experienced engineers,
challenge assumptions constructively, facilitate decisions, and influence
outcomes without direct authority.
- Ability to collaborate effectively with SREs,
architects, application teams, and specialists across multiple IT domains.
Preferred
Qualifications
- AWS certification, preferably AWS Certified Solutions
Architect Professional, AWS Certified DevOps Engineer Professional, or a
relevant specialty certification.
- Kubernetes or OpenShift certification, such as CKA,
CKS, or Red Hat Certified OpenShift Administrator.
- Experience integrating enterprise platforms such as
Dynatrace, AWS, ServiceNow, GitLab, Kubernetes, and ROSA.
- Experience developing self-service reliability
capabilities, internal developer platforms, or enterprise engineering
products.
- Experience with performance engineering, capacity
planning, resilience testing, and chaos engineering.
- Experience establishing observability and reliability
controls for AI agents, LLM-enabled applications, or AI-driven workflows.
- Experience working with large-scale, highly available,
business-critical enterprise systems.