Title: Site reliability Engineer
Location: Dallas, TX
Duration: 6 Months and extendable
Description:
Are you the next Service Reliability Engineer we are looking for?
- You
will accelerate Application teams’ ability to reliably and consistently
deliver applications by developing standardized automation to control,
build, artifact and deploy managed services, integrated into loosely
coupled toolchains, to form a common continuous deployment pipeline for
application development teams.
- You
will be responsible for Capacity Planning, Change Management, Problem
Management, Incident Management, Release Management, Performance
improvement and automation as well as tool development.
- You
will be expected to excel under pressure, work well with others, be
self-motivated and able to manage short- and long-term projects.
Implementing automation for kickstarting, monitoring, management and
support will be a key component of the position.
- You
will actively interface with software developers, network engineers,
system support, storage, project management and database administrators
on pro-jects and providing second tier on call support.
- You
will have to identify the root cause, troubleshoot and resolve issues
quickly and effectively, sometimes under pressure. Good communication
and teamwork are extremely important.
- You will be participating in the 24x7 on-call rotation of the team
In this role you’ll:
- Support an ultra-highly available cloud-based applicative platform for our customers.
- Support application deployments, building new systems, upgrading and patching existing ones.
- Develop automation to quickly and rapidly deploy instances from blue-printed applications or golden images.
- Develop
and use monitoring tools to find problems, resolve and/or escalate to
development and ensure that we exceed our SLAs.
- Build and manage development and testing environments, assisting developers in debugging application issues using tools.
- Participate in the building of tools and processes to support the infrastructure.
- Leverage scripting to build required automation and tools on an ad-hoc basis.
- Operate the platform within our security and privacy guidelines.
- Learning on the job and explore new technologies with little supervision.
- Ability to use a wide variety of open-source technologies and tools.
- Experience with systems and IT operations.
- Comfort with frequent, incremental code testing and deployment.
- A strong focus on business outcomes.
- Strong sense of collaboration, open communication and reaching beyond functional borders.
- Provide hands-on engineering, administration and technical support.
- Troubleshoot issues across the entire stack - hardware, software, application and network.
- Document current and future configuration processes and policies.
- Proactive thought leadership for creative and efficient technology solutions.
- Drive continuous improvement to the service delivered solutions to customers (agility, stability)
- Process reengineering and optimization.
- Drive
the enforcement and definition of operational requirements and
non-functional requirements in collaboration with application owners and
middleware organizations.
About the ideal candidate:
- Education: requires a bachelor’s degree (or foreign equivalent) in Computer Science, Engineering, or a related field
- Relevant Work Experience:
- Minimum 3-5 years in systems administration/Software Engineering/DevOps, networking in a large environment.
- Minimum 3 years' experience of application build and release engineering in SOA architectures.
- Business Understanding:
- Excellent
understanding of Software Engineering methodologies and development
cycle (Open-Source development), including Version Control systems (GIT
and Sub-Version) as well as Continuous Integration and testing methods
(Jenkins)
- Strong knowledge on Service Oriented Architecture design patterns
- Good knowledge in Networking is including:
- Communication Protocols (TCP/IP, DNS, SSH, HTTP/S)
- Load balancing techniques, traffic routing, and caching for distributed applications, scalability
- Identifying, troubleshooting, and resolving system level issues on large, busy networks.
- Proficiency
in deployment and infrastructure configuration management tools (such
as Maven, Capistrano, Puppet, NPM, etc.), especially Ansible
- Excellent knowledge in Linux operating system administration (RHEL or SLES)
- Working under Linux
- Good understanding of Linux Containers deployment technologies (Docker or LXC)
- Good knowledge of C, C++ or Java, also Shell, Perl, GO or Python
- Understanding of monitoring tools and concepts (Kibana, Elasticsearch), especially Grafana
- Very
good understanding of Cloud concepts and Cloud computing and related
ecosystems (Cloud Stack, OpenStack, AWS API, Azure, OpenShift/Kubernetes
etc.)
- Virtualization Technology (such as EC2, Xen, KVM, OpenStack)
- Very good knowledge in relational DB (Oracle, MySQL, MariaDB) and NoSQL technology (Hadoop, MongoDB, Couchbase) a plus
- Good understanding of security information and event management technologies
- Good exposure to Agile methodologies is a plus
- Curious to learn new stuff every day
- Willing to automate his/her work
- Feeling responsible about his/her platform
- Excellent written and verbal communication skills
- Conflict resolution-oriented Team Player
- Support Engineering
- Skills:
- Languages: English is a MUST, French would be appreciated
- Specific Knowledge: EDIFACT protocol, Pricing
- Other:
- Communication skills, teamwork, Software Development Methodologies (Agile and Waterfall), Coordination Skills, Leadership.
- Experience with customer and premium support in premises of an airline are mandatory
--
Thanks & Regards
Chai
NAVA Software Solutions LLC
Phone: 860-780-7890
E-Verified Company | Certified MBE