Skip to content
GetuJobs

Site Reliability Engineering Lead_Truist

  • Infosys Limited
  • Bangalore
  • 5 - 9 Years
  • Full Time
  • Ansible
  • Kubernetes
  • Python - OpenSystem
  • Site Reliability Engineering(SRE)

Posted August 18, 2026 applications close September 17, 2026


Job Description

Responsibilities

Reliability Engineering & Automation

Architect and deliver automation solutions that eliminate toil, reduce MTTR, and increase service resilience. Experience in Ansible, Puppet or Chef is a plus.

Implement intelligent alerting, anomaly detection, and event correlation leveraging AI and AIOps tools.

Guide and enforce SLO/SLI adoption across product teams, ensuring metrics inform decision-making and prioritization.

Utilize Infrastructure-as-Code (IaC) tools for automating deployment of assets within cloud tenants.

Observability & Operational Excellence

Ensure operational readiness of applications and platforms through resiliency testing, chaos engineering, and failure-mode validation.

Cross-Functional Leadership & Influence

Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution.

Standardization & Documentation

Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns.

Follow best practices and internal processes for Non-Functional requirements to improve resiliency and reliability.

Mentorship & Technical Development

Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline.

Provide thought leadership in SRE methodologies, cloud-native operational patterns, and automated reliability engineering.

Incident Leadership & Production Operations

Lead P1/P0 incident bridges and direct technical investigation efforts.

Perform hands-on triage using logs, traces, metrics, and application telemetry.

Drive mitigation, recovery, RCA development, and follow-through remediation.

Provide executive communications during major incidents.

Build operational automation based on recurring production issues.

Establish credibility through technical leadership during live service disruptions.

Additional Responsibilities

Experience enabling large-scale SRE transformations or modernization initiatives.

Demonstrated proficiency with GitLab Duo, or similar AI technologies.

Familiarity with chaos engineering, resilience assessments, and service failure modeling.

Exposure to hybrid-cloud and multi-cloud operational frameworks.

Experience contributing to or leading Center for Enablement functions or Communities of Practice.

Expertise with highly regulated industries preferred.

Technical and Professional Requirements

7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations.

Deep hands‑on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling.

Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible).

Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring.

Proven leadership in major incident management and cross-team technical coordination.

Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns.

Excellent communication skills, including executive-level situational awareness during critical incidents.

Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.

Preferred Skills

  • Kubernetes
  • Ansible
  • Site Reliability Engineering(SRE)
  • Python – OpenSystem

Educational Requirements

Bachelor Of Comp. Applications,Bachelor Of Computer Science,Bachelor Of Science,Bachelor of Engineering,Bachelor Of Technology

Site Reliability Engineering Lead_Truist Sign in to apply

Browse Job Vacancies