Site Reliability Engineering Lead_Truist
-
Infosys Limited
- Bangalore
- 5 - 9 Years
- Full Time
- Ansible
- Kubernetes
- Python - OpenSystem
- Site Reliability Engineering(SRE)
Posted August 18, 2026 applications close September 17, 2026
Please sign in or register for free to apply.
Job Description
Responsibilities
Reliability Engineering & Automation
Architect and deliver automation solutions that eliminate toil, reduce MTTR, and increase service resilience. Experience in Ansible, Puppet or Chef is a plus.
Implement intelligent alerting, anomaly detection, and event correlation leveraging AI and AIOps tools.
Guide and enforce SLO/SLI adoption across product teams, ensuring metrics inform decision-making and prioritization.
Utilize Infrastructure-as-Code (IaC) tools for automating deployment of assets within cloud tenants.
Observability & Operational Excellence
Ensure operational readiness of applications and platforms through resiliency testing, chaos engineering, and failure-mode validation.
Cross-Functional Leadership & Influence
Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution.
Standardization & Documentation
Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns.
Follow best practices and internal processes for Non-Functional requirements to improve resiliency and reliability.
Mentorship & Technical Development
Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline.
Provide thought leadership in SRE methodologies, cloud-native operational patterns, and automated reliability engineering.
Incident Leadership & Production Operations
Lead P1/P0 incident bridges and direct technical investigation efforts.
Perform hands-on triage using logs, traces, metrics, and application telemetry.
Drive mitigation, recovery, RCA development, and follow-through remediation.
Provide executive communications during major incidents.
Build operational automation based on recurring production issues.
Establish credibility through technical leadership during live service disruptions.
Additional Responsibilities
Experience enabling large-scale SRE transformations or modernization initiatives.
Demonstrated proficiency with GitLab Duo, or similar AI technologies.
Familiarity with chaos engineering, resilience assessments, and service failure modeling.
Exposure to hybrid-cloud and multi-cloud operational frameworks.
Experience contributing to or leading Center for Enablement functions or Communities of Practice.
Expertise with highly regulated industries preferred.
Technical and Professional Requirements
7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations.
Deep hands‑on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling.
Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible).
Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring.
Proven leadership in major incident management and cross-team technical coordination.
Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns.
Excellent communication skills, including executive-level situational awareness during critical incidents.
Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.
Preferred Skills
- Kubernetes
- Ansible
- Site Reliability Engineering(SRE)
- Python – OpenSystem
Educational Requirements
Bachelor Of Comp. Applications,Bachelor Of Computer Science,Bachelor Of Science,Bachelor of Engineering,Bachelor Of Technology