Lead Site Reliability Engineer
Date:
27 Jul 2026
Lead Site Reliability Engineer
Company:
IT & Digital Solutions
Job Purpose
To lead the reliability, availability, performance, and continuous improvement of Air Arabia's mission-critical Java-based applications and airline reservation systems by implementing Site Reliability Engineering (SRE) best practices. Responsible for driving operational excellence through proactive monitoring, automation, incident management, scalability improvements, and collaboration with cross-functional teams, while ensuring compliance with organizational policies, industry standards, and applicable regulatory requirements.
Key Result Responsibilities
- Lead Site Reliability Engineering (SRE) initiatives for Java-based microservices and monolithic applications supporting mission-critical airline operations and Passenger Service Systems (PSS).
- Establish, monitor, and continuously improve system reliability by defining and managing Service Level Agreements (SLAs), Service Level Objectives (SLOs), and error budgets.
- Lead and facilitate root cause analysis (RCA) for complex production incidents, ensuring timely resolution and implementation of preventive measures to minimize recurrence.
- Drive architectural enhancements to improve system performance, scalability, resilience, availability, and operational efficiency through optimization techniques such as caching and distributed system design.
- Mentor and provide technical guidance to team members by promoting best practices in incident management, production support, troubleshooting, and operational excellence.
- Collaborate closely with software engineering, infrastructure, and cross-functional teams to enhance application design, improve system reliability, and ensure production readiness.
- Own and define CI/CD pipeline standards, GitOps practices, and infrastructure automation strategy, setting the framework that engineers across the team execute against.
- Set the containerization and orchestration strategy (Docker, Kubernetes) for the team, defining standards for scalability, resilience, and high availability that other engineers implement.
- Own the on-call escalation framework, ensuring adequate coverage and clear escalation paths, and lead post-incident reviews.
- Evaluate new tools and technologies to strengthen reliability and represent SRE in capacity planning and release readiness reviews.
- Own reliability reporting to management, translating SLA/SLO performance and incident trends into clear business updates.
Qualifications (Academic, training, languages)
- Bachelor’s Degree in Computer Engineering/ Computer Science/ Information Technology.
- Fluent in English Language.
- Strong expertise in Java, Spring Boot, and troubleshooting complex production issues within enterprise environments.
- Airline, aviation, or travel industry experience, particularly with Passenger Service Systems (PSS), is preferred.
- Proficiency in MS Office.
- Strong knowledge of SQL with experience in database design, optimization, and performance tuning; experience with Oracle Database is an advantage.
- Solid understanding of distributed systems, system architecture, scalability, high availability, and resilient application design.
- Proficiency in monitoring, logging, and observability tools such as Prometheus, Grafana, Elasticsearch, or Datadog, used to drive proactive incident detection and reliability strategy.
Work Experience
- With 6-8 years of experience in Software Engineering, SRE, or Production Support (Java Applications).
- Hands-on experience designing, developing, and supporting both microservices and monolithic application architectures.
- Strong hands-on experience with Docker and Kubernetes for containerization, orchestration, and production deployment (mandatory).
- Experience with JBoss Application Server or similar enterprise Java application servers is an added advantage.
- Proven experience designing and governing CI/CD pipeline standards, GitOps practices, and Git-based workflows across multiple teams.
- Hands-on experience with caching technologies (e.g., Redis) and messaging platforms to improve application performance and reliability.