Apply now »

Site Reliability Engineer

Date:  14 Sept 2026

Site Reliability Engineer

Company:  IT & Digital Solutions

Job Purpose

To support the reliability, availability, and performance of enterprise applications and mission-critical airline systems by monitoring production environments, resolving incidents, troubleshooting application issues, and implementing reliability improvements. Contributes to maintaining stable and resilient software services through collaboration with Software Engineering and Site Reliability Engineering teams while ensuring compliance with organizational policies, industry standards, and applicable regulatory requirements.

Key Result Responsibilities

  • Investigate day-to-day production incidents by analyzing application logs, system metrics, and monitoring data to identify root causes and escalate complex issues appropriately.
  • Troubleshoot, debug, and support Java-based applications and services under the guidance of senior engineers to ensure day-to-day system stability and availability.
  • Develop, implement, and validate bug fixes and application enhancements under the direction of senior engineers, following established coding and reliability standards.
  • Operate and maintain existing application monitoring, alerting, and observability tools as configured by senior team members, flagging gaps in coverage.
  • Execute day-to-day CI/CD pipeline runs, deployment activities, and release tasks according to established processes and runbooks.

Key Result Responsibilities-Continued

  • Operate and support existing containerized applications on Docker and Kubernetes within pre-defined production configurations, escalating architectural changes to senior engineers.
  • Participate in the on-call rotation, responding to production alerts within defined SLAs and escalating unresolved issues to senior engineers or the Lead SRE.
  • Document incident resolutions, troubleshooting steps, and known issues in runbooks and the team knowledge base to support faster future resolution.
  • Perform routine system health checks, log reviews, and capacity/performance monitoring to proactively flag potential issues before they impact production.
  • Execute scheduled maintenance activities, including patching, backups, and routine configuration changes, following approved change procedures.

Qualifications (Academic, training, languages)

  • Bachelor’s degree in computer engineering/computer science/information technology. 
  • Fluent in English Language.
  • Good knowledge of Java, Spring Boot, and enterprise application development principles.
  • Understanding of microservices and monolithic application architectures and their deployment models.
  • Basic debugging and troubleshooting skills with the ability to diagnose and resolve application and production issues.
  • Familiarity with CI/CD tools (e.g., Jenkins), Git-based version control, and automated deployment practices.
  • Strong knowledge of SQL with experience in database querying and troubleshooting; experience with Oracle Database is an advantage.
  • Familiarity with Linux operating systems, command-line tools, and networking fundamentals.
  • Proficient in MS Office.

Work Experience

  • With 2–4 years of experience in Java development or support.
  • Experience with JBoss Application Server or similar enterprise Java application servers is an added advantage.
  • Hands-on experience with Docker and Kubernetes for deploying and supporting containerized applications.

Apply now »