Site Reliability Engineer (SRE) / Production Support Engineer - #1170651
EXPRESS PTE. LTD.
Job Summary
We are seeking an experienced Site Reliability Engineer (SRE) / Production Support Engineer with 10+ years of experience in enterprise infrastructure, production support, system administration, and SRE operations. The ideal candidate will have strong hands-on experience in AIX, Linux/RHEL, application production support, monitoring, automation, incident management, and cloud/DevOps environments.
The candidate will be responsible for ensuring high availability, reliability, performance, and security of critical enterprise applications and infrastructure in a 24x7 production support environment. The role requires close collaboration with development, infrastructure, middleware, security, and business teams to identify and resolve production issues and continuously improve service reliability.
Key Responsibilities
Site Reliability Engineering
- Define and establish SLIs, SLOs, SLAs, error budgets, MTTD, and MTTR for enterprise applications and services.
- Monitor application availability, latency, performance, capacity, and overall reliability.
- Identify and implement Golden Signals and observability best practices across applications and infrastructure.
- Analyze production incidents and recurring issues to identify root causes and implement permanent remediation.
- Develop automation and engineering solutions to reduce manual operational activities and improve reliability.
- Participate in 24x7x365 production support and on-call operations.
- Perform capacity planning and proactively identify infrastructure and application resource requirements.
- Support highly available and resilient application architectures.
- Participate in disaster recovery planning, testing, and implementation.
Production & Infrastructure Support
- Provide L2/L3 production and infrastructure support for critical enterprise applications.
- Perform administration, troubleshooting, configuration, maintenance, and performance tuning of IBM AIX and RHEL/Linux servers.
- Support server patching, upgrades, maintenance, and DR activities.
- Troubleshoot application, middleware, operating system, connectivity, and infrastructure-related issues.
- Perform middleware administration and restart activities for WebSphere, JBoss, IBM HTTP Server, and Apache Tomcat.
- Support firewall changes, SSL certificate renewals, SCP configuration, and system-to-system key exchanges.
- Support application deployment activities across Blue/Green environments.
- Coordinate application-related changes and deployments across development, infrastructure, middleware, and business teams.
Monitoring & Observability
- Configure and maintain monitoring and alerting for infrastructure and application services.
- Grafana
- Centreon
- AppDynamics
- ELK / Kibana
- Application and infrastructure logging platforms
- Analyze system and application metrics, logs, and performance trends.
- Develop appropriate alerts to proactively identify service degradation and failures.
- Promote observability practices and help development teams implement effective monitoring.
Incident & Change Management
- Manage and resolve incidents within defined SLA/OLA timelines.
- Participate in major incident and emergency response activities.
- Perform incident investigation, troubleshooting, root-cause analysis, and problem management.
- Raise and manage Change Requests (CRs) for application deployments, BAU fixes, infrastructure changes, middleware changes, and maintenance activities.
- Create and manage service requests and incident tickets.
- Prepare RCA and corrective/preventive action plans for recurring and high-priority incidents.
- Follow ITIL-based Incident, Change, Problem, and Service Request Management processes.
Automation & DevOps
- Develop automation scripts using Python and Shell scripting to improve operational efficiency.
- Automate repetitive infrastructure and production support activities.
- Git / GitHub / Bitbucket
- Jenkins
- Ansible
- Chef
- Docker
- Maven
- JFrog / Nexus Repository
- SonarQube
- Fortify / Nexus IQ
- Support CI/CD pipelines and application deployment processes.
- Collaborate with development teams to integrate monitoring, logging, and reliability controls into deployment pipelines.
Cloud & Application Support
- Provide support for applications hosted on AWS and Pivotal Cloud Foundry (PCF) environments.
- IBM WebSphere
- IBM HTTP Server
- JBoss
- Apache Tomcat
- Java applications
- Support microservices-based applications and their associated infrastructure.
- Assist with application migration, optimization, and adoption of new technologies where required.
Security & Compliance
- Apply system security best practices across AIX and Linux environments.
- Support security hardening, patching, access management, and vulnerability remediation.
- Work with security and access management tools such as IBM Tivoli Access Manager.
- Support SSL certificate management and renewal activities.
- Prepare quarterly operational reports and documentation required for audit, risk, and compliance activities.
- Ensure production changes and operational activities comply with organizational security and governance standards.
Required Technical Skills
Operating Systems
- IBM AIX 5.x / 6.x / 7.x
- Red Hat Enterprise Linux (RHEL)
- Linux
- Windows Server
SRE / Monitoring
- SLI / SLO / SLA
- Error Budgets
- MTTD / MTTR
- Observability
- Golden Signals
- Grafana
- Centreon
- AppDynamics
- ELK / Kibana
Cloud & Containers
- AWS
- Pivotal Cloud Foundry (PCF)
- Docker
Middleware & Application Technologies
- IBM WebSphere Application Server
- IBM HTTP Server
- JBoss
- Apache Tomcat
- Java
Databases
- DB2 UDB
- SQL
- MySQL / MariaDB
Automation & Scripting
- Python
- Shell Scripting
- Groovy
DevOps / CI-CD
- Git / GitHub
- Bitbucket
- Jenkins
- Ansible
- Chef
- Maven
- JFrog / Nexus Repository
- SonarQube
- Fortify
- Nexus IQ
ITSM / Ticketing
- ServiceNow
- BMC Remedy
- IBM ISM
- JIRA
Soft Skills
- Strong analytical and troubleshooting skills.
- Excellent incident management and problem-solving capabilities.
- Ability to work effectively in a high-pressure 24x7 production environment.
- Strong communication and stakeholder management skills.
- Ability to collaborate effectively with Development, Infrastructure, Middleware, Security, and Business teams.
- Strong ownership and accountability for production services.
- Ability to prioritize critical issues and coordinate resolution during major incidents.
Education
- Bachelor's degree in engineering / technology / computer science or equivalent.
- B.Tech / B.E. preferred.
Experience
- 10+ years of overall IT experience in SRE, Production Support, Infrastructure Support, or System Administration.
- Strong hands-on experience with AIX and Linux/RHEL administration.
- Experience supporting enterprise applications in banking, financial services, or other mission-critical environments is preferred.
- Experience working in 24x7x365 production support and on-call environments.
Preferred Profile
The ideal candidate will be a hands-on SRE/Production Support professional with strong expertise in AIX/Linux administration, enterprise application support, observability, incident management, automation, DevOps, and cloud technologies, with a proven ability to improve application reliability, reduce operational effort, and maintain high service availability.
How to apply
To apply for this job you need to authorize on our website. If you don't have an account yet, please register.
Post a resumeSimilar jobs
Manager, Attractions Designer
Senior Specialist / Assistant Manager, Marketing and Corporate Communications
Trust Administrator