Site Reliability Engineer
POSITION SUMMARY
- We are seeking a Site Reliability Engineer (SRE) with 6 to 7 years of experience in tech operations and application support, with a strong focus on production reliability, monitoring, and operational ownership. The engineer will own incident resolution, maintain secrets and credentials lifecycle, and drive tech hygiene and operational improvements across business critical applications hosted on AWS.
POSITION RESPONSIBILITIES
- Minimum of 6 to 7 years of experience in SRE, Application Support, or TechOps roles.
- Incident Management: Experience owning end-to-end incident lifecycle including triage, resolution, root cause analysis, and post-incident documentation following ITIL practices.
- Monitoring and Observability: Hands-on experience configuring monitoring, alerting, health checks, and escalation policies using Datadog and PagerDuty. Ability to proactively identify monitoring gaps, reduce alert noise, and ensure correct alert routing and ownership across applications and environments.
- Secrets and credentials management: Experience managing end-to-end secrets lifecycle including rotation scheduling, expiry monitoring, and coordination with application and DBA teams using AWS Secrets Manager, HashiCorp Vault, and Beacon Vault.
- Aws and cloud operations: Working knowledge of AWS S3 and IAM for operational tasks such as
- key management, access provisioning, and permission-related requests.
- Scripting for operational automation: Ability to write and maintain scripts using Bash or equivalent for operational tasks such as alert automation, health checks, and monitoring scripts. AI-assisted scripting tools may be used to support delivery.
- Version control and pipeline onboarding: Basic working knowledge of GitHub and GitHub Actions
- to onboard repositories onto standardised CI/CD workflows provided by the platform team by importing templates and configuring repository-specific variables.
- Documentation and coordination: Ability to create and maintain runbooks, wikis, and operational procedures, and to coordinate effectively with development, platform, infrastructure, and business teams to drive deliverables to closure.
EXPERIENCE AND REQUIRED SKILL SETS
- Minimum of 6 to 7 years of experience in SRE, Application Support, or TechOps roles.
- The engineer will coordinate closely with development, platform, and infrastructure teams on
- operational deliverables, where tasks require skills or capacity beyond the SRE's scope, the
- expectation is to actively coordinate and follow through to closure.
- The role is SRE-first, with DevOps responsibilities introduced progressively based on bandwidth and team need.