At Betsson Group, we strive to deliver the best customer experience in the industry. We are one of the market leaders in iGaming, offering online gaming products in a number of markets, both through our own gaming sites as well as several partner brands.
###
###
Responsibilities
• Incident & Problem Management: Investigate system incidents, drive Root Cause Analysis (RCAs), and execute long-term remedial fixes. Proactively reduce the number of incidents caused by system changes.
• Observability & Metrics: Define and enforce Service Level Agreements (SLAs), Service Level Objectives (SLOs), and success metrics for new initiatives. Build and maintain comprehensive dashboards to achieve observability excellence.
• Performance & Capacity: Identify and help resolve performance bottlenecks. Optimize infrastructure and code to maintain fast service, and conduct capacity planning to forecast future hardware or cloud resource requirements.
• Availability & Change Management: Guarantee the Platform components remain highly reachable and functional for users. Oversee deployments to ensure new code does not disrupt the existing system.
###
Requirements
• Observability & Monitoring: Deep experience building dashboards and tracking SLAs/SLOs using tools like Prometheus, Grafana, Coralogix, Splunk, or Loki.
• Programming & Automation: Proficiency in scripting and coding to automate manual tasks (eliminate "toil") and build reliability tools. Strong skills in .NET, Python, Powershell or Bash are highly preferred.
• Infrastructure as Code (IaC) & Cloud: Experience provisioning and managing infrastructure using Terraform or Ansible, along with a solid understanding of cloud platforms (AWS, GCP, or Azure).
• Containerization & Orchestration: Hands-on experience scaling and managing distributed systems using Kubernetes (K8s) and Docker.
• CI/CD & Change Management: Familiarity with deployment pipelines (GitLab CI, GitHub Actions, Team City, Octopus) to ensure safe, automated rollouts that don't cause incidents.
• Core Competencies: Strong analytical skills for Root Cause Analysis (RCA), a calm approach to incident response, and the ability to lead blameless post-mortems.
• AWS Cloud infrastructure, CDNs, and other various systems running in multiple data centres and environments
• Cloud Application Load Balancer, preferably with experience on AWS ALB
• Cloud DNS support such as AWS Route 53, GCP Cloud DNS, or Azure DNS
• Experience with Microsoft SQL databases, PostgreSQL, and Couchbase is considered an asset.