Auf einen Blick
Site Reliability Engineer bei Unlimit zur Sicherung von Platform Reliability und Operational Excellence. Verantwortung für Infrastruktur, Automation und Incident Management.
💰 $110.000–170.000/Jahr
📊 Senior
🕒 Vollzeit
🌍 Remote
🗺️ EMEA
- 5+ Jahre Linux/Cloud Infrastructure
- 3+ Jahre Kubernetes Production
- Terraform Master
- On-Call Experience
Linux
Kubernetes
AWS
Terraform
Prometheus
Incident Response
✅ Geeignet für
- Senior Linux/Cloud Engineers
- Persons mit Incident Response Background
- Automation-fokussierte Infrastruktur-Profis
🚫 Weniger geeignet
- Personen ohne On-Call Experience
- Junior DevOps ohne Production Kubernetes
- Personen mit reiner Development Background
💡 Gut zu wissen
- Belgrade-basiert
- On-Call Rotation ist Pflicht
- Platform Reliability ist Business-kritisch
Über das Unternehmen
Unlimit betreibt globale Financial Infrastructure mit Fokus auf Zuverlässigkeit und Skalierung.
Deine Aufgaben
- Platform Reliability und Operations: Sicherung von Verfügbarkeit, Resilienz und Performance
- Incident Management mit Troubleshooting, Eskalation und Compliance zu SLAs
- On-Call Rotation für Production Systems
- Infrastructure Engineering: Linux-basierte Systemarchitektur Design und Deployment
- AWS und Kubernetes Setup, Deployment und Management
- Automation mit Infrastructure as Code: Terraform, Ansible
- CI/CD Pipeline Management über 20+ Repositories
- Observability: Monitoring, Logging und Alerting mit Prometheus, Grafana, Splunk
- Dokumentation und Runbooks erstellen
Deine Voraussetzungen
- 5+ Jahre Linux Systems Engineering oder Cloud Infrastructure
- 3+ Jahre Kubernetes und AWS Production Experience
- Terraform und Infrastructure as Code Expertise
- Monitoring und Observability Tools (Prometheus, Grafana, Splunk)
- On-Call und Incident Response Experience
- Starke Linux Kernel Kenntnisse