Talent.com
Stellar Cyber
Staff Site Reliability EngineerStellar Cyber • Banyoles, Cataluña, Spain
Buscar otras ofertas
Staff Site Reliability Engineer

Staff Site Reliability Engineer

Stellar Cyber • Banyoles, Cataluña, Spain
Hace 12 días
Descripción del trabajo
  • We are seeking a highly skilled Staff Site Reliability Engineer (SRE) to join our team and drive reliability, scalability, and efficiency across our production systems.
  • The ideal candidate will have deep expertise in cloud infrastructure, Kubernetes administration, observability, and incident management, with a proven track record of building and maintaining highly available and resilient platforms.
  • As a senior member of the SRE team, you will not only operate complex distributed systems but also influence architecture, tooling, and best practices to ensure operational excellence
  • Administer and maintain container orchestration platforms and containerized workloads
  • Monitor and troubleshoot production systems, participating in on-call rotations to ensure reliability
  • Drive observability improvements by enhancing monitoring, logging, and alerting capabilities across systems and data platforms
  • Administer and optimize cloud-based environments across multiple providers
  • Manage and support distributed data platforms and real-time processing systems
  • Develop and maintain continuous integration and delivery pipelines for efficient and reliable deployments
  • Own and implement Infrastructure as Code (IaC) practices to ensure consistency and scalability
  • Automate and orchestrate infrastructure using programming and scripting languages
  • Perform system administration and networking tasks to support internal and external environments
  • Collaborate effectively with engineers and stakeholders across different time zones

Benefits

  • Pre-IPO Stock Options
  • Medical, Dental & Vision care
  • 401(k)
  • Employee Assistance Program
  • Employee Discount Program
  • Life Insurance
  • Paid time off
  • Referral Program
  • Rewards and Recognition Program
  • Strong technical background in distributed systems, databases, networking, and Linux administration
  • Deep understanding of Infrastructure as Code (Terraform, Helm)
  • Proven success leading large-scale production systems in cloud environments (AWS, GCP, Azure, or OCI)
  • Strong experience with production on-call operations and incident management
  • Strong programming and automation skills in Python and Bash
  • 5+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles
  • Expertise in operating data platforms (Elasticsearch, MongoDB, Spark, Kafka, Redis)
  • Advanced proficiency in Kubernetes administration and troubleshooting
  • Knowledge in chat-based operations interfaces and/or auto-remediation controllers using AI agentic framework
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field
  • Excellent problem-solving, communication, and leadership abilities
  • Understanding of AI agents for Auto-triaging alerts, correlate signals and suggest/root-cause hypotheses
  • Demonstrated leadership in driving incident response, on-call best practices, and reliability-focused culture
  • Proficiency with public cloud services (AWS, Azure, GCP, or OCI)
  • Experience with CI/CD pipelines (GitHub Actions, Bitbucket, ArgoCD)
  • Hands-on experience with observability tools: Prometheus, Grafana, Loki, and Alertmanager
  • Certifications in AWS, GCP, Observability, Linux or Kubernetes are a plus

#J-18808-Ljbffr

Crear una alerta de empleo para esta búsqueda

Staff Site Reliability Engineer • Banyoles, Cataluña, Spain