Kubernetes Infrastructure Reliability Engineer en España - Jobeax
Descripción de la vacante
Kubernetes Infrastructure Reliability Engineer en España
PresencialTrabaja desde la oficina
Tiempo completoJornada semanal estándar
Full-time hours
Spain
Kubernetes Infrastructure Reliability Engineer en España is listed on Jobeax. Browse 80,000+ vacancies available.
What you'll do
Service Reliability and Optimization: Focus on capacity planning and launch reviews for services before they go live. Perform blameless postmortems and proactive identification of potential outages to foster iterative improvements
Accountability/Problem Solving: Resolves complex problems in a global Kubernetes-based infrastructure through in-depth evaluation of variable factors, including inter-organizational impact, balanced with effective consultative engagement of key stakeholders. Leads end-to-end design of infrastructure solutions and maintains component standards. Evaluates promising solutions via Proof of Concept (PoCs) and feasibility studies across multiple areas, and serves as an internal escalation point for major incidents
Stakeholder Management: Acts as a bridge between engineering and operations. Communicates and presents complex information and potential solutions to cross-functional teams and the business in non-technical terms. Represents the organization as a prime contact on initiatives and interacts with senior internal and external personnel. Uses deep knowledge to influence IT infrastructure vendor product evaluations and collaborates with multiple IT partners (e.g. Enterprise Architects, Solution Owners) to integrate feedback. Mentors and shares DevOps culture, guiding developers on how to create and deploy cloud-native applications
Impact/Strategy: Provides technical leadership and direction for small-to-medium sized initiatives (projects, lifecycle work, PoCs). Ensures solutions comply with Quality/Regulatory standards and that designs adhere to the organization's Technical Architecture Framework (TAF) policies and directions. Assists in planning technology projects, estimating engineering resources, dependencies, risks and timelines for successful delivery
Business / Technical ability: Applies extensive cloud native technical expertise, acting as a recognized expert in Kubernetes and maintaining in-depth knowledge across related cloud native technologies (containers, AWS, etc.). Demonstrates a detailed understanding of how IT infrastructure impacts respective Roche business processes and outcomes
Qualifications
Education & certifications
Without Degree:4-7 years of relevant experience
Bachelor's Degree:2-5 years of relevant experience
Master's Degree:1-3 years of relevant experience
At least 1 year of experience working in a multinational environment; healthcare industry experience is a plus
Excellent problem-solving skills, decision-making ability, and sound judgment
A strong team-oriented mindset with the ability to function independently with low supervision and navigate ambiguity.
Highly fluent oral and written English communication skills are required.
Ability to work across multiple time zones
Strong customer & delivery focus
*** 24/7 on-call rotation is required for this role****
About the company & team
A healthier future drives us to innovate. Together, more than 100'000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.
Let's build a healthier future, together.
Roche is an Equal Opportunity Employer.
Job description
At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.
The Position
The Kubernetes Infrastructure Reliability Engineer is a highly skilled expert responsible for solving complex business problems using advanced cloud native technologies. The engineer will build and maintain a Kubernetes-based infrastructure, enabling the modernization of business applications and processes.
This role combines software and systems engineering to optimize systems, increase efficiency, and eliminate operational work through automation.
You will be part of the global CaaS infrastructure team at a leading healthcare company, working with members across different regions. The team's mandate is to deliver, maintain, and continuously improve a highly available Kubernetes platform across hybrid cloud deployments, including on-premise data centers and public clouds like AWS. In this role, you will apply software engineering principles to operations to build and run massively distributed, fault-tolerant systems, focusing heavily on automation, security, and observability.
Technical Skills
Kubernetes & Containers: Strong hands-on experience navigating, managing, and hardening Kubernetes clusters and containers, including knowledge of distributed storage. A Certified Kubernetes Administrator (CKA) certification is a strong plus. Knowledge of tools like Rancher or Portworx is beneficial
Infrastructure as Code (IaC): Hands-on experience delivering and managing infrastructure automation using tools like Ansible and Terraform
Scripting & Software Engineering: Proficiency in scripting and programming languages, primarily Python, Bash, or Go, including experience with test automation (e.g., pytest) and APIs deployment and management
CI/CD Tools: Expert knowledge of implementing software delivery pipelines using tools (e.g., Jenkins, Rundeck, or GitLab)
Systems & Networking: Strong understanding of Linux operating systems and core networking principles, including DNS, load balancing, firewalls, routing, and service meshes.
Observability: Experience configuring logging, metrics, and monitoring tools, specifically focusing on setting up alerts based on symptoms rather than waiting for system outages
Cloud Infrastructure: Experience with public cloud platforms, with a strong preference for AWS, specifically involving managed services for compute, networking, security, and identity (e.g., EKS, VPC, IAM)
General and Operational Knowledge
Proven experience applying best practices in an always-up, always-available service environment utilizing Scrum and Agile methodologies
Deep understanding of Technical Architecture Frameworks (TAF) and Quality/Regulatory compliance standards
Acerca del rol Estamos buscando un Senior Cloud Infrastructure Engineer con amplia experiencia en AWS y Kubernetes para unirse a nuestro equipo de infraestructura en expansión. En esta posición, serás responsable de diseñar, implementar y mantener soluciones cloud escalables y robustas para nuestros clientes del sector ...
... Elastic's complete, cloud-based solutions for search, security, and observability help organizations deliver on the promise of AI. As a Principal Platform Engineer focused on capacity, you will play a crucial role in managing and optimizing our compute resources, ensuring that our Elastic Cloud Hosted and Serverless workloads ...
... distributed systems at scale, including consistency, fault tolerance, and performance optimization Modern cloud development practices - Familiarity with infrastructure-as-code, GitOps workflows, service mesh technologies, and cloud-native development patterns Infrastructure automation - Experience with infrastructure-as-code ...
... proactive and innovative professional that is keen to join as a Senior Site Reliability Engineer! The mission of Nexthink's SRE team is to strengthen our infrastructure and enhance our ability to deploy, monitor, and scale systems effectively and reliably. They work closely with over 50 Product Engineering teams that develop ...
About the role We are looking for a Site Reliability Engineer (SRE) passionate about infrastructure reliability, automation, and the development of scalable production systems. What you'll do - Own and improve production infrastructure reliability and stability - Prepare, execute, and support deployments and infrastructure ...
... of the technical support team providing senior level technical knowledge on our customer issues. - Work with and foster relationships with the development engineering team to ensure tracking on engineering escalations, bugs and feature requests. - Develop technical troubleshooting sessions, trainings and guides for internal ...
... Engineering, Telecommunications Engineering, or a related field. - 3+ years of experience in telecommunications or related fields. - Experience working in Site Reliability Engineering, DevOps, Infrastructure Operations, or similar roles. - Strong experience with SIM/eSIM technologies. - Hands-on experience with Remote SIM Provisioning ...
About the role Senior Site Reliability Engineer - Observability We are looking for a Senior Site Reliability Engineer with deep observability expertise to strengthen the reliability and visibility of Shopmonkey's production infrastructure. This is a hands-on contract role for an experienced SRE who has built production-grade ...
... credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or compounding interest. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to "Operate What They Own" with excellence to protect their customers' ...
Want your engineering skills to enable real scientific breakthroughs? We’re looking for a Senior Site Reliability Engineer to join our highly skilled IT Operations team at EMBL-EBI. At EMBL-EBI, our IT & Technical Services department underpins groundbreaking research that improves human and planetary health. As part of ...
... mejora continua de plataformas complejas (M2M), queremos conocerte. ¿Te unes a nuestro #WelcomeHome? ¿Qué buscamos? Experiencia sólida de 3 a 6 años como Site Reliability Engineer (SRE), DevOps Engineer o en roles similares , trabajando sobre entornos productivos críticos. Experiencia en soporte técnico y funcional de aplicaciones ...
... practical SRE standards, and improve the reliability of critical infrastructure and delivery workflows. What you'll do - Define meaningful SLIs, SLOs, and reliability targets for the platform. - Collaborate with the software engineering teams to define and achieve the best practices for software observability, SLIs, SLOs ...
... sustainable rotation health. What we're looking for - Bachelor's degree in Computer Engineering or a similar discipline. - 5+ years of experience as a Site Reliability Engineer or in a similar role. - 3+ years of experience with AWS services including strong knowledge of container orchestration. - 2+ years of Kubernetes ...
About the role We are seeking a Site Reliability Engineer to join the Observability group inside our Platform Engineering domain. Platform Engineering’s goal is to provide easy to use, self-service platforms to enable other segments to easily build, deploy and monitor their business applications. And Observability’s role ...
... solid set of product ideas lined up ready for innovative engineers to tackle. And of course we have big plans to take over the taxi app service industry! Site Reliability Engineers at Cabify work on improving all aspects of our platform and have an impact across the whole organisation. They are a blend of systems engineers ...
... everyone's job), and help the infra team grow toward shared ownership and fewer single-person dependencies. What we're looking for - You've worked as an Infrastructure Engineer, SRE, or in a DevOps or platform engineering role, on systems real people depended on. - You've run production Kubernetes. You've authored and maintained ...
... and improve monitoring and observability systems (e.g., Prometheus, Grafana) — not just react to alerts - Collaborate closely with internal teams (CX, CS, Engineering) to ensure system reliability and performance - Work hands-on with modern DevOps and infrastructure tools including Kubernetes, Terraform, CI/CD pipelines, ...
... Internal careers site Locations : Madrid, Spain Required Qualifications - 3+ years of experience as a Site Reliability Engineer, DevOps Engineer, or Cloud Infrastructure Engineer. - Deep expertise in Linux administration and core networking fundamentals. - Hands-on experience building and managing infrastructure in AWS and ...
... hands-on Kubernetes: operating and scaling production clusters and container tooling (Docker) and its ecosystem. - Experience building and managing cloud infrastructure on AWS (or similar). - Strong infrastructure-as-code practice with Terraform. - Experience with reliability frameworks: SLOs, SLIs, error budgets, alerting ...