Senior Site Reliability Engineer en Spain, España - Jobeax
Descripción de la vacante
Senior Site Reliability Engineer en Spain, España
Lodgify
RemotoTrabaja desde cualquier lugar
Tiempo completoJornada semanal estándar
ContratoTemporal o freelance
Spain
Senior Site Reliability Engineer en Spain, España is listed on Jobeax. Browse 80,000+ vacancies available.
About the role
Are you a systems-minded engineer who cares deeply about reliability, scalability, and production excellence? Join Lodgify as a Senior Site Reliability Engineer and help our engineering teams build and operate services that are reliable, observable, scalable, and resilient by design. In this role, you will improve the reliability of our shared infrastructure and product services while helping teams own their systems in production. You will work in the Platform team to strengthen observability, reduce operational toil, improve incident response, define practical SRE standards, and improve the reliability of critical infrastructure and delivery workflows.
What you'll do
Define meaningful SLIs, SLOs, and reliability targets for the platform.
Collaborate with the software engineering teams to define and achieve the best practices for software observability, SLIs, SLOs and reliability.
Strengthen production readiness by improving service ownership, observability, alerting, runbooks, scaling assumptions, rollback paths, and failure-mode preparedness.
Improve the reliability, scalability, and performance of cloud, Kubernetes, and shared infrastructure, including how systems scale during growth, traffic spikes, and dependency failures.
Build actionable observability using metrics, logs, traces, and golden signals, with tools such as Datadog, Prometheus, and Grafana.
Implement operational and security best practices through guidelines, policies and automation.
Reduce alert noise and improve signal quality so teams can detect, understand, and resolve issues quickly.
Automate repetitive operational work using Python or other languages, turning recurring manual work into safer automation and clearer runbooks.
Implement self-service Internal Developer Platform features via APIs and Kubernetes operators.
Improve deployment safety, rollbackability, and release observability.
Improve reliability of critical stateful systems such as databases, caches, queues, and streaming platforms.
Participate in on-call, troubleshoot, and coordinate incident response, and facilitate blameless post-incident reviews that turn into concrete improvements.
Execute disaster recovery drills and analyse cloud/platform usage to identify cost and resource-efficiency gains without compromising reliability.
What we're looking for
You have 7+ years of production experience operating Kubernetes-based platforms and cloud infrastructure.
You understand and apply SRE practices: SLIs, SLOs, error budgets, production readiness, incident response, post-incident learning, toil reduction, scalability, capacity planning, high availability, backups, and disaster recovery.
You can design and improve observability and alerting for critical systems using metrics, logs, traces, and golden signals, and are comfortable troubleshooting complex distributed systems to identify systemic reliability improvements.
You can write maintainable software to automate operational tasks and reduce manual intervention.
You have experience with stateful production systems such as relational databases, caches, queues, or streaming platforms.
You know how to balance reliability, performance, cost, and delivery speed pragmatically.
You are comfortable working in a transitional environment where SRE practices are being introduced while critical infrastructure and delivery systems still need hands-on reliability support.
You collaborate effectively with Engineering, Platform, Security, and Product stakeholders.
You communicate clearly, document well, and enjoy coaching teams toward stronger production ownership.
You model initiative and accountability, raising risks early and driving improvements through to completion.
What you'll get
*
🏠 Remote Flexibility: The freedom to work from home any day that works for you.
🌴 Time to Recharge: 25 working days of paid vacation and Jornada Intensiva in August..
💊 Alan Health Insurance: Premium health, dental, and mental health support via Alan. Pre-existing conditions are covered.
😋 Meal Perk:€150/month allowance on your Alan card + 50% off Ametller Origen prepared dishes at the office.
💸 Tax-Free Savings: Increase your take-home pay by using Flexible Remuneration for extra meal costs (up to €70/mo) and public transport (up to €136/mo).
🖥️ Home Office Gear: We provide a table, ergonomic chair, and monitor for your home setup.
🇪🇸 Language Learning: Free Spanish classes.
🤑 Referrals: Cash rewards for bringing in new talent.
🌟 Social Life: Daily office breakfast and monthly team events
🎯 Dynamic Hub: A high-energy, inclusive environment designed for collaboration and connection with a team that represents over 60 countries.
*Benefits offered may differ based on the type of contract that is issued
So, what are you waiting for?
All applications and CVs must be submitted in English 😉
About the company & team
⭐
Lodgify is a fast-growing scale-up company leading the vacation rental industry. Backed by $30M in funding, our platform empowers property owners and managers worldwide to efficiently manage and grow their business through technology.
Headquartered in sunny Barcelona, we're now a team of 380+ people representing over 60 nationalities, united by a passion for transforming the future of short-term rentals.
You'll be part of a growing, dynamic company with a truly international team. At Lodgify, we are full of contagious energy, hard work, and passion for what we do. We celebrate diversity and are proud to acknowledge a variety of backgrounds, perspectives and skills in our team; committed to creating a workplace where everyone is heard and feels a sense of belonging.
⭐ What does success look like?
Critical services have clear owners, meaningful SLIs/SLOs, actionable alerts, dashboards, runbooks, and production readiness coverage.
Reliability targets are consistently met across critical infrastructure and services.
Operational toil and manual intervention are measurably reduced through automation and safer workflows.
MTTR improves through reduced alert noise, better signal quality, stronger observability, and clear incident response playbooks and escalation paths.
Post-incident actions are tracked, completed, and used to reduce repeat incidents.
Disaster recovery exercises validate that critical services and infrastructure can recover within agreed expectations.
Cloud and infrastructure resources are optimised without sacrificing performance, elasticity, or resilience.
... seamless, resilient, and scalable platform around the clock. We are looking for an experienced, proactive and innovative professional that is keen to join as a Senior Site Reliability Engineer! The mission of Nexthink's SRE team is to strengthen our infrastructure and enhance our ability to deploy, monitor, and scale systems ...
Want your engineering skills to enable real scientific breakthroughs? We’re looking for a Senior Site Reliability Engineer to join our highly skilled IT Operations team at EMBL-EBI. At EMBL-EBI, our IT & Technical Services department underpins groundbreaking research that improves human and planetary health. As part of ...
... layers or bureaucracy. We're fully remote by design, genuinely global, and united by a shared mission to make travel simpler for everyone. We are looking for a Senior Site Reliability Engineer to join our growing engineering team. We are a company that values SRE principles and practices. We believe in empowering our SREs ...
... solid set of product ideas lined up ready for innovative engineers to tackle. And of course we have big plans to take over the taxi app service industry! Site Reliability Engineers at Cabify work on improving all aspects of our platform and have an impact across the whole organisation. They are a blend of systems engineers ...
About the role As a Senior SRE at Remote, you'll work with a high degree of autonomy on complex reliability and platform problems, owning the plan and execution of features and projects within our SRE/Platform domain. You'll contribute to the platform's architecture and reliability strategy, translating ambiguous requirements ...
... users worldwide. Our commitment to reliability is a key foundation of our product and our dedication to exceeding customer availability expectations is a core engineering focus. As a Senior Site Reliability Engineer, you'll join our SRE team based in Europe to ensure our production systems are not only operational but also ...
About the role We're looking for a Senior Site Reliability Engineer to join the Software Logistics Team (CI/CD) within our Platform Engineering Segment. Platform Engineering's mission is to provide easy-to-use, self-service platforms that enable other segments to build, deploy, and monitor their business applications with ...
... reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or compounding interest. Site Reliability Engineering at Affirm is a small, yet crucial, team that helps our Engineering partners to "Operate What They Own" with excellence to protect their ...
... Elastic's complete, cloud-based solutions for search, security, and observability help organizations deliver on the promise of AI. As a Principal Platform Engineer focused on capacity, you will play a crucial role in managing and optimizing our compute resources, ensuring that our Elastic Cloud Hosted and Serverless workloads ...
About the role We're hiring a Senior Site Reliability Engineer to join our Platform team and help us build secure, reliable, and scalable systems as Maze grows. This role sits at the intersection of cloud infrastructure, reliability, and security . You'll work across our platform to improve how we operate production systems, ...
... resilience - Persistent in identifying systemic issues, understanding failure modes in distributed systems, and driving solutions to completion Performance and reliability focus - Analytical mindset toward observing end-to-end service performance, system health, and user impact in production environments Dynamic environment ...
About the role We are looking for a Site Reliability Engineer (SRE) passionate about infrastructure reliability, automation, and the development of scalable production systems. What you'll do - Own and improve production infrastructure reliability and stability - Prepare, execute, and support deployments and infrastructure ...
... Migrations - Analyze and plan complex migrations. What are we looking for? - Bachelor's degree in Computer Engineering, Electronics Engineering, Telecommunications Engineering, or a related field. - 3+ years of experience in telecommunications or related fields. - Experience working in Site Reliability Engineering, DevOps, Infrastructure ...
... evolución y mejora continua de plataformas complejas (M2M), queremos conocerte. ¿Te unes a nuestro #WelcomeHome? ¿Qué buscamos? Experiencia sólida de 3 a 6 años como Site Reliability Engineer (SRE), DevOps Engineer o en roles similares , trabajando sobre entornos productivos críticos. Experiencia en soporte técnico y funcional ...
About the role We are seeking a Site Reliability Engineer to join the Observability group inside our Platform Engineering domain. Platform Engineering’s goal is to provide easy to use, self-service platforms to enable other segments to easily build, deploy and monitor their business applications. And Observability’s role ...
... grow together. You will be part of a culture that values trust, accountability, and shared success where your work truly matters. Your Impact Join a team of senior engineers operating in a large-scale, multi-cloud production environment supporting tens of thousands of enterprise customers worldwide. This is not a typical ...
... are invited, and ideas matter. A team where everyone makes play happen. Reports to: Technical Director Hiring process External careers site, Internal careers site Locations : Madrid, Spain Required Qualifications - 3+ years of experience as a Site Reliability Engineer, DevOps Engineer, or Cloud Infrastructure Engineer. ...
... place for you? What you'll do - You will report to the Manager of Customer Reliability, work as an important member of the technical support team providing senior level technical knowledge on our customer issues. - Work with and foster relationships with the development engineering team to ensure tracking on engineering ...
About the role Senior Site Reliability Engineer - Observability We are looking for a Senior Site Reliability Engineer with deep observability expertise to strengthen the reliability and visibility of Shopmonkey's production infrastructure. This is a hands-on contract role for an experienced SRE who has built production-grade ...