Senior Inference Engineer en España is listed on Jobeax. Browse 80,000+ vacancies available.
About the job
As a Senior Inference Engineer at vCluster Labs, you are the first engineer we're hiring to own inference. You'll partner directly with our CTO to build the platform's inference layer from the ground up, taking models and turning them into a production-grade, query-to-response pipeline running at scale. From there, you will help lead the engineering direction of inference at vCluster, partnering with Product to shape what we build next as the space evolves.
As a Senior Inference Engineer, your role will include:
Deploying models to production: Take LLMs and put them into production across one or more machines on GPU infrastructure, owning the full pipeline from a customer's query to the served response.
Serving frameworks: Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM.
Optimizing for scale: Apply quantization, batching, caching, and routing to keep latency and cost in check as traffic grows.
Programming, not just configuring: Build real infrastructure in Python or Golang - this is an engineering role, not a research or data-science one.
Owning the roadmap: Build the first iteration alongside our CTO, then take the lead on the inference platform and partner with Product to decide what we build next.
This role could be a fit for you if you bring
Production LLM serving experience: You've deployed and served LLMs using vLLM, SGLang, or TensorRT-LLM, ideally at a company built around inference at scale.
Inference optimization know-how: Hands-on experience with quantization, batching, caching, and routing, not just familiarity with the terms.
Hands-on programming experience: Strong engineering skills in Python or Golang, with real production code experience.
Communication: Strong communication skills, explaining technical concepts clearly to both engineers and non-technical stakeholders.
Bonus points for
Familiarity with containerized environments (Docker, Kubernetes)
Hands-on generative AI experience with common ML frameworks (PyTorch, Transformers)
Good understanding of the GPU stack: CUDA, NCCL, drivers, and related libraries
Knowledge of model architectures and fine-tuning approaches
Experience with NVIDIA Dynamo
About vCluster Labs
We're the #1 platform for AI infrastructure, trusted by the world's fastest-growing AI cloud builders. We're a venture-backed startup that's raised over $28M from top-tier investors including Khosla Ventures (first investor in OpenAI, GitLab, Stripe, and DoorDash), and we're in a hyper-growth phase looking for motivated people to join our team. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe with a remote-first culture.
We give AI Cloud providers and AI factories a hyperscaler-like experience on their own GPU infrastructure. Our platform runs the full stack an operator needs, from bare metal provisioning and node lifecycle management up through managed Kubernetes, Slurm, Ray, and inference clusters, so they can turn raw GPUs into cluster products they can sell in days instead of spending 12+ months building it themselves. Today we power over 100,000 GPUs and 1 million CPUs across 50+ AI clouds and Fortune 500 companies, backed by a team of 40+ infrastructure engineers who build alongside our customers rather than just shipping them software.
We're the company behind vCluster, the open source technology for tenant isolation on Kubernetes, with 11,000+ GitHub stars and 40M+ tenant clusters created since 2021. Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI, a Kubernetes-native framework built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.
Benefits
We offer the following benefits:
Competitive Salary: We offer a competitive compensation package, including equity.
Platinum-Level Insurance: Health, dental, vision, and life Insurance, including plans for you and eligible dependents (benefits vary depending on country).
Flexible Working Schedule: You have a doctor's appointment or need to head to the supermarket to get groceries at 2pm? We won't have an issue with that. To us, results matter more than clocking in and out at the same time every day.
Workplace Flexibility: We're very flexible about where you work. We know things can change in life and we're happy to adjust the work environment for you along the way.
Culture & Values
At vCluster Labs, we value and stand for:
Make it Happen: We have a relentless bias for action and the grit to push through obstacles. We do whatever it takes to figure it out, put in the work, and ruthlessly prioritize the actions that drive measurable impact for the business.
Own the Outcome: We understand that our responsibility doesn't end when a task is checked off; it ends when the value is delivered. We connect our daily individual actions to the broader success of the company and our customers.
Create Wow: We measure success by the experience we generate, both inside and outside the company. For our customers, this means impressive speed and intuitive experiences. For our team, this means going the extra mile to support one another and to continuously drive each other to new heights.
Open Source, Open Mind: We are actively contributing to and maintaining open-source projects. Internally, we foster meritocracy - the strongest ideas win, no matter who or where they come from.
Build Tomorrow's Standards, Intentionally: We don't just ship software; we define the state-of-the-art of tomorrow. We are fearless in tearing down old approaches to build something better, but we are disciplined in how we do it because we know our users rely on our technology to run mission-critical infrastructure platforms.
... MS or PhD in Computer Science, Engineering, High-Performance Computing, or equivalent professional experience. - 5+ years of experience in Neural Networks inference optimization. - Solid understanding of transformers inference optimization: quantization, disaggregated inference, speculative decoding, continuous batching, ...
... infrastructure engineers, and executive team members. Ways to stand out from the crowd - Experience with NVIDIA's inference stack, including TensorRT-LLM, Triton Inference Server, NIM, and NVIDIA Dynamo. - Experience with GPU orchestration on Kubernetes. - You have operated inference at scale inside a frontier AI lab or hyperscale's ...
About the role NVIDIA is hiring an NCX Senior Engineer who is passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join our DSX team. This role involves working closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities essential for running large-scale NVIDIA accelerated ...
About the role We are looking for Senior CV Engineer . What you'll do Your main tasks will be: - Train and fine-tune generative models (text2image, image2image, video2image, IP-adapters) to produce photorealistic and stylized visuals - Work with image classifiers, rankers, and image/video captioning models that understand ...
About the role NVIDIA is seeking a Senior MLOps Engineer to join our DSX Enablement team, collaborating closely with strategic customers to implement and enhance groundbreaking AI workloads. We partner with the world's most innovative AI companies and open-source communities to address their most challenging technical problems. ...
About the role We are seeking senior engineers to pioneer new methodologies for accurately assessing the performance and capabilities of ground-breaking deep learning models, including LLMs, RAG, agents, and vision models. You will collaborate across the organization to bring the latest flagship models from our community ...
... AI infrastructure, and hardware/software co-design , this is your opportunity to shape technology that protects AI models, data, and workloads at scale. As a Senior Secure AI Engineer , you'll be at the intersection of AI, security, cloud-native technologies, and next-generation silicon , designing solutions that make confidential ...
... a position to a new member, the Open Home Foundation aims to provide a total compensation package that matches the 75th percentile for the new hire's role, seniority, and local market rates. For a Software Engineer in our primary operating countries, the approximate yearly compensation will be the following: - Netherlands: ...
About the role We are seeking a highly experienced and technically adept Senior Machine Learning Engineer to join our team. In this pivotal role, you will own the entire lifecycle of MLOps from optimization to deployment, with a strong focus on LLMOps. This position demands a strong blend of hands-on technical expertise, ...
About the role Reference 393_26_AII_AI_RE3 Job title Senior Research Engineer (RE3) About BSC The Barcelona Supercomputing Center - Centro Nacional de Supercomputación (BSC-CNS) is the leading supercomputing center in Spain. It houses MareNostrum, one of the most powerful supercomputers in Europe, was a founding and hosting ...
... Geometric deep learning, or graph neural networks applied to spatial data. - Vector graphics, CAD, GIS or trajectory modelling in an industrial setting. - Inference optimization and quantization. - Published or open‑sourced work we can read. - Spanish or Catalan. This is not an LLM or prompt‑engineering role, and it is ...
About the role Reference 393_26_AII_AI_RE3T1 Job title Senior Research Engineer (RE3-T1) About BSC The Barcelona Supercomputing Center - Centro Nacional de Supercomputación (BSC-CNS) is the leading supercomputing center in Spain. It houses MareNostrum, one of the most powerful supercomputers in Europe, was a founding and ...
What you'll do As a Senior Machine Learning Engineer, these will be some of the key activities in your day-to-day role: Design, train and deploy scalable machine learning models using Getnet data to solve high-impact payment and business challenges across geographies. Build and maintain end-to-end machine learning pipelines, ...
... reliable, high-quality output across the entire company. You'll join the AI & DevEx team. Our mission is to raise the performance of Fever's whole Product Engineering organization, not just individual engineers, but entire squads (product, design, and engineering together) and everything that surrounds their work: CI/CD, ...
... Architectures: Proven experience training, fine-tuning, and deploying modern transformer-based models and encoding methods for complex text tasks. - MLOps & Production Engineering: Strong track record of optimizing models for low-latency environments (e.g., batch inference, model distillation, hardware acceleration) and deploying them ...
... pipelines and high-concurrency inference workloads. - Ensure inference infrastructure scales efficiently across thousands of GPUs. - Partner with platform and inference teams to ensure storage design supports evolving inference architectures and avoids becoming a bottleneck in throughput, latency, or concurrency. AI Data Path ...
... solutions, ensuring accuracy, performance, security, and scalability. - Implement and maintain end-to-end AI/ML pipelines - from data ingestion and feature engineering through to model development, validation, and deployment with guidance from senior engineers on complex architectural decisions - Instrument AI/ML services ...
... support that's both relevant and compliant wherever you are. ★ OUR PROCESS What to expect as you move through our hiring process: - HR interview with Claire, Senior Talent Acquisition Partner - Technical test to complete remotely - Technical interview with Gwennaëlle, Staff ML Engineer and Pedro, Senior ML Engineer, - Manager ...
... are leaders in digital technology services, delivering large-scale technological solutions for some of the world's leading organizations. We are looking for a Senior DevOps Engineer AWS + AI to join our team. This role is focused on the design, automation, deployment, and operation of cloud infrastructures on AWS, with a ...