... mediante estrategias adecuadas de benchmarking, validación y análisis de generalización. - Desarrollar pipelines reproducibles y escalables para ejecución en HPC y/o cloud. - Seleccionar metodologías apropiadas y comunicar supuestos y limitaciones. - Colaborar estrechamente con equipos experimentales para traducir resultados ...
Spain, Pamplona
Ver vacanteFully remote
... Computer Science, Engineering, or equivalent experience. - Over 7 years of practical experience in distributed AI training, including direct involvement with HPC and/or AI environments with multi-node GPU clusters. - Solid understanding of training infrastructure and how it affects efficiency and scalability. - Strong ...
... to develop scripting to deploy system monitoring and other metrics based tools to integrate with customer infrastructure - Support AI/ML, data‑intensive, and HPC workloads running at scale in on‑prem, hybrid, and cloud‑adjacent environments - Work closely with customers to optimize the their AI applications to better work ...
- Remoto
- Híbrido
- Tiempo completo
- Contrato
Full-time hours
... Infrastructure networking solutions at scale please do not hesitate to apply if your experience does not match the full scope of the position. What you'll do AI Fabric & HPC Networking - Design and operate high-performance GPU networking fabrics supporting distributed AI workloads. - Architect large-scale RoCE fabrics optimized for ...
About the role Join our multidisciplinary team and help build and improve GPU and CPU accelerated data processing software libraries. Projects like DALI or nvImageCodec are used in all kinds of processing workflows and support NVIDIA's vision and growth.
Fully remote
... with a strong emphasis on datacentre power, cooling, and MEP (Mechanical, Electrical, and Plumbing) requirements to support the deployment of advanced AI and HPC GPU infrastructure. Are you keen to join a team that brings GenAI, AI, and ML hardware and software technologies into real-world production, with a particular ...
Fully remote
... Experience with Ethernet or InfiniBand networking, particularly 100+Gb networks - Experience with performance monitoring/tuning Linux systems - Experience with HPC environments, including knowledge of HPC cluster management tools - Experience with Parallel Filesystems (Lustre, GPFS/SpectrumScale, DDN) - Familiar with SDS ...
- Remoto
- Híbrido
- Tiempo completo
- Contrato
de 67000 a 77000 €/año
... Platform (GCP) . - Knowledge of APIs, cloud architecture, data platforms, SQL and modern data engineering concepts. - Experience with High Performance Computing (HPC) environments or HPC to cloud migrations. Experience with scientific computing platforms or life sciences software such as Benchling or Schrödinger . Cloud certifications ...
Spain, Vitoria
Ver vacanteHybrid
... Enterprises, Public Administrations, Supercomputing Centers, and AI Factories. These environments include infrastructure for Traditional IT, High-Performance Computing (HPC), Accelerated Infrastructure for AI Inference and Training, and Quantum Computing (QC). The creation and management of new types of advanced infrastructure require ...
- Remoto
- Tiempo completo
- Contrato
Permanent contract
... for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population. - GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour ...