... with a strong emphasis on datacentre power, cooling, and MEP (Mechanical, Electrical, and Plumbing) requirements to support the deployment of advanced AI and HPC GPU infrastructure. Are you keen to join a team that brings GenAI, AI, and ML hardware and software technologies into real-world production, with a particular ...
... mediante estrategias adecuadas de benchmarking, validación y análisis de generalización. - Desarrollar pipelines reproducibles y escalables para ejecución en HPC y/o cloud. - Seleccionar metodologías apropiadas y comunicar supuestos y limitaciones. - Colaborar estrechamente con equipos experimentales para traducir resultados ...
... Experience with Ethernet or InfiniBand networking, particularly 100+Gb networks - Experience with performance monitoring/tuning Linux systems - Experience with HPC environments, including knowledge of HPC cluster management tools - Experience with Parallel Filesystems (Lustre, GPFS/SpectrumScale, DDN) - Familiar with SDS ...
... Platform (GCP) . - Knowledge of APIs, cloud architecture, data platforms, SQL and modern data engineering concepts. - Experience with High Performance Computing (HPC) environments or HPC to cloud migrations. Experience with scientific computing platforms or life sciences software such as Benchling or Schrödinger . Cloud certifications ...
... Computer Science, Engineering, or equivalent experience. - Over 7 years of practical experience in distributed AI training, including direct involvement with HPC and/or AI environments with multi-node GPU clusters. - Solid understanding of training infrastructure and how it affects efficiency and scalability. - Strong ...
... Enterprises, Public Administrations, Supercomputing Centers, and AI Factories. These environments include infrastructure for Traditional IT, High-Performance Computing (HPC), Accelerated Infrastructure for AI Inference and Training, and Quantum Computing (QC). The creation and management of new types of advanced infrastructure require ...
... for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population. - GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour ...
... INT4, quantization-aware settings) and understanding how calibration and sparsity interact with benchmark results. - Comfort running large-scale workloads on HPC/Slurm clusters, including reproducible experiment management (MLflow, W&B) and compute cost optimization across hundreds of benchmark runs. What you'll get Your ...
... delivery teams to understand real incident-response needs, not just technical specs Preferred Skills & Experience - Experience building observability for GPU/HPC infrastructure or other specialized, high-performance compute environments - Experience with eBPF-based observability tooling - Familiarity with AIOps/ML-based ...
... internal teams. About you - Cloud depth. You are already strong in at least one of AWS, Azure, or GCP. Multi-cloud expertise is a real advantage. Exposure to HPC schedulers such as SLURM or Grid Engine is welcome. - Networking and security. You are confident across networking, TLS and certificate management (public and ...
About the role Join our multidisciplinary team and help build and improve GPU and CPU accelerated data processing software libraries. Projects like DALI or nvImageCodec are used in all kinds of processing workflows and support NVIDIA's vision and growth.
... Infrastructure networking solutions at scale please do not hesitate to apply if your experience does not match the full scope of the position. What you'll do AI Fabric & HPC Networking - Design and operate high-performance GPU networking fabrics supporting distributed AI workloads. - Architect large-scale RoCE fabrics optimized for ...
... access patterns. - Experience optimizing storage for GPU-accelerated workloads. - WEKA Data Platform (an enterprise high-performance storage system for AI and HPC) - Familiarity with Kubernetes storage integrations such as CSI. - Experience operating large-scale storage clusters. - Experience owning both architecture and ...
... founding and hosting member of the former European HPC infrastructure PRACE (Partnership for Advanced Computing in Europe), and is now hosting entity for EuroHPC JU, the Joint Undertaking that leads large-scale investments and HPC provision in Europe. The mission of BSC is to research, develop and manage information technologies ...
... industries. - Ability to collaborate effectively with scientists, engineers, and business stakeholders. Nice to have - Experience with High Performance Computing (HPC) environments or HPC to cloud migrations. - Knowledge of MLOps, AI and Machine Learning platforms. - Experience with scientific computing platforms or life sciences ...
... equivalente). - Implementar y operar entornos basados en contenedores para la ejecución de cargas computacionales (Docker, Kubernetes, colas de procesos, entornos HPC). - Definir y mantener pipelines de CI/CD para despliegue automatizado de servicios e infraestructura. - Gestionar infraestructura como código (Terraform o herramientas ...
... performance test results, providing stakeholders with clear, actionable insights and visualizations. ð What We’re Looking ForAdvanced degree in Computer Science, HPC, or related field.5+ years of practical experience in performance modeling, testing, analysis, and optimization for critical software systems. Experience with ...
... organoids. - Writing reports, attending international meetings, and supervising related work. Required - Experience in big-data metagenomic analysis. - Fluency in HPC and Linux environments, processing pipelines and genomic data analysis (all technical work is bioinformatic). Valued - Knowledge/experience in the human microbiome, ...
... headquartered in Barcelona, Spain. We aim to democratize the usage of Chips by developing Systems on chips, SOCs, that combines RISC-V and accelerated chiplets for AI and HPC, everything interconnected with UCIe open interfaces. Our technologies will provide value in multiple fields as Artificial Intelligence, Security and Privacy ...
... investigación actuales son las siguientes: - Analítica: BigData, Inteligencia Artificial - IoT - Edge/Fogcomputing Sistemas distribuidos - Machine / DeepLearning - HPC - Virtualización y orquestación Nos gusta hacer las cosas bien tanto en tecnología como en la en forma de trabajar, por eso toda nuestra metodología de desarrollo ...