Machine Learning Engineer, Speech - Joint Audio-Video Modeling en Remote (U.S. or Europe), España
Cantina
RemotoTrabaja desde cualquier lugar
HíbridoCombinación de oficina y remoto
PresencialTrabaja desde la oficina
Tiempo completoJornada semanal estándar
US$200,000 - US$220,000 a year
Spain
Machine Learning Engineer, Speech - Joint Audio-Video Modeling en Remote (U.S. or Europe), España is listed on Jobeax. Browse 80,000+ vacancies available.
About the role
We're looking for a Research / ML Engineer to join our Speech Team to build state-of-the-art speech and audio generation systems end-to-end from data specs through production inference with a focus on joint audio-video modeling.
You'll own the audio side of multimodal generation: the representations (audio VAEs, neural codecs), the generative backbone (diffusion / flow-matching transformers), and the conditioning and alignment machinery that makes characters speak, sing, and emote in sync with what's on screen. That includes voice cloning and multi-speaker conditioning inside joint AV models, cinematic dialogue with music and sound design, and adjacent speech tasks (controllable TTS, voice conversion) that feed the same stack.
You'll drive the model ↔ data ↔ eval flywheel, partnering closely with research, video, data, and infra to ship fast, reliable, and cost-aware models. In this role you'll work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems.
You will thrive in this role if you:
See research and engineering as two sides of the same coin and enjoy owning work end-to-end.
Are excited to work across modalities and collaborate closely with a video generation team rather than staying inside audio.
Are results-oriented, flexible, and willing to pick up whatever moves the needle.
Like collaborating closely with infra, data, and product to ship measurable improvements.
Enjoy designing experiments, listening tests, and metrics that correlate with user-perceived quality.
Are eager to learn every day, and to find and solve unique large-scale problems.
What you'll do
Audio Representations: Design, train, and improve the audio VAEs, neural codecs, and vocoders our generative models sit on top of latent design, reconstruction and perceptual objectives, compression-vs-fidelity tradeoffs.
Model Building: Architect, implement, pre-train, fine-tune, and post-train/alignment (e.g., GRPO/DPO) diffusion and flow-matching transformers for large-scale audio and video generation.
Joint Audio-Video Modeling: Design the audio conditioning and cross-modal alignment inside joint AV models, audio latents alongside video latents, reference-audio and multi-speaker conditioning, multi shot generation audio/video modeling.
Experimental Design: Design, run, and analyze scientific experiments to advance our understanding of the models.
Data Ownership: Define data requirements and collaborate on acquisition, curation, AV-sync and quality filtering, annotation quality, and synthetic data strategies for paired audio-video and speech corpora.
Rigorous Evaluation: Design automated objective/subjective evaluations audio fidelity and intelligibility metrics, AV-sync, listening and viewing tests, robustness & bias checks, and red-team studies.
Inference Efficiency: Drive distillation, step-count reduction, quantization, and kernel/memory optimization to meet interactive latency and cost targets.
Pipeline Delivery: Harden the training → evaluation → inference pipeline; profile latency, memory, and cost; and meet production SLAs with robust monitoring and rollback.
GPU Scaling: Partner with infrastructure to run distributed training/inference on cloud fleets and productionize models with reliability and observability.
Project Leadership: Independently lead small research projects while collaborating on larger team initiatives, including cross-team work with video generation.
Tool Development: Develop and improve dev tooling to enhance team productivity.
Safety & Responsibility: Contribute to safety/consent guardrails, watermarking, and misuse/abuse mitigation for responsible voice and likeness technology.
What you'll get
The anticipated annual base salary range for this role is between $200,000-$220,000 (€170,000-€190,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.
Competitive salary and generous company equity
Medical, dental, and vision insurance - 99.99% of premiums covered by Cantina
42 days of paid time off, including:
15 PTO days
10 sick days
15 company holidays
2 floating holidays
Generous parental leave & fertility support
401(k) retirement savings plan
Lifestyle spending account - $500/month to use however you'd like
Complimentary lunch and snacks for in-office employees
One Medical membership, and more!
About Cantina
Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.
If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!
What You'll Bring
Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data).
Deep hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation.
Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders latent/tokenizer design, reconstruction and perceptual objectives, adversarial training.
Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent).
Strong software engineering skills with a proven track record of building complex systems.
Strong with PyTorch and performance work (profiling, CUDA/Triton/C++ as needed) and writing reliable production-quality code.
Shipped large-scale speech/audio or multimodal generative models to production.
Background in working with large-scale ML data, and the ability to iterate on data and triangulate quality using both subjective and objective signals.
Experience with voice cloning, speech control/steerability, or expressive speech generation.
Notable publications and/or open-source contributions in speech/audio/ML.
Strongly preferred:
Experience with multimodal audio-video modeling: joint AV generation of multi-shot, multi-speaker scenes with dialogue, music, and sound design generated jointly with video, and the cross-modal alignment that keeps them in sync.
Experience with video generation: video diffusion/flow-matching transformers, video VAEs, conditioned and multi-shot generation, building data pipelines for video models.
Buscamos un/a Data Scientist / ML Engineer con sólida experiencia en IA tradicional para incorporarse a un equipo de datos en expansión. La persona seleccionada será responsable de diseñar, entrenar y poner en producción modelos predictivos orientados a pricing dinámico, personalización de oferta y forecast de demanda en ...
About the role We are seeking an experienced Senior Machine Learning Engineer to join our innovative AI team in Barcelona. You will be responsible for designing, developing, and deploying cutting-edge machine learning solutions for our leading automotive and manufacturing clients. This is a Full-time, on-site position where ...
... Anywhere! - We are a remote-first company with a globally distributed team. You can find your productive zone and work from there. What you'll be doing As a Machine Learning Engineer at Sardine, you'll own the systems that make real-time fraud detection possible. Our data science team builds custom models for our clients, ...
¿Te gustaría participar en proyectos de innovación tecnológica , trabajando con ingeniería de datos, Cloud, Inteligencia Artificial y Machine Learning aplicados a entornos industriales? En Keapps buscamos un/a Ingeniero/a de Datos y Machine Learning para incorporarse a un proyecto estable dentro de un Departamento de Innovación ...
EData Scientist / Machine Learning Engineer Ayesa Digital Cequelinos, galicia, Spain Descripción del trabajo ¡En Ayesa Digital crecemos Contigo! Desplácese hacia abajo para encontrar una descripción detallada de este trabajo y lo que se espera de los candidatos. Cada profesional de nuestra empresa es importante para nosotros, ...
... record of delivering production-grade ML systems. - Strong expertise in NLP, particularly transformer-based models, along with solid understanding of classical machine learning techniques. - Advanced Python skills and hands-on experience with deep learning frameworks (PyTorch preferred). - Strong experience with SQL (e.g., ...
... https://jobeax.com/link/8om4ubga9gvjbUNE a mejorar su posición competitiva mediante la gestión avanzada de la información. ¡Únete a nuestro equipo! Actualmente buscamos Machine Learning Operations Engineer con experiencia comprobable en Vertex AI (GCP) y liderazgo para trabajar en un proyecto líder del sector retail. Requisitos obligatorios: ...
... implementation of cutting-edge machine learning models for real-world applications. You'll work on challenging projects involving image recognition, object detection, and video analysis, while mentoring junior engineers and contributing to our technical strategy. Who are we looking for We're looking for a highly experienced machine ...
Sobre el rol Buscamos un ingeniero de Machine Learning senior altamente cualificado para unirse a nuestro equipo de transformación digital en Banco Santander. En esta posición, serás responsable de diseñar, desarrollar e implementar soluciones avanzadas de machine learning que impulsen la innovación en el sector financiero. ...
... and other nationally funded research projects. The team is seeking a Machine Learning Engineer with experience in speech technologies, particularly in deep learning and model development for tasks such as speech recognition, speech synthesis, and LLMs. The successful candidate will join the Speech Team, work within a highly ...
... metodologías DevOps. Habilidades y competencias - Experiencia avanzada en Python, SQL y manipulación de datos con pandas y NumPy - Dominio de frameworks de deep learning como TensorFlow, PyTorch o Keras - Conocimiento sólido en algoritmos de machine learning, estadística y análisis de datos - Experiencia con herramientas de ...
Data Scientist – Machine Learning / CRM Analytics Desde Métrica estamos buscando un/a Data Scientist / Machine Learning para incorporarse a un proyecto dentro del área de CRM & Data/AI , participando en el desarrollo de soluciones avanzadas de análisis y Machine Learning orientadas a cliente. Modalidad: 100% remoto. SBA: ...
About the job About the opportunity We are seeking a data-driven Technical Product Manager for Applied Machine Learning to join our Intelligent Operations Platforms (IOP) segment. In this role, you will build and deploy Machine Learning solutions that empower internal teams and enable N26 to scale operations efficiently ...
Data Scientist – Machine Learning / CRM Analytics Desde Métrica estamos buscando un/a Data Scientist / Machine Learning para incorporarse a un proyecto dentro del área de CRM & Data/AI , participando en el desarrollo de soluciones avanzadas de análisis y Machine Learning orientadas a cliente. Modalidad: 100% remoto. SBA: ...
... capacidad para resolver problemas complejos y una mentalidad orientada a la calidad. Valoramos especialistas con experiencia en implementación de soluciones de machine learning integradas en aplicaciones web. Tu capacidad para adaptarte rápidamente a nuevas tecnologías, tu pensamiento analítico y tu excelente comunicación ...
Newton Colmore Consulting is partnering with a venture-backed deep tech company to develop novel intelligent systems. The Machine Learning Engineer will own advanced modelling, ML and predictive analytics across product platforms, tackling challenging problems with autonomy and leadership exposure. You will join a talented ...
Acerca del rol Buscamos un Senior Machine Learning Engineer experimentado para unirse a nuestro equipo de innovación en Barcelona. En esta posición híbrida, serás responsable de diseñar, desarrollar e implementar soluciones de machine learning de alto impacto que transformen nuestros productos y servicios. Trabajarás en ...
Original Advert - Job ID: 365686 - Date posted: 11/09/2026 What you'll need to have • Bachelor's degree in Mathematics, Computer Engineering, Physics, Science, or equivalent. Preferred: Master's degree in Data Science, Machine Learning, Big Data, or Advanced Analytics • Python, PySpark, in Jupyter Lab, Big Query, DataBircks ...
... mismo objetivo en 10 países. Si eres una persona entusiasta y buscas un nuevo reto profesional, ¡Este es tu sitio! Buscamos ampliar nuestro equipo con un/a AI / Machine Learning Specialist para unirse a un equipo de profesionales multidisciplinar. ¿Qué vas a hacer? Participarás en proyectos de Inteligencia Artificial, trabajando ...
... entre Barcelona, Madrid y centros remotos en toda Europa. Proceso de selección Nuestro proceso incluye: entrevista inicial con RRHH (30 min), prueba técnica de machine learning (2-3 horas), entrevista técnica con senior engineers (45 min), y entrevista final con responsable de equipo (30 min). Total: 1-2 semanas aproximadamente.