- Remoto
- Tiempo completo
€75,000 - €145,000 a year
... your role will include: - Deploying models to production: Take LLMs and put them into production across one or more machines on GPU infrastructure, owning the full pipeline from a customer's query to the served response. - Serving frameworks: Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM. ...
España