Software Engineer: Applied NLP/ML and Data Systems… - Jobeax
Descripción de la vacante
Software Engineer: Applied NLP/ML and Data Systems (Mid-career / Senior) en Bellprat, España
Theia Insights
HíbridoCombinación de oficina y remoto
ContratoTemporal o freelance
Bellprat, España
About the role
Theia Insights builds foundational financial intelligence products, including industry classification, knowledge graphs and factor risk models, for institutional investors. We serve some of the largest asset managers, hedge funds, index providers and sell-side banks.
As an engineer on the Data Products team you'll own the pipelines that ingest NLP and financial data from public equities around the world to produce the Theia Insights Industry Classification (TIIC) and the datasets behind our Thematic Factor Risk Models (TFM).
The Data Products team owns the data that underpins everything we sell. It's a small, senior group that values correctness and reproducibility over volume, and it sits close to the product leads who shape the methodology. We value durability and good judgement over familiarity with the flashiest tools.
What you’ll do
Build and maintain pipelines that classify global public equities across our five-level taxonomy, sector, industry, sub-industry, major theme and micro theme, by extracting information from filings and web content, and assigning thematic exposures based on this information.
Run large-scale NLP and LLM inference (entity extraction, classification, knowledge graph construction) over company documents, with cost- and throughput-aware batch execution.
Ingest market data and publish datasets to external distributors.
Own schema and contract evolution for datasets with real downstream consumers.
Work with economists and engineers to turn modelling decisions into reliable production data.
Essential
Strong production Python.
Datasets in pandas and Parquet/Arrow, plus an analytical engine, e.g. DuckDB, or a warehouse such as Snowflake.
Orchestrated batch pipelines you've operated, not just written: Dagster or Airflow, S3-based data flows, and a habit of testing outputs for correctness rather than only for exceptions.
Applied ML in production: embeddings and semantic similarity, clustering, or operationalising models (not necessarily training from scratch).
AWS fluency and CI/CD discipline.
Nice to have
Practical LLM engineering: prompting, batch inference, and cost and throughput trade-offs across providers.
SageMaker, Bedrock, or comparable managed ML tooling.
Financial and equities domain knowledge: classification taxonomies, factor models, index construction. Valuable but learnable.
Infrastructure as code (AWS CDK or Terraform) and Docker.
Experience we’re looking for
We care more about what you’ve owned than years on a CV. If you’ve built a pipeline that runs on a schedule against real volume, and you were the person who got paged when it broke, you’re in scope. More senior candidates will typically have made the cost and throughput trade-offs, and set the standard for how a team tests data quality.
Competitive salary plus share options
25 working days holiday, plus Spanish public holidays
... learning and NLP techniques to deliver products, with a solid understanding of ML concepts relevant to the Swedish language. * Strong skills in object-oriented software design and programming, and experience with working on large code bases. * Ability to analyze data and make data-driven decisions to improve user experience ...
... projects for a significant European Agency. The role requires strong analytics, ML/NLP expertise, and proficiency in Python, R, and SQL. You will collaborate with UX and product teams to design, implement and monitor data-driven solutions. Join a team that emphasizes reproducibility, scalable data pipelines, and robust analytics ...
A leading technology company in Barcelona is seeking a Senior Applied Data Scientist to help shape scalable data infrastructure and support AI model development. The ideal candidate will have over 5 years of relevant experience, strong proficiency in Python and SQL, and a master's degree in a related field. This role involves ...
... leading-edge tech stack at the intersection of software engineering, cloud architecture, data platforms, and AI - Learn from the best - access world-class mentorship and training, working with renowned leaders and academics in machine learning, data, and software engineering - Grow your career - discover diverse opportunities ...
... advanced machine learning models to complement and enhance our large-scale, high-precision detection pipeline. In this role, you will work closely with our core engineering, data, and technical leadership teams to ensure the seamless integration of new ML technologies with our existing systems. The core objective of your role ...
Ayesa Digital busca un Data Scientist / Machine Learning Engineer para unirse a un proyecto tecnológico donde participarás en todo el ciclo de vida del dato y en el diseño, entrenamiento e implementación de ML en entornos productivos. Se valorarán 3–5 años de experiencia, dominio de Python y SQL, y experiencia con PySpark ...
... role for an engineer who wants technical depth, customer impact, and a direct voice in the evolution of AI products. Accountabilities - Build polished prototypes and technical demonstrations across serverless inference, databases, MLflow, MLOps, and applied AI use cases, including Physical AI and healthcare and life sciences. ...
EY Global Delivery Services (GDS) in Malaga seeks a Senior Data Scientist with 3+ years in Data Science and AI, focusing on NLP, GenAI, and MLOps. You will design and implement AI models, work with cloud platforms (Azure/AWS/GCP), and collaborate with software engineering teams to operationalize solutions. Ideal candidates ...
... and walk away; you'll keep them accurate as the data and the business shift. And you won't be heads‑down in isolation either. You'll co‑create with engineers and product people, and you'll grow a genuine understanding of how accountants actually think about a document. - Build the models: Train and evaluate classic ML ...
... do we value? - 2–3 years of experience in Data Engineering, Machine Learning, or similar roles. - Advanced proficiency in Python, including libraries such as Pandas, NumPy, and Scikit-learn. - Strong SQL skills for data querying, transformation, and analysis. - Hands-on experience with PySpark for large-scale data processing. ...