• Data Modeling and Pipelines development with Spark and Scala to ingest and transform data from different sources (Kafka topics, APIs, HDFS, structured databases, files...) into HDFS, IBM Cloud Storage (generally in parquet format) or SQL/NOSQL databases
followign complex business rules.
• Manage big data storage solutions in the platform (HDFS, IBM Cloud Storage, structured and non-structured databases).
• Data Transformation and Quality: implement data transformation and quality control processes to ensure data consistency and accuracy. Utilize programming languages such as Scala and SQL. And libraries like Spark for data transformation and enrichment
operations.
• CI/CD Pipeline Implementation: set up CI/CD pipelines to automate deployment, unit testing, and development management.
• Infrastructure Migration: migrate the existing Hadoop infrastructure to cloud infrastructure on Kubernetes Engine, Object Storage (IBM Cloud storage), Spark as a service on Scala (to build the data pipelines), and Airflow as a service (to orchestrate and
schedule the data pipelines).
• Implementation of schemas,queries, and views in SQL/NOSQL databases like Oracle, Postgres or MongoDB.
• Develop and configure scheduling of data pipelines with a combination of shell scripting and AirFlow as a service.
• Validation Testing: conduct unit and validation tests to ensure accuracy and integrity.
• Documentation: write technical documentation (specifications, operational documents) to ensure knowledge capitalization.
• Code Improvement: modify the existing code as per business requirements and continuously improve for better performance and maintainability.
• Configure Dremio Data Virtualization to interface with Parquet or as a way to expose the data in the different data products.
• Performance Optimization and Security: ensure the performance and security of the data infrastructure and follow the best practices of Data engineering.
IT Tools:
• Spark on Scala as legacy data pipeline development language.
• Spark as a service on Scala as data pipeline development platform.
• Experience in the design and development of streaming procesess using Spark Streaming, Spark Structure Streaming and Apache Kafka.
• Management of legacy big data storage solutions (HDFS).
• Management of big data storage solutions (IBM Cloud Object Storage and parquet format).
• Implementation of SQL/NO SQL database schemas,queries and views (MongoDB, Oracle, Postgres).
• Shell scripting and Airflow as data pipelie scheduling solution,
• Dremio as data virtualization tool.