Description:
• Data Pipeline Development: Designing and building
scalable, efficient, and reliable data pipelines to extract, transform, and
load (ETL) data from various sources into data storage systems, such as
data lakes or data warehouses.
• Provide leadership on the migration from Hadoop to Azure and on how we
leverage data in multiple portfolios
• Data Modeling: Developing and implementing data models that optimize data
storage, retrieval, and analysis, considering factors such as data volume,
variety, and velocity.
• Data Integration: Integrating data from different sources, including
databases, APIs, streaming platforms, and third-party systems, to ensure
data consistency and accuracy.
• Data Transformation and Cleansing: Applying data transformation techniques
to clean, preprocess, and enrich data, ensuring its quality and readiness
for analysis and reporting.
• Data Warehousing and Storage: Designing and maintaining data warehouses or
data lakes to efficiently store and manage large volumes of structured and
unstructured data.
• Data Governance and Security: Implementing data governance practices and
security measures to protect sensitive data, ensure data privacy, and
comply with regulations such as GDPR or CCPA.
• Performance Optimization: Monitoring and optimizing data processing and
query performance, identifying and resolving bottlenecks to improve data
retrieval and analysis efficiency.
• Collaboration with Data Analysts and Scientists: Collaborating with data
analysts and data scientists to understand their requirements and provide
them with reliable, accurate, and well-organized data sets for analysis and
modeling.
Key Skillsets :
• Strong Programming and Scripting
Skills: Proficiency in programming
languages such as Python, SQL, or Scala, as well as experience with
scripting and automation tools, for data manipulation, ETL processes, and
pipeline development.
• Experience with Azure cloud, Azure data bricks, SQL , Spark, Scala or
Python
• Experience in Jira, Github, DevOps tools
• Data Integration and ETL Tools: Familiarity with tools and frameworks for
data integration and ETL processes, such as Apache Kafka, Apache Spark,
Apache Airflow, or Talend.
• Data Modeling and Database Design: Knowledge of data modeling concepts,
dimensional modeling, and schema design to create efficient and scalable
database structures.
• Big Data Technologies: Experience with big data platforms and technologies
like Apache Hadoop, Apache Hive, Apache HBase, or Apache Cassandra for
handling large volumes of data.
• Cloud Platforms: Familiarity with cloud platforms and related services for
data storage, processing, and analytics.
• Data Governance and Security: Understanding of data governance principles,
data security best practices, and compliance requirements to ensure data
integrity, privacy, and regulatory compliance.
• Problem-solving and Analytical Thinking: Strong problem-solving skills,
attention to detail, and analytical thinking to identify and resolve
data-related issues and optimize data pipelines.
• Communication and Collaboration: Excellent communication skills to
effectively collaborate with cross-functional teams, including data
analysts, data scientists, and stakeholders, and to explain complex
technical concepts to non-technical stakeholders.