Big Data Engineer at SmartHub Innovation Pvt Ltd · Bengaluru (Bangalore) · 5 - 7 years · ₹15L - ₹20L / yr · Raised funding · Posted 10 Jul 2023

JD Code: SHI-LDE-01
Version#: 1.0
Date of JD Creation: 27-March-2023
Position Title: Lead Data Engineer
Reporting to: Technical Director
Location: Bangalore Urban, India (on-site)
SmartHub.ai (www.smarthub.ai) is a fast-growing Startup headquartered in Palo Alto, CA, and with offices in Seattle and Bangalore. We operate at the intersection of AI, IoT & Edge Computing. With strategic investments from leaders in infrastructure & data management, SmartHub.ai is redefining the Edge IoT space. Our “Software Defined Edge” products help enterprises rapidly accelerate their Edge Infrastructure Management & Intelligence. We empower enterprises to leverage their Edge environment to increase revenue, efficiency of operations, manage safety and digital risks by using Edge and AI technologies.
SmartHub is an equal opportunity employer and will always be committed to nurture a workplace culture that supports, inspires and respects all individuals, encourages employees to bring their best selves to work, laugh and share. We seek builders who hail from a variety of backgrounds, perspectives and skills to join our team.
Summary
This role requires the candidate to translate business and product requirements to build, maintain, optimize data systems which can be relational or non-relational in nature. The candidate is expected to tune and analyse the data including from a short and long-term trend analysis and reporting, AI/ML uses cases.
We are looking for a talented technical professional with at least 8 years of proven experience in owning, architecting, designing, operating and optimising databases that are used for large scale analytics and reports.
Responsibilities
- Provide technical & architectural leadership for the next generation of product development.
- Innovate, Research & Evaluate new technologies and tools for a quality output.
- Architect, Design and Implement ensuring scalability, performance and security.
- Code and implement new algorithms to solve complex problems.
- Analyze complex data, develop, optimize and transform large data sets both structured and unstructured.
- Ability to deploy and administrator the database and continuously tuning for performance especially container orchestration stacks such as Kubernetes
- Develop analytical models and solutions Mentor Junior members technically in Architecture, Designing and robust Coding.
- Work in an Agile development environment while continuously evaluating and improvising engineering processes
Required
- At least 8 years of experience with significant depth in designing and building scalable distributed database systems for enterprise class products, experience of working in product development companies.
- Should have been feature/component lead for several complex features involving large datasets.
- Strong background in relational and non-relational database like Postgres, MongoDB, Hadoop etl.
- Deep exp database optimization, tuning ertise in SQL, Time Series Databases, Apache Drill, HDFS, Spark are good to have
- Excellent analytical and problem-solving skill sets.
- Experience in for high throughput is highly desirable
- Exposure to database provisioning in Kubernetes/non-Kubernetes environments, configuration and tuning in a highly available mode.
- Demonstrated ability to provide technical leadership and mentoring to the team

About SmartHub Innovation Pvt Ltd
About
Similar jobs (10)
At Mitratech, we are a team of technocrats focused on building world-class products that simplify operations in the Legal, Risk, Compliance, and HR functions. We are a close-knit, globally dispersed team that thrives in an ecosystem that supports individual excellence and takes pride in its diverse and inclusive work culture centered around great people practices, learning opportunities, and having fun! Our culture is the ideal blend of entrepreneurial spirit and enterprise investment, enabling the chance to move at a rapid pace with some of the most complex, leading-edge technologies available.
For over 35 years, the experts at Mitratech have been focused on solving the complex needs. Today, we serve 20,000 client companies of all sizes globally, representing 30% of the Fortune 500 and over 500,000 users in over 160 countries.
As we continue to grow, we’re always looking for resourceful, enthusiastic, and fresh perspectives. Join our global team and see what makes Mitratech a truly exceptional place to work!
Job Overview
Principal Data Engineer
About Engineering at Mitratech Legal Solutions
Mitratech's engineering organization is a collaborative and dynamic environment where engineers are empowered to drive technical direction and innovation. Our engineers are passionate about delivering high-quality products and solutions that meet the evolving needs of our customers, and we're committed to fostering a culture of continuous learning and growth.
About the Role
Mitratech is a fast-paced and dynamic environment, and this role requires someone who is adaptable, resilient, and able to thrive in a rapidly changing landscape. If you’re a seasoned engineer with a passion for technical leadership, innovation, and collaboration — including building the data foundations that power trusted reporting and agentic AI-driven products — we’d love to hear from you.
What You Will Do
• Drive technical direction for a significant product domain or platform capability, ensuring alignment with business objectives and customer needs
• Design and maintain data pipelines and reporting models that power trusted business metrics and increasingly feed agentic AI systems (e.g., RAG ingestion, embeddings, vector stores, AI agent workflows)
• Use AI-assisted and agentic engineering tools (e.g., Claude Code, Copilot, Cursor, AI agents) as part of your own workflow, and help other engineers adopt agentic development practices effectively
• Reduce systemic complexity by identifying and leading architectural debt remediation, and developing strategies for ongoing technical debt management
• Partner with Product and Engineering leadership to inform multi-quarter roadmap feasibility, and provide technical guidance and oversight to ensure successful implementation
• Elevate engineering craft across multiple teams through RFCs, mentorship, and knowledge sharing, and develop training programs to improve engineering skills and knowledge
• Represent Mitratech’s technical capabilities externally, including speaking at conferences, contributing to open-source projects, and engaging with industry peers and thought leaders
What We Are Looking For
To be successful in this role, you will need:
• 10+ years of experience in software engineering, with a focus on technical leadership and architecture
• Deep understanding of data engineering principles, including data modeling, data warehousing, reporting, and data governance
• Strong technical expertise in SQL, PostgreSQL, ETL/ELT pipelines, BI tools, and analytics platforms
• Practical experience with AI/LLM-adjacent and agentic AI data work — e.g., RAG ingestion pipelines, embedding generation, vector store management, or building/operating AI agent workflows over data — using AI coding assistants (Claude Code, Copilot, Cursor, or similar) as a regular part of the engineering workflow
• Working knowledge of modern cloud platforms such as AWS
• Experience with BI, reporting, dashboards, and customer-facing analytics
• Experience leading cross-functional initiatives with product, engineering, analytics, and business teams
Nice to Have
• Working knowledge of Ruby on Rails and React
• Experience with a semantic or metrics layer (e.g., dbt Semantic Layer, headless BI)
• Understanding of CI/CD, Git-based workflows, and infrastructure-as-code
The Stack Context
• Modern data stack: Fivetran, Airbyte, dbt, Snowflake, GitHub, Terraform, or similar tools
• Application context (nice to have): Ruby on Rails, React, or similar backend/frontend frameworks
• Data modeling: SQL, analytics models, documentation, testing, naming standards, and version control
• Infrastructure: cloud-based data infrastructure, infrastructure-as-code, CI/CD, monitoring, and cloud storage
• Data workflows: ingestion, transformation, orchestration, reporting, deployment, and change management
• Reporting focus: trusted metrics, scalable reporting models, dashboards, exports, and data quality
• AI surface: data pipelines and quality practices supporting AI/LLM and agentic AI use cases (RAG, embeddings, vector stores, AI agents) alongside traditional BI
Why This Role
This role offers a unique opportunity to drive technical direction and innovation at a rapidly growing company, while also mentoring and coaching engineers to improve their craft. As a Principal Data Engineer at Mitratech, you will have the chance to work on complex and challenging problems spanning trusted reporting and agentic AI systems, collaborate with cross-functional teams, and represent the company's technical capabilities externally. If you're looking for a role that offers a mix of technical leadership, data and reporting depth, agentic AI innovation, and collaboration, this could be the perfect fit for you.
We are an equal-opportunity employer that values diversity at all levels. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, national origin, age, sexual orientation, gender identity, disability, or veteran status.
Job Summary
The Technical Lead will be responsible for overseeing and leading projects related to Azure Data Factory (ADF), Azure Databricks, SQL, Oracle PL/SQL, and Python. The role involves designing, developing, and implementing data solutions while ensuring they meet the business requirements and align with best practices. (1.) Key Responsibilities
1. Lead and manage end-to-end data engineering projects using azure data factory, azure databricks, sql, oracle pl/sql, and python.
2. Collaborate with stakeholders to gather and understand requirements for data pipelines and analytics solutions.
3. Design and develop etl processes, data models, and data integration solutions.
4. Provide technical guidance and mentorship to the team members.
5. Ensure data quality, data governance, and data security standards are maintained throughout the project lifecycle.
6. Troubleshoot and optimize data pipelines and processes for performance and efficiency.
7. Stay updated on the latest trends and technologies in data engineering and contribute to continuous improvement efforts.
Skill Requirements
1. Proficiency in azure data factory (adf) and azure databricks for building and managing data pipelines.
2. Strong experience with sql and oracle pl/sql for data querying and manipulation.
3. Advanced programming skills in python for scripting and data processing tasks.
4. Knowledge of data modeling, data warehousing concepts, and database design principles.
5. Ability to work in a collaborative team environment and communicate effectively with stakeholders.
6. Strong analytical and problem-solving skills with attention to detail.
7. Experience in data visualization tools and techniques is a plus.
Certifications: Relevant certifications in Azure Data Factory, Azure Databricks, SQL, Oracle PL/SQL, or Python are advantageous.
Skill (Primary)
Data Fabric-Azure-Azure Data Factory (ADF)
Job Summary
Role Overview
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, Advanced SQL, CI/CD, DevOps, and Data Analytics. The ideal candidate should have hands-on experience designing and developing scalable data pipelines, transforming large datasets, and supporting data-driven applications.
Experience with Google Cloud Platform (GCP) will be an added advantage.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Python and SQL.
- Develop complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain reliable data integration workflows across multiple data sources.
- Perform data cleansing, validation, transformation, and quality checks.
- Analyze data and provide insights to support business and technical requirements.
- Implement and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps practices and tools to automate deployments, monitoring, and infrastructure processes.
- Troubleshoot data pipeline failures, performance issues, and production incidents.
- Optimize data processing workflows for performance, scalability, and reliability.
- Collaborate with Data Analysts, Data Scientists, Developers, and other stakeholders.
- Follow best practices for version control, testing, documentation, and deployment.
- Contribute to cloud-based data engineering initiatives, preferably on GCP.
Required Skills
- 5–7 years of hands-on experience in Data Engineering.
- Strong programming skills in Python.
- Strong expertise in Advanced SQL and database concepts.
- Hands-on experience with ETL/ELT processes and data pipelines.
- Good understanding of Data Warehousing and Data Modeling concepts.
- Experience with CI/CD practices and tools.
- Strong understanding of DevOps principles, automation, and deployment processes.
- Strong data analytics and problem-solving skills.
- Experience working with large datasets and performance optimization.
- Good understanding of Git/version control and software development best practices.
Good to Have
- Hands-on experience with Google Cloud Platform (GCP).
- Exposure to GCP data services such as BigQuery, Cloud Storage, Dataflow, Composer, or Pub/Sub.
- Experience with containerization/orchestration technologies such as Docker/Kubernetes.
- Experience with workflow orchestration tools such as Airflow.
- Knowledge of cloud-based data architecture and distributed data processing.
Preferred Candidate Profile
- Strong analytical and problem-solving abilities.
- Good communication and stakeholder management skills.
- Ability to work independently as well as in a collaborative team environment.
- Strong ownership of data pipelines and production systems.
- Candidates who can join at short notice are preferred.
Mandatory Skills
Data Engineer, Python , ETL, GCP, Advanced SQL, Strong Data Analytics skills, CICD, Devops
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Summary
We are seeking a motivated Data Engineer with strong skills in SQL, Python, and Linux to design, build, and maintain scalable data pipelines and support data-driven decision-making. The ideal candidate should have experience working with large datasets, ETL processes, and relational databases while ensuring data quality and performance.
Key Responsibilities
- Design, develop, and maintain ETL/ELT data pipelines.
- Write optimized SQL queries, stored procedures, and database objects.
- Develop Python scripts for data extraction, transformation, and automation.
- Work in Linux environments to manage scripts, cron jobs, and system processes.
- Monitor and troubleshoot data pipeline failures.
- Ensure data integrity, consistency, and quality across systems.
- Collaborate with data analysts, software engineers, and business stakeholders.
- Optimize database performance and query execution.
- Participate in code reviews and follow best engineering practices.
Required Skills
- Strong proficiency in SQL (joins, subqueries, window functions, CTEs, indexing, query optimization).
- Good programming experience in Python.
- Hands-on experience with Linux commands and shell scripting.
- Understanding of ETL/ELT concepts and data warehousing.
- Knowledge of relational databases such as PostgreSQL, MySQL, Oracle, or SQL Server.
- Familiarity with Git for version control.
- Strong problem-solving and analytical skills.
Roles & Responsibilities
- Design, develop, and deliver scalable end-to-end data pipelines using Azure Data Factory, ensuring robust integration
of enterprise-wide data from diverse sources
• Build and optimize data engineering workflows using Databricks and PySpark
• Write efficient, high-performance SQL for data transformation and analysis
• Work with the Azure Cloud platform and associated services, applying strong understanding of data warehousing,
data models, and pipelines
• Provide technical leadership to a team of developers, including code reviews and enforcing best practices across the
development lifecycle
• Oversee CI/CD implementation using Azure DevOps, managing deployments across development, QA, and production
environments with proper change control processes
• Collaborate with cross-functional teams to translate business requirements into scalable data solutions
• Ensure data quality, reliability, and performance across all pipelines and platforms
Ideal Candidate
1Strong Azure Databricks Engineer / Senior Data Engineer Profile
2Mandatory (Experience 1) – Must have minimum 8+ years of overall experience in Data Engineering, Data Development, or related data technology roles, with strong hands-on experience in enterprise data pipeline development.
3Mandatory (Experience 2) – Must have strong hands-on experience with Azure Databricks, including development and optimization of scalable data engineering workflows using Databricks and PySpark.
4Mandatory (Experience 3) – Must have strong hands-on proficiency in PySpark/Python and SQL, with proven experience developing complex data transformations, processing workflows, and performance-optimized queries.
5Mandatory (Experience 4) – Must have hands-on experience with Azure Data Factory (ADF) for designing, developing, and orchestrating end-to-end data pipelines and integrating data from multiple sources.
6Mandatory (Experience 5) – Must have strong experience working on the Azure Cloud platform and associated data services, with solid understanding of data warehousing, data modeling, pipeline architecture, and enterprise data solutions.
7Mandatory (Experience 6) – Must have hands-on experience implementing CI/CD using Azure DevOps, including deployment and release management across development, QA, and production environments.
8Mandatory (Experience 7) – Must have proven technical leadership experience, including code reviews, enforcing development best practices, mentoring developers, and providing technical guidance to a data engineering team.
9Mandatory (Notice Period) – Immediate joiners or candidates who can join within 15 days.
10Mandatory (Note) - The position is open across all Cognizant offices pan India. Candidates must be willing to attend the F2F interview at the nearest Cognizant office location.
Job Title : Senior Data Engineer – Databricks
Experience : 14 to 20 Years
Location : HSR Layout, Bangalore
Work Mode : Hybrid – 3 Days WFO
Shift : 11:30 AM – 07:30 PM IST
Positions : 2
Notice Period : Immediate Joiners Only
Interview : 1 Technical Round + 2 Client Rounds
Role Overview :
We are looking for a Senior Data Engineer to build and lead enterprise-scale data platforms for a Switzerland-based commodity client.
The role requires a strong hands-on Data Engineering professional with expertise in Databricks, PySpark, Python, SQL, and AWS, along with technical leadership and stakeholder management experience.
Must-Have Skills :
- 14 to 20 years of Data Engineering experience
- Databricks & Apache Spark / PySpark
- Python & SQL
- AWS Cloud
- Lakehouse Architecture
- ETL / ELT & Distributed Data Processing
- Batch & Streaming Pipelines
- Data Pipeline Optimization & Data Modeling
- CDC & Incremental Processing
- Git, CI/CD & Testing
- Data Quality, Monitoring & Observability
- Technical Leadership & Stakeholder Management
Key Responsibilities :
- Design and build scalable data pipelines using Databricks, PySpark, Python, SQL, and AWS.
- Own data products from design through production.
- Develop batch / streaming pipelines and reusable ETL / ELT frameworks.
- Optimize pipelines for performance, scalability, reliability, and cost.
- Design scalable data architectures and data models.
- Implement data quality, monitoring, lineage, and CI/CD practices.
- Lead technical discussions and mentor engineering teams.
- Collaborate with business stakeholders, architects, product owners, and engineering teams.
- Remain hands-on while providing technical leadership.
Ideal Candidate :
A 14 to 20 years experienced, hands-on Data Engineering leader with strong Databricks + PySpark + AWS expertise, excellent communication, stakeholder management, and experience delivering enterprise-scale data platforms.
🔴 Super Urgent : Only Bangalore-based immediate joiners.
Key Responsibilities
- Lead end-to-end data migration initiatives, including assessment, planning, mapping, transformation, validation, and reconciliation.
- Define and implement data governance frameworks, standards, policies, and processes.
- Design and manage data solutions using Microsoft Azure Data Services.
- Lead development of BI dashboards, reports, KPIs, and analytics solutions.
- Work with business and technical stakeholders to understand reporting and data requirements.
- Develop and maintain data models, data pipelines, ETL/ELT processes, and reporting architecture.
- Ensure data quality, consistency, integrity, security, and compliance throughout migration and reporting processes.
- Identify data risks, dependencies, gaps, and migration challenges and drive their resolution.
- Establish data validation and reconciliation mechanisms to ensure migration accuracy.
- Provide technical leadership and guidance to data engineers, BI developers, and other project team members.
- Collaborate with application, cloud, infrastructure, and business teams during project implementation.
- Monitor data migration and BI deliverables against project timelines, quality standards, and business objectives.
- Prepare technical documentation, data dictionaries, mapping documents, governance guidelines, and project reports.
Required Skills & Experience
- 7+ years of experience in data, migration, BI, or related technology roles.
- Strong experience working on software/IT projects and managing data-related workstreams.
- Hands-on experience with Microsoft Azure Data Services.
- Strong understanding of data migration methodologies, ETL/ELT, data transformation, and reconciliation.
- Experience in Data Governance, Data Quality, Master Data, Metadata Management, and Data Security.
- Strong experience with BI reporting and dashboard development.
- Good understanding of SQL and relational databases.
- Experience with Power BI and data visualization is highly desirable.
- Knowledge of Azure services such as Azure Data Factory, Azure Data Lake, Azure Synapse Analytics, Azure SQL Database, or equivalent.
- Strong understanding of data architecture and data lifecycle management.
- Excellent stakeholder management, communication, analytical, and problem-solving skills.
Must-Have Skills
- Minimum 3 years of experience in Data Engineering / Analytics Engineering / Fintech Data roles
- Must have worked on SMS Parsing, intelligent platform, converting RAW customer SMS data into structured actionable financial signals and enabling downstream usage of SMS derived variables
- Must have established a continuous learning cycle to expand parser coverage
- Experience in Lending / NBFC / Fintech domain
- Experience working with Bureau, SMS, Device, or Banking data
- Strong Python and SQL (production level)
- Experience handling unstructured data (SMS, logs, JSON, APIs)
- Experience building data pipelines, schedulers, and cron jobs
- Strong database design and data modelling skills
- Ability to work in a startup environment with high ownership
- Familiarity with modern platforms like AWS, Snowflake, Google BigQuery, Redshift
Good to Have
- Experience in STPL, especially less than 25K ticket size
- Experience with streaming (Kafka/Kinesis) and orchestration (Airflow or Step Functions)
- Experience with feature stores and risk analytics datasets
- Knowledge of regex, NLP basics for SMS parsing
- Experience supporting real-time decision engines/underwriting systems
Role Summary
This role will be responsible for owning the end-to-end data-structuring layer across the organisation. The individual will transform large volumes of raw, unstructured, and semi-structured data (such as SMS, device, bureau, and app data) into clean, standardised, and analysis-ready datasets. These structured datasets will directly power risk analytics, fraud detection, marketing insights, collections strategy, and policy decisioning.
Key Objective of the Role
Ensure all raw lending data (SMS, Bureau, Device, AA, App logs) is captured, parsed, structured, and stored in a clean analytics-ready format inside databases (PostgreSQL, DynamoDB, AWS stack) so that the Risk and Data Science team can directly use it for feature creation, policy building, and portfolio monitoring.
Core Responsibilities
- End-to-End Data Ownership
- Design, build, and maintain end-to-end data pipelines (batch + streaming) using AWS native services (Glue, Lambda, Step Functions, Kinesis, S3, Athena, Redshift, EMR/Spark, etc.): ingestion
→ parsing → structuring → storage
- Work closely with Tech, Product, and Data Science to define what data should be captured
- Maintain data documentation, data dictionaries, and schema governance
- Ensure data quality, consistency, and version control
- Unstructured Data Processing (Highest Priority)
- Parse raw SMS dumps and categorise into salary, EMI, loan apps, collections, credits, debits, OTP, etc.
- Process device fingerprint, behavioural logs, and vendor data (FinBox, AA, Bureau APIs)
- Convert JSON, logs, and raw API responses into structured feature tables
- Build regex/keyword-based parsers for financial SMS classification
- Feature Implementation (From Risk & Data Science Team)
- Implement feature creation logic provided by Risk/Data Science team
- Translate business and policy logic into SQL/Python pipelines
- Create reusable feature layers for underwriting, fraud, collections, and monitoring
- Maintain a feature store for consistent model and policy usage
- Lending Data Understanding (Domain-Specific Requirement)
- Work with Bureau data
- Structure SMS-derived financial variables (income, stress, EMI signals)
- Work with Account Aggregator and bank transaction datasets
- Understand fintech alternate data used in underwriting and fraud detection
- Data Pipelines & Automation
- Build and maintain ETL/ELT pipelines using Python & SQL
- Create cron jobs for automated data ingestion and feature refresh
- Automate vendor data pulls (Bureau, SMS SDK, AA, device data)
- Ensure low-latency pipelines for real-time underwriting use cases
- Database Structuring & Storage Architecture
- Structure clean datasets in PostgreSQL (analytics layer)
- Manage raw data storage in DynamoDB / S3 data lake
- Design normalized and denormalised tables for risk analytics
- Optimise database performance for large-scale query workloads
- Dashboards & Readable Data Layer
- Create analytics-ready datasets, implement & write Metabase queries and convert into dashboards (Metabase / Power BI)
- Enable self-serve data access for Risk, Business, and Founders
- Support ad-hoc analysis requirements from leadership
- Cross-Functional Collaboration (Very Important)
- The role requires close collaboration with data science, tech, product, and business teams to ensure reliable data pipelines, well-defined schemas, API integrations, logging architecture and high data quality, enabling faster and more accurate decision-making across lending workflows.
Tech Stack (Current Environment)
- AWS Services
- PostgreSQL (Primary analytics DB)
- DynamoDB (Raw/NoSQL storage)
- Python (Pandas, NumPy, ETL frameworks)
- Advanced SQL
- APIs, JSON, and Log Data Handling
Job Title : Data Engineer – Databricks
Experience : 6+ Years
Location : Noida / Hyderabad / Chennai / Pune / Bengaluru (Hybrid)
Shift : IST (Normal Shift)
Job Summary :
We are seeking an experienced Data Engineer with strong expertise in Databricks, Snowflake, Python, and Spark to build and optimize scalable data pipelines and support AI/ML model deployments. The ideal candidate should have experience working with cloud-based data platforms and preferably possess exposure to the Healthcare domain.
Required Skills :
- Databricks (Preferred)
- Snowflake
- Python
- Apache Spark
- SQL
- Azure Cloud
- Kubernetes
- Apache Airflow
- GitHub & CI/CD Pipelines
- AI/ML Model Deployment
- Data Analytics
Preferred :
- Experience in the Healthcare domain.
- Strong understanding of scalable data engineering architectures and best practices.
Job Summary
We are seeking a skilled Data Engineer to design, build, and maintain scalable data pipelines and infrastructure. The ideal candidate should have strong expertise in SQL, Python, Linux, and modern data engineering practices to support data integration, transformation, and analytics.
Key Responsibilities
- Design, develop, and maintain ETL/ELT data pipelines.
- Write efficient and optimized SQL queries for data extraction, transformation, and reporting.
- Develop automation scripts using Python for data processing and workflow optimization.
- Work with Linux environments for deployment, monitoring, and troubleshooting.
- Ensure data quality, integrity, and reliability across data platforms.
- Collaborate with data analysts, software engineers, and business stakeholders to deliver data solutions.
- Monitor, troubleshoot, and optimize data pipelines for performance and scalability.
- Implement best practices for data security, governance, and documentation.
Required Skills
- Strong experience in Data Engineering concepts and ETL/ELT processes.
- Proficiency in SQL, including query optimization and database design.
- Strong programming skills in Python.
- Hands-on experience with Linux commands, shell scripting, and system administration basics.
- Experience with relational databases such as PostgreSQL, MySQL, SQL Server, or Oracle.
- Familiarity with Git/version control.
- Strong analytical and problem-solving skills.
Preferred Skills
- Experience with cloud platforms (AWS, Azure, or GCP).
- Knowledge of Apache Spark, Airflow, Kafka, or similar data engineering tools.
- Experience with data warehousing solutions and big data technologies.
- Understanding of CI/CD pipelines and containerization (Docker/Kubernetes).
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Relevant certifications in cloud or data engineering are an added advantage.






