Data Engineering Lead at Healthtech Startup · Bengaluru (Bangalore) · 6 - 10 years · ₹20L - ₹30L / yr · Posted 2 Nov 2023

Description:
As a Data Engineering Lead at Company, you will be at the forefront of shaping and managing our data infrastructure with a primary focus on Google Cloud Platform (GCP). You will lead a team of data engineers to design, develop, and maintain our data pipelines, ensuring data quality, scalability, and availability for critical business insights.
Key Responsibilities:
1. Team Leadership:
a. Lead and mentor a team of data engineers, providing guidance, coaching, and performance management.
b. Foster a culture of innovation, collaboration, and continuous learning within the team.
2. Data Pipeline Development (Google Cloud Focus):
a. Design, develop, and maintain scalable data pipelines on Google Cloud Platform (GCP) using services such as BigQuery, Dataflow, and Dataprep.
b. Implement best practices for data extraction, transformation, and loading (ETL) processes on GCP.
3. Data Architecture and Optimization:
a. Define and enforce data architecture standards, ensuring data is structured and organized efficiently.
b. Optimize data storage, processing, and retrieval for maximum
performance and cost-effectiveness on GCP.
4. Data Governance and Quality:
a. Establish data governance frameworks and policies to maintain data quality, consistency, and compliance with regulatory requirements. b. Implement data monitoring and alerting systems to proactively address data quality issues.
5. Cross-functional Collaboration:
a. Collaborate with data scientists, analysts, and other cross-functional teams to understand data requirements and deliver data solutions that drive business insights.
b. Participate in discussions regarding data strategy and provide technical expertise.
6. Documentation and Best Practices:
a. Create and maintain documentation for data engineering processes, standards, and best practices.
b. Stay up-to-date with industry trends and emerging technologies, making recommendations for improvements as needed.
Qualifications
● Bachelor's or Master's degree in Computer Science, Data Engineering, or related field.
● 5+ years of experience in data engineering, with a strong emphasis on Google Cloud Platform.
● Proficiency in Google Cloud services, including BigQuery, Dataflow, Dataprep, and Cloud Storage.
● Experience with data modeling, ETL processes, and data integration. ● Strong programming skills in languages like Python or Java.
● Excellent problem-solving and communication skills.
● Leadership experience and the ability to manage and mentor a team.

Similar jobs (10)
Job Summary
We are seeking a highly skilled GCP Data Engineer with strong expertise in Google Cloud Platform (GCP), Python, ETL, and modern data engineering technologies. The ideal candidate should have hands-on experience designing and building scalable data pipelines using BigQuery, Dataflow, Pub/Sub, Airflow, and modern data lake technologies such as Apache Iceberg or Delta Lake.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines on Google Cloud Platform.
- Build and optimize data processing workflows using Python and Google Cloud Dataflow (Apache Beam).
- Develop and manage large-scale analytical data models in BigQuery.
- Implement event-driven data ingestion using Google Cloud Pub/Sub.
- Create, schedule, and monitor workflows using Apache Airflow and Autosys.
- Design and implement modern data lake architectures using Apache Iceberg or Delta Lake.
- Optimize query performance, storage, and compute costs in GCP.
- Ensure data quality, governance, security, and compliance across data platforms.
- Collaborate with Data Scientists, Analysts, and Application teams to deliver scalable data solutions.
- Troubleshoot production issues and continuously improve pipeline reliability and performance.
Mandatory Skills
- Strong hands-on experience with Google Cloud Platform (GCP).
- Proficiency in Python programming.
- Experience in designing and implementing ETL/ELT pipelines.
- Strong knowledge of BigQuery.
- Experience with Google Cloud Dataflow (Apache Beam).
- Experience with Google Cloud Pub/Sub.
- Hands-on experience with Apache Airflow.
- Experience in job scheduling using Autosys.
- Experience with modern table formats such as Apache Iceberg or Delta Lake.
- Strong SQL and data modeling skills.
Preferred Skills
- Experience with Cloud Storage, Dataproc, Cloud Composer, and Cloud Functions.
- Knowledge of CI/CD pipelines and DevOps practices.
- Experience with Docker and Kubernetes.
- Familiarity with Git and Agile/Scrum methodologies.
- Knowledge of data warehousing and dimensional modeling.
- Exposure to streaming and real-time data processing.
Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
- 4–8+ years of experience in Data Engineering with hands-on expertise in GCP technologies.
Required Experience
- Strong experience in developing enterprise-grade data pipelines using Python and GCP.
- Hands-on experience with BigQuery, Dataflow, Pub/Sub, and Airflow.
- Experience scheduling and monitoring batch workflows using Autosys.
- Experience implementing modern data lake architectures using Apache Iceberg or Delta Lake.
- Strong understanding of ETL best practices, performance tuning, and data optimization.
- Excellent analytical, troubleshooting, and problem-solving skills.
Mandatory Skills
- Google Cloud Platform (GCP)
- Python
- ETL
- BigQuery
- Autosys
- Apache Airflow
- Google Cloud Pub/Sub
- Google Cloud Dataflow (Apache Beam)
- Apache Iceberg / Delta Lake
- SQL & Data Modeling
Role Overview
We are looking for a GCP Data Engineer with 10+ years of experience to design, develop, and optimize scalable cloud-based data solutions. The ideal candidate will have strong hands-on expertise in GCP, BigQuery, and advanced SQL, with experience building data pipelines and working with large-scale datasets.
Key Responsibilities
- Design and develop scalable data pipelines and ETL/ELT processes on GCP.
- Build, optimize, and maintain data solutions using Google BigQuery.
- Develop complex SQL queries for data transformation, aggregation, and analysis.
- Design efficient data models and optimize pipelines for performance, scalability, and cost.
- Integrate data from multiple sources and ensure data quality, reliability, and availability.
- Troubleshoot pipeline and data issues and drive continuous improvement.
- Collaborate with data architects, analysts, application teams, and business stakeholders.
- Follow best practices for cloud security, data governance, testing, and documentation.
Required Skills
- 8+ years of Data Engineering experience
- Strong hands-on experience with GCP, Django, and MongoDB
- Extensive experience with BigQuery
- Advanced SQL skills
- Strong understanding of ETL/ELT and data pipeline development
- Data modeling and data warehousing experience
- Experience handling large-scale datasets and performance optimization
- Strong problem-solving and communication skills
Good to Have
- GCP services such as Cloud Storage, Dataflow, Pub/Sub, Cloud Composer, or Cloud Functions
- Python or other data engineering languages
- Experience with data governance and security
- Agile development experience
Experience: 5+ Years
Employment Type: Full-Time
Role Overview
We are looking for an experienced GCP Data Engineer with 5+ years of experience in data engineering and strong hands-on expertise in Google BigQuery, Google Cloud Storage (GCS), Airflow/Cloud Composer, Python, and Vertex AI. The candidate should be capable of designing, developing, and maintaining scalable data pipelines and cloud-based data solutions on Google Cloud Platform.
Key Skills – Mandatory
- BigQuery – Strong hands-on experience in data warehousing, SQL, optimization, and performance tuning.
- Google Cloud Storage (GCS) – Experience with data storage, file management, and integration with data pipelines.
- Airflow / Cloud Composer – Experience in developing, scheduling, monitoring, and managing data workflows.
- Python – Strong programming skills for data engineering, ETL/ELT development, automation, and pipeline implementation.
- Vertex AI – Experience working with ML/AI workflows, model integration, or data pipelines supporting AI/ML solutions.
Good to Have / Added Advantage
- Dataproc – Experience with distributed data processing and Spark-based workloads.
- Cloud Data Fusion – Experience in building and managing data integration pipelines.
- Cloud Run – Understanding of deploying and running containerized applications/services on GCP.
- Experience with ETL/ELT processes and data pipeline development.
- Knowledge of GCP data architecture and cloud-native services.
- Experience in data quality, validation, monitoring, and troubleshooting.
Responsibilities
- Design, develop, and maintain scalable GCP-based data pipelines.
- Build and optimize data solutions using BigQuery and Cloud Storage.
- Develop and manage workflows using Airflow / Cloud Composer.
- Write efficient and reusable Python code for data processing and automation.
- Support Vertex AI integrations and AI/ML data workflows.
- Monitor pipeline performance and troubleshoot data processing issues.
- Work with cross-functional teams to understand data requirements and deliver reliable solutions.
- Implement best practices for data security, quality, scalability, and performance.
You must have :
- 5+ years of overall experience in Data Engineering.
- Strong hands-on experience with BigQuery, GCS, Airflow/Cloud Composer, Python, and Vertex AI.
- Strong understanding of data engineering concepts, ETL/ELT, data pipelines, and cloud technologies.
- Dataproc, Data Fusion, and Cloud Run experience will be an added advantage.
Job Summary
Role Overview
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, Advanced SQL, CI/CD, DevOps, and Data Analytics. The ideal candidate should have hands-on experience designing and developing scalable data pipelines, transforming large datasets, and supporting data-driven applications.
Experience with Google Cloud Platform (GCP) will be an added advantage.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Python and SQL.
- Develop complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain reliable data integration workflows across multiple data sources.
- Perform data cleansing, validation, transformation, and quality checks.
- Analyze data and provide insights to support business and technical requirements.
- Implement and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps practices and tools to automate deployments, monitoring, and infrastructure processes.
- Troubleshoot data pipeline failures, performance issues, and production incidents.
- Optimize data processing workflows for performance, scalability, and reliability.
- Collaborate with Data Analysts, Data Scientists, Developers, and other stakeholders.
- Follow best practices for version control, testing, documentation, and deployment.
- Contribute to cloud-based data engineering initiatives, preferably on GCP.
Required Skills
- 5–7 years of hands-on experience in Data Engineering.
- Strong programming skills in Python.
- Strong expertise in Advanced SQL and database concepts.
- Hands-on experience with ETL/ELT processes and data pipelines.
- Good understanding of Data Warehousing and Data Modeling concepts.
- Experience with CI/CD practices and tools.
- Strong understanding of DevOps principles, automation, and deployment processes.
- Strong data analytics and problem-solving skills.
- Experience working with large datasets and performance optimization.
- Good understanding of Git/version control and software development best practices.
Good to Have
- Hands-on experience with Google Cloud Platform (GCP).
- Exposure to GCP data services such as BigQuery, Cloud Storage, Dataflow, Composer, or Pub/Sub.
- Experience with containerization/orchestration technologies such as Docker/Kubernetes.
- Experience with workflow orchestration tools such as Airflow.
- Knowledge of cloud-based data architecture and distributed data processing.
Preferred Candidate Profile
- Strong analytical and problem-solving abilities.
- Good communication and stakeholder management skills.
- Ability to work independently as well as in a collaborative team environment.
- Strong ownership of data pipelines and production systems.
- Candidates who can join at short notice are preferred.
Mandatory Skills
Data Engineer, Python , ETL, GCP, Advanced SQL, Strong Data Analytics skills, CICD, Devops
Experience: 6+ years overall Data Engineering experience.
Must-have — candidates should have hands-on experience in ALL of these:
- GCP (Google Cloud Platform) – strong hands-on experience
- Python – data engineering/ETL development
- SQL – advanced SQL, query optimization, data transformation
- BigQuery – strong hands-on experience with development, optimization and data warehousing
- Data Engineering / ETL – building and maintaining data pipelines
- GCP data services – preferably Cloud Storage, Dataflow, Pub/Sub, Composer/Airflow, etc.
- Data warehousing / dimensional modeling
- Design, build, and maintain scalable ETL/ELT pipelines for batch and real-time data ingestion and transformation.
- Develop and optimize data lake and data warehouse architectures (e.g., Snowflake, BigQuery, Redshift).
- Work with cloud platforms GCP, Azure to manage data infrastructure.
- GCP as mandatory skills
- Collaborate with analytics and product teams to understand data needs and deliver solutions.
- Ensure data quality, reliability, security, and compliance across all data systems.
- Mentor junior data engineers and contribute to best practices and code reviews.
- Monitor and troubleshoot data pipeline performance and resolve data-related issues.
- Automate data validation, monitoring, and alerting processes.
- 8+ years of experience in data engineering or software engineering with a data focus.
- Proficient in SQL and at least one programming language (e.g., Python, Scala, Java).
- Experience with modern data warehousing tools (e.g., Snowflake, Redshift, BigQuery).
- Strong understanding of data modeling, data lakes, and ETL/ELT design.
- Hands-on experience with orchestration tools like Airflow, dbt, or similar.
- Solid experience with cloud data platforms (AWS/GCP/Azure).
- Familiarity with CI/CD pipelines, containerization (Docker/Kubernetes), and version control (Git).
- Experience working in a DevOps or DataOps environment.
- Knowledge of data governance, lineage, and cataloging tools (e.g., Collibra, Alation).
- Familiarity with streaming technologies (Kafka, Spark Streaming, Flink).
- Experience supporting machine learning workflows and data science initiatives.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Description – Lead Data Engineer (AWS + Big Data)
Location: Pune - Hybrid
Experience: 8+ Years
Role Overview
We are looking for a Lead Data Engineer with strong hands-on expertise in AWS-based Big Data platforms. The ideal candidate should have extensive implementation experience in designing and building scalable data pipelines, mentoring engineering teams, and driving technical delivery. Databricks exposure is mandatory, while deep implementation experience in Databricks is not essential. Candidates with HBase experience will be preferred.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines and ETL/ELT solutions on AWS.
- Build and optimize Big Data applications using Spark/PySpark, HBase, Hive, Kafka, and related technologies.
- Work with AWS services such as S3, Glue, Lambda, IAM, and CloudWatch for cloud-native data engineering.
- Lead technical implementation, mentor engineers, conduct code reviews, and drive engineering best practices.
- Ensure data quality, performance optimization, CI/CD adoption, and production support.
Required Skills
- Strong hands-on experience with AWS (S3, Glue, Lambda, IAM, CloudWatch)
- Apache Spark (PySpark/Scala) and Big Data ecosystem
- Databricks exposure (mandatory)
- HBase (strongly preferred)
- Python or Java, Advanced SQL
- ETL/ELT development and Data Warehousing concepts
- Apache Airflow or similar orchestration tools
- Git, CI/CD, Performance Tuning, and Data Quality
Good to Have
- Kafka / Spark Structured Streaming
- Hive, Impala, Hadoop ecosystem
- Delta Lake / Lakehouse concepts
- Snowflake or other modern cloud data platforms
Experience Required
- 8–14 years of Data Engineering experience with strong AWS implementation expertise.
- Proven experience leading technical delivery and mentoring engineering teams.
- Strong understanding of enterprise-scale data platforms and Big Data architectures.
- Ability to collaborate with architects, stakeholders, and cross-functional teams to deliver scalable solutions.
NOTE: Final Technical round is mandatory to be taken F2F from Pune, office.
Role Summary
The GCP Delivery Lead is responsible for driving the technical vision, architecture, and delivery of enterprise-scale solutions on Google Cloud Platform (GCP). This role combines deep technical expertise with strong delivery leadership and customer engagement capabilities, working closely with clients, Google teams, and internal stakeholders.
Key Responsibilities
Technical Leadership
- Define and own GCP solution architecture and technical strategy for client engagements
- Design scalable, secure, and resilient cloud-native architectures
- Establish architecture standards, best practices, and reusable frameworks
- Provide governance and oversight across multiple programs
Delivery Leadership
- Ensure successful delivery of cloud transformation and modernization initiatives
- Conduct architecture reviews, risk assessments, and quality assurance
- Guide engineering teams through complex implementations
- Ensure solutions meet performance, security, and scalability standards
Google Partnership & Pre-Sales
- Act as primary technical interface with Google Cloud teams
- Support joint solutioning, co-selling, and client workshops
- Contribute to proposals, solution design, and executive presentations
AI & Data Solutions
- Lead enterprise AI, GenAI, and data platform architecture on GCP
- Drive adoption of Vertex AI, BigQuery, and modern AI/ML platforms
- Advise clients on AI strategy, transformation, and innovation
Practice Development
- Build and mentor high-performing architecture and engineering teams
- Drive certifications, capability building, and hiring initiatives
- Create accelerators, reusable assets, and reference architectures
- Contribute to practice growth and go-to-market strategies
Required Qualifications
Experience
- 12+ years in cloud architecture, software engineering, or consulting
- 5+ years of hands-on GCP architecture and delivery experience
- Proven track record in large-scale cloud transformation programs
- Experience managing multi-client and multi-program delivery
Job Title: Data Engineer – PySpark | Oracle | GCP
Experience: 5–7 Years
Location: Hyderabad
Notice Period: Immediate Joiners Preferred
Job Summary
We are seeking an experienced Data Engineer with strong expertise in PySpark, Oracle, and Google Cloud Platform (GCP) to design, develop, and optimize scalable data pipelines. The ideal candidate should have hands-on experience in ETL development, data integration, and cloud-based data engineering solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL/data pipelines using PySpark.
- Extract, transform, and load data from Oracle databases into GCP environments.
- Build and optimize batch data processing workflows for high performance and reliability.
- Develop data engineering solutions using GCP services.
- Ensure data quality through validation, monitoring, and troubleshooting.
- Optimize SQL queries and ETL jobs for performance and scalability.
Required Skills
- 5–7 years of experience as a Data Engineer.
- Strong hands-on experience with PySpark.
- Solid experience with Oracle Database and advanced SQL.
- Hands-on experience with Google Cloud Platform (GCP).
- Strong understanding of ETL processes and data warehousing concepts.
Work Location: Hyderabad
Notice Period: Immediate Joiners Preferred
At Mitratech, we are a team of technocrats focused on building world-class products that simplify operations in the Legal, Risk, Compliance, and HR functions. We are a close-knit, globally dispersed team that thrives in an ecosystem that supports individual excellence and takes pride in its diverse and inclusive work culture centered around great people practices, learning opportunities, and having fun! Our culture is the ideal blend of entrepreneurial spirit and enterprise investment, enabling the chance to move at a rapid pace with some of the most complex, leading-edge technologies available.
For over 35 years, the experts at Mitratech have been focused on solving the complex needs. Today, we serve 20,000 client companies of all sizes globally, representing 30% of the Fortune 500 and over 500,000 users in over 160 countries.
As we continue to grow, we’re always looking for resourceful, enthusiastic, and fresh perspectives. Join our global team and see what makes Mitratech a truly exceptional place to work!
Job Overview
Principal Data Engineer
About Engineering at Mitratech Legal Solutions
Mitratech's engineering organization is a collaborative and dynamic environment where engineers are empowered to drive technical direction and innovation. Our engineers are passionate about delivering high-quality products and solutions that meet the evolving needs of our customers, and we're committed to fostering a culture of continuous learning and growth.
About the Role
Mitratech is a fast-paced and dynamic environment, and this role requires someone who is adaptable, resilient, and able to thrive in a rapidly changing landscape. If you’re a seasoned engineer with a passion for technical leadership, innovation, and collaboration — including building the data foundations that power trusted reporting and agentic AI-driven products — we’d love to hear from you.
What You Will Do
• Drive technical direction for a significant product domain or platform capability, ensuring alignment with business objectives and customer needs
• Design and maintain data pipelines and reporting models that power trusted business metrics and increasingly feed agentic AI systems (e.g., RAG ingestion, embeddings, vector stores, AI agent workflows)
• Use AI-assisted and agentic engineering tools (e.g., Claude Code, Copilot, Cursor, AI agents) as part of your own workflow, and help other engineers adopt agentic development practices effectively
• Reduce systemic complexity by identifying and leading architectural debt remediation, and developing strategies for ongoing technical debt management
• Partner with Product and Engineering leadership to inform multi-quarter roadmap feasibility, and provide technical guidance and oversight to ensure successful implementation
• Elevate engineering craft across multiple teams through RFCs, mentorship, and knowledge sharing, and develop training programs to improve engineering skills and knowledge
• Represent Mitratech’s technical capabilities externally, including speaking at conferences, contributing to open-source projects, and engaging with industry peers and thought leaders
What We Are Looking For
To be successful in this role, you will need:
• 10+ years of experience in software engineering, with a focus on technical leadership and architecture
• Deep understanding of data engineering principles, including data modeling, data warehousing, reporting, and data governance
• Strong technical expertise in SQL, PostgreSQL, ETL/ELT pipelines, BI tools, and analytics platforms
• Practical experience with AI/LLM-adjacent and agentic AI data work — e.g., RAG ingestion pipelines, embedding generation, vector store management, or building/operating AI agent workflows over data — using AI coding assistants (Claude Code, Copilot, Cursor, or similar) as a regular part of the engineering workflow
• Working knowledge of modern cloud platforms such as AWS
• Experience with BI, reporting, dashboards, and customer-facing analytics
• Experience leading cross-functional initiatives with product, engineering, analytics, and business teams
Nice to Have
• Working knowledge of Ruby on Rails and React
• Experience with a semantic or metrics layer (e.g., dbt Semantic Layer, headless BI)
• Understanding of CI/CD, Git-based workflows, and infrastructure-as-code
The Stack Context
• Modern data stack: Fivetran, Airbyte, dbt, Snowflake, GitHub, Terraform, or similar tools
• Application context (nice to have): Ruby on Rails, React, or similar backend/frontend frameworks
• Data modeling: SQL, analytics models, documentation, testing, naming standards, and version control
• Infrastructure: cloud-based data infrastructure, infrastructure-as-code, CI/CD, monitoring, and cloud storage
• Data workflows: ingestion, transformation, orchestration, reporting, deployment, and change management
• Reporting focus: trusted metrics, scalable reporting models, dashboards, exports, and data quality
• AI surface: data pipelines and quality practices supporting AI/LLM and agentic AI use cases (RAG, embeddings, vector stores, AI agents) alongside traditional BI
Why This Role
This role offers a unique opportunity to drive technical direction and innovation at a rapidly growing company, while also mentoring and coaching engineers to improve their craft. As a Principal Data Engineer at Mitratech, you will have the chance to work on complex and challenging problems spanning trusted reporting and agentic AI systems, collaborate with cross-functional teams, and represent the company's technical capabilities externally. If you're looking for a role that offers a mix of technical leadership, data and reporting depth, agentic AI innovation, and collaboration, this could be the perfect fit for you.
We are an equal-opportunity employer that values diversity at all levels. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, national origin, age, sexual orientation, gender identity, disability, or veteran status.







