Cutshort logo
For Employers
Team Geek Solutions logo
Senior Data Engineer
Senior Data Engineer

Senior Data Engineer at Team Geek Solutions · Bengaluru (Bangalore), Pune, Hyderabad, Chennai · 7 - 10 years · ₹16L - ₹24L / yr · Bootstrapped · Posted 19 Aug 2026

Team Geek Solutions's logo

Senior Data Engineer

Mamta K's profile picture
Posted by Mamta K
7 - 10 yrs
₹16L - ₹24L / yr
Bengaluru (Bangalore), Pune, Hyderabad, Chennai
Skills
skill iconPython
Data engineering
ETL
SQL
Microsoft fabric
Data Factory
Azure Synapse
Data modeling

Job Description:

Position: Senior Data Engineer

Location: Chennai / Pune / Bangalore / Hyderabad

Working Type: WFO

Shift: UK Shift (2:00 – 11:00 PM)

Experience : 7+ years overall

Interviews: Assessment || 2 Interview rounds.


Notice Period: Immediate Joiner



Key Responsibilities


Implement ingestion, transformation, and optimization of enterprise data sources into Microsoft Fabric Lakehouse environments.

Configure and optimize Fivetran connectors (Oracle, SQL DB, etc.)

Manage large-volume ingestion and backfill operations

Implement Bronze to Silver transformation pipelines

Develop incremental load and CDC logic

Optimize Lakehouse performance and storage patterns

Implement monitoring (record counts, load duration, failure tracking)

Support Dev/Test/Prod promotion processes



Required Qualifications

7+ years of data engineering experience

Hands-on experience with Microsoft Fabric or Azure Synapse/Data Factory

Strong experience with Fivetran or similar ELT tools

Experience handling high-volume datasets (hundreds of millions of records)

Proficiency in SQL, Python, and data modeling concepts

Strong understanding of Medallion architecture.

Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About Team Geek Solutions

Founded :
2018
Type :
Services
Size
Stage :
Bootstrapped

About

Team Geek Solutions (TGS) is a global technology partner based in Leander, Texas, with a regional office in Pune, India. Founded in 2018 and established as a private limited company in 2020, TGS has rapidly grown into a USD $40 million company with over 150 team members worldwide. The company specializes in AI and Generative AI solutions, custom software development, and talent optimization, offering managed remote teams and scalable solutions tailored to various industries. TGS provides a wide range of services, including AI/ML application development, cloud migration, custom software development, and cybersecurity audits. They also offer specialized resources in Generative AI and automated hiring solutions. TGS supports businesses in sectors such as banking, telecom, fintech, healthcare, and manufacturing, helping them enhance operational efficiency and drive innovation. The company is committed to delivering custom software solutions and cloud-based platforms that empower organizations to operate more effectively.

Read more

Company social profiles

linkedin

Similar jobs (10)

company logo
shruthi k
Posted by shruthi k
Remote, Pune
7 - 18 yrs
₹1L - ₹35L / yr (ESOP available)
microsoft fabric,dataengineering,one lake

Data Engineer – Microsoft Fabric

Location: Pune, India

Work Mode: Hybrid

Experience: 6+ Years

Employment Type: Full-time contactor

Compensation: As per market standards, commensurate with experience and expertise

Shift Timings: 2:00 PM – 11:00 PM IST

Notice Period: 0 – 15 days

About the Role

Jade Business Services (JBS) is seeking a Data Engineer – Microsoft Fabric to join our Pune team and work on enterprise-scale data transformation and analytics initiatives.

We are looking for a hands-on Data Engineer with strong experience in Microsoft Fabric, SQL, Python/PySpark and modern data engineering practices. The candidate will be responsible for building scalable data pipelines, implementing Lakehouse and Warehouse solutions, developing data models and supporting governed, reliable and AI-ready data platforms.

The ideal candidate should be comfortable working with architects, engineering teams and client stakeholders to translate business requirements into scalable and production-ready data solutions.

Roles and Responsibilities

  • Design and develop data solutions using Microsoft Fabric, including OneLake, Lakehouse, Warehouse and Data Factory pipelines.
  • Build and maintain scalable ETL/ELT pipelines for batch and incremental data processing.
  • Develop data ingestion and transformation pipelines using Fabric Data Factory, SQL, Python and/or PySpark.
  • Implement Medallion Architecture using Bronze, Silver and Gold layers.
  • Work with Lakehouse and Fabric Warehouse for enterprise data processing and analytics.
  • Develop and maintain data models, tables, views and optimized SQL queries.
  • Build and support semantic models for Power BI and analytical workloads.
  • Implement data quality, validation, monitoring and error-handling mechanisms.
  • Work with metadata, lineage and governance requirements using Microsoft Purview.
  • Implement data security, access controls and role-based permissions across data platforms.
  • Support Data Product and domain-oriented data architecture principles.
  • Follow DataOps practices including CI/CD, deployment, monitoring and production support.
  • Troubleshoot pipeline failures, performance issues and data quality problems.
  • Optimize data pipelines, queries and storage for performance and cost efficiency.
  • Work closely with Data Architects and business stakeholders to understand requirements and implement technical solutions.
  • Participate in technical design discussions, code reviews and architecture reviews.
  • Maintain technical documentation, data flow diagrams and pipeline documentation.
  • Support production deployments, incident resolution and SLA-driven data platform operations.
  • Identify opportunities for automation and AI-assisted improvements across data engineering processes.

Qualifications and Skills

  • 6+ years of experience in Data Engineering, Data Integration or Data Platform development.
  • Strong hands-on experience with Microsoft Fabric.
  • Experience with:
  • Microsoft Fabric Lakehouse
  • Fabric Warehouse
  • OneLake
  • Fabric Data Factory / Pipelines
  • Semantic Models
  • Strong understanding of Lakehouse and Medallion Architecture.
  • Strong SQL development and query optimization skills.
  • Hands-on experience with Python and/or PySpark.
  • Experience developing enterprise ETL/ELT and data integration pipelines.
  • Experience with batch and incremental data processing.
  • Understanding of data modelling concepts including dimensional modelling.
  • Knowledge of data quality, metadata, lineage and data governance.
  • Working knowledge of Microsoft Purview.
  • Understanding of Data Mesh and Data Product concepts.
  • Experience with CI/CD, version control, monitoring and DataOps practices.
  • Understanding of cloud security, access controls and data privacy.
  • Good troubleshooting and problem-solving skills.
  • Strong communication skills and ability to work with distributed and client-facing teams.

Preferred Skills

  • Microsoft Fabric or Azure Data certifications.
  • Experience migrating workloads from Azure Synapse, SQL Server, Databricks or other data platforms to Microsoft Fabric.
  • Experience implementing Medallion Architecture on Microsoft Fabric.
  • Experience with Power BI and semantic modelling.
  • Exposure to AI/ML, Generative AI or Agentic AI use cases on enterprise data platforms.
  • Experience working with Data Products or domain-oriented data solutions.
  • Experience in Energy & Utilities, Healthcare, Financial Services or Insurance.
  • Experience working with US or international enterprise clients.

What We Expect

The ideal candidate should be hands-on first and capable of independently building, troubleshooting and optimizing Fabric data solutions. You should be able to explain the technical decisions behind your implementation and work effectively with architects and engineering teams to deliver production-ready solutions.

 

Read more
company logo
Lata Deepak
Posted by Lata Deepak
Bengaluru (Bangalore)
8 - 15 yrs
₹10L - ₹27L / yr
Python,data lake ,data pipeline, Azure Databricks,

Azure Data Factory and Azure Databricks, processing 5 million+ records/week from 4+ source systems into a governed Lakehouse. 


Read more
Service Based Company
Service Based Company
Agency job
via by Chandra M
Delhi, Gurugram, Noida, Ghaziabad, Faridabad, Chennai, Hyderabad, Pune, Kolkata
7 - 10 yrs
₹10L - ₹15L / yr
databricks
Azure Data Factory
Data engineering

Dear Candidate,


Greeting from NAM Info Pvt Ltd.


We have a role for Data Engineer position with NAM Info.


This role will be permanent with NAM info and deploy to client

location NEW DELHI~CHENNAI~HYDERABAD~PUNE~KOLKATA.


Work Mode: WORK FROM OFFICE

A decent hike can be provided based on current CTC

Interview Mode: Virtual

Role Descriptions:

Exp Range: 7 - 10 years

City Locations: NEW DELHI~CHENNAI~HYDERABAD~PUNE~KOLKATA

Key Responsibilities*

Role: Data Engineer


Location: ~NEW DELHI~CHENNAI~HYDERABAD~PUNE~KOLKATA

Skills: Digital: Databricks, Azure Data Factory

Experience Required: 8-10

Descriptions:

Good information and sound knowledge in Azure Synapse Analytics Azure Data Factory (ADF)Big Data technologies and data processing frameworks Azure Data Warehouse and associated Azure data platform services Data integration| data modelling| and performance optimization


Desire candidate

  • Candidate should have valid PF.


Regards,

NAM Info 

Read more
leading global IT services company offering enterprise digit
leading global IT services company offering enterprise digit
Agency job
via by Priyadharshini Periyasamy
icon

The recruiter has not been active on this job recently. You may apply but please expect a delayed response.

Bengaluru (Bangalore), Pune, Hyderabad, Chennai
11 - 14 yrs
₹25L - ₹35L / yr
databricks
Azure

Hi 

Position: Senior Data Engineer 

Experience: 5–7Years

Location: Bangalore / Chennai / Hyderabad / Pune

Work Model: Hybrid (Mandatory 2 Days Work From Office)

Job Description

  • Design and build Bronze, Silver, and Gold data pipelines using Microsoft Fabric and Databricks.
  • Implement Cascade 1.0 & 2.0 merge logic in the Silver layer, including handling new and nulled columns.
  • Migrate legacy SQL Server tables, views, and stored procedures into Microsoft Fabric Notebooks and Pipelines.
  • Develop scalable PySpark transformations and Delta Lake models while optimizing performance and cost.
  • Establish data quality checks, lineage, and source-to-target reconciliation.
  • Collaborate with Report Engineers and QA teams to deliver governed semantic models.

Mandatory Skills

  • Microsoft Fabric (Lakehouse, OneLake, Pipelines)
  • PySpark
  • SQL / T-SQL
  • Delta Lake
  • Azure Data Factory (ADF)
  • Power BI (Datasets / Semantic Models)
  • Databricks
  • Medallion Architecture (Bronze / Silver / Gold)

Preferred Skills

  • SQL Server
  • SSIS / SQL Server ETL Migration
  • Lakehouse Architecture
  • Data Quality & Validation
  • Performance Tuning

Candidate Requirements

  • 5-7 years of overall IT experience with a minimum of 5+ years in Data Engineering.
  • Hands-on expertise in Microsoft Fabric and/or Databricks.
  • Strong experience in PySpark, SQL/T-SQL, Delta Lake, and Azure Data Factory.
  • Experience migrating legacy SQL Server/SSIS ETL workloads to modern Lakehouse platforms is highly preferred.
  • Good understanding of Power BI datasets and semantic models.
  • Strong communication and stakeholder management skills.

Interview & Work Model

  • Face-to-Face HR interview is mandatory.
  • Candidate should be flexible to travel to the Chennai office whenever required for project-related work.
Read more
company logo
SaiSruthi Nuthanpati
Posted by SaiSruthi Nuthanpati
Tirupati, Chennai
5 - 10 yrs
Best in industry
SQL
skill iconPython
Stored Procedures
skill iconAmazon Web Services (AWS)
Microsoft Windows Azure
+10 more

About Us:

The QX Impact was launched with a mission to make A.I accessible and affordable and deliver AI Products/Solutions at scale for the enterprises by bringing the power of Data, AI, and Engineering to drive digital transformation. We believe without insights; businesses will continue to face challenges to better understand their customers and even lose them. Secondly, without insights businesses won't’ be able to deliver differentiated products/services; and finally, without insights, businesses can’t achieve a new level of “Operational Excellence” is crucial to remain competitive, meeting rising customer expectations, expanding markets, and digitalization.


Job Summary:

We are looking for a Senior Data Engineer who is creative, collaborative, and adaptable to join our agile team of data scientists, engineers, and UX developers. The role focuses on building and maintaining robust data pipelines to support advanced analytics, data science, and BI solutions.

As a Senior Data Engineer, you will work with internal and external data, collaborate with data scientists, and contribute to the design, development, and deployment of innovative solutions.


Key Responsibilities:

  • Design, develop, test, and maintain optimal data pipeline and ETL architectures.
  • Map out data systems and define/design required integrations, ETL, BI, and AI systems/processes.
  • Prepare and optimize data for predictive and prescriptive modeling.
  • Collaborate with teams to integrate ERP data into the enterprise data lake, ensuring seamless flow and quality.
  • Enhance cloud data infrastructure on AWS or Azure for scalability and performance.
  • Utilize big data tools and frameworks to optimize data acquisition and preparation.
  • Build architectures to move data to/from data lakes and data warehouses for advanced analytics.
  • Develop and curate data models for analytics, dashboards, and reports.
  • Conduct code reviews, maintain production-level code, and implement testing approaches.
  • Monitor, troubleshoot, and resolve data ingestion workflows to maintain reliability and uptime.
  • Drive innovation and implement efficient new approaches to data engineering tasks.


Must-Have Skills:

  • Bachelor’s degree in Computer Science, Mathematics, Engineering, or a related field.
  • 5+ years of experience working with enterprise data platforms, including building and managing data lakes.
  • 3–5 years of experience designing and implementing data warehouse solutions.
  • Expertise in SQL, including developing stored procedures (SP) and applying advanced data design concepts.
  • Proficiency in Spark (Python/Scala) and Spark Streaming for real-time data pipelines.
  • Experience with AWS or Azure services (e.g., AWS Glue, Azure Data Factory, Redshift, Snowflake).
  • Familiarity with big data tools such as Apache Kafka, Apache Spark, or Flink.
  • Hands-on experience with orchestration tools (e.g., Apache Airflow, Prefect).
  • Knowledge of CI/CD processes, version control (e.g., Git, Jenkins), and deployment automation.
  • Strong problem-solving, communication, and collaboration skills.


Good-to-Have Skills:

  • Experience in integrating ERP data into data lakes.
  • Experience with traditional ETL tools (e.g., Talend, Pentaho).


Competencies:

  • Tech Savvy - Anticipating and adopting innovations in business-building digital and technology applications.
  • Self-Development - Actively seeking new ways to grow and be challenged using both formal and informal development channels.
  • Action Oriented - Taking on new opportunities and tough challenges with a sense of urgency, high energy, and enthusiasm.
  • Customer Focus - Building strong customer relationships and delivering customer-centric solutions.
  • Optimize Work Processes - Knowing the most effective and efficient processes to get things done, with a focus on continuous improvement.


Why Join Us?

  • Be part of a collaborative and agile team driving cutting-edge AI and data engineering solutions.
  • Work on impactful projects that make a difference across industries.
  • Opportunities for professional growth and continuous learning.
  • Competitive salary and benefits package.


Application Details

Ready to make an impact? Apply today and become part of the QX Impact team!


Read more
company logo
Bengaluru (Bangalore), Mumbai, Pune, Hyderabad, Noida, Kolkata
8 - 15 yrs
₹13L - ₹20L / yr
Azure Data Factory
Azure Databricks
PySpark
skill iconPython
SQL
+4 more

Roles & Responsibilities

  • Design, develop, and deliver scalable end-to-end data pipelines using Azure Data Factory, ensuring robust integration

of enterprise-wide data from diverse sources

• Build and optimize data engineering workflows using Databricks and PySpark

• Write efficient, high-performance SQL for data transformation and analysis

• Work with the Azure Cloud platform and associated services, applying strong understanding of data warehousing,

data models, and pipelines

• Provide technical leadership to a team of developers, including code reviews and enforcing best practices across the

development lifecycle

• Oversee CI/CD implementation using Azure DevOps, managing deployments across development, QA, and production

environments with proper change control processes

• Collaborate with cross-functional teams to translate business requirements into scalable data solutions

• Ensure data quality, reliability, and performance across all pipelines and platforms

Ideal Candidate

1Strong Azure Databricks Engineer / Senior Data Engineer Profile

2Mandatory (Experience 1) – Must have minimum 8+ years of overall experience in Data Engineering, Data Development, or related data technology roles, with strong hands-on experience in enterprise data pipeline development.

3Mandatory (Experience 2) – Must have strong hands-on experience with Azure Databricks, including development and optimization of scalable data engineering workflows using Databricks and PySpark.

4Mandatory (Experience 3) – Must have strong hands-on proficiency in PySpark/Python and SQL, with proven experience developing complex data transformations, processing workflows, and performance-optimized queries.

5Mandatory (Experience 4) – Must have hands-on experience with Azure Data Factory (ADF) for designing, developing, and orchestrating end-to-end data pipelines and integrating data from multiple sources.

6Mandatory (Experience 5) – Must have strong experience working on the Azure Cloud platform and associated data services, with solid understanding of data warehousing, data modeling, pipeline architecture, and enterprise data solutions.

7Mandatory (Experience 6) – Must have hands-on experience implementing CI/CD using Azure DevOps, including deployment and release management across development, QA, and production environments.

8Mandatory (Experience 7) – Must have proven technical leadership experience, including code reviews, enforcing development best practices, mentoring developers, and providing technical guidance to a data engineering team.

9Mandatory (Notice Period) – Immediate joiners or candidates who can join within 15 days.

10Mandatory (Note) - The position is open across all Cognizant offices pan India. Candidates must be willing to attend the F2F interview at the nearest Cognizant office location.

Read more
company logo
Netra Shettigar
Posted by Netra Shettigar
Pune, Bengaluru (Bangalore), Hyderabad, Chennai, Noida, Coimbatore, Navi Mumbai
5 - 15 yrs
₹8L - ₹16L / yr
Microsoft Windows Azure
Apache Synapse
synapse
SQL Azure
PL/SQL
+1 more

Job Description

We are looking for an experienced Azure Synapse Data Engineer with strong hands-on expertise in Azure data engineering, data warehousing, ETL/ELT, and SQL performance optimization.

Key Responsibilities

  • Design, develop, and optimize solutions using Azure Synapse Analytics.
  • Build and maintain data integration pipelines using Azure Data Factory (ADF) and Synapse Pipelines.
  • Design and implement scalable Azure Data Lake Storage Gen2 (ADLS Gen2) architectures.
  • Develop robust and high-performance ETL/ELT frameworks for enterprise data platforms.
  • Monitor and troubleshoot large-scale data ingestion, transformation, and processing workloads.
  • Identify and resolve performance bottlenecks across data movement and transformation processes.
  • Ensure data quality, availability, reliability, and consistency across analytics platforms.
  • Perform advanced SQL query analysis and optimization.
  • Analyze execution plans and optimize complex analytical workloads.
  • Configure and tune Synapse Dedicated SQL Pools.
  • Design efficient partitioning strategies for large datasets.
  • Optimize indexing, statistics, and workload distribution to improve query performance.

Mandatory Skills

  • Azure Synapse Analytics
  • Azure Data Factory (ADF)
  • Synapse Pipelines
  • ADLS Gen2
  • ETL/ELT
  • Strong SQL
  • Synapse Dedicated SQL Pool
  • SQL query/execution plan optimization
  • Partitioning, indexing, and statistics
  • Azure data warehouse / data lake architecture
  • Performance tuning and troubleshooting

Ideal Candidate

A strong Azure Data Engineer with hands-on experience in Synapse + ADF + ADLS Gen2 + Advanced SQL + Dedicated SQL Pool performance tuning.


Read more
company logo
Tanisha Gupta
Posted by Tanisha Gupta
Bengaluru (Bangalore)
6 - 9 yrs
₹20L - ₹26L / yr
Data-flow analysis
DevOps
Data integration
Data Structures
ETL
+4 more

Job Description

• Design and Implement Data Solutions: Lead the design, development, and implementation of scalable and secure Azure-

based data platforms, ensuring integration with various data sources and business systems. Deliver at least 2 major

projects every year with a focus on data engineering best practices.

• Optimize Data Pipelines: Build and optimize end-to-end data pipelines using Azure Data Factory, Azure Databricks, and

Azure Synapse, with an emphasis on automating data workflows. Achieve a 20% reduction in pipeline execution times

within the first 6 months.

• Cloud Infrastructure Management: Manage and maintain the Azure data environment, ensuring high availability, disaster

recovery, and cost optimization. Track and improve system uptime to exceed 99.9% reliability.

• Collaborate with Cross-Functional Teams: Partner with data scientists, data analysts, and business stakeholders to translate

business requirements into efficient data solutions. Facilitate at least 3 collaborative sessions per quarter to address key

business use cases.

• Ensure Data Security & Compliance: Implement data security measures, ensuring compliance with industry regulations

(GDPR, HIPAA, etc.) and company policies. Achieve and maintain full compliance in all data environments within the first

quarter of onboarding.

• Continuous Learning & Knowledge Sharing: Stay up-to-date with emerging Azure technologies, and mentor junior

engineers to promote knowledge sharing. Complete 1 Azure certification annually and conduct at least 2 internal

knowledge-sharing sessions per year.

• Sound knowledge of data governance practices, data quality management, and data security principles.

• Play a pivotal role in shaping our organization's data-driven journey, driving innovation through data analytics and insights.

• Optimize data storage, processing and retrieval mechanisms for performance, cost, and scalability using data storage

services (such as Azure Data Lake Storage, Azure SQL Database, etc.), data processing services (such as Azure Data Bricks,

Azure Synapse, etc.) and data visualization (PowerBI, Qlik, etc.) & integration services (Data API builder, logic apps, etc.)

• Monitor and troubleshoot data platform performance, identify and resolve issues, and provide recommendations for

continuous improvement.

• Collaborate with DevOps teams to automate deployment, configuration, and monitoring processes using Azure DevOps,

PowerShell, or other relevant tools.

• Stay up to date with the latest trends and advancements in cloud data services and provide recommendations on adopting

new technologies or features to enhance the data platform.

• Document technical designs, procedures, and guidelines for data platform engineering and operations

Knowledge, Skills & Experience

Job Experience • Bachelor's degree in Computer Science, Engineering, or a related field. Advanced

degree preferred.

• Proven 6-10 years experience in playing platform engineer or admin role

• Experience with big data technologies such as Apache Spark, Hadoop, or similar

frameworks.

• Solid understanding of cloud computing concepts and experience with cloud

infrastructure management and provisioning.

• Solid understanding of network security concepts and technologies (such as

firewalls, VPNs, intrusion detection/prevention systems, etc.) and data security

concepts and technologies (such as access controls, encryption, observability,

privacy laws/regulations, etc.)

• Experience in a Retail setup is preferred.

Required Skills The position will require someone with the following:

• Strategic Planning

Public

• Communication and Collaboration

• Problem Solving Skills A/B testing & experimentation

• SQL, BI tools, and storytelling with data

Read more
company logo
Mayank Choudhary
Posted by Mayank Choudhary
Chennai
10 - 15 yrs
₹27L - ₹32L / yr
Data engineering
databricks
skill iconPython
SQL
Apache Spark

Strong Databricks Architect Profile with end-to-end Lakehouse ownership

2

Mandatory (Experience 1) – Must have 10+ years of software engineering experience with atleast 5+ years in Data Engineering with hands on exposure to Databricks and strong ownership of end-to-end data pipeline development.

3

Mandatory (Experience 2) – Must have atleast 5+ years of expertise across the Databricks ecosystem — Delta Lake, Delta Live Tables, Autoloader, Structured Streaming, Workflows, Unity Catalog

4

Mandatory (Tech skill 1) – Must have worked at architecture level, owning end-to-end design through deployment

5

Mandatory (Tech skill 2) – Must have strong experience with Python and SQL for data processing and Apache Spark for performance tuning & scalability

6

Mandatory (Tech skill 3) – Must have experience in large-scale data warehousing & advanced data modeling (3NF and dimensional) across batch and real-time systems

7

Mandatory (AI Exposure) – Must have at least a basic working understanding of how AI services or tools work

8

Mandatory (Communication) – Must have strong stakeholder management & requirement-gathering experience with US or UK clients

9

Mandatory (Company) – Must come from a B2B IT services or IT consulting background

10

Mandatory (Note) – CTC is inclusive of 5% variable

11

Preferred (Tech skill 1) – Azure Databricks or Azure data services experience (project runs on Azure DevOps)

12

Preferred (Tech skill 2) – MLflow or MLOps practices and AI use cases (RAG, AI/BI)

13

Preferred (Tech skill 3) – CI/CD, Databricks Asset Bundles (DABs) or equivalent packaging, Terraform or IaC, reusable deployment templates

14

Preferred (Integrations) – ServiceNow or enterprise system integrations

15

Preferred (Certifications) – Databricks (Data Engineer Associate or Professional, ML or GenAI tracks), Azure or AWS cloud certifications

Read more
company logo
Sri Priyanka
Posted by Sri Priyanka
Remote only
8 - 17 yrs
Best in industry
Data engineering
Medallion
lakehouse
ETL
skill iconAmazon Web Services (AWS)
+4 more

Data Engineer

Data Lakehouse & Platform Engineering 


About the Role

We are hiring Data Engineer to own the lifecycle of our enterprise Data Lakehouse platform. We are looking for engineers who think in systems, make platform-level design decisions, and can build and operate a production-grade, multi-source lakehouse from the ground up, covering ingestion through consumption across a complex, multi-cloud source landscape.

You will be the technical authority for a platform that consolidates data from 18+ enterprise products (Costpoint, GovWin, Specpoint, Vantagepoint, and others) into a governed, medallion-architected data lake on AWS S3 with Apache Iceberg table format, orchestrated via AWS Step Functions, and queryable through AWS Athena and Trino. This role is end-to-end: you own ingestion, transformation, quality, orchestration, ML data supply, and BI consumption.


Key Responsibilities

•      Architect and evolve the full medallion lakehouse — Bronze, Silver, and Gold layers — on AWS S3 with Apache Iceberg; own schema design, partitioning, compaction, and retention policies.

•      Design and implement scalable Glue ETL (PySpark) pipelines for bronze_to_silver and silver_to_gold transformations, incorporating dbt for SQL-layer transformations where appropriate.

•      Own and extend CDC ingestion via Fivetran; manage schema evolution, connector health, and sync reliability across 18+ source products.

•      Build and maintain AWS Step Functions state machines and EventBridge schedules for end-to-end pipeline orchestration; implement Lambda-based quality and drift monitors.

•      Govern the Glue Catalog and Lake Formation policies; enforce column-level security, row-level access controls, and audit logging to meet SOC2 and regulatory requirements.

•      Architect the query layer — optimize Athena workgroups and partition pruning; plan and execute Trino-on-EKS deployment for sub-second analytics workloads.

•      Partner with data science teams on SageMaker data supply: feature engineering pipelines, training dataset preparation, and model registry integration.

•      Implement real-time and near-real-time streaming solutions using Kafka or Kinesis where sub-13-minute latency is required.

•      Lead platform modernization initiatives: evaluate emerging formats (Iceberg vs. Delta Lake vs. Hudi), tooling, and cost optimization strategies.

•      Establish and enforce data engineering best practices: code reviews, CI/CD for pipeline code, IaC (Terraform / CloudFormation), and incident response runbooks.

•      Mentor and level up junior and mid-level data engineers; define team standards for pipeline design, testing, and documentation.


 

Required Qualifications

•      Software or data engineering experience, with at least 4 years in an architect or technical lead capacity designing large-scale cloud data platforms.

•      Deep, hands-on expertise with AWS data services: S3, Glue (PySpark ETL), Athena, Step Functions, Lambda, EventBridge, Lake Formation, SageMaker, and CloudWatch.

•      Production experience with Apache Iceberg (or Delta Lake / Hudi) table formats — compaction, snapshot management, schema evolution, and time travel.

•      Strong PySpark and Python skills; ability to write, review, and optimize distributed data processing jobs at scale.

•      Hands-on experience with CDC-based ingestion platforms (Fivetran, Debezium, or equivalent) across heterogeneous source systems.

•      Proven experience designing and implementing medallion (Bronze/Silver/Gold) or equivalent multi-hop lakehouse architectures.

•      Experience with data pipeline orchestration: AWS Step Functions, Apache Airflow, or equivalent; event-driven pipeline design patterns.

•      Strong SQL skills; experience with Athena, Trino, Presto, or equivalent query engines for large-scale analytical workloads.

•      Familiarity with data governance tooling: catalog management (Glue Catalog, Apache Polaris/Iceberg REST), data lineage, access controls, and audit frameworks.

•      Experience with Infrastructure as Code (Terraform or CloudFormation) for data platform provisioning and drift management.

•      Solid understanding of dimensional modeling, schema design (star/snowflake), and data normalization for BI and analytics workloads.

•      Bachelor's degree in Computer Science, Engineering, or a related field; or equivalent professional experience.


Preferred Qualifications

•      Experience operating Trino or PrestoDB on Kubernetes (EKS); tuning for sub-second query latency and multi-tenant workloads.

•      Familiarity with streaming platforms (Kafka, Kinesis, or Pub/Sub) and real-time lakehouse patterns.

•      Experience with Apache Polaris or other Iceberg REST catalog implementations.

•      Exposure to SageMaker MLOps pipelines, Model Registry, and feature store patterns for ML data supply.

•      Experience with dbt (data build tool) for SQL-layer transformation and documentation in lakehouse environments.

•      Government contracting or ERP domain knowledge (Costpoint, Deltek, Oracle, or similar enterprise platforms) is a strong plus.

•      AWS certifications: Data Engineer Associate, Solutions Architect Professional, or equivalent.


What You Will Build

You will be a founding architect of a strategic, cross-product data platform that serves 18+ enterprise applications and their analytics, ML, and AI workloads. The decisions you make on schema, storage format, query layer, governance, and orchestration will shape the data foundation of the company for years. This is a high-impact, high-ownership role with direct visibility to senior leadership.

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos