Cutshort logo
For Employers

50+ PySpark Jobs in India

Apply to 50+ PySpark Jobs on CutShort.io. Find your next job, effortlessly. Browse PySpark Jobs and apply today!

icon
Straatix Partners

at Straatix Partners

2 candid answers
Harshita Jamwal
Posted by Harshita Jamwal
Bengaluru (Bangalore), Pune
8 - 15 yrs
₹25L - ₹35L / yr
PySpark
SQL

CAP-190 : Senior Data Analyst

📍Location : Bangalore / Hyderabad / Chennai / Pune / NCR

🧠Experience : 7 - 15 Years

🆔Job Code : CAP-190

🏢Work Type : Hybrid - 3 Days in a Week


About the Client (CODE: CAP)

CAP operates at the forefront of the financial services industry, providing consulting, technology, and digital transformation solutions to leading organizations worldwide. With a focus on innovation, CAP empowers clients to navigate complex regulatory landscapes, optimize operations, and drive business growth. The company fosters a collaborative and agile culture, encouraging continuous learning and excellence.


Key Responsibilities

● Define and obtain source data required to deliver insights and use cases.

● Determine data mapping and join multiple data sets across various sources.

● Develop methods to highlight and report data inconsistencies for user review.

● Propose and assist with suitable data migration sets for stakeholders.

● Support teams in processing data migration sets and coordinating migration activities.

● Plan, track, and coordinate the data migration team and migration run-book.

● Collaborate with stakeholders to avoid negative customer and business impacts.

● Ensure robust communication and escalation mechanisms across project portfolios.

● Implement strategic solutions and avoid short-term workarounds.

● Maintain strong control and compliance standards in data handling.


Required Skills

● Minimum 7+ years of experience as a Data Analyst, preferably in financial services.

● Strong expertise in Pyspark, Python, and SQL.

● Experience with big data programs and data models in banking or financial markets.

● Ability to write SQL queries and navigate databases such as Hive, CMD, Putty, and Note++.

● Excellent analytical skills and commercial acumen.

● Strong verbal and written communication skills.

● Proven ability to manage multiple priorities and deliver within tight deadlines.

● Business analysis skills, including defining and understanding requirements.

● Familiarity with SDLC, Agile processes, and a bias towards TDD.

● Attention to detail and a proactive, problem-solving mindset.


Nice to Have

● Knowledge and experience in Data Quality & Governance.

● Working experience with Spark Scala or Java for Spark.

● Proven track record of managing small, delivery-focused data teams (for senior roles).

● Experience with market data vendors and domains such as Party/Client, Trade, Settlements, Payments, Instrument and Pricing, Market and/or Credit Risk.


Why Join CAP (Code Name)

Join CAP to work on impactful data initiatives within the financial services sector, tackling complex technical challenges and driving meaningful business outcomes. You'll collaborate with talented professionals in a dynamic, agile environment that values innovation and continuous improvement. CAP offers opportunities for professional growth, skill development, and the chance to contribute to high-visibility projects that shape the future of financial technology.


About the Employment Model

Direct Hire (Client Payroll) : For this role, you’ll be hired directly by the client and be part of their internal team. Straatix supports the hiring process, but your employment, payroll, and benefits are all managed by the client.

Read more
NeoGenCode Technologies Pvt Ltd
Bengaluru (Bangalore)
14 - 25 yrs
₹50L - ₹70L / yr
Data engineering
databricks
Apache Spark
PySpark
skill iconPython
+19 more

Job Title : Senior Data Engineer – Databricks

Experience : 14 to 20 Years

Location : HSR Layout, Bangalore

Work Mode : Hybrid – 3 Days WFO

Shift : 11:30 AM – 07:30 PM IST

Positions : 2

Notice Period : Immediate Joiners Only

Interview : 1 Technical Round + 2 Client Rounds


Role Overview :

We are looking for a Senior Data Engineer to build and lead enterprise-scale data platforms for a Switzerland-based commodity client.

The role requires a strong hands-on Data Engineering professional with expertise in Databricks, PySpark, Python, SQL, and AWS, along with technical leadership and stakeholder management experience.


Must-Have Skills :

  • 14 to 20 years of Data Engineering experience
  • Databricks & Apache Spark / PySpark
  • Python & SQL
  • AWS Cloud
  • Lakehouse Architecture
  • ETL / ELT & Distributed Data Processing
  • Batch & Streaming Pipelines
  • Data Pipeline Optimization & Data Modeling
  • CDC & Incremental Processing
  • Git, CI/CD & Testing
  • Data Quality, Monitoring & Observability
  • Technical Leadership & Stakeholder Management


Key Responsibilities :

  • Design and build scalable data pipelines using Databricks, PySpark, Python, SQL, and AWS.
  • Own data products from design through production.
  • Develop batch / streaming pipelines and reusable ETL / ELT frameworks.
  • Optimize pipelines for performance, scalability, reliability, and cost.
  • Design scalable data architectures and data models.
  • Implement data quality, monitoring, lineage, and CI/CD practices.
  • Lead technical discussions and mentor engineering teams.
  • Collaborate with business stakeholders, architects, product owners, and engineering teams.
  • Remain hands-on while providing technical leadership.


Ideal Candidate :

A 14 to 20 years experienced, hands-on Data Engineering leader with strong Databricks + PySpark + AWS expertise, excellent communication, stakeholder management, and experience delivering enterprise-scale data platforms.

🔴 Super Urgent : Only Bangalore-based immediate joiners.

Read more
Quantiphi

at Quantiphi

3 candid answers
1 video
Nikita Sinha
Posted by Nikita Sinha
Mumbai, Bengaluru (Bangalore)
4 - 7 yrs
Best in industry
Data Transformation Tool (DBT)
skill iconAmazon Web Services (AWS)
PySpark
iceberg
SQL

As a Senior Data Engineer, you will be responsible for designing, developing, and optimizing scalable cloud-native data platforms on AWS. You will build high-performance batch and analytical data pipelines using DBT (Cloud + Core), AWS Glue, SMUS, Redshift, Python, MWAA, APIs, and other AWS services while ensuring reliability, scalability, and operational excellence.

The role requires strong expertise in data engineering, SQL optimization, distributed data processing, and cloud-native architectures.


Must Have Primary Skills

Data Engineering

  • 3+ years of hands-on experience building large-scale AWS Data Engineering solutions.
  • Strong experience designing scalable batch and streaming data pipelines on AWS.
  • Experience implementing monitoring, logging, alerting, and observability for production data pipelines.
  • Strong understanding of data quality, troubleshooting, debugging, and production support.

DBT & Data Warehousing

  • Strong hands-on experience with DBT (Cloud + Core) (Highly Recommended).
  • Hands-on expertise with Amazon Redshift for Data Warehousing and Analytics.

Programming & Querying

  • Strong expertise in:
  • SQL (Analytical Queries, Window Functions, Stored Procedures)
  • Python
  • Spark / PySpark
  • Strong experience in Spark/PySpark performance tuning and optimization.

Lakehouse & Apache Iceberg

  • Hands-on experience designing and implementing Lakehouse solutions using Apache Iceberg.
  • Experience optimizing Apache Iceberg and Amazon Aurora PostgreSQL for performance and scalability.
  • Strong understanding of:
  • Data Modeling
  • Partitioning
  • File Formats
  • Lakehouse / Data Lake Architectures

AWS Cloud Services

Hands-on experience with AWS services including:

  • AWS Glue
  • Amazon MWAA
  • Amazon EMR
  • Amazon Redshift
  • Amazon S3
  • Amazon Athena
  • AWS Glue Catalog
  • Amazon Aurora PostgreSQL
  • AWS Lambda
  • Amazon EC2
  • Amazon DynamoDB
  • Amazon API Gateway
  • Amazon CloudWatch
  • Amazon SQS
  • Amazon SNS
  • Amazon EventBridge
  • AWS Secrets Manager
  • AWS IAM

Development Practices

  • Experience with CI/CD.
  • Strong knowledge of Git.
  • Familiarity with Data Engineering best practices.

Soft Skills

  • Strong communication skills.
  • Excellent problem-solving ability.
  • Stakeholder management experience.

Preferred Certification

  • AWS Data Engineer Associate or equivalent AWS Certification.

Good to Have Skills

  • AWS Step Functions
  • Terraform / Infrastructure as Code (IaC)
  • Kafka
  • Hive
  • HDFS
  • Other Big Data technologies
  • Experience using GenAI-assisted development tools such as:
  • GitHub Copilot
  • Cursor
  • Kiro
  • Similar AI coding assistants
  • ClickHouse Database
  • Telecom / Mobile Network domain knowledge
  • AtScale design and development
  • Docker
  • Kubernetes 


Read more
Straatix Partners

at Straatix Partners

2 candid answers
Arushi Jamwal
Posted by Arushi Jamwal
Bengaluru (Bangalore), Pune, Hyderabad, Mumbai
7 - 14 yrs
₹25L - ₹40L / yr
skill iconPython
SQL
PySpark

CAP-190 : Senior Data Analyst

📍Location : Bangalore / Hyderabad / Chennai / Pune / NCR

🧠Experience : 7 - 15 Years

🆔Job Code : CAP-190

🏢Work Type : Hybrid - 3 Days in a Week

📨Application Link : https://drive.google.com/drive/folders/1bApWyMvoFHvI6krgw9CPbGZHaqYU0kSZ?usp=sharing


Your Experience at a Glance

We’re hiring a Senior Data Analyst for our client (Code Name:CAP) - CAP is a global consulting firm specializing in digital transformation and technology solutions for the financial services industry. As a Data Analyst / Senior Data Analyst at CAP, you will play a pivotal role in delivering actionable insights and supporting data migration initiatives within the financial services domain. You will collaborate with cross-functional teams to define data requirements, ensure data quality, and drive strategic solutions. This senior-level position requires strong technical expertise, stakeholder management, and a proactive approach to delivering change in a dynamic environment.


About the Client (CODE: CAP )

CAP operates at the forefront of the financial services industry, providing consulting, technology, and digital transformation solutions to leading organizations worldwide. With a focus on innovation, CAP empowers clients to navigate complex regulatory landscapes, optimize operations, and drive business growth. The company fosters a collaborative and agile culture, encouraging continuous learning and excellence.


Key Responsibilitie

● Define and obtain source data required to deliver insights and use cases.

● Determine data mapping and join multiple data sets across various sources.

● Develop methods to highlight and report data inconsistencies for user review.

● Propose and assist with suitable data migration sets for stakeholders.

● Support teams in processing data migration sets and coordinating migration activities.

● Plan, track, and coordinate the data migration team and migration run-book.

● Collaborate with stakeholders to avoid negative customer and business impacts.

● Ensure robust communication and escalation mechanisms across project portfolios.

● Implement strategic solutions and avoid short-term workarounds.

● Maintain strong control and compliance standards in data handling.


Required Skills

● Minimum 7+ years of experience as a Data Analyst, preferably in financial services.

● Strong expertise in Pyspark, Python, and SQL.

● Experience with big data programs and data models in banking or financial markets.

● Ability to write SQL queries and navigate databases such as Hive, CMD, Putty, and Note++.

● Excellent analytical skills and commercial acumen.

● Strong verbal and written communication skills.

● Proven ability to manage multiple priorities and deliver within tight deadlines.

● Business analysis skills, including defining and understanding requirements.

● Familiarity with SDLC, Agile processes, and a bias towards TDD.

● Attention to detail and a proactive, problem-solving mindset.


Nice to Have

● Knowledge and experience in Data Quality & Governance.

● Working experience with Spark Scala or Java for Spark.

● Proven track record of managing small, delivery-focused data teams (for senior roles).

● Experience with market data vendors and domains such as Party/Client, Trade, Settlements, Payments, Instrument and Pricing, Market and/or Credit Risk.


Why Join CAP (Code Name)

Join CAP to work on impactful data initiatives within the financial services sector, tackling complex technical challenges and driving meaningful business outcomes. You'll collaborate with talented professionals in a dynamic, agile environment that values

innovation and continuous improvement. CAP offers opportunities for professional growth, skill development, and the chance to contribute to high-visibility projects that shape the future of financial technology.


About the Employment Model

Direct Hire (Client Payroll) : For this role, you’ll be hired directly by the client and be part of their internal team. Straatix supports the hiring process, but your employment, payroll, and benefits are all managed by the client.

Apply here-

https://drive.google.com/drive/folders/1bApWyMvoFHvI6krgw9CPbGZHaqYU0kSZ?usp=sharing

Read more
ProofofSkill
Hyderabad
5 - 7 yrs
₹15L - ₹20L / yr
skill iconPython
SQL
PySpark
Data Warehouse (DWH)
Amazon Redshift
+1 more

Location – Hyderabad (Hybrid)

Work Experience – 5 to 7 years

CTC – upto 20 LPA


Roles & Responsibilities:

· We are looking for a Senior Data Engineering who will be majorly responsible for designing, building and maintaining ETL/ ELT pipelines.

· Integration of data from multiple sources or vendors to provide the holistic insights from data.

· You are expected to build and manage Data warehouse solutions, designing data models, creating ETL processes, implementing data quality mechanisms etc.

· Performs EDA (exploratory data analysis) required to troubleshoot data related issues and assist in the resolution of data issues.

· Should have experience in client interaction.

· Experience in mentoring juniors and providing required guidance.

Required Technical Skills

 

· Extensive hands on experience in Python, Pyspark, SQL, Dataiku.

· Strong experience in Data Warehouse, ETL, Data Modelling, building ETL Pipelines, Snowflake database.

· Working knowledge in Databricks, Redshift, ADF etc.

· Hands-on experience in cloud services like Azure, AWS- S3, Glue, Lambda, CloudWatch, Athena.

· Sound knowledge in end-to-end Data management, Data ops, quality and data governance.

· Familiar with SFDC, Waterfall/ Agile methodology.

· Strong domain knowledge in Pharma domain/ life sciences commercial data operations.

 

Qualifications

 

· Bachelor’s or master’s Engineering/ MCA or equivalent degree.

· 5-7 years of relevant industry experience as Data Engineer.

· Experience working on Pharma syndicated data such as IQVIA, Veeva, Symphony; Claims, CRM, Sales etc.

· High motivation, good work ethic, maturity, self-organized and personal initiative.

· Ability to work collaboratively and providing the support to the team.

· Excellent written and verbal communication skills.

· Strong analytical and problem-solving skills. 

Read more
Wissen Technology

at Wissen Technology

4 recruiters
Robin Silverster
Posted by Robin Silverster
Bengaluru (Bangalore)
7 - 10 yrs
₹15L - ₹40L / yr
Data engineering
skill iconPython
PySpark
Data Transformation Tool (DBT)
Apache Airflow
+2 more

About the Role

We are looking for a Senior Data Engineer with strong hands-on expertise in Databricks, Python, PySpark, and SQL to build scalable, high-performance data engineering solutions. You’ll architect and develop large scale, high-performance data pipelines capable of handling massive real-time and batch data volumes across multiple business systems. Databricks is the core enterprise data and processing platform for this role. You will also use Apache Airflow for workflow orchestration and dbt for ELT transformations, and will contribute to designing reliable, secure, and governed data platforms that enable analytics, reporting, and AI-driven use cases.

Key Responsibilities

  • Design and implement large-scale data pipelines using Python/PySpark, Databricks, and Microsoft Fabric.
  • Develop and optimize data processing workloads in Databricks using PySpark and Spark SQL, with a strong focus on scalability, reliability, performance, and maintainability.
  • Develop and maintain dbt models including layered architecture, incremental models, snapshots, macros, testing, and documentation.
  • Design, develop, and maintain Apache Airflow DAGs for orchestrating reliable, scalable, and observable data pipelines.
  • Design and implement data quality, observability, and governance frameworks, including automated testing, monitoring, lineage, access control, and data privacy standards.
  • Partner with analytics, product, and business stakeholders to turn requirements into trustworthy datasets, and raise the engineering bar through design discussions, code reviews, and mentoring junior engineers.

Required Skills

  • Strong expertise in Python for developing scalable, modular, and production-ready data engineering applications.
  • Strong expertise in PySpark, including DataFrame API, Spark SQL, Structured Streaming, partitioning strategies, joins, caching, handling data skew, and Spark performance optimization.
  • Strong hands-on experience with Databricks for data ingestion, transformation, processing, and optimization, including Delta Lake, Unity Catalog, Databricks Workflows, notebooks, jobs, and Databricks-native data engineering capabilities.
  • Strong experience in Databricks/Spark performance tuning, including query and job optimization, partitioning, file sizing, caching, join optimization, handling data skew, and efficient use of compute resources.
  • Hands-on experience with Delta Lake, including transactional data processing, schema management, incremental data processing, and reliable batch and streaming data pipelines.
  • Hands-on experience in developing dbt projects using layered architecture, incremental models, snapshots, macros/Jinja, testing, documentation, and deployment best practices.
  • Expertise in advanced SQL and data modelling — dimensional modeling, slowly changing dimensions, schema evolution, and query optimization.
  • Hands-on experience in developing and managing Apache Airflow DAGs, scheduling workflows, dependency management, retries, backfills, and operational monitoring.
  • Hands-on experience with at least one major cloud platform (AWS, Azure or GCP).
  • Strong problem-solving skills and the ability to work independently with business and analytics stakeholders.

Nice to Have

  • Hands-on exposure to Microsoft Fabric for data integration and analytics.
  • Experience using AI coding assistants (e.g. Claude Code, GitHub Copilot) as part of a development workflow.
  • Familiarity with modern DevOps practices, including CI/CD pipelines, Infrastructure as Code (IaC), and containerization (Docker/Kubernetes).
  • Domain expertise in financial services.


Read more
Straatix Partners

at Straatix Partners

2 candid answers
Ashwini Ranganath
Posted by Ashwini Ranganath
Bengaluru (Bangalore), Delhi, Gurugram, Noida, Ghaziabad, Faridabad
7 - 15 yrs
₹20L - ₹40L / yr
PySpark

Define and obtain source data required to deliver insights and use cases.

● Determine data mapping and join multiple data sets across various sources.

● Develop methods to highlight and report data inconsistencies for user review.

● Propose and assist with suitable data migration sets for stakeholders.

● Support teams in processing data migration sets and coordinating migration activities.

● Plan, track, and coordinate the data migration team and migration run-book.

● Collaborate with stakeholders to avoid negative customer and business impacts.

● Ensure robust communication and escalation mechanisms across project portfolios.

● Implement strategic solutions and avoid short-term workarounds.

● Maintain strong control and compliance standards in data handling.

Required Skills

● Minimum 7+ years of experience as a Data Analyst, preferably in financial services.

● Strong expertise in Pyspark, Python, and SQL.

● Experience with big data programs and data models in banking or financial markets.

● Ability to write SQL queries and navigate databases such as Hive, CMD, Putty, and Note++.

● Excellent analytical skills and commercial acumen.

● Strong verbal and written communication skills.

● Proven ability to manage multiple priorities and deliver within tight deadlines.

● Business analysis skills, including defining and understanding requirements.

● Familiarity with SDLC, Agile processes, and a bias towards TDD.

● Attention to detail and a proactive, problem-solving mindset.

Nice to Have

● Knowledge and experience in Data Quality & Governance.

● Working experience with Spark Scala or Java for Spark.

● Proven track record of managing small, delivery-focused data teams (for senior roles).

● Experience with market data vendors and domains such as Party/Client, Trade, Settlements, Payments, Instrument and Pricing, Market and/or Credit Risk.

Read more
Bengaluru (Bangalore)
9 - 18 yrs
₹4L - ₹35L / yr
Azure Databricks
Azure data factory
skill iconPython
PySpark
Apache Spark
+1 more

Role Overview

We are looking for experienced Azure Databricks Data Engineers with strong hands-on expertise in Apache Spark/PySpark, Python, SQL, and Azure Data Factory (ADF). The candidate will be responsible for designing, developing, and optimizing scalable data engineering solutions on the Azure cloud platform.

Mandatory Skills

  • Strong hands-on experience with Azure Databricks
  • Strong knowledge of Apache Spark and/or PySpark
  • Proficiency in Python
  • Strong SQL development and query optimization skills
  • Hands-on experience with Azure Data Factory (ADF)
  • Experience developing and maintaining scalable ETL/ELT data pipelines
  • Good understanding of Azure data engineering concepts and cloud-based data platforms

Key Responsibilities

  • Design, develop, and maintain data pipelines using Azure Databricks and ADF
  • Develop efficient data processing solutions using PySpark/Spark and Python
  • Write complex SQL queries for data transformation, validation, and analysis
  • Build scalable ETL/ELT workflows for large datasets
  • Optimize Spark jobs, Databricks notebooks, and data pipelines for performance
  • Integrate data from multiple sources into Azure-based data platforms
  • Implement data quality, validation, error handling, and monitoring mechanisms
  • Troubleshoot production issues and provide timely resolutions
  • Collaborate with data architects, analysts, developers, and business stakeholders
  • Follow best practices for code quality, security, performance, and maintainability

Experience Requirements

For 5+ Years

  • 5+ years of overall experience in data engineering
  • Strong hands-on experience in Azure Databricks, PySpark/Spark, Python, SQL, and ADF
  • Experience working on enterprise-scale data pipelines

For 9+ Years

  • 9+ years of overall experience in data engineering
  • Strong expertise in Azure Databricks and modern Azure data engineering
  • Proven experience designing and optimizing large-scale data pipelines
  • Ability to lead technical discussions and mentor junior engineers

Preferred Skills

  • Azure Data Lake Storage (ADLS)
  • Delta Lake
  • Databricks Workflows
  • Azure DevOps / CI-CD
  • Git
  • Data warehousing concepts
  • Experience with Agile/Scrum methodologies
Read more
IKLAVYA INDIA TECH PRIVATE LIMITED
Bengaluru (Bangalore)
8 - 18 yrs
₹5L - ₹18L / yr
ELT
SQL
PySpark
skill iconAmazon Web Services (AWS)
NOSQL Databases

Design, develop, and maintain ETL pipelines involving large-scale data.

Develop data processing and analytics applications primarily using PySpark and Python.

Build scalable and distributed data processing solutions using Apache Spark.

Develop and deploy data applications on AWS cloud.

Work with AWS services related to storage, compute, ETL, data warehousing, analytics, and streaming.

Implement distributed storage and processing solutions capable of handling high-volume datasets.

Design data processing applications with a focus on performance, scalability, reliability, and optimization.

Work with both SQL and NoSQL databases for data storage, processing, and analytics.

Write, optimize, and analyze SQL, HQL, and NoSQL queries.

Troubleshoot data pipeline and processing issues and ensure data quality and reliability.

Collaborate with data engineers, analysts, architects, and other technical teams to deliver data-driven solutions.

Read more
Wissen Technology

at Wissen Technology

4 recruiters
Anisha Jindal
Posted by Anisha Jindal
Bengaluru (Bangalore), Mumbai
5 - 14 yrs
Best in industry
Data engineering
skill iconPython
PySpark
DAX
PowerBI

Job Summary

We are looking for a skilled and experienced Data Engineer to join our growing data team. The ideal candidate will have strong expertise in Python, PySpark, Data Modeling, and Power BI, with hands-on experience in designing, developing, and optimizing scalable data solutions. The role requires working closely with business stakeholders, data architects, and analytics teams to build robust data pipelines and semantic models that enable data-driven decision-making.


Technical Skills

  • Strong hands-on experience in Python and PySpark development.
  • Expertise in building and optimizing Data Engineering solutions and ETL pipelines.
  • Strong understanding of Data Modeling concepts (Star Schema, Snowflake Schema, Dimensional Modeling).
  • Experience with Power BI Data Modeling and Semantic Layer development.
  • Proficiency in DAX (Data Analysis Expressions).
  • Experience designing and managing Semantic Models in Power BI.
  • Strong SQL skills and experience working with large datasets.
  • Knowledge of data warehousing concepts and best practices.


Preferred Skills

  • Experience with cloud platforms such as Azure, AWS, or GCP.
  • Exposure to modern data platforms like Databricks.
  • Understanding of data governance and data quality frameworks.
Read more
TalentXO
Bengaluru (Bangalore), Mumbai, Pune, Hyderabad, Noida, Kolkata
8 - 15 yrs
₹13L - ₹20L / yr
Azure Data Factory
Azure Databricks
PySpark
skill iconPython
SQL
+4 more

Roles & Responsibilities

  • Design, develop, and deliver scalable end-to-end data pipelines using Azure Data Factory, ensuring robust integration

of enterprise-wide data from diverse sources

• Build and optimize data engineering workflows using Databricks and PySpark

• Write efficient, high-performance SQL for data transformation and analysis

• Work with the Azure Cloud platform and associated services, applying strong understanding of data warehousing,

data models, and pipelines

• Provide technical leadership to a team of developers, including code reviews and enforcing best practices across the

development lifecycle

• Oversee CI/CD implementation using Azure DevOps, managing deployments across development, QA, and production

environments with proper change control processes

• Collaborate with cross-functional teams to translate business requirements into scalable data solutions

• Ensure data quality, reliability, and performance across all pipelines and platforms

Ideal Candidate

1Strong Azure Databricks Engineer / Senior Data Engineer Profile

2Mandatory (Experience 1) – Must have minimum 8+ years of overall experience in Data Engineering, Data Development, or related data technology roles, with strong hands-on experience in enterprise data pipeline development.

3Mandatory (Experience 2) – Must have strong hands-on experience with Azure Databricks, including development and optimization of scalable data engineering workflows using Databricks and PySpark.

4Mandatory (Experience 3) – Must have strong hands-on proficiency in PySpark/Python and SQL, with proven experience developing complex data transformations, processing workflows, and performance-optimized queries.

5Mandatory (Experience 4) – Must have hands-on experience with Azure Data Factory (ADF) for designing, developing, and orchestrating end-to-end data pipelines and integrating data from multiple sources.

6Mandatory (Experience 5) – Must have strong experience working on the Azure Cloud platform and associated data services, with solid understanding of data warehousing, data modeling, pipeline architecture, and enterprise data solutions.

7Mandatory (Experience 6) – Must have hands-on experience implementing CI/CD using Azure DevOps, including deployment and release management across development, QA, and production environments.

8Mandatory (Experience 7) – Must have proven technical leadership experience, including code reviews, enforcing development best practices, mentoring developers, and providing technical guidance to a data engineering team.

9Mandatory (Notice Period) – Immediate joiners or candidates who can join within 15 days.

10Mandatory (Note) - The position is open across all Cognizant offices pan India. Candidates must be willing to attend the F2F interview at the nearest Cognizant office location.

Read more
Codnatives
Agency job
via VY SYSTEMS PRIVATE LIMITED by Ajeethkumar s
Hyderabad, Bengaluru (Bangalore)
5 - 10 yrs
₹4L - ₹16L / yr
skill iconPython
ETL
PySpark
Data engineering
skill iconAmazon Web Services (AWS)
+2 more

Skills Referential (Required knowledge, skills and abilities)

Technical Skills:

Python

Pyspark

SQL

ETL Aws, Azure, gcp

Read more
Bengaluru (Bangalore)
1 - 3 yrs
₹3.5L - ₹4.5L / yr
skill iconPython
skill iconKubernetes
skill iconDocker
TensorFlow
PySpark
+22 more

Key Responsibilities  

  • Design, build, and optimize scalable data pipelines for AI/ML applications.
  • Develop, train, evaluate, and deploy Machine Learning and Deep Learning models.
  • Build production-ready LLM applications using Retrieval-Augmented Generation (RAG), prompt engineering, and vector databases.
  • Fine-tune open-source and foundation models using domain-specific datasets.
  • Develop and maintain end-to-end MLOps pipelines for model deployment, monitoring, and lifecycle management.
  • Perform data preprocessing, feature engineering, exploratory data analysis (EDA), and model evaluation.
  • Develop APIs and AI services for production deployment.
  • Collaborate with cross-functional teams to deliver scalable AI-driven solutions.
  • Monitor model performance, troubleshoot production issues, and maintain technical documentation.


Required Skills  

Mandatory  

  • 1–3 years of experience in Data Science, Data Engineering, or AI/ML development.
  • Strong programming skills in Python and SQL.
  • Hands-on experience with Machine Learning frameworks such as PyTorch, TensorFlow, or Scikit-learn.
  • Experience building LLM-powered applications using RAG, Prompt Engineering, and Embeddings.
  • Hands-on experience with LangChain, LlamaIndex, CrewAI, or n8n for LLM orchestration and AI workflow automation.
  • Experience in LLM fine-tuning and working with Hugging Face models.
  • Knowledge of MLOps concepts including model deployment, monitoring, versioning, and CI/CD.
  • Experience with Git, REST APIs, Linux environments, and data processing libraries.


Preferred  

  • Experience with vector databases such as Pinecone, Chroma, Milvus, or Weaviate.
  • Familiarity with Docker, Kubernetes, and MLflow.
  • Exposure to Apache Spark or Airflow for data engineering workflows.
  • Experience with cloud platforms (AWS, Azure, or GCP).


Primary Technology Stack  

  • Languages & Data Processing: Python, SQL, Pandas, NumPy, Apache Spark
  • AI & Machine Learning: PyTorch, TensorFlow, Scikit-learn
  • Application Frameworks: LangChain, LlamaIndex, CrewAI, n8n
  • Core Methodologies: Retrieval-Augmented Generation (RAG), Model Fine-Tuning, Prompt Engineering, Embeddings
  • Models & Infrastructure: OpenAI APIs, Hugging Face Ecosystem, Embedding Models
  • Vector Databases: Pinecone, Chroma, Milvus, Weaviate
  • Databases: PostgreSQL, MongoDB
  • MLOps & DevOps: Docker, Kubernetes, MLflow, CI/CD, Git
  • Cloud Platforms: AWS, Azure, GCP


Experience: 1–3 Years

Domain: Data Science | Data Engineering | Machine Learning | Generative AI | MLOps

Read more
MNC

MNC

Agency job
via VY SYSTEMS PRIVATE LIMITED by aafia parveen
Hyderabad
5 - 8 yrs
₹2L - ₹20L / yr
Data engineering
Google Cloud Platform (GCP)
Oracle
PySpark
ETL

Job Title: Data Engineer – PySpark | Oracle | GCP


Experience: 5–7 Years

Location: Hyderabad

Notice Period: Immediate Joiners Preferred


Job Summary

We are seeking an experienced Data Engineer with strong expertise in PySpark, Oracle, and Google Cloud Platform (GCP) to design, develop, and optimize scalable data pipelines. The ideal candidate should have hands-on experience in ETL development, data integration, and cloud-based data engineering solutions.

Key Responsibilities


  • Design, develop, and maintain scalable ETL/data pipelines using PySpark.
  • Extract, transform, and load data from Oracle databases into GCP environments.
  • Build and optimize batch data processing workflows for high performance and reliability.
  • Develop data engineering solutions using GCP services.
  • Ensure data quality through validation, monitoring, and troubleshooting.
  • Optimize SQL queries and ETL jobs for performance and scalability.


Required Skills

  • 5–7 years of experience as a Data Engineer.
  • Strong hands-on experience with PySpark.
  • Solid experience with Oracle Database and advanced SQL.
  • Hands-on experience with Google Cloud Platform (GCP).
  • Strong understanding of ETL processes and data warehousing concepts.


Work Location: Hyderabad

Notice Period: Immediate Joiners Preferred

Read more
Bengaluru (Bangalore)
9 - 16 yrs
₹4L - ₹17L / yr
Azure Data Engineer
Azure Databricks
skill iconPython
PySpark
SQL

Job Summary

We are seeking a skilled Azure Data Engineer with hands-on experience in Azure Data Services, Azure Databricks, Python, PySpark, and SQL. The ideal candidate will be responsible for designing, developing, and optimizing scalable data pipelines and data processing solutions to support business intelligence, analytics, and reporting requirements.

Key Responsibilities

  • Design, develop, and maintain scalable ETL/ELT pipelines using Azure Databricks and PySpark.
  • Build and optimize data processing workflows using Python and SQL.
  • Develop and manage data ingestion pipelines from multiple structured and unstructured data sources.
  • Work with Azure Data Factory (ADF) to orchestrate and schedule data pipelines.
  • Implement data transformation and cleansing logic using PySpark.
  • Optimize SQL queries and Spark jobs for performance and scalability.
  • Collaborate with data architects, analysts, and business stakeholders to understand data requirements.
  • Ensure data quality, integrity, and governance across the data platform.
  • Monitor, troubleshoot, and resolve production data pipeline issues.
  • Follow coding standards, version control, and CI/CD best practices.

Required Skills

  • Strong experience with Microsoft Azure cloud services.
  • Hands-on experience with Azure Databricks.
  • Strong programming skills in Python.
  • Expertise in PySpark for large-scale data processing.
  • Strong SQL skills, including query optimization and performance tuning.
  • Experience with Azure Data Factory (ADF).
  • Knowledge of Delta Lake, Spark SQL, and Databricks notebooks.
  • Experience with Git or Azure DevOps for source code management.
  • Understanding of data warehousing concepts and ETL/ELT processes.

Preferred Skills

  • Experience with Azure Synapse Analytics.
  • Knowledge of Delta Live Tables (DLT).
  • Experience with Azure Data Lake Storage (ADLS Gen2).
  • Familiarity with Unity Catalog and data governance.
  • Exposure to CI/CD pipelines and infrastructure-as-code.
  • Experience working in Agile/Scrum environments.
Read more
Quantiphi

at Quantiphi

3 candid answers
1 video
Nikita Sinha
Posted by Nikita Sinha
Mumbai, Bengaluru (Bangalore)
6 - 12 yrs
Best in industry
skill iconAmazon Web Services (AWS)
skill iconPostgreSQL
skill iconPython
PySpark
SQL

As an Associate Technical Architect - Data, you will lead the end-to-end design, architecture, and implementation of enterprise-scale AWS data platforms and Lakehouse solutions. You will provide technical leadership across the entire project lifecycle—from solution architecture and technology selection to implementation, optimization, deployment, and production support.

The role requires deep hands-on expertise in modern AWS data engineering technologies, strong architectural skills, and the ability to mentor engineering teams while collaborating with business and technical stakeholders to deliver scalable, secure, and high-performance data solutions.


Must Have Skills:

  • 8+ years of experience designing and delivering enterprise-scale Data Lake, Lakehouse, or Data Warehouse solutions on AWS.
  • Proven experience leading end-to-end implementation of cloud-native data platforms, including architecture, design, development, deployment, and production support.
  • Strong hands-on expertise in SQL (analytical queries, window functions, stored procedures), Spark/PySpark, and Python.
  • Strong hands-on experience designing and implementing Lakehouse architectures using Apache Iceberg.
  • Strong knowledge of AWS services including EMR, S3, Athena, Glue Catalog, Aurora PostgreSQL, Lambda, CloudWatch, SQS, SNS, EventBridge, IAM, and related AWS data services.
  • Experience designing and implementing scalable batch and streaming data pipelines using AWS native services.
  • Strong expertise in Spark/PySpark performance tuning and optimization.
  • Hands-on experience optimizing Apache Iceberg and Aurora PostgreSQL for performance, scalability, and cost efficiency.
  • Strong understanding of data modeling, distributed data processing, partitioning strategies, file formats, and Lakehouse/Data Lake architectures.
  • Strong understanding of AWS architecture principles, including security, networking, disaster recovery, scalability, resiliency, and cost optimization.
  • Experience designing orchestration workflows using Apache Airflow or AWS Step Functions.
  • Ability to define cloud data platform architectures, evaluate technology choices, and articulate architectural trade-offs and best practices.
  • Experience leading globally distributed engineering teams, mentoring developers, conducting architecture/code reviews, and driving engineering best practices.
  • Excellent communication, stakeholder management, problem-solving, and technical leadership skills with the ability to translate business requirements into scalable technical solutions.
  • AWS Solution Architect Associate/Professional or AWS Data Engineer Associate certification is preferred.


Good to Have Skills:

  • Experience with ClickHouse, including performance tuning and query optimization.
  • Experience with Infrastructure as Code using Terraform or CloudFormation.
  • Experience implementing CI/CD pipelines for data engineering workloads.
  • Exposure to Kafka, Hive, HDFS, or other Big Data technologies.
  • Hands-on experience using GenAI-assisted development tools such as Kiro, GitHub Copilot, Cursor, or similar AI coding assistants to improve engineering productivity.
  • Experience integrating with data governance and metadata management tools such as Collibra.
  • Experience integrating with data virtualization platforms such as Denodo.
  • Telecom/Mobile Network domain knowledge is preferred but not mandatory.
Read more
offrolls

offrolls

Agency job
via TalentXO by tabbasum shaikh
Chennai
6 - 10 yrs
₹40L - ₹50L / yr
Agile Software Development
Angular
BigQuery
CI/CD
DevOps
+10 more

Job Description: Software Engineer / Core Engineer - Java Full Stack (GCP)

Location: Sholinganallur, Chennai - Work From Office (WFO)only chennai candidate

Experience: 6+ years of IT experience with 4+ years in Software Development [Ideal: 6-10 years]

Education: Bachelor's Degree [B.E./B.Tech or equivalent]

Employment Type: Full Time

About the Role:

This is a Java Full Stack Developer role with a strong GCP requirement. We are looking for a hands-on Software Engineer / Core Engineer to design, build, and scale cloud-native full-stack applications. Java, Spring Boot, Angular, and GCP are mandatory skills for this position.

Mandatory Skills:

  • Full Stack Java Development
  • Java
  • Spring Boot
  • Angular
  • Google Cloud Platform (GCP)
  • REST APIs & Microservices
  • Agile Software Development

Preferred Skills:

  • Data & Analytics on GCP: BigQuery, Dataflow, Dataproc, Data Fusion, Cloud SQL, PostgreSQL
  • Big Data & Orchestration: Apache Airflow, PySpark, Python
  • IaC & DevOps: Terraform, Tekton, Maven, Gradle, SonarQube
  • Frameworks & Tools: Spring Cloud, APIGEE, Swagger/OpenAPI
  • UI/UX Collaboration

Roles & Responsibilities:

  1. Design, develop, test, and deploy robust full-stack applications using Java, Spring Boot, and Angular.
  2. Build and maintain scalable REST APIs and microservices-based architecture.
  3. Develop responsive, high-performance Angular-based UI applications and collaborate with UI/UX teams for seamless user experience.
  4. Design software architecture and end-to-end cloud-native solutions on Google Cloud Platform.
  5. Work on GCP deployments, infrastructure setup, and cloud migration initiatives.
  6. Follow and advocate Agile, TDD, CI/CD, and DevOps best practices throughout the SDLC.
  7. Collaborate effectively with Product Owners, Architects, UI/UX teams, and cross-functional stakeholders.
  8. Ensure application security, performance optimization, scalability, and high code quality with SonarQube, code reviews, and testing.

Ideal Candidate Profile:

  • 6 to 10 years of total IT experience with strong hands-on development background.
  • Strong expertise in Java, Spring Boot, Angular, and GCP - must have built production-grade applications.
  • Hands-on experience with Microservices, REST APIs, CI/CD pipelines, and Cloud Deployment.
  • Experience with Cloud SQL, PostgreSQL, and GCP data services is a plus.
  • Good understanding of software architecture, distributed systems, and performance tuning.
  • Excellent communication skills and proven experience working in Agile teams.
  • Willingness to work from Sholinganallur, Chennai office (WFO).

What We Offer:

  • Opportunity to work on cloud-native, large-scale platforms
  • Agile and collaborative work culture
  • Exposure to modern GCP data stack and DevOps ecosystem
Read more
VDart Digital
Marimuthu J
Posted by Marimuthu J
Bengaluru (Bangalore)
12 - 15 yrs
Best in industry
Microsoft Azure
Azure Databricks
PySpark
Azure Data Factory
Data modeling
  • Job Title: Lead Data Engineer
  • Location: Marathahalli, Bangalore 
  • Job Mode: Hybrid


Key Responsibilities:

·   Create and maintain data sources inside our data lake on Azure

·   Assemble and analyze large, complex data sets that meet both functional and non-functional business requirements

·   Identify, design, and implement internal process improvements: automating manual processes; optimizing data delivery; re-designing infrastructure for greater scalability.

·   Work with stakeholders including the Executive, Tech, Product, Data and Design teams to optimize data driven decisions

·   Keep our data separated and secure by understanding and applying best practice access policies

·   Create tools for analytics and data scientist team members that assist them in building and optimizing our product into an innovative industry leader

·   Ensure data integrity with thorough data quality test coverage and data validation

·   Improve, create and implement CI/CD processes and tooling inside the data lake

 

Required Qualifications

-     Join immediately

-     Databricks

-     Databricks hands on experience (esp. Pipelines PySpark & SQL, Delta Lake, Unity Catalog, DLT*)

-     Ingesting & transforming data using Databricks - Min 3 years continuous hands-on

-     Advanced SQL and Python knowledge

-     Data Lake housing and Modelling experience (e.g. medallion architecture & query performance optimization)

-     Experience with backend systems ERP, CRM (e.g. SAP, Oracle, Salesforce)

-     Exposure of working with Azure (ADLS, ADF)

-     Good Communication skills

Read more
Service Co

Service Co

Agency job
via Vikash Technologies by Rishika Teja
Pune
10 - 15 yrs
₹34L - ₹38L / yr
Data engineering
databricks
PySpark
SparkSQL
skill iconPython
+8 more

Hiring for Engineering Manager – Data Engineering


Exp : 10 - 15 yrs

Work Location : Pune WFO


Must Have Skills :


10+ years of experience in Data Engineering.

3+ years of experience leading Data Engineering teams.

Must be a Manager

Strong hands-on experience with Databricks on Azure. 

Strong expertise in PySpark, Spark SQL, Python, and SQL.

Hands-on experience with Delta Lake, Delta Live Tables (DLT), Unity Catalog, Databricks Workflows, and Auto Loader.

Experience designing and building enterprise-scale data platforms.

Strong knowledge of ETL/ELT, Medallion Architecture, and batch & streaming data pipelines.



Read more
VY SYSTEMS PRIVATE LIMITED
Chennai, Tamil Nadu
5 - 10 yrs
₹4L - ₹24L / yr
skill iconPython
PySpark
SQL
ETL

We are looking for a skilled Python & PySpark Developer with strong expertise in Big Data technologies, Spark, SQL/PL-SQL, and REST API development using Flask or Django. The ideal candidate should have experience building scalable data pipelines, processing large datasets, developing APIs, and working with distributed computing frameworks.

Key Responsibilities

  • Develop, optimize, and maintain scalable data pipelines using PySpark and Apache Spark.
  • Design, develop, and optimize complex SQL and PL/SQL queries, stored procedures, functions, and database objects.
  • Build and maintain RESTful APIs using Flask or Django.
  • Develop robust Python applications for data engineering and backend services.
  • Process and analyze large-scale datasets using Big Data technologies.
  • Optimize Spark jobs for performance, scalability, and reliability.
  • Integrate APIs with internal and external systems.
  • Collaborate with cross-functional teams including Data Engineers, Data Scientists, and Application Developers.
  • Troubleshoot production issues and implement performance improvements.
  • Follow coding standards, version control, and CI/CD best practices.

Mandatory Skills

  • Strong proficiency in Python programming.
  • Hands-on experience with PySpark and Apache Spark.
  • Strong SQL coding skills.
  • Experience with PL/SQL development.
  • Experience in Big Data ecosystem.
  • REST API development using Flask or Django.
  • Experience in developing and consuming Python APIs.
  • Knowledge of data processing, ETL, and distributed computing.
  • Experience with Git/version control.

Preferred Skills

  • Experience with Hadoop ecosystem (Hive, HDFS, YARN).
  • Exposure to cloud platforms such as AWS, Azure, or GCP.
  • Knowledge of Airflow or other workflow orchestration tools.
  • Experience with Docker and Kubernetes.
  • Familiarity with Kafka or other streaming technologies.
  • Understanding of CI/CD pipelines.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Information Technology, or a related field.
  • 4–8+ years of experience in Python and Big Data development (can be adjusted based on the role).

Required Experience

  • Strong hands-on experience in Python, PySpark, and Apache Spark.
  • Extensive experience writing optimized SQL and PL/SQL code.
  • Experience developing REST APIs using Flask or Django.
  • Experience working with large-scale data processing and ETL pipelines.
  • Strong analytical, debugging, and problem-solving skills.

Mandatory Skills: Python, PySpark, SQL Coding, Apache Spark, Big Data, Flask/Django (REST API), PL/SQL, Python APIs.

Read more
Global IT Consulting Company

Global IT Consulting Company

Agency job
via BIG IT JOBS by Meetesh Soni
Bengaluru (Bangalore)
12 - 18 yrs
₹30L - ₹40L / yr
skill iconPython
PySpark
Architecture
Data architecture
Snowflake
+8 more

Job Summary

We are looking for an experienced AI Data Architect to design and build an enterprise AI-ready data platform that serves as the single source of truth for AI applications, including RAG, Agentic AI, Conversational AI, ML models, and analytics.

Key Responsibilities

  • Design enterprise AI data platform and Lakehouse architecture.
  • Build batch & real-time data pipelines.
  • Develop semantic models, knowledge graphs, and vector databases.
  • Architect RAG and LLMOps infrastructure.
  • Implement data governance, security, and AI observability.
  • Modernize legacy data platforms to cloud-native architectures.

Required Skills

  • Python, SQL, PySpark
  • Databricks, Delta Lake, Snowflake
  • Kafka, Spark Structured Streaming
  • AWS / Azure
  • LangChain, LlamaIndex
  • OpenAI, Claude, Bedrock
  • Pinecone, FAISS, ChromaDB, Neo4j
  • MLflow, Docker, Kubernetes, Terraform
  • FastAPI, GitHub Actions, Jenkins
  • Data Governance, RBAC, CI/CD

Requirements

  • 12+ years in Data Engineering/Data Architecture.
  • Experience with AI/ML, RAG, LLMOps, and enterprise AI platforms.
  • Strong expertise in Lakehouse, Data Mesh, Cloud, and Vector Databases.
  • Hands-on experience with enterprise-scale AI data architecture and governance. 


Read more
Credilio Financial Technologies Pvt. Ltd.
Munjal Dhamecha
Posted by Munjal Dhamecha
Mumbai
3 - 8 yrs
Best in industry
skill iconAmazon Web Services (AWS)
PySpark
Data engineering
ClickHouse

We are looking for a hands-on Data Engineer to help build and manage our data platform for reporting, analytics, and future data science use cases.


The role will involve working with SQL, Python, PySpark, AWS, ETL/ELT pipelines, data warehouses, and BI/reporting tools. Our architecture may use technologies such as ClickHouse, Redshift, AWS Glue, Airflow, Step Functions, Lambda, S3, and CDC-based replication tools based on scale, cost, and operational needs.


Key Responsibilities

·     Design, build, and maintain scalable ETL/ELT data pipelines from multiple databases, applications, and external systems.

·     Build raw, cleaned, and business-ready data layers to support reporting, analytics, and future data science use cases.

·     Write efficient SQL, Python, and PySpark jobs for data ingestion, transformation, validation, and processing.

·     Implement workflow orchestration using Apache Airflow, AWS Glue, Step Functions, or similar tools.

·     Work with data warehouses such as ClickHouse, Redshift, Snowflake, BigQuery, or similar.

·     Support cross-service reporting as the architecture moves towards independent microservice databases.

·     Build reusable reporting tables, aggregates, summaries, and basic data marts.

·     Monitor, troubleshoot, and optimize data pipelines, warehouse queries, and processing jobs.

·     Implement data quality checks for freshness, completeness, consistency, duplicates, and reconciliation.

·     Support BI/reporting needs through tools such as Power BI, Metabase, Superset, Redash, or similar.

·     Apply data governance, access control, security, and PII-handling best practices.

·     Collaborate with engineering, DevOps, product, business, finance, risk, and support teams.


Required Skills

·     Strong expertise in SQL for joins, aggregations, window functions, query optimization, and analytical reporting.

·     Hands-on experience with Python for data processing, automation, validation, and scripting.

·     Working experience with PySpark / Apache Spark for processing large datasets.

·     Good understanding of ETL/ELT pipelines, data warehousing, and data modelling concepts.

·     Experience with workflow orchestration using Apache Airflow, AWS Step Functions, AWS Glue, or similar tools.

·     Experience with AWS data services such as S3, Glue, Lambda, Step Functions, Redshift, DMS, CloudWatch, or similar.

·     Experience with any data warehouse such as ClickHouse, Redshift, Snowflake, BigQuery, or similar.

·     Understanding of relational databases, preferably PostgreSQL.

·     Ability to debug data mismatches, failed pipelines, slow queries, and data quality issues.

·     Exposure to BI tools such as Power BI, Metabase, Superset, Redash, or similar.

Good ownership, problem-solving, communication, and collaboration skills.

Read more
Hiyrnow
Durga A
Posted by Durga A
Bengaluru (Bangalore)
5 - 10 yrs
₹5L - ₹26L / yr
skill iconData Science
skill iconMachine Learning (ML)
skill iconPython
SQL
databricks
+1 more

Senior Data Scientist (ML) – 5+ Years

Required Skills:

  • 5+ years of experience in Data Science/ML
  • Strong knowledge of Probability & Statistics
  • Strong educational background in Statistics or Mathematics (degree/course specialization). 
  • Good understanding of core Machine Learning algorithms
  • Proficiency in Python, PySpark, and SQL
  • Experience with Databricks, Azure ML, SageMaker, or Vertex AI
  • Hands-on experience in Natural Language Processing (NLP)
  • Understanding of Transformer-based models and architectures
  • Ability to build, deploy, and optimize ML solutions at scale
Read more
Wissen Technology

at Wissen Technology

4 recruiters
Annie Varghese
Posted by Annie Varghese
Pune
7 - 10 yrs
Best in industry
skill iconAmazon Web Services (AWS)
databricks
ETL
PySpark
Apache Spark

Data Engineer – Databricks & AWS (7+ Years)

Location: Baner, Pune

Work Model:5 days from office


Required Skills

  • 7+ years of experience in Data Engineering with strong expertise in Databricks, PySpark, Apache Spark, and SQL.
  • Hands-on experience building scalable ETL/ELT pipelines and Data Lake/Lakehouse solutions using Delta Lake.
  • Experience with AWS services including S3, Glue, Lambda, IAM, EMR, and Redshift.
  • Strong knowledge of Apache Airflow, Git, CI/CD, data quality, performance tuning, and production support.
  • Experience with Kafka/Spark Streaming and Banking, Financial Services, or Credit Bureau domains is preferred.

Roles & Responsibilities

  • Designed and developed scalable data pipelines using Databricks, PySpark, and AWS to process large-scale credit bureau, customer, loan, and repayment datasets.
  • Built ETL/ELT workflows and Delta Lake-based data models to support credit risk analytics, regulatory reporting, and customer profiling.
  • Developed and orchestrated batch and near real-time data processing pipelines using Airflow, Kafka, and Spark, ensuring data quality and reliability.
  • Optimized Spark workloads through performance tuning techniques, improving processing efficiency and reducing execution time.
  • Collaborated with business stakeholders, data architects, and risk teams to deliver data solutions while supporting production environments and operational excellence.

NOTE: One technical round is mandatory to be taken F2F from Pune office.

Read more
Heavy Manufacturing, Automotive/Commercial Vehicles, Industr

Heavy Manufacturing, Automotive/Commercial Vehicles, Industr

Agency job
via Recro by Alok Singh
Bengaluru (Bangalore)
4 - 6 yrs
₹20L - ₹30L / yr
PySpark
databricks
Delta Lake
ADF
Apache Kafka

About the Role

The Smart Operations (Smart Ops) division is seeking a seasoned Senior Data Engineer / Lead Data Engineer to design and scale our next-generation industrial data platform. In this role, you will architect the real-time data backbone that processes high-frequency iHistorian telemetry, IoT sensor feeds, and complex manufacturing data across global operations.

You will be the core technical owner of our Databricks Lakehouse ecosystem, responsible for building optimized, cost-efficient streaming and batch pipelines that deliver high-trust data to our Advanced Analytics, MLOps, and Business Intelligence teams.


Key Responsibilities

  • End-to-End Lakehouse Architecture: Design, implement, and scale a robust Medallion Architecture (Bronze-Silver-Gold) on Azure Databricks and Delta Lake.
  • Real-Time IoT Ingestion: Build fault-tolerant, low-latency streaming pipelines to ingest high-frequency telemetry and machine sensor data.
  • Enterprise Governance: Enforce data security, fine-grained access controls, and comprehensive lineage tracking using Unity Catalog.
  • Performance Tuning & Cost Optimization: Actively profile, tune, and optimize Spark workloads (utilizing partitioning, liquid clustering, AQE, and Photon) to drastically minimize cloud compute costs and pipeline runtimes.
  • CI/CD Automation: Modernize data platform deployments across Dev/QA/Prod environments using Infrastructure as Code (IaC) and Databricks Asset Bundles (DABs).
  • Cross-Functional Collaboration: Partner closely with manufacturing stakeholders, supply chain analysts, and data scientists to provide curated datasets for predictive maintenance, anomaly detection, and operational KPI reporting.


Technical Skills Breakdown

Mandatory Skills (Must-Haves)

  • Domain Context: Minimum 5+ years of Data Engineering experience with direct exposure to Heavy Manufacturing, Automotive/Commercial Vehicles, Industrial Automation, or Global Logistics data environments.
  • Databricks Ecosystem: Deep architectural mastery of Azure Databricks, PySpark, Spark SQL, Delta Lake, and Unity Catalog governance.
  • Streaming & CDC: Proven hands-on experience engineering real-time data pipelines using Apache Kafka, Azure Event Hubs, or Spark Structured Streaming, alongside Change Data Capture (CDC) tools (e.g., Debezium).
  • Cloud Infrastructure: Strong proficiency within the Microsoft Azure stack—specifically Azure Data Factory (ADF), Azure Data Lake Storage Gen2 (ADLS Cores), Azure Key Vault, and Azure Monitor.
  • Advanced SQL & Python: Expert-level Python programming and complex SQL optimisation (advanced indexing, partition pruning, performance tuning of massive datasets).
  • DevOps Automation: Hands-on experience setting up production CI/CD pipelines using Git, Azure DevOps, GitHub Actions, or Databricks Asset Bundles (DABs).


Good to Have Skills (Nice-to-Haves)

  • Professional Certifications: Databricks Certified Data Engineer Professional (highly preferred), Databricks Certified Associate Developer for Apache Spark, or Microsoft Certified: Azure Data Engineer Associate (DP-203).
  • Industrial & ERP Systems: Familiarity with industrial data systems (SCADA, PLC, iHistorian), heavy vehicle telemetry standards (MDF/MF4 files), or deep integration experience with enterprise ERPs (SAP S/4HANA, SAP BW).
  • MLOps & GenAI: Exposure to working alongside Machine Learning pipelines, dataset feature engineering, or implementing Generative AI frameworks (RAG, Vector Databases like FAISS/Weaviate) for semantic query self-service.
  • Infrastructure as Code (IaC): Basic proficiency with Terraform for provisioning cloud data infrastructure.


Education & Qualifications

Bachelor’s or Master’s Degree in Computer Science, Information Technology, Electrical/Electronics Engineering, or a closely related technical field.

Read more
venanalytics

at venanalytics

2 candid answers
Rincy jain
Posted by Rincy jain
Pune
3 - 4 yrs
₹7L - ₹13L / yr
PySpark
Data modeling
skill iconPython
SQL
ETL

About the Role:


We are looking for a highly skilled Data Engineer with a strong foundation in Power BI, SQL, Python, and Big Data ecosystems to help design, build, and optimize end-to-end data solutions. The ideal candidate is passionate about solving complex data problems, transforming raw data into actionable insights, and contributing to data-driven decision-making across the organization.


Key Responsibilities:


  • Data Modelling & Visualization
  • Build scalable and high-quality data models in Power BI using best practices.
  • Define relationships, hierarchies, and measures to support effective storytelling.
  • Ensure dashboards meet standards in accuracy, visualization principles, and timelines.
  • Data Transformation & ETL
  • Perform advanced data transformation using Power Query (M Language) beyond UI-based steps.
  • Design and optimize ETL pipelines using SQL, Python, and Big Data tools.
  • Manage and process large-scale datasets from various sources and formats.
  • Business Problem Translation
  • Collaborate with cross-functional teams to translate complex business problems into scalable, data-centric solutions.
  • Decompose business questions into testable hypotheses and identify relevant datasets for validation.
  • Performance & Troubleshooting
  • Continuously optimize performance of dashboards and pipelines for latency, reliability, and scalability.
  • Troubleshoot and resolve issues related to data access, quality, security, and latency, adhering to SLAs.
  • Analytical Storytelling
  • Apply analytical thinking to design insightful dashboards—prioritizing clarity and usability over aesthetics.
  • Develop data narratives that drive business impact.
  • Solution Design
  • Deliver wireframes, POCs, and final solutions aligned with business requirements and technical feasibility.

Required Skills & Experience:

  • Minimum 3+ years of experience as a Data Engineer or in a similar data-focused role.
  • Strong expertise in Power BI: data modeling, DAX, Power Query (M Language), and visualization best practices.
  • Hands-on with Python and SQL for data analysis, automation, and backend data transformation.
  • Deep understanding of data storytelling, visual best practices, and dashboard performance tuning.
  • Familiarity with DAX Studio and Tabular Editor.
  • Experience in handling high-volume data in production environment.
  • Exposure to Big Data technologies such as:
  • PySpark (must have)
  • Hadoop
  • Hive / HDFS
  • Spark Streaming (optional but preferred)

Why Join Us?

  • Work with a team that's passionate about data innovation.
  • Exposure to modern data stack and tools.
  • Flat structure and collaborative culture.
  • Opportunity to influence data strategy and architecture decisions.


Read more
Wissen Technology

at Wissen Technology

4 recruiters
Bipasha Rath
Posted by Bipasha Rath
Bengaluru (Bangalore)
5 - 8 yrs
Best in industry
Google Cloud Platform (GCP)
skill iconPython
PySpark
SQL
Data engineering

Location: Bangalore (Hybrid/Onsite)

Experience: 5–8 Years

Work location -Manyata Tech park

Job Description

We are seeking a skilled GCP Data Engineer with 5–8 years of experience in designing, developing, and maintaining scalable data pipelines and cloud-based data solutions. The ideal candidate should have strong expertise in Google Cloud Platform (GCP), Python, PySpark, SQL, and Data Engineering concepts.

Key Responsibilities

  • Design, build, and optimize scalable ETL/ELT data pipelines.
  • Develop and maintain data processing solutions using Python and PySpark.
  • Work with large-scale structured and unstructured datasets.
  • Implement data ingestion, transformation, and data quality frameworks.
  • Build and manage data solutions on Google Cloud Platform (GCP).
  • Develop and optimize complex SQL queries, stored procedures, and data models.
  • Collaborate with business stakeholders, data analysts, and cross-functional teams to understand data requirements.
  • Monitor, troubleshoot, and improve data pipeline performance and reliability.
  • Ensure data governance, security, and compliance standards are followed.
  • Support data warehousing and analytics initiatives.

Required Skills

  • 5–8 years of experience in Data Engineering.
  • Strong programming experience in Python.
  • Hands-on experience with PySpark and distributed data processing.
  • Strong expertise in SQL and database performance tuning.
  • Experience with Google Cloud Platform (GCP) services such as:
  • BigQuery
  • Cloud Storage
  • Dataflow
  • Dataproc
  • Cloud Composer
  • Pub/Sub
  • Experience in designing ETL/ELT workflows.
  • Knowledge of data warehousing concepts and dimensional modeling.
  • Experience with version control tools such as Git.
  • Strong problem-solving and analytical skills.

Preferred Skills

  • Experience with CI/CD pipelines and DevOps practices.
  • Exposure to orchestration tools like Airflow/Cloud Composer.
  • Experience working in Agile/Scrum environments.
  • Knowledge of streaming data processing and real-time data pipelines.

Educational Qualification

  • Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
Read more
UK Based

UK Based

Agency job
via Visionyle Solutions by Abinayasri Tamilmani
Mumbai, Pune
12 - 30 yrs
₹10L - ₹40L / yr
Snow flake schema
snowflake
Data architecture
databricks
skill iconPython
+1 more

Visionyle is Hiring for UK Based Company- Lead Data Architect 


Role:


· Lead Data Architect Snowflake


Job Location: Mumbai, Pune

Experience: 15-25 yrs

Client: Capita

Job Type: FTE


Job Description: 

· Mandatory Skills:

Data Architecture, Microsoft Azure, Databricks, Snowflake, Azure Synapse Analytics, Data Modelling, Application Integration, Presales.

· Good to Have Skills:

Microsoft Fabric, Data Strategy / Governance, Data Visualisation, Data Science / AI, DevOps / Cloud Infrastructure

· Experience in end-to-end project all aspects of our delivery right from initiating a project from technology side, initiating a project meaning understanding the requirements from stakeholder, designing architecture in it right up to production and production support organization, support, pre-sales, estimations, costing.


Certifications:

Candidates possessing one or more of the following certifications will be highly preferred:

• Databricks Certified Data Engineer Associate or Professional

• Databricks Partner Solution Architect Champion

• SnowPro Advanced: Data Engineer

• SnowPro Advanced: Architect

• Microsoft Azure Certifications in Data Engineering, Artificial Intelligence (AI), or Solution Architecture.


Applicants holding any of the above certifications are encouraged to apply.

 


Roles & Responsibilities:

· The Lead Data Architect is responsible for providing strategic leadership and technical expertise in data architecture to drive innovation and excellence in solution delivery.

· With a focus in client satisfaction and project success, lead a team of data professionals, providing mentorship, guidance and support to ensure the successful execution of projects

Read more
UK Based

UK Based

Agency job
via Visionyle Solutions by Abinayasri Tamilmani
Mumbai
12 - 30 yrs
₹10L - ₹40L / yr
Snow flake schema
snowflake
databricks
skill iconPython
PySpark
+1 more

Visionyle is Hiring Lead Data Architect 


Role:


· Lead Data Architect Snowflake


Job Location: Mumbai, Pune

Experience: 15-25 yrs

Client: Capita

Job Type: FTE


Job Description: 

· Mandatory Skills:

  • Data Architecture, Microsoft Azure, Databricks, Snowflake, Azure Synapse Analytics, Data Modelling, Application Integration, Presales.

· Good to Have Skills:

  • Microsoft Fabric, Data Strategy / Governance, Data Visualisation, Data Science / AI, DevOps / Cloud Infrastructure
  • Experience in end-to-end project all aspects of our delivery right from initiating a project from technology side, initiating a project meaning understanding the requirements from stakeholder, designing architecture in it right up to production and production support organization, support, pre-sales, estimations, costing.


Certifications:

Candidates possessing one or more of the following certifications will be highly preferred:

• Databricks Certified Data Engineer Associate or Professional

• Databricks Partner Solution Architect Champion

• SnowPro Advanced: Data Engineer

• SnowPro Advanced: Architect

• Microsoft Azure Certifications in Data Engineering, Artificial Intelligence (AI), or Solution Architecture.


Applicants holding any of the above certifications are encouraged to apply.

 


Roles & Responsibilities:

  • The Lead Data Architect is responsible for providing strategic leadership and technical expertise in data architecture to drive innovation and excellence in solution delivery.
  • With a focus in client satisfaction and project success, lead a team of data professionals, providing mentorship, guidance and support to ensure the successful execution of projects


Read more
The supreme consultancy
Pune
4 - 7 yrs
₹23L - ₹38L / yr
SQL
skill iconPython
skill iconMachine Learning (ML)
skill iconDeep Learning
Image Embeddings
+17 more

Role & Responsibilities


Responsibilities


• Contribute to the development and optimization of enterprise-wide search systems and models.

• Design and implement algorithms to improve indexing, query relevance, and search accuracy.

• Support taxonomy, ontology, and metadata model creation for better search outcomes.

• Collaborate with business units (Loans, Insurance, Investments) to build AI-enabled search features.

• Conduct analysis of user behavior and system metrics to refine search performance.

• Work with engineers, product managers, and designers to deliver integrated search solutions.

• Develop production-grade ML systems for ranking, personalization, and recommendations.

• Participate in proof-of-concept initiatives with internal and external partners.

• Follow best practices in software engineering including CI/CD, testing, and monitoring.

• Keep abreast of emerging developments in AI/ML to apply them in practical solutions.


Ideal Candidate


Strong Data Scientist / AI Engineer / Machine Learning Engineer profiles.

Mandatory (Experience 1) – Must have minimum 5+ years of hands-on experience in Data Science, Machine Learning, Applied AI, NLP, Deep Learning, or Generative AI solutions.

Mandatory (Experience 2) – Must have strong hands-on experience in Python programming, SQL, data analysis, feature engineering, model development, and production-grade ML applications.

Mandatory (Experience 3) – Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, Scikit-learn, or equivalent.

Mandatory (Experience 4) – Must have hands-on experience working on NLP, embeddings, semantic search, text classification, document understanding, recommendation systems, or similar AI/ML use cases.

Mandatory (Experience 5) – Must have experience working with Large Language Models (LLMs) such as GPT, Llama, Mistral, Claude, Gemini, Phi, or similar foundation models.

Mandatory (Experience 6) – Must have hands-on experience building or implementing RAG (Retrieval Augmented Generation) systems, vector search, knowledge retrieval, embeddings, chunking, indexing, or semantic retrieval solutions.

Mandatory (Experience 7) – Must have experience working with Git, CI/CD practices, production environments, and scalable AI/ML systems.

Mandatory (CTC) – The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.

Mandatory (Age) - Candidate's Age should be below 30 Years

Preferred (Experience 1) – Experience with MLFlow, Kubeflow, Airflow, Prefect, Feature Stores, Model Registry, or MLOps/LLMOps frameworks.

Preferred (Experience 2) – Experience working with Vector Databases, Spark, PySpark, distributed ML pipelines, large-scale data processing, or real-time ML systems..

Preferred (Experience 3) – Familiarity with Docker, Kubernetes, Azure, AWS, GCP, cloud-native AI deployments, and scalable ML architecture.

Preferred (Company) – Candidates from AI-first startups, Fintech, Banking, Lending, Fraud Analytics, Risk Analytics, Product Companies, SaaS organizations, or data-driven technology companies.


Kindly provide the following details while sending your CV: (Mandatory details)


1) Date of Birth

2) Current Location-

3) Current CTC-

4) Expected CTC-

5) Notice Period-

6) Ready to relocate to Pune?



Regards,

The Supreme Consultancy

Website- https://lnkd.in/eawfxfxU

Read more
Deqode

at Deqode

1 recruiter
purvisha Bhavsar
Posted by purvisha Bhavsar
Pune, Gurugram, Kolkata, Bengaluru (Bangalore), Chennai
5 - 10 yrs
₹5L - ₹25L / yr
Google Cloud Platform (GCP)
PySpark
skill iconPython
SQL
ETL
+1 more

Hiring: GCP Data Engineer (FTE)

📍 Location: Bangalore | Chennai | Pune | Gurgaon | Kolkata

💼 Employment Type: Full-Time

Notice Period - Immediate Joiner ( Serving Notice Period )

Work Mode - Hybrid


We are looking for experienced GCP Data Engineers with strong expertise in building scalable cloud data solutions.


Required Skills:

✔ GCP Data Engineering Experience

✔ BigQuery

✔ SQL & Python

✔ PySpark / Apache Spark

✔ Apache Beam / Dataflow

✔ ETL / ELT Pipeline Development

✔ Airflow / Cloud Composer


Responsibilities:

  • Design and develop scalable ETL/ELT pipelines on GCP
  • Build and optimize BigQuery solutions
  • Process large-scale structured/unstructured data using Spark
  • Develop automated workflows and cloud-native data pipelines
  • Work with GCP services like Dataflow, Dataproc, Cloud Storage, etc.

Good to Have:

  • Google Cloud Certifications (Professional Data Engineer / Solution Architect)
  • Experience with BigTable, Cloud SQL, Spanner, NoSQL databases


🎓 Qualification: Bachelor’s / Master’s in CS, Engineering, or related field

Read more
10XScale.ai
Naveen Balne
Posted by Naveen Balne
Hyderabad, Bengaluru (Bangalore)
8 - 16 yrs
₹10L - ₹50L / yr
Google Cloud Platform (GCP)
skill iconPython
PySpark
Google BigQuery
Data engineering
+4 more

GCP Data Engineer

Experience: 8 – 15 yrs

Grade: C2/D1

Skill: GCP + Python/Pyspark

NP: Immediate joiners


  • 8+ years of hands-on experience in Python.
  • 8+ years of hands-on experience in Data Engineering.
  • 5+ Years of hands-on experience in GCP Big Query.
  • Experience in building scalable data pipelines and automation frameworks.
  • Experience migrating data and pipelines from SQL Server to GCP Big Query.
  • Familiarity with CI/CD tools and Agile methodologies.
  • Good understanding on Data Governance, Data Quality, Metadata, Lineage
  • Expertise in Data Model design in Big query.
  • Expertise in writing optimal Big query SQL and Stored Proc.
  • Expertise in GCS Cloud Storage, Pub Sub, Cloud Composer, DAG, Apache Airflow, Data Flow, Data Proc, Data Plex, Cloud Run.
  • Expertise in Vertex AI and Feature Store.
  • Expertise in Spark and Apache Beam is desirable.


Email resumes to: naveenkb @ 10xscale.ai

Read more
Wissen Technology

at Wissen Technology

4 recruiters
Shrutika SaileshKumar
Posted by Shrutika SaileshKumar
Bengaluru (Bangalore)
5 - 8 yrs
Best in industry
databricks
PySpark
SQL
Spark
skill iconPython

Experience: 5 to 8 years 

 

Job Summary: 

  • We are looking for an experienced Databricks Developer with a strong background in Spark/pyspark to design, develop, and optimize data processing applications. The ideal candidate will build scalable data engineering solutions and ensure seamless data flow across platforms. 
  • Key Responsibilities: 
  • Design, develop, and maintain scalable data processing applications using Python, PySpark and Spark. 
  • Collaborate with data engineers, data scientists, and other stakeholders to understand business requirements and deliver high-quality solutions. 
  • Ensure data integrity, performance, and reliability across all data processing pipelines. 
  • Write clean, maintainable, and efficient code following best practices and coding standards. 
  • Perform data analysis and implement data validation to ensure quality. 
  • Monitor and troubleshoot performance issues in data workflows and pipelines. 
  • Implement and manage CI/CD pipelines for automated testing, integration, and deployment. 

 

  • Required Skills & Qualifications: 
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field. 
  • Proven experience as a Databricks Developer, with strong expertise in Spark/Pyspark
  • Proficiency in SQL and hands-on experience with relational databases
  • Strong programming experience in Python and PySpark
  • Familiarity with version control systems such as Git
  • Strong analytical and problem-solving abilities. 
  • Excellent communication and collaboration skills. 

 

Read more
Everforth Quinnox
Bengaluru (Bangalore), Mumbai
7 - 10 yrs
₹20L - ₹25L / yr
PowerBI
ADF
PySpark
Spark
SQL Queries
+3 more

            Azure Database Developer

 

Mandatory Skills: Data Bricks (Pyspark/Spark), Unity Catalog, T-SQL, Azure Data Factory, Power BI.

Power BI, & Azure Devops

·        Very strong understanding on ETL and ELT

·        Very strong understanding on Medallion architecture.

·        Very strong knowledge in Pyspark and Spark architecture.

·        Very strong knowledge in Unity Catalog architecture using databricks.

·        Very strong knowledge in T-SQL 

·        Good knowledge in Azure data lake architecture and access controls.

·        Good knowledge in data pipeline creation and monitoring using Azure data factory

·        Good knowledge in CI /CD process using Azure devops

·        Power BI

Read more
NA
Remote only
17 - 20 yrs
₹30L - ₹55L / yr
Data modeling
ADF
databricks
PySpark
SQL
+1 more

Company Description

VDart is a global leader in digital solutions, product development, and professional services. Headquartered in Atlanta, GA, USA, the company has a robust global presence across North America, Europe, the Middle East, and Asia. VDart Digital specializes in delivering cutting-edge digital transformation solutions, leveraging technologies like AI/ML, blockchain, cloud computing, IoT, and data analytics. Its innovative product portfolio includes offerings such as TestSamurAI, LendSmartAI, IDocLens, and more, which are designed to optimize operations and drive business growth.


Role Description

We are looking for a seasoned data leader to design, build, and own enterprise-scale data platforms on Azure. This role goes beyond development — it requires end-to-end accountability for architecture, data pipelines, transformation frameworks, and production readiness.

You will act as the critical link between business stakeholders, data engineering teams, and analytics functions, ensuring scalable and high-performance data solutions are delivered and maintained.


Key Responsibilities:

  • Design and implement robust data pipelines using Azure Data Factory (ADF), including integration with REST APIs and external data sources
  • Build scalable data transformation workflows using Databricks (PySpark), handling complex and nested JSON datasets
  • Architect and implement Delta Lake-based data platforms, including fact and dimension models (star schema)
  • Define and enforce best practices for data modeling, performance optimization, and cost efficiency
  • Own end-to-end data platform lifecycle — from architecture and deployment to monitoring and operational support
  • Establish production readiness frameworks, including logging, alerting, and data quality checks
  • Collaborate closely with business and analytics teams to translate requirements into scalable technical solutions
  • Mentor engineering teams and drive architectural governance across projects 


Required Experience & Skills:


•                Experience building pipelines with Azure Data Factory 

•                Experience connecting to REST API sources using Azure Data Factory 

•                Experience building transformations with Databricks using PySpark 

•                Experience handling complex nested JSON files using PySpark 

•                Experience designing dimensional models/star schema 

•                Experience implementing facts and dimension tables in Databricks Delta Lake 

•                Around 15-20 years of solid experience in building, managing, and optimizing enterprise data platforms with at least 5 years in Azure cloud data services

•                Act as a bridge between business, data engineering, and analytics teams to ensure requirements are clearly understood and implemented correctly

•                Own end-to-end production readiness of the data platform, including architectural design, deployment patterns, monitoring strategy and operational support.

Read more
Eassy Onboard
Pawan Beesetti
Posted by Pawan Beesetti
Remote only
4 - 8 yrs
₹25L - ₹50L / yr
PySpark
SQL
Spark
databricks
Google Cloud Platform (GCP)
+2 more

Company Description


Eassy Onboard LLP is a team of Databricks Certified Data Engineers committed to empowering businesses through data-driven solutions. Specializing in automated workflows, scalable architectures, optimized data pipelines, and AI solutions, we help organizations reduce manual effort, optimize costs, and achieve reliable insights. We ensure secure data operations with robust validation processes and strong data integrity. As an Employer of Record (EOR), we assist global companies in hiring top Indian talent, managing payroll, compliance, and regulatory requirements. Our mission is to accelerate enterprise transformation and enable companies to build future-ready, compliant teams.


Role Description


We are seeking a Senior Data Engineer with deep expertise in Spark/PySpark/SQL to join our data team.

This is a hands-on technical role for someone passionate about building scalable data systems, mentoring engineers, and shaping data strategy.

You will architect systems that power high-performance data processing, enable advanced analytics, and accelerate AI initiatives.


What You'll Do

  • Design and evolve scalable, distributed data infrastructure across cloud platforms including GCP and AWS.
  • Build and maintain real-time and batch data processing pipelines supporting AI/ML workloads, consumer applications, and analytics.
  • Develop and manage integrations with third-party e-commerce platforms to expand the data ecosystem.
  • Ensure data availability, reliability, and quality through monitoring and automated auditing.
  • Partner with engineering, AI, and product teams on data solutions for business-critical needs.
  • Mentor and support data engineers, establishing best practices and code quality standards.



Qualifications

  • Bachelor's degree in Computer Science or a related field, or equivalent practical experience.
  • 5+ years of software development and data engineering experience with ownership of production-grade data infrastructure.
  • Deep expertise scaling Spark, PySpark, and SQL in production, including Databricks or DataProc on GCP.
  • Strong understanding of distributed computing and modern data modeling for scalable systems.
  • Proficient in Python with experience implementing software engineering best practices.
  • Hands-on experience with both relational and NoSQL systems including MySQL, MongoDB, and Elasticsearch.
  • Strong communicator with experience influencing cross-functional stakeholders.


Nice to Have

  • Experience with job orchestration and containerization tools such as Airflow and Docker.
  • Experience working with vector stores and knowledge graphs.
  • Experience working in early-stage, high-growth environments.
  • Familiarity with MLOps pipelines and integrating ML models into data workflows.
  • A proactive, problem-solving mindset with a passion for innovative solutions.
Read more
VDart Digital

at VDart Digital

2 recruiters
Martin Raj
Posted by Martin Raj
Remote only
15 - 20 yrs
₹40L - ₹55L / yr
Microsoft Windows Azure
databricks
datafactory
Data modeling
PySpark
+2 more

Company Description:

VDart Digital is a global leader in digital solutions, product development, and professional services. Headquartered in Atlanta, GA, USA, the company has a robust global presence across North America, Europe, the Middle East, and Asia. VDart Digital specializes in delivering cutting-edge digital transformation solutions, leveraging technologies like AI/ML, blockchain, cloud computing, IoT, and data analytics. Its innovative product portfolio includes offerings such as TestSamurAI, LendSmartAI, IDocLens, and more, which are designed to optimize operations and drive business growth.

Role Description

  • Design, build, and own enterprise-scale data platforms on Azure
  • Take end-to-end accountability for architecture, data pipelines, transformation frameworks, and production readiness
  • Act as a bridge between business stakeholders, data engineering, and analytics teams to deliver scalable, high-performance solutions

Key Responsibilities

  • Design and implement data pipelines using Azure Data Factory (ADF), including REST API and external source integration
  • Build data transformation workflows using Databricks (PySpark) for complex and nested JSON datasets
  • Architect and implement Delta Lake-based platforms with fact and dimension models (star schema)
  • Define and enforce best practices for data modeling, performance optimization, and cost efficiency
  • Own the full data platform lifecycle: architecture, deployment, monitoring, and support
  • Establish production readiness frameworks including logging, alerting, and data quality checks
  • Collaborate with business and analytics teams to translate requirements into scalable solutions
  • Mentor engineering teams and drive architectural governance

Required Experience & Skills

  • Experience building pipelines using Azure Data Factory
  • Experience integrating REST APIs using Azure Data Factory
  • Experience building transformations using Databricks (PySpark)
  • Experience handling complex and nested JSON datasets using PySpark
  • Experience designing dimensional models (star schema)
  • Experience implementing fact and dimension tables in Databricks Delta Lake
  • 15–20 years of experience in enterprise data platforms, with at least 5 years in Azure data services
  • Ability to bridge business, data engineering, and analytics teams
  • Experience owning end-to-end production readiness including architecture, deployment, monitoring, and support


Read more
Bengaluru (Bangalore)
5 - 10 yrs
₹1L - ₹10L / yr
databricks
PySpark
Apache Spark
ETL
CI/CD
+10 more

Profile - Databricks Developer

Experience- 5+ years

Location- Bangalore (On site)

PF & BGV is Mandatory


Job Description: -

* Design, build, and optimize data pipelines and ETL/ELT workflows using Databricks and

Apache Spark (PySpark).

* Develop scalable, high performance data solutions using Spark distributed processing.

* Lead engineering initiatives focused on automation, performance tuning, and platform

modernization.

* Implement and manage CI/CD pipelines using Git-based workflows and tools such as

GitHub Actions or Jenkins.

* Collaborate with cross-functional teams to translate business needs into technical

solutions.

* Ensure data quality, governance, and security across all processes.

* Troubleshoot and optimize Spark jobs, Databricks clusters, and workflows.

* Participate in code reviews and develop reusable engineering frameworks.

* Should have knowledge of utilizing AI tools to improve productivity and support daily

engineering activities.

* Strong knowledge and hands-on experience in Databricks Genie, including prompt

engineering, workspace usage, and automation.

Required Skills & Experience:

* 5+ years of experience in Data Engineering or related fields.

* Strong hands-on expertise in Databricks (notebooks, Delta Lake, job orchestration).

* Deep knowledge of Apache Spark (PySpark, Spark SQL, optimization techniques).

* Strong proficiency in Python for data processing, automation, and framework

development.

* Strong proficiency in SQL, including complex queries, performance tuning, and analytical

functions.

* Strong knowledge of Databricks Genie and leveraging it for engineering workflows.

* Strong experience with CI/CD and Git-based development workflows.

* Proficiency in data modeling and ETL/ELT pipeline design.


* Experience with automation frameworks and scheduling tools.

* Solid understanding of distributed systems and big data concepts

Read more
Global MNC serving 40+ Fortune 500 Companies

Global MNC serving 40+ Fortune 500 Companies

Agency job
Bengaluru (Bangalore)
5 - 8 yrs
₹20L - ₹26L / yr
Generative AI
Retrieval Augmented Generation (RAG)
skill iconMachine Learning (ML)
LangGraph
langchain
+11 more

Want to work on exciting GenAI projects for Fortune 500 companies across multiple sectors? Then read on..


About Company:

CSG is a multi-national company having a presence in 20 countries with 1600+ Engineers. Company works with more than 40 Fortune 500 customers such as Sony, Samsung, ABB, Thyssenkrup, Toyota, Mitsubishi and many more.


Job Description:

We are looking for a talented Generative AI Developer to join our dynamic AI/ML team. This position offers an exciting opportunity to leverage cutting-edge Generative AI (GenAI) technologies to drive innovation to solve real world problems. You will be responsible for developing and optimizing GenAI-based applications, implementing advanced techniques like Retrieval-Augmented Generation (RAG), RIG (Retrieval Interleaved Generation), Agentic Frameworks and vector databases. This is a collaborative role where you will work directly with customers cross-functional teams to design, implement, and optimize AI-driven solutions. Exposure to cloud-native AI platforms such as Amazon Bedrock and Microsoft Azure OpenAI is highly desirable.


Key Responsibilities

Generative AI Application Development:

Design, develop, and deploy GenAI-driven applications to address complex industrial challenges.

Implement Retrieval-Augmented Generation (RAG) and Agentic frameworks


Data Management & Optimization:

Design and optimize document chunking strategies tailored to specific datasets and use cases.

Build, manage, and optimize data embeddings for high-performance similarity searches across vector databases.


Collaboration & Integration:

Work closely with data engineers and scientists to integrate AI solutions into existing pipelines.

Collaborate with cross-functional teams to ensure seamless AI implementation.


Cloud & AI Platform Utilization:

Explore and implement best practices for utilizing cloud-native AI platforms, such as Amazon Bedrock and Azure OpenAI, to enhance solution delivery.

Continuous Learning & Innovation:

Stay updated with the latest trends and emerging technologies in the GenAI and AI/ML fields, ensuring our solutions remain cutting-edge.


Requirements:

The ideal candidate will have strong experience in Generative AI technologies, particularly in the areas of RAG, document chunking, and vector database management. They will be able to quickly adapt to evolving AI frameworks and leverage cloud-native platforms to create efficient, scalable solutions. You will be working in a fast-paced and collaborative environment, where innovation and the ability to learn and grow are key to success.

- 3 to 5 years of overall experience in software development, with 3 years focused on AI/ML.

- Minimum 2 years of experience specifically working with Generative AI (GenAI) technologies.

- Python, PySpark and SQL knowledge is necessary for tasks

- Proven ability to work in a collaborative, fast-paced, and innovative environment.


Technical Skills:

- Generative AI Frameworks & Technologies:

- Expertise in Generative AI frameworks, including prompt engineering, fine-tuning, and few-shot learning.

- Familiarity with frameworks such as T5 (Text-to-Text Transfer Transformation), LangChain, Lang Graph, Open-source tech stalk Ollama, Mistral, DeepSeek.

- Strong knowledge of Retrieval-Augmented Generation (RAG) for combining LLMs with external data retrieval systems.


Data Management:

- Experience in designing chunking strategies for different datasets.

- Expertise in data embedding techniques and experience with vector databases like Pinecone, ChromaDB etc

- Programming & AI/ML Libraries:

- Strong programming skills in Python.

- Experience with AI/ML libraries such as TensorFlow, PyTorch, and Hugging Face Transformers.


Cloud Platforms & Integration:

- Familiarity with cloud services for AI/ML workloads (AWS, Azure).

- Experience with API integration for AI services and building scalable applications.

- Certifications (Optional but Desirable):

- Certification in AI/ML (e.g., TensorFlow, AWS Certified Machine Learning Specialty).

- Certification or coursework in Generative AI or related technologies.

Read more
Bengaluru (Bangalore)
5 - 10 yrs
₹1L - ₹8L / yr
databricks
ETL
PySpark
Apache Spark
CI/CD
+7 more

Profile - Databricks Developer

Experience- 5+ years

Location- Bangalore (On site)

PF & BGV is Mandatory


Job Description: -


* Design, build, and optimize data pipelines and ETL/ELT workflows using Databricks and Apache Spark (PySpark).

* Develop scalable, high performance data solutions using Spark distributed processing.

* Lead engineering initiatives focused on automation, performance tuning, and platform modernization.

* Implement and manage CI/CD pipelines using Git-based workflows and tools such as GitHub Actions or Jenkins.

* Collaborate with cross-functional teams to translate business needs into technical solutions.

* Ensure data quality, governance, and security across all processes.

* Troubleshoot and optimize Spark jobs, Databricks clusters, and workflows.

* Participate in code reviews and develop reusable engineering frameworks.

* Should have knowledge of utilizing AI tools to improve productivity and support daily engineering activities.

* Strong knowledge and hands-on experience in Databricks Genie, including prompt engineering, workspace usage, and automation


. Required Skills & Experience:

* 5+ years of experience in Data Engineering or related fields.

* Strong hands-on expertise in Databricks (notebooks, Delta Lake, job orchestration).

* Deep knowledge of Apache Spark (PySpark, Spark SQL, optimization techniques).

* Strong proficiency in Python for data processing, automation, and framework development.

* Strong proficiency in SQL, including complex queries, performance tuning, and analytical functions.

* Strong knowledge of Databricks Genie and leveraging it for engineering workflows.

* Strong experience with CI/CD and Git-based development workflows. * Proficiency in data modeling and ETL/ELT pipeline design.

* Experience with automation frameworks and scheduling tools.

* Solid understanding of distributed systems and big data concepts

Read more
Searce Inc

at Searce Inc

3 recruiters
Jatin Gereja
Posted by Jatin Gereja
Bengaluru (Bangalore), Mumbai, Pune
10 - 18 yrs
Best in industry
Google Cloud Platform (GCP)
skill iconAmazon Web Services (AWS)
Enterprise Data Warehouse (EDW)
Data modeling
Big Data
+9 more

Director - Data engineering


What are we looking for

real solver?

Solver? Absolutely. But not the usual kind. We're searching for the architects of the audacious & the pioneers of the possible. If you're the type to dismantle assumptions, re-engineer ‘best practices,’ and build solutions that make the future possible NOW, then you're speaking our language.


Your Responsibilities

what you will wake up to solve.

1. Delivery & Tactical Rigor

  • Methodology Implementation: Implement and manage a unified, 'DataOps-First' methodology for data engineering delivery (ETL/ELT pipelines, Data Modeling, MLOps, Data Governance) within assigned business units. This ensures predictable outcomes and trusted data integrity by reducing architecture variability at the project level.
  • Operational Stewardship: Drive initiatives to optimize team utilization and enhance operational efficiency within the practice. You manage the commercial success of your squads, ensuring data delivery models (from migration to modern data stack implementation) are executed profitably, scalably, and cost-effectively.
  • Execution & Technical Resolution
  • Technical Escalation: Serve as the primary escalation point for delivery issues, personally leading the resolution of complex data integration bottlenecks and pipeline failures to protect client timelines and data reliability standards.
  • Quality Enforcement
  • Quality Oversight: Execute and monitor technical data quality standards, ensuring engineering teams adhere to strict policies regarding data lineage, automated quality checks (observability), security/privacy compliance (GDPR/CCPA/PII), and active catalog management.

2. Strategic Growth & Practice Scaling

  • Talent & Scaling Execution: Execute the strategy for data engineering talent acquisition and development within your business units. Implement objective metrics to assess and grow the 'Data-Native' DNA of your teams, ensuring squads are consistently equipped to handle petabyte-scale environments and high-impact delivery.
  • Offerings Alignment: Drive the adoption of standardized regional offerings (e.g., Modern Data Platform, Data Mesh, Lakehouse Implementation). Ensure your teams leverage the profitable frameworks defined by the practice to accelerate time-to-insight and eliminate architectural fragmentation in client environments.
  • Innovation & IP Development: Lead the practical integration of Vector Databases and LLM-ready architectures into project delivery. Champion the hands-on development of IP and reusable accelerators (e.g., automated ingestion engines) that improve delivery speed and enhance data availability across your portfolio.

3. Leadership & Unit Management

  • Unit Leadership: Directly lead, mentor, and manage the Engineering Managers and Lead Architects within your business unit. Hold your teams accountable for project-level operational consistency, technical talent development, and strict adherence to the practice's data governance standards.
  • Stakeholder Communication: Clearly articulate the business unit’s operational performance, technical quality metrics, and delivery progress to the C-suite Stakeholders and regional client leadership, bridging the gap between technical execution and business value.
  • Ecosystem Alignment: Maintain strong technical relationships with key partner contacts (Snowflake, Databricks, AWS/GCP). Align team delivery capabilities with current product roadmaps and ensure squad-level participation in training, certifications, and partner-led enablement opportunities.


Welcome to Searce

The ‘process-first’, AI-native modern tech consultancy that's rewriting the rules.

We don’t do traditional.

As an engineering-led consultancy, we are dedicated to relentlessly improving the real business outcomes. Our solvers co-innovate with clients to futurify operations and make processes smarter, faster & better.


Functional Skills

1. Delivery Management & Operational Excellence

  • Methodology Execution: Expert capability in implementing and enforcing a unified delivery methodology (DataOps, Agile, Mesh Principles) within specific business units. Proven track record of auditing squad-level adherence to ensure consistency across the project lifecycle.
  • Operational Performance: High proficiency in managing day-to-day operational metrics, including squad utilization, resource forecasting, and productivity tracking. Skilled at optimizing team performance to meet profitability and efficiency targets.
  • SOW & Risk Mitigation: Proven experience in operationalizing Statement of Work (SOW) requirements and identifying technical delivery risks early. Expert at mitigating scope creep and data-specific bottlenecks (e.g., latency, ingestion gaps) before they impact client outcomes.
  • Technical Escalation Leadership: Demonstrated ability to lead "war room" efforts to resolve complex pipeline failures or data integrity issues. Skilled at providing clear, rapid remediation plans and communicating technical status directly to regional stakeholders.

2. Architectural Implementation & Technical Oversight

  • Modern Stack Proficiency: Deep, hands-on expertise in implementing Cloud-Native architectures (Lakehouse, Data Mesh, MPP) on Snowflake, Databricks, or hyperscalers. Ability to conduct deep-dive architectural reviews and course-correct design decisions at the squad level to ensure scalability.
  • Operationalizing Governance: Proven experience in embedding data quality and observability (completeness, freshness, accuracy) directly into the CI/CD pipeline. Responsible for technical enforcement of regulatory compliance (GDPR/PII) and maintaining the integrity of data catalogs across active projects.
  • Applied Domain Expertise: Practical experience leading the delivery of high-growth solutions, specifically Generative AI infrastructure (RAG, Vector DBs), Real-Time Streaming, and large-scale platform migrations with a focus on zero-downtime execution.
  • DataOps & Engineering Standards: Expert-level mastery of DataOps, including the setup and management of orchestration frameworks (Airflow, Dagster) and Infrastructure as Code (IaC). You ensure that automation is a baseline requirement, not an afterthought, for all delivery teams.

3. Unit Management & Commercial Execution

  • Unit & Team Management: Proven success in leading and mentoring Engineering Managers and Lead Architects. Responsible for the operational metrics, technical output, and career development of the business unit's talent pool.
  • Offerings Implementation & Scoping: Expertise in translating service offerings (e.g., Data Maturity Assessments, Lakehouse Builds) into accurate project scopes, technical estimates, and resource plans to ensure delivery is both profitable and competitive.
  • Talent Growth & Mentorship: Functional ability to implement growth frameworks for data engineering roles. Focus on hands-on coaching and scaling high-performance technical talent to meet the demands of complex, petabyte-scale environments.
  • Partner Enablement: Functional competence in managing regional technical relationships with major partners (Snowflake, Databricks, GCP/AWS). Drives squad-level certifications, joint technical enablement, and alignment with partner product roadmaps.

Tech Superpowers

  • Modern Data Architect – Reimagines business with the Modern Data Stack (MDS) to deliver data mesh implementations, insights, & real value to clients.
  • End-to-End Ecosystem Thinker – Builds modular, reusable data products across ingestion, transformation (ETL/ELT), governance, and consumption layers.
  • Distributed Compute Savant – Crafts resilient, high-throughput architectures that survive petabyte-scale volume and data skew without breaking the bank.
  • Governance & Integrity Guardian – Embeds data quality, complete lineage, and privacy-by-design (GDPR/PII) into every table, view, and pipeline.
  • AI-Ready Orchestrator – Engineers pipelines that bridge structured data with Unstructured/Vector stores, powering RAG models and Generative AI workflows.
  • Product-Minded Strategist – Balances architectural purity with time-to-insight; treats every dataset as a measurable "Data Product" with clear ROI.
  • Pragmatic Stack Curator – Chooses the simplest tools that compound reliability; fluent in SQL, Python, Spark, dbt, and Cloud Warehouses.
  • Builder @ Heart – Writes, reviews, and optimizes queries daily; proves architectures with cost-performance benchmarks, not slideware. Business-first, data-second, outcome focused technology leader.

Experience & Relevance

  • Executive Experience: Minimum 10+ years of progressive experience in data engineering and analytics, with at least 3 years in a Senior Manager or Director -level role managing multiple technical teams and owning significant operational and efficiency metrics for a large data service line.
  • Delivery Standardization: Demonstrated success in defining and implementing globally consistent, repeatable delivery methodologies (DataOps/Agile Data Warehousing) across diverse teams.
  • Architectural Depth: Must retain deep, current expertise in Modern Data Stack architectures (Lakehouse, MPP, Mesh) and maintain the ability to personally validate high-level architectural and data pipeline design decisions.
  • Operational Leadership: Proven expertise in managing and scaling large professional services organizations, demonstrated ability to optimize utilization, resource allocation, and operational expense.
  • Domain Expertise: Strong background in Enterprise Data Platforms, Applied AI/ML, Generative AI integration, or large-scale Cloud Data Migration.
  • Communication: Exceptional executive-level presentation and negotiation skills, particularly in communicating complex operational, data quality, and governance metrics to C-level stakeholders.

Join the ‘real solvers’

ready to futurify?

If you are excited by the possibilities of what an AI-native engineering-led, modern tech consultancy can do to futurify businesses, apply here and experience the ‘Art of the possible’. Don’t Just Send a Resume. Send a Statement.

Read more
Searce Inc

at Searce Inc

3 recruiters
Tejashree Kokare
Posted by Tejashree Kokare
Bengaluru (Bangalore), Pune, Mumbai
6 - 15 yrs
Best in industry
Google Cloud Platform (GCP)
Data engineering
Data warehouse architecture
Data architecture
Data modeling
+6 more

Solutions Architect - Data Engineering


Modern tech solutions advisory & 'futurify' consulting as a Searce lead fds (‘forward deployed solver’) architecting scalable data platforms and robust data engineering solutions that power intelligent insights and fuel AI innovation.

If you’re a tech-savvy, consultative seller with the brain of a strategist, the heart of a builder, and the charisma of a storyteller — we’ve got a seat for you at the front of the table.

You're not a sales lead. You're the transformation driver.


What are we looking for

real solver?

Solver? Absolutely. But not the usual kind. We're searching for the architects of the audacious & the pioneers of the possible. If you're the type to dismantle assumptions, re-engineer ‘best practices,’ and build solutions that make the future possible NOW, then you're speaking our language.

  • Improver. Solver. Futurist.
  • Great sense of humor.
  • ‘Possible. It is.’ Mindset.
  • Compassionate collaborator. Bold experimenter. Tireless iterator.
  • Natural creativity that doesn’t just challenge the norm, but solves to design what’s better.
  • Thinks in systems. Solves at scale.


This Isn’t for Everyone. But if you’re the kind who questions why things are done a certain way— and then identifies 3 better ways to do it — we’d love to chat with you.


Your Responsibilities

what you will wake up to solve.


You are not just a Solutions Architect; you are a futurifier of our data universe and the primary enabler of our AI ambitions. With a deep-seated passion for data engineering, you will architect and build the foundational data infrastructure that powers the customers entire data intelligence ecosystem.

As the Directly Responsible Individual (DRI) for our enterprise-grade data platforms, you own the outcome, end-to-end. You are the definitive solver for our customer's most complex data challenges, leveraging a powerful tech stack including Snowflake, Databricks, etc. and core GCP & AWS services (BigQuery, Spanner, Airflow, Kafka). This is a hands-on-keys role where you won't just design solutions—you'll build them, break them, and perfect them.


  • Solution Design & Pre-sales Excellence:Collaborate with cross-functional teams, including sales, engineering, and operations, to ensure successful project delivery.
  • Design Core Data Engineering: Master data modeling, architecting high-performance data ingestion pipelines and ensuring data quality and governance throughout the data lifecycle.
  • Enable Cloud & AI: Design and implement solutions utilizing core GCP data services, building foundational data platforms that efficiently support advanced analytics and AI/ML initiatives.
  • Optimize Performance & Cost: Continuously optimize data architectures and implementations for performance, efficiency, and cost-effectiveness within the cloud environment.
  • Bridge Business & Tech: Translate complex business requirements into clear technical designs, providing technical leadership and guidance to data engineering teams.
  • Stay Ahead of the Curve: Continuously research and evaluate new data technologies, architectural patterns, and industry trends to keep our data platforms at the cutting edge.


Functional Skills:


  • Enterprise Data Architecture Design: Expert ability to design holistic, scalable, and resilient data architectures for complex enterprise environments.
  • Cloud Data Platform Strategy: Proven capability to strategize, design, and implement cloud-native data platforms.
  • Pre-Sales & Technical Storyteller: Crafts compelling, client-ready proposals, architectural decks, and technical demonstrations. Doesn't just present; shapes the strategic technical narrative behind every proposed solution.
  • Advanced Data Modelling: Mastery in designing various data models for analytical, operational, and transactional use cases.
  • Data Ingestion & Pipeline Orchestration: Strong expertise in designing and optimizing robust data ingestion and transformation pipelines.
  • Stakeholder Communication: Exceptional skills in articulating complex technical concepts and architectural decisions to both technical and non-technical stakeholders.
  • Performance & Cost Optimization: Adept at optimizing data solutions for performance, efficiency, and cost within a cloud environment.


Tech Superpowers:


  • Cloud Data Mastery: You're a wizard at leveraging public cloud data services, with deep expertise in GCP (BigQuery, Spanner, etc.) and expert proficiency in modern data warehouse solutions like Snowflake.
  • Data Engineering Core: Highly skilled in designing, implementing, and managing data workflows using tools like Apache Airflow and Apache Kafka. You're also an authority on advanced data modeling and ETL/ELT patterns.
  • AI/ML Data Foundation: You instinctively design data pipelines and structures that efficiently feed and empower Machine Learning and Artificial Intelligence applications.
  • Programming for Data: You have a strong command over key programming languages (Python, SQL) for scripting, automation, and building data processing applications.


Experience & Relevance:


  • Architectural Leadership (8+ Years): You bring extensive experience (7+ years) specifically in a Solutions Architect role, focused on data engineering and platform building.
  • Cloud Data Expertise: You have a proven track record of designing and implementing production-grade data solutions leveraging major public cloud platforms, with significant experience in Google Cloud Platform (GCP).
  • Data Warehousing & Data Platform: Demonstrated hands-on experience in the end-to-end design, implementation, and optimization of modern data warehouses and comprehensive data platforms.
  • Databricks & BigQuery Mastery: You possess significant practical experience with Databricks as a core data warehouse and GCP BigQuery for analytical workloads.
  • Data Ingestion & Orchestration: Proven experience designing and implementing complex data ingestion pipelines and workflow orchestration using tools like Airflow and real-time streaming technologies like Kafka.
  • AI/ML Data Enablement: Experience in building data foundations specifically geared towards supporting Machine Learning and Artificial Intelligence initiatives.


Join the ‘real solvers’

ready to futurify?

If you are excited by the possibilities of what an AI-native engineering-led, modern tech consultancy can do to futurify businesses, apply here and experience the ‘Art of the possible’.


Don’t Just Send a Resume. Send a Statement.


So, If you are passionate about tech, future & what you read above (we really are!), apply here to experience the ‘Art of Possible’

Read more
AI Industry

AI Industry

Agency job
via Peak Hire Solutions by Dharati Thakkar
Mumbai, Bengaluru (Bangalore), Hyderabad, Gurugram
6 - 10 yrs
₹32L - ₹42L / yr
ETL
SQL
Google Cloud Platform (GCP)
Data engineering
ELT
+17 more

Role & Responsibilities:

We are looking for a strong Data Engineer to join our growing team. The ideal candidate brings solid ETL fundamentals, hands-on pipeline experience, and cloud platform proficiency — with a preference for GCP / BigQuery expertise.


Responsibilities:

  • Design, build, and maintain scalable data pipelines and ETL/ELT workflows
  • Work with Dataform or DBT to implement transformation logic and data models
  • Develop and optimize data solutions on GCP (BigQuery, GCS) or AWS/Azure
  • Support data migration initiatives and data mesh architecture patterns
  • Collaborate with analysts, scientists, and business stakeholders to deliver reliable data products
  • Apply data governance and quality best practices across the data lifecycle
  • Troubleshoot pipeline issues and drive proactive monitoring and resolution


Ideal Candidate:

  • Strong Data Engineer Profile
  • Must have 6+ years of hands-on experience in Data Engineering, with strong ownership of end-to-end data pipeline development.
  • Must have strong experience in ETL/ELT pipeline design, transformation logic, and data workflow orchestration.
  • Must have hands-on experience with any one of the following: Dataform, dbt, or BigQuery, with practical exposure to data transformation, modeling, or cloud data warehousing.
  • Must have working experience on any cloud platform: GCP (preferred), AWS, or Azure, including object storage (GCS, S3, ADLS).
  • Must have strong SQL skills with experience in writing complex queries and optimizing performance.
  • Must have programming experience in Python and/or SQL for data processing.
  • Must have experience in building and maintaining scalable data pipelines and troubleshooting data issues.
  • Exposure to data migration projects and/or data mesh architecture concepts.
  • Experience with Spark / PySpark or large-scale data processing frameworks.
  • Experience working in product-based companies or data-driven environments.
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.


NOTE:

  • There will be an interview drive scheduled on 28th and 29th March 2026, and if shortlisted, they will be expected to be available on these Interview dates. Only Immediate joiners are considered.
Read more
Deqode

at Deqode

1 recruiter
purvisha Bhavsar
Posted by purvisha Bhavsar
Pune
5 - 6 yrs
₹4L - ₹10L / yr
Windows Azure
skill iconPython
PySpark
ADF
databricks
+2 more

🚀 Hiring: Data Engineer ( Azure ) at Deqode

⭐ Experience: 5+ Years

📍 Location: Pune, Bhopal, Jaipur, Gurgaon, Delhi, Banglore,

⭐ Work Mode:- Hybrid

⏱️ Notice Period: Immediate Joiners

(Only immediate joiners & candidates serving notice period)


⭐ Hiring: Databricks Data Engineer – Lakeflow | Streaming | DBSQL | Data Intelligence

We are looking for a Databricks Data Engineer ( Azure ) to build reliable, scalable, and governed data pipelines powering analytics, operational reporting, and the Data Intelligence Layer.


🔹 Key Responsibilities

✅ Build optimized batch pipelines using Delta Lake (partitioning, OPTIMIZE, Z-ORDER, VACUUM)

✅ Implement incremental ingestion using Databricks Autoloader with schema evolution & checkpointing

✅ Develop Structured Streaming pipelines with watermarking, late data handling & restart safety

✅ Implement declarative pipelines using Lakeflow

✅ Design idempotent, replayable pipelines with safe backfills

✅ Optimize Spark workloads (AQE, skew handling, shuffle & join tuning)

✅ Build curated datasets for Databricks SQL (DBSQL), dashboards & downstream applications

✅ Package and deploy using Databricks Repos & Asset Bundles (CI/CD)

Ensure governance using Unity Catalog and embedded data quality checks


✅ Mandatory Skills (Must Have)

👉 Databricks & Delta Lake (Advanced Optimization & Performance Tuning)

👉 Structured Streaming & Autoloader Implementation

👉 Databricks SQL (DBSQL) & Data Modeling for Analytics

Read more
The Blue Owls Solutions

at The Blue Owls Solutions

2 candid answers
Apoorvo Chakraborty
Posted by Apoorvo Chakraborty
Pune
2 - 5 yrs
₹10L - ₹18L / yr
PySpark
SQL
skill iconPython
Data engineering
ETL

Blue Owls Solutions is looking for a mid-level Azure Data Engineer with approximately 4 years of hands-on experience to join our growing data team. In this role, you will design, build, and maintain scalable data pipelines and architectures that power business-critical analytics and reporting. You'll work closely with cross-functional teams to transform raw data into reliable, high-quality datasets that drive decision-making across the organization.

Required Skills

  • 4+ years of professional experience as a Data Engineer or in a similar data-focused role
  • Strong proficiency in SQL for data manipulation, querying, and performance optimization
  • Hands-on experience with PySpark for large-scale data processing and transformation
  • Solid working knowledge of the Microsoft Azure ecosystem (Azure Data Factory, Azure Data Lake, Azure Synapse, etc.)
  • Experience with Microsoft Fabric for end-to-end data analytics workflows
  • Ability to design and implement robust data architectures including data warehouses, lakehouses, and ETL/ELT frameworks
  • Strong coding and scripting skills with Python
  • Proven problem-solving ability with a knack for debugging complex data issues and optimizing pipeline performance
  • Understanding of data modeling concepts, dimensional modeling, and data governance best practices


Interview Process

  • Take-Home Assessment
  • 60-Minute Technical Interview
  • Culture Fit Round


Preferred Skills & Certifications

  • Microsoft Certified: Fabric Analytics Engineer Associate (DP-600)
  • Microsoft Certified: Fabric Data Engineer Associate (DP-700)
  • Experience with CI/CD practices for data pipelines
  • Familiarity with version control systems such as Git
  • Exposure to real-time streaming data solutions
  • Experience working in Agile or Scrum environments
  • Strong communication skills with the ability to translate technical concepts for non-technical stakeholders

What We Offer

  • Competitive salary and performance-based bonuses
  • Flexible hybrid options
  • Opportunities for professional development, training, and certification sponsorship
  • A collaborative, innovation-driven team culture
  • Paid time off and company holidays
Read more
TVARIT GmbH

at TVARIT GmbH

2 candid answers
DrSoumya Sahadevan
Posted by DrSoumya Sahadevan
Pune
7 - 15 yrs
₹20L - ₹30L / yr
skill iconAmazon Web Services (AWS)
Windows Azure
Google Cloud Platform (GCP)
PySpark
databricks
+2 more

About TVARIT

TVARIT GmbH specializes in developing and delivering cutting-edge artificial intelligence (AI) solutions for the metal industry, including steel, aluminum, copper, cast iron, and more. Our software products empower customers to make intelligent, data-driven decisions, driving advancements in Predictive Quality (PsQ), Predictive Maintenance (PdM), and Energy Consumption Reduction (PsE), etc. With a strong portfolio of renowned reference customers, state-of-the-art technology, a talented research team from prestigious universities, and recognition through esteemed awards such as the EU Horizon 2020 AI Prize, TVARIT is recognized as one of the most innovative AI companies in Germany and Europe. We are seeking a self-motivated individual with a positive "can-do" attitude and excellent oral and written communication skills in English to join our team.


Job Description: We are looking for a Senior Data Engineer with strong expertise in Azure Databricks, PySpark, and distributed computing to develop and optimize scalable ETL pipelines for manufacturing analytics. The role involves working with high-frequency industrial data to enable real-time and batch data processing.


Key Responsibilities · Build scalable real-time and batch processing workflows using Azure Databricks, PySpark, and Apache Spark.

· Perform data pre-processing, including cleaning, transformation, deduplication, normalization, encoding, and scaling to ensure high-quality input for downstream analytics.

· Design and maintain cloud-based data architectures, including data lakes, lakehouses, and warehouses, following Medallion Architecture.

· Deploy and optimize data solutions on Azure (preferred), AWS, or GCP with a focus on performance, security, and scalability.

· Develop and optimize ETL/ELT pipelines for structured and unstructured data from IoT, MES, SCADA, LIMS, and ERP systems. · Automate data workflows using CI/CD and DevOps best practices, ensuring security and compliance with industry standards

· Monitor, troubleshoot, and enhance data pipelines for high availability and reliability.

· Utilize Docker and Kubernetes for scalable data processing.

· Collaborate with automation team, data scientists and engineers to provide clean, structured data for AI/ML models.


Desired Skills and Qualifications · Bachelor’s or Master’s degree in Computer Science, Information Technology, or a related field.

· 7+ years of experience in core data engineering, with a strong focus on cloud platforms such as Azure (preferred), AWS, or GCP · Proficiency in PySpark, Azure Databricks, Python and Apache Spark, etc.

. 2 years of team handling experience.

· Expertise in relational databases (e.g., SQL Server, PostgreSQL), time series databases (e.g. Influx DB), and NoSQL databases (e.g., MongoDB, Cassandra) · Experience in containerization (Docker, Kubernetes).

· Strong analytical and problem-solving skills with attention to detail.

· Good to have MLOps, DevOps including model lifecycle management

· Excellent communication and collaboration skills, with a proven ability to work effectively as a team player.

· Comfortable working in a dynamic, fast-paced startup environment, adapting quickly to changing priorities and responsibilities.

Read more
PhotonMatters
Human Resource
Posted by Human Resource
Remote only
4 - 13 yrs
₹8L - ₹20L / yr
skill iconPython
ETL
Spark
skill iconAmazon Web Services (AWS)
ELT
+2 more

 

 

 

Job Title: Data Engineer

Experience: 4–14 Years

Work Mode: Remote

Employment Type: Full-Time

 

Position Overview:

We are looking for highly experienced Senior Data Engineers to design, architect, and lead scalable, cloud-based data platforms on AWS. The role involves building enterprise-grade data pipelines, modernizing legacy systems, and developing high-performance scoring engines and analytics solutions and collaborate closely with architecture, analytics, risk, and business teams to deliver secure, reliable, and scalable data solutions.

 

Key Responsibilities:

·      Design and build scalable data pipelines for financial and customer data

·      Build and optimize scoring engines (credit, risk, fraud, customer scoring)

·      Design, develop, and optimize complex ETL/ELT pipelines (batch & real-time)

·      Ensure data quality, governance, reliability, and compliance standards

·      Optimize large-scale data processing using SQL, Spark/PySpark, and cloud technologies

·      Lead cloud data architecture, cost optimization, and performance tuning initiatives

·      Collaborate with Data Science, Analytics, and Product teams to deliver business-ready datasets

·      Mentor junior engineers and establish best practices for data engineering

 

Key Requirements:

·      Strong programming skills in Python and advanced SQL

·      Experience building scalable scoring or rule-based decision engines

·      Hands-on experience with Big Data technologies (Spark/PySpark/Kafka)

·      Strong expertise in designing ETL/ELT pipelines and data modeling

·      Experience with cloud platforms (AWS/Azure) and modern data architectures

·      Solid understanding of data warehousing, data lakes, and performance tuning

·      Knowledge of CI/CD, version control (Git), and production support best practices

Read more
Global Digital Transformation Solutions Provider

Global Digital Transformation Solutions Provider

Agency job
via Peak Hire Solutions by Dharati Thakkar
Hyderabad
5 - 7 yrs
₹15L - ₹21L / yr
skill iconPython
Terraform
PySpark
skill iconAmazon Web Services (AWS)

Job Details

Job Title: Lead I - Data Engineering (Python, AWS Glue, Pyspark, Terraform)

Industry: Global digital transformation solutions provider

Domain - Information technology (IT)

Experience Required: 5-7 years

Employment Type: Full Time

Job Location: Hyderabad

CTC Range: Best in Industry

 

Job Description

Data Engineer with AWS, Python, Glue, Terraform, Step function and Spark

 

Skills: Python, AWS Glue, Pyspark, Terraform - All are mandatory

 

******

Notice period - 0 to 15 days only

Job stability is mandatory

Location: Hyderabad 

Read more
Global digital transformation solutions provider.

Global digital transformation solutions provider.

Agency job
via Peak Hire Solutions by Dharati Thakkar
Hyderabad
5 - 8 yrs
₹11L - ₹20L / yr
PySpark
Apache Kafka
Data architecture
skill iconAmazon Web Services (AWS)
EMR
+32 more

JOB DETAILS:

* Job Title: Lead II - Software Engineering - AWS, Apache Spark (PySpark/Scala), Apache Kafka

* Industry: Global digital transformation solutions provider

* Salary: Best in Industry

* Experience: 5-8 years

* Location: Hyderabad

 

Job Summary

We are seeking a skilled Data Engineer to design, build, and optimize scalable data pipelines and cloud-based data platforms. The role involves working with large-scale batch and real-time data processing systems, collaborating with cross-functional teams, and ensuring data reliability, security, and performance across the data lifecycle.


Key Responsibilities

ETL Pipeline Development & Optimization

  • Design, develop, and maintain complex end-to-end ETL pipelines for large-scale data ingestion and processing.
  • Optimize data pipelines for performance, scalability, fault tolerance, and reliability.

Big Data Processing

  • Develop and optimize batch and real-time data processing solutions using Apache Spark (PySpark/Scala) and Apache Kafka.
  • Ensure fault-tolerant, scalable, and high-performance data processing systems.

Cloud Infrastructure Development

  • Build and manage scalable, cloud-native data infrastructure on AWS.
  • Design resilient and cost-efficient data pipelines adaptable to varying data volume and formats.

Real-Time & Batch Data Integration

  • Enable seamless ingestion and processing of real-time streaming and batch data sources (e.g., AWS MSK).
  • Ensure consistency, data quality, and a unified view across multiple data sources and formats.

Data Analysis & Insights

  • Partner with business teams and data scientists to understand data requirements.
  • Perform in-depth data analysis to identify trends, patterns, and anomalies.
  • Deliver high-quality datasets and present actionable insights to stakeholders.

CI/CD & Automation

  • Implement and maintain CI/CD pipelines using Jenkins or similar tools.
  • Automate testing, deployment, and monitoring to ensure smooth production releases.

Data Security & Compliance

  • Collaborate with security teams to ensure compliance with organizational and regulatory standards (e.g., GDPR, HIPAA).
  • Implement data governance practices ensuring data integrity, security, and traceability.

Troubleshooting & Performance Tuning

  • Identify and resolve performance bottlenecks in data pipelines.
  • Apply best practices for monitoring, tuning, and optimizing data ingestion and storage.

Collaboration & Cross-Functional Work

  • Work closely with engineers, data scientists, product managers, and business stakeholders.
  • Participate in agile ceremonies, sprint planning, and architectural discussions.


Skills & Qualifications

Mandatory (Must-Have) Skills

  1. AWS Expertise
  • Hands-on experience with AWS Big Data services such as EMR, Managed Apache Airflow, Glue, S3, DMS, MSK, and EC2.
  • Strong understanding of cloud-native data architectures.
  1. Big Data Technologies
  • Proficiency in PySpark or Scala Spark and SQL for large-scale data transformation and analysis.
  • Experience with Apache Spark and Apache Kafka in production environments.
  1. Data Frameworks
  • Strong knowledge of Spark DataFrames and Datasets.
  1. ETL Pipeline Development
  • Proven experience in building scalable and reliable ETL pipelines for both batch and real-time data processing.
  1. Database Modeling & Data Warehousing
  • Expertise in designing scalable data models for OLAP and OLTP systems.
  1. Data Analysis & Insights
  • Ability to perform complex data analysis and extract actionable business insights.
  • Strong analytical and problem-solving skills with a data-driven mindset.
  1. CI/CD & Automation
  • Basic to intermediate experience with CI/CD pipelines using Jenkins or similar tools.
  • Familiarity with automated testing and deployment workflows.

 

Good-to-Have (Preferred) Skills

  • Knowledge of Java for data processing applications.
  • Experience with NoSQL databases (e.g., DynamoDB, Cassandra, MongoDB).
  • Familiarity with data governance frameworks and compliance tooling.
  • Experience with monitoring and observability tools such as AWS CloudWatch, Splunk, or Dynatrace.
  • Exposure to cost optimization strategies for large-scale cloud data platforms.

 

Skills: big data, scala spark, apache spark, ETL pipeline development

 

******

Notice period - 0 to 15 days only

Job stability is mandatory

Location: Hyderabad

Note: If a candidate is a short joiner, based in Hyderabad, and fits within the approved budget, we will proceed with an offer

F2F Interview: 14th Feb 2026

3 days in office, Hybrid model.

 


Read more
Wissen Technology

at Wissen Technology

4 recruiters
Janane Mohanasankaran
Posted by Janane Mohanasankaran
Mumbai, Pune
3 - 6 yrs
Best in industry
skill iconPython
PySpark
pandas
SQL
ADF
+2 more

* Python (3 to 6 years): Strong expertise in data workflows and automation

* Spark (PySpark): Hands-on experience with large-scale data processing

* Pandas: For detailed data analysis and validation

* Delta Lake: Managing structured and semi-structured datasets at scale

* SQL: Querying and performing operations on Delta tables

* Azure Cloud: Compute and storage services

* Orchestrator: Good experience with either ADF or Airflow

Read more
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Why apply via Cutshort?
Connect with actual hiring teams and get their fast response. No spam.
Find more jobs
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort