Cutshort logo
For Employers
PySpark Jobs in Hyderabad

60 PySpark Jobs in Hyderabad | PySpark Job openings in Hyderabad

Apply to 60+ PySpark Jobs in Hyderabad on CutShort.io. Explore the latest PySpark Job opportunities across top companies like Google, Amazon & Adobe.

icon

Bengaluru (Bangalore), Hyderabad · 5 - 8 years · ₹2L - ₹15L / yr · Profitable · Posted 7 Oct 2026

skill iconData Science
skill iconPython
skill iconMachine Learning (ML)
Statistics/ML Fundamentals
GenAI/LLM
+8 more

Role Overview

We are looking for an experienced Data Scientist – Agentic AI with strong expertise in Python, Machine Learning, Generative AI, Large Language Models (LLMs), RAG and Agentic AI.

The ideal candidate should have hands-on experience in developing, fine-tuning, evaluating and deploying machine learning and GenAI solutions. The candidate should be comfortable working with open-source LLMs, LangChain/LangGraph, PySpark and AI observability/tracing frameworks.

The role involves building intelligent AI systems that can reason, use tools, retrieve information and execute multi-step tasks using Agentic AI architectures.

Mandatory Technical Skills

1. Data Science & Python

  • 5+ years of experience in Data Science / Machine Learning / AI.
  • Strong programming experience in Python.
  • Strong understanding of data analysis, feature engineering and statistical techniques.
  • Experience with Python ML and data science libraries such as:
  • NumPy
  • Pandas
  • Scikit-learn
  • Matplotlib / Seaborn
  • Good understanding of data preprocessing, exploratory data analysis and experimentation.

2. Machine Learning & Statistics

  • Strong understanding of Machine Learning fundamentals.
  • Experience with supervised and unsupervised learning techniques.
  • Knowledge of:
  • Regression
  • Classification
  • Clustering
  • Feature Engineering
  • Model Selection
  • Hyperparameter Tuning
  • Cross-validation
  • Strong understanding of Statistics / ML fundamentals.
  • Ability to interpret model performance and statistical results.

3. Generative AI / LLM

  • Strong hands-on experience with Generative AI and Large Language Models (LLMs).
  • Understanding of Transformer architecture and modern LLM-based applications.
  • Experience working with commercial or open-source LLMs.
  • Strong understanding of:
  • Prompt Engineering
  • Context Management
  • Embeddings
  • Tokenization
  • LLM inference
  • Hallucination mitigation

4. RAG – Retrieval Augmented Generation

  • Strong hands-on experience developing RAG applications.
  • Experience with:
  • Document ingestion
  • Chunking
  • Embeddings
  • Vector search
  • Semantic search
  • Retrieval pipelines
  • Context retrieval
  • Reranking
  • Ability to optimize RAG pipelines for relevance, accuracy and latency.
  • Experience integrating LLMs with enterprise knowledge sources.

5. Agentic AI

  • Hands-on experience building Agentic AI / AI Agent solutions.
  • Understanding of agent architecture and multi-step reasoning workflows.
  • Experience with:
  • AI Agents
  • Multi-Agent systems
  • Tool Calling
  • Function Calling
  • Agent orchestration
  • Planning and reasoning workflows
  • Memory
  • Workflow automation
  • Ability to build agents that can interact with tools, APIs, databases and external systems.

6. LangChain / LangGraph

  • Strong hands-on experience with LangChain and/or LangGraph.
  • Experience building LLM workflows and agent-based applications.
  • Understanding of:
  • Chains
  • Agents
  • Tools
  • State management
  • Graph-based workflows
  • Agent orchestration
  • Retrieval workflows
  • Experience designing scalable Agentic AI workflows.

7. LLM Fine-Tuning

  • Hands-on experience with LLM fine-tuning.
  • Understanding of techniques such as:
  • Supervised Fine-Tuning (SFT)
  • Parameter-Efficient Fine-Tuning
  • LoRA
  • QLoRA
  • Experience preparing datasets for fine-tuning.
  • Ability to evaluate fine-tuned models against baseline models.
  • Understanding of model optimization and inference considerations.

8. BERT / LLaMA / Open-Source LLMs

Experience working with one or more open-source / transformer-based models such as:

  • BERT
  • LLaMA / Llama
  • Mistral
  • Gemma
  • Qwen
  • Other open-source LLMs

Candidate should understand model loading, inference, fine-tuning and evaluation.

9. PySpark

  • Strong experience with PySpark for large-scale data processing.
  • Experience working with large datasets and distributed data processing.
  • Knowledge of:
  • Data transformations
  • Data cleaning
  • Aggregations
  • Joins
  • Spark SQL
  • Performance optimization
  • Ability to build scalable data processing pipelines.

10. Model Validation & Evaluation

  • Experience validating and evaluating ML and GenAI models.
  • Understanding of traditional ML evaluation metrics.
  • Experience evaluating LLM/RAG applications using relevant quality metrics.
  • Ability to compare model performance and identify areas for improvement.
  • Experience with:
  • Accuracy
  • Precision
  • Recall
  • F1 Score
  • ROC-AUC
  • Retrieval metrics
  • LLM response quality
  • Groundedness / relevance
  • Experience designing evaluation datasets and test cases is preferred.

11. AI Tracing / Observability

  • Experience with AI/LLM tracing and observability.
  • Ability to monitor AI applications in production.
  • Experience tracking:
  • LLM requests/responses
  • Latency
  • Token usage
  • Errors
  • Retrieval performance
  • Agent/tool execution
  • Model performance
  • Exposure to tools/frameworks such as LangSmith, OpenTelemetry, Arize Phoenix, MLflow or similar is preferred.

12. Model Deployment

  • Experience deploying ML/LLM/GenAI solutions into production.
  • Exposure to cloud and/or on-premise model deployment.
  • Experience with model serving, APIs and production inference.
  • Knowledge of deployment environments such as:
  • AWS
  • Azure
  • GCP
  • On-premise infrastructure
  • Experience with Docker, APIs and CI/CD is an advantage.

Key Responsibilities

  • Design, develop and deploy Data Science, Machine Learning and GenAI solutions.
  • Build production-ready RAG and Agentic AI applications.
  • Develop intelligent agents capable of tool calling, reasoning and multi-step task execution.
  • Build LLM-powered applications using LangChain/LangGraph.
  • Work with open-source LLMs including BERT, LLaMA and other transformer-based models.
  • Fine-tune LLMs for specific business use cases.
  • Develop scalable data processing pipelines using PySpark.
  • Perform data analysis, feature engineering and statistical modeling.
  • Develop and maintain model validation and evaluation frameworks.
  • Evaluate ML and LLM models using appropriate performance and quality metrics.
  • Implement AI tracing, monitoring and observability for production GenAI systems.
  • Deploy models and AI applications in cloud or on-premise environments.
  • Optimize model performance, response quality, latency and cost.
  • Troubleshoot issues related to model inference, retrieval, agents and LLM workflows.
  • Collaborate with Data Scientists, ML Engineers, Software Engineers and business stakeholders.
  • Convert business requirements into scalable AI/ML solutions.

Good to Have

  • Experience with Vector Databases such as:
  • FAISS
  • Pinecone
  • Weaviate
  • Milvus
  • Chroma
  • Azure AI Search
  • Experience with MLflow or similar ML lifecycle tools.
  • Experience with Docker/Kubernetes.
  • Experience with REST APIs / FastAPI.
  • Knowledge of cloud AI/ML services.
  • Experience with MLOps / LLMOps.
  • Experience with multi-agent frameworks other than LangChain/LangGraph.
  • Experience working with enterprise GenAI applications.

Ideal Candidate Profile

The ideal candidate should be a Data Scientist / ML Engineer with strong GenAI and Agentic AI experience, rather than a pure Python developer.

A strong candidate would typically have:

Data Science + Python + ML + Statistics + GenAI/LLM + RAG + Agentic AI + LangChain/LangGraph + LLM Fine-Tuning + Open-Source LLMs + PySpark + Model Evaluation + AI Observability + Model Deployment.

Core Mandatory Skills

Data Science, Python, Machine Learning, Statistics/ML Fundamentals, GenAI/LLM, RAG, Agentic AI, LangChain/LangGraph, LLM Fine-Tuning, BERT/LLaMA/Open-Source LLMs, PySpark, Model Validation/Evaluation, AI Tracing/Observability, Cloud/On-Prem Model Deployment.

Read more
Browse more Data Science Jobs in Hyderabad | Data Science Job openings in Hyderabad →
VY SYSTEMS PRIVATE LIMITED

Bengaluru (Bangalore), Hyderabad · 5 - 10 years · ₹18L - ₹28L / yr · Profitable · Posted 5 Oct 2026

skill iconPython
PySpark
SQL

1st virtual , 2nd round F2F

Python pyspark, SQL, data engineer

5+yrs

9+yrs

Bangalore/Hyderabad

immediate to 15days.

Read more
Browse more Python Jobs in Hyderabad | Python Job openings in Hyderabad →
VY SYSTEMS PRIVATE LIMITED

Bengaluru (Bangalore), Hyderabad · 9 - 15 years · ₹8L - ₹28L / yr · Profitable · Posted 5 Oct 2026

skill iconPython
PySpark
SQL

Job Description – Azure Data Engineer

Role: Azure Data Engineer

Experience: 9+ Years

Location: Bangalore / Hyderabad

Notice Period: Immediate to 15 Days

Interview Process: 1st Round – Virtual | 2nd Round – F2F

Mandatory Skills

  • Python
  • PySpark
  • SQL
  • Azure Data Engineering

Job Description

We are looking for an experienced Azure Data Engineer with 9+ years of experience and strong hands-on expertise in Python, PySpark, SQL, and Azure Data Engineering.

Key Responsibilities

  • Develop and maintain scalable data engineering solutions using Azure.
  • Build and optimize data processing pipelines using PySpark and Python.
  • Write complex SQL queries for data extraction and transformation.
  • Work with Azure data services and cloud-based data platforms.
  • Perform data processing, transformation, and integration.
  • Troubleshoot data pipeline and production issues.
  • Collaborate with technical and business teams to deliver data solutions.

Preferred: Immediate to 15 Days joiners.

Read more
Browse more Python Jobs in Hyderabad | Python Job openings in Hyderabad →
VY SYSTEMS PRIVATE LIMITED
Santhanalakshmi A
Posted by Santhanalakshmi A

Bengaluru (Bangalore), Hyderabad · 5 - 12 years · ₹8L - ₹16L / yr · Profitable · Posted 3 Oct 2026

skill iconPython
PySpark
SQL

1st virtual , 2nd round F2F

Python pyspark, SQL, data engineer

5+yrs

Bang/hyderabad

immediate to 15days.

Read more
Browse more Python Jobs in Hyderabad | Python Job openings in Hyderabad →
Envisioned Strategy and Consulting

Hyderabad, Chennai, Pune · 6 - 10 years · ₹16L - ₹20L / yr · Profitable · Posted 30 Sep 2026

PySpark
databricks
Snowflake
ETL
ELT
+4 more

Responsibilities and JD

Job Description: We are looking for a Senior Developer with strong expertise in PySpark, Databricks, and Snowflake to build scalable data engineering solutions and enterprise data platforms.

Key Responsibilities:

  1. Design, develop, and maintain ETL/ELT pipelines using PySpark, Databricks, and Snowflake.
  2. Develop batch and real-time data processing solutions for structured and semi-structured data.
  3. Build and optimize Databricks notebooks, workflows, and Delta Lake solutions.
  4. Design and implement Snowflake databases, schemas, views, stored procedures, tasks, and streams.
  5. Develop scalable data models, data marts, and data warehouse solutions.
  6. Optimize PySpark jobs, Databricks workloads, and Snowflake queries for performance and cost efficiency.
  7. Implement data quality, validation, governance, and security controls.
  8. Collaborate with business stakeholders, architects, and cross-functional teams to deliver data solutions.
  9. Manage source control and CI/CD deployments using Git and Azure DevOps.
  10. Troubleshoot production issues, perform root cause analysis, and ensure pipeline reliability.
  11. Mentor junior team members and participate in code reviews and technical design discussions.

 

Required Skills: PySpark, Databricks, Snowflake, Python, SQL.

 

Experience: 5+ years of Data Engineering experience with strong hands-on expertise in PySpark, Databricks, and Snowflake.

Read more
Browse more PySpark Jobs in India →
VY SYSTEMS PRIVATE LIMITED

Hyderabad, Bengaluru (Bangalore) · 9 - 15 years · ₹10L - ₹25L / yr · Profitable · Posted 25 Sep 2026

Google Cloud Platform (GCP)
skill iconPython
PySpark
PL/SQL

Job Summary: GCP Data Engineering Lead

Experience: 9+ Years

Location: Bangalore / Hyderabad

Notice Period: Immediate to 15 Days

Key Skills:

  • Strong experience in GCP Data Engineering
  • Proven Technical Lead / Lead experience
  • Strong programming skills in Python
  • Hands-on experience with PySpark
  • Strong expertise in SQL / PL-SQL
  • Good understanding of GCP data services and data engineering concepts
  • Experience in designing and developing scalable data pipelines
  • Strong problem-solving and technical leadership skills

Roles & Responsibilities:

  • Lead the design and development of scalable GCP data engineering solutions.
  • Develop and optimize data pipelines using Python, PySpark and SQL/PL-SQL.
  • Design data processing solutions and ensure performance and scalability.
  • Provide technical leadership, conduct code reviews, and mentor team members.
  • Collaborate with business and technical teams to understand requirements and deliver data solutions.
  • Troubleshoot issues and ensure quality across the data engineering lifecycle.
Read more
Browse more Google Cloud Platform (GCP) Jobs in Hyderabad | Google Cloud Platform (GCP) Job openings in Hyderabad →
VY SYSTEMS PRIVATE LIMITED

Bengaluru (Bangalore), Hyderabad · 5 - 10 years · ₹7L - ₹18L / yr · Profitable · Posted 24 Sep 2026

skill iconPython
PySpark
SQL

🚨 WE ARE HIRING – SENIOR DATA ENGINEER | BANGALORE 🚨


Looking for experienced Senior Data Engineers with strong expertise in PySpark, Oracle SQL/PLSQL, and Data Modeling!


🔹 Role: Senior Data Engineer

🔹 Experience: 7+ Years

🔹 Location: Bangalore

🔹 CTC: Up to 26 LPA

🔹 Notice Period: Immediate to 10 Days Preferred


💻 Mandatory Skills


✅ PySpark

✅ Oracle SQL / PL/SQL

✅ Data Modeling & Design / Modernization

✅ Python

✅ ETL

✅ Data Pipelines


⚙️ Good to Have / Ecosystem Skills


✅ Kafka

✅ Hadoop

✅ AWS / Azure / GCP

✅ Git

✅ JIRA


📌 Key Responsibilities


• Develop and maintain scalable data pipelines using PySpark

• Work extensively with Oracle SQL/PLSQL for data processing and transformation

• Design and implement data models and data architecture

• Work on data design and modernization initiatives

• Develop and optimize ETL processes and data pipelines

• Handle large volumes of data using PySpark and Python

• Work with technologies such as Kafka, Hadoop, and Cloud platforms

• Collaborate with cross-functional teams to deliver scalable data solutions

• Use Git and JIRA for version control and project tracking


📩 Interested candidates can share their updated CV along with:


Total Experience:

Relevant PySpark Experience:

Relevant Oracle SQL/PLSQL Experience:

Data Modeling Experience:

Current Location:

Notice Period:

Current CTC:

Expected CTC:


#Hiring #DataEngineer #SeniorDataEngineer #PySpark #Oracle #PLSQL #DataModeling #Python #ETL #DataPipelines #Kafka #Hadoop #AWS #Azure #GCP #BangaloreJobs #ITJobs #TechJobs #ImmediateJoiners

Read more
Browse more Python Jobs in Hyderabad | Python Job openings in Hyderabad →
VY SYSTEMS PRIVATE LIMITED

Bengaluru (Bangalore), Hyderabad · 9 - 15 years · ₹8L - ₹24L / yr · Profitable · Posted 22 Sep 2026

GCP
PL/SQL
skill iconPython
PySpark


🚨 Hiring: GCP Data Engineer

We are looking for experienced GCP Data Engineers to join our team!

🔹 Experience: 9+ Years

🔹 Relevant Experience: 4+ Years in GCP Data Engineering

🔹 Required Skills: GCP, Oracle PL/SQL, Python, PySpark

🔹 Location: Bangalore / Hyderabad

🔹 Notice Period: Immediate to 10 Days Preferred

Key Skills:

🔹 Strong hands-on experience in GCP Data Engineering

🔹 Good experience with PySpark & Python

🔹 Strong knowledge of Oracle PL/SQL

🔹 Experience in data processing, ETL, and data pipelines

🔹 Good understanding of cloud-based data engineering

📩 Interested candidates can share their updated CV via DM.

#Hiring #GCPDataEngineer #GCP #DataEngineering #PySpark #Python #OraclePLSQL #DataEngineer #BangaloreJobs #HyderabadJobs #ImmediateJoiner #TechJobs #ITJobs #HiringNow

Read more
Browse more PL/SQL Jobs in Hyderabad | PL/SQL Job openings in Hyderabad →
VY SYSTEMS PRIVATE LIMITED

Bengaluru (Bangalore), Hyderabad · 9 - 15 years · ₹10L - ₹28L / yr · Profitable · Posted 22 Sep 2026

Google Cloud Platform (GCP)
skill iconPython
PySpark
Oracle
skill iconLeadership

GCP Data Engineering Lead

Experience: 9+ Years

Lead Experience: 2+ Years

Location: Bangalore / Hyderabad

Key Skills:

  • Strong experience in GCP Data Engineering and BigQuery
  • Hands-on experience with Oracle Exadata / PL-SQL
  • Strong knowledge of PySpark / Scala and Python
  • Experience with GoldenGate, Kafka and CDC
  • Hands-on experience with Apache Airflow
  • Good experience in CI/CD and DevOps practices
  • Experience with Terraform / Infrastructure as Code
  • Exposure to AI/LLM technologies and GenAI solutions
  • Strong understanding of data architecture, ETL/ELT and data pipelines

Roles & Responsibilities:

  • Lead the design and development of scalable GCP data engineering solutions.
  • Design and implement batch and real-time data pipelines using BigQuery, PySpark, Kafka/CDC and Airflow.
  • Work with Oracle Exadata/PL-SQL and GoldenGate for data integration and migration.
  • Implement CI/CD pipelines and infrastructure automation using Terraform.
  • Explore and integrate AI/LLM capabilities into data engineering solutions.
  • Lead technical discussions, code reviews, solution design and mentor team members.
  • Collaborate with business and technical teams to deliver high-quality data solutions.


Read more
Browse more Google Cloud Platform (GCP) Jobs in Hyderabad | Google Cloud Platform (GCP) Job openings in Hyderabad →
VY SYSTEMS PRIVATE LIMITED

Hyderabad, Bengaluru (Bangalore) · 5 - 9 years · ₹9L - ₹20L / yr · Profitable · Posted 21 Sep 2026

PySpark
SQL
skill iconPython
ETL

Data Engineer Hiring Post


🚨 Hiring: Data Engineer | PySpark + Python + SQL

We are looking for experienced Data Engineers to join our team!

🔹 Experience: 5 to 9 Years

🔹 Locations: Bangalore / Hyderabad

🔹 Interview Process:

• 1st Round – Virtual

• 2nd Round – Face-to-Face (Karat Test)

🔑 Key Skills:

✅ PySpark

✅ SQL

✅ Python

✅ ETL

📩 Interested candidates can share their updated resume.

#Hiring #DataEngineer #PySpark #Python #SQL #ETL #BangaloreJobs #HyderabadJobs #TechHiring #ImmediateHiring

Read more
Browse more PySpark Jobs in India →
VY SYSTEMS PRIVATE LIMITED

Hyderabad, Pune · 5 - 9 years · ₹18L - ₹20L / yr · Profitable · Posted 18 Sep 2026

PySpark
SQL
skill iconPython

Data Engineer Short Hiring Post


🚨 Hiring: Data Engineer

🔹 Experience: 5–9 Years

🔹 Location: Bangalore / Hyderabad

🔹 Skills: PySpark, Python, SQL, ETL, CI/CD, Data Modeling

🔹 Process: L1 Virtual → L2 F2F Karat Test

🔹 F2F: Bangalore / Hyderabad Location

🔹 Positions: Immediate requirement

⚠️ Note: Candidates must be available for F2F Karat immediately after L1.

#Hiring #DataEngineer #PySpark #Python #SQL #BangaloreJobs #HyderabadJobs #Mphasis #ImmediateJoiners

Read more
Browse more PySpark Jobs in India →
ProofofSkill

Hyderabad · 5 - 7 years · ₹15L - ₹20L / yr · Raised funding · Posted 20 Aug 2026

skill iconPython
SQL
PySpark
Data Warehouse (DWH)
Amazon Redshift
+1 more

Location – Hyderabad (Hybrid)

Work Experience – 5 to 7 years

CTC – upto 20 LPA


Roles & Responsibilities:

· We are looking for a Senior Data Engineering who will be majorly responsible for designing, building and maintaining ETL/ ELT pipelines.

· Integration of data from multiple sources or vendors to provide the holistic insights from data.

· You are expected to build and manage Data warehouse solutions, designing data models, creating ETL processes, implementing data quality mechanisms etc.

· Performs EDA (exploratory data analysis) required to troubleshoot data related issues and assist in the resolution of data issues.

· Should have experience in client interaction.

· Experience in mentoring juniors and providing required guidance.

Required Technical Skills

 

· Extensive hands on experience in Python, Pyspark, SQL, Dataiku.

· Strong experience in Data Warehouse, ETL, Data Modelling, building ETL Pipelines, Snowflake database.

· Working knowledge in Databricks, Redshift, ADF etc.

· Hands-on experience in cloud services like Azure, AWS- S3, Glue, Lambda, CloudWatch, Athena.

· Sound knowledge in end-to-end Data management, Data ops, quality and data governance.

· Familiar with SFDC, Waterfall/ Agile methodology.

· Strong domain knowledge in Pharma domain/ life sciences commercial data operations.

 

Qualifications

 

· Bachelor’s or master’s Engineering/ MCA or equivalent degree.

· 5-7 years of relevant industry experience as Data Engineer.

· Experience working on Pharma syndicated data such as IQVIA, Veeva, Symphony; Claims, CRM, Sales etc.

· High motivation, good work ethic, maturity, self-organized and personal initiative.

· Ability to work collaboratively and providing the support to the team.

· Excellent written and verbal communication skills.

· Strong analytical and problem-solving skills. 

Read more
Browse more Python Jobs in Hyderabad | Python Job openings in Hyderabad →
TalentXO

Bengaluru (Bangalore), Mumbai, Pune, Hyderabad, Noida, Kolkata · 8 - 15 years · ₹13L - ₹20L / yr · Profitable · Posted 10 Aug 2026

Azure Data Factory
Azure Databricks
PySpark
skill iconPython
SQL
+4 more

Roles & Responsibilities

  • Design, develop, and deliver scalable end-to-end data pipelines using Azure Data Factory, ensuring robust integration

of enterprise-wide data from diverse sources

• Build and optimize data engineering workflows using Databricks and PySpark

• Write efficient, high-performance SQL for data transformation and analysis

• Work with the Azure Cloud platform and associated services, applying strong understanding of data warehousing,

data models, and pipelines

• Provide technical leadership to a team of developers, including code reviews and enforcing best practices across the

development lifecycle

• Oversee CI/CD implementation using Azure DevOps, managing deployments across development, QA, and production

environments with proper change control processes

• Collaborate with cross-functional teams to translate business requirements into scalable data solutions

• Ensure data quality, reliability, and performance across all pipelines and platforms

Ideal Candidate

1Strong Azure Databricks Engineer / Senior Data Engineer Profile

2Mandatory (Experience 1) – Must have minimum 8+ years of overall experience in Data Engineering, Data Development, or related data technology roles, with strong hands-on experience in enterprise data pipeline development.

3Mandatory (Experience 2) – Must have strong hands-on experience with Azure Databricks, including development and optimization of scalable data engineering workflows using Databricks and PySpark.

4Mandatory (Experience 3) – Must have strong hands-on proficiency in PySpark/Python and SQL, with proven experience developing complex data transformations, processing workflows, and performance-optimized queries.

5Mandatory (Experience 4) – Must have hands-on experience with Azure Data Factory (ADF) for designing, developing, and orchestrating end-to-end data pipelines and integrating data from multiple sources.

6Mandatory (Experience 5) – Must have strong experience working on the Azure Cloud platform and associated data services, with solid understanding of data warehousing, data modeling, pipeline architecture, and enterprise data solutions.

7Mandatory (Experience 6) – Must have hands-on experience implementing CI/CD using Azure DevOps, including deployment and release management across development, QA, and production environments.

8Mandatory (Experience 7) – Must have proven technical leadership experience, including code reviews, enforcing development best practices, mentoring developers, and providing technical guidance to a data engineering team.

9Mandatory (Notice Period) – Immediate joiners or candidates who can join within 15 days.

10Mandatory (Note) - The position is open across all Cognizant offices pan India. Candidates must be willing to attend the F2F interview at the nearest Cognizant office location.

Read more
Browse more PySpark Jobs in Mumbai | PySpark Job openings in Mumbai →
Codnatives
Agency job
via VY SYSTEMS PRIVATE LIMITED by Ajeethkumar s

Hyderabad, Bengaluru (Bangalore) · 5 - 10 years · ₹4L - ₹16L / yr · Profitable · Posted 8 Aug 2026

skill iconPython
ETL
PySpark
Data engineering
skill iconAmazon Web Services (AWS)
+2 more

Skills Referential (Required knowledge, skills and abilities)

Technical Skills:

Python

Pyspark

SQL

ETL Aws, Azure, gcp

Read more
Browse more Backend Developer Jobs in Hyderabad | Backend Developer Job openings in Hyderabad →
MNC

MNC

Agency job
via VY SYSTEMS PRIVATE LIMITED by aafia parveen

Hyderabad · 5 - 8 years · ₹2L - ₹20L / yr · Posted 7 Aug 2026

Data engineering
Google Cloud Platform (GCP)
Oracle
PySpark
ETL

Job Title: Data Engineer – PySpark | Oracle | GCP


Experience: 5–7 Years

Location: Hyderabad

Notice Period: Immediate Joiners Preferred


Job Summary

We are seeking an experienced Data Engineer with strong expertise in PySpark, Oracle, and Google Cloud Platform (GCP) to design, develop, and optimize scalable data pipelines. The ideal candidate should have hands-on experience in ETL development, data integration, and cloud-based data engineering solutions.

Key Responsibilities


  • Design, develop, and maintain scalable ETL/data pipelines using PySpark.
  • Extract, transform, and load data from Oracle databases into GCP environments.
  • Build and optimize batch data processing workflows for high performance and reliability.
  • Develop data engineering solutions using GCP services.
  • Ensure data quality through validation, monitoring, and troubleshooting.
  • Optimize SQL queries and ETL jobs for performance and scalability.


Required Skills

  • 5–7 years of experience as a Data Engineer.
  • Strong hands-on experience with PySpark.
  • Solid experience with Oracle Database and advanced SQL.
  • Hands-on experience with Google Cloud Platform (GCP).
  • Strong understanding of ETL processes and data warehousing concepts.


Work Location: Hyderabad

Notice Period: Immediate Joiners Preferred

Read more
Browse more Data engineering Jobs in Hyderabad | Data engineering Job openings in Hyderabad →
VY SYSTEMS PRIVATE LIMITED

Hyderabad · 5 - 12 years · ₹4L - ₹18L / yr · Profitable · Posted 18 Jul 2026

Google Cloud Platform (GCP)
PySpark

Job Summary

We are seeking a highly skilled GCP Data Engineer with strong expertise in Google Cloud Platform (GCP), Python, ETL, and modern data engineering technologies. The ideal candidate should have hands-on experience designing and building scalable data pipelines using BigQuery, Dataflow, Pub/Sub, Airflow, and modern data lake technologies such as Apache Iceberg or Delta Lake.

Key Responsibilities

  • Design, develop, and maintain scalable ETL/ELT data pipelines on Google Cloud Platform.
  • Build and optimize data processing workflows using Python and Google Cloud Dataflow (Apache Beam).
  • Develop and manage large-scale analytical data models in BigQuery.
  • Implement event-driven data ingestion using Google Cloud Pub/Sub.
  • Create, schedule, and monitor workflows using Apache Airflow and Autosys.
  • Design and implement modern data lake architectures using Apache Iceberg or Delta Lake.
  • Optimize query performance, storage, and compute costs in GCP.
  • Ensure data quality, governance, security, and compliance across data platforms.
  • Collaborate with Data Scientists, Analysts, and Application teams to deliver scalable data solutions.
  • Troubleshoot production issues and continuously improve pipeline reliability and performance.

Mandatory Skills

  • Strong hands-on experience with Google Cloud Platform (GCP).
  • Proficiency in Python programming.
  • Experience in designing and implementing ETL/ELT pipelines.
  • Strong knowledge of BigQuery.
  • Experience with Google Cloud Dataflow (Apache Beam).
  • Experience with Google Cloud Pub/Sub.
  • Hands-on experience with Apache Airflow.
  • Experience in job scheduling using Autosys.
  • Experience with modern table formats such as Apache Iceberg or Delta Lake.
  • Strong SQL and data modeling skills.

Preferred Skills

  • Experience with Cloud Storage, Dataproc, Cloud Composer, and Cloud Functions.
  • Knowledge of CI/CD pipelines and DevOps practices.
  • Experience with Docker and Kubernetes.
  • Familiarity with Git and Agile/Scrum methodologies.
  • Knowledge of data warehousing and dimensional modeling.
  • Exposure to streaming and real-time data processing.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
  • 4–8+ years of experience in Data Engineering with hands-on expertise in GCP technologies.

Required Experience

  • Strong experience in developing enterprise-grade data pipelines using Python and GCP.
  • Hands-on experience with BigQuery, Dataflow, Pub/Sub, and Airflow.
  • Experience scheduling and monitoring batch workflows using Autosys.
  • Experience implementing modern data lake architectures using Apache Iceberg or Delta Lake.
  • Strong understanding of ETL best practices, performance tuning, and data optimization.
  • Excellent analytical, troubleshooting, and problem-solving skills.

Mandatory Skills

  • Google Cloud Platform (GCP)
  • Python
  • ETL
  • BigQuery
  • Autosys
  • Apache Airflow
  • Google Cloud Pub/Sub
  • Google Cloud Dataflow (Apache Beam)
  • Apache Iceberg / Delta Lake
  • SQL & Data Modeling
Read more
Browse more Google Cloud Platform (GCP) Jobs in Hyderabad | Google Cloud Platform (GCP) Job openings in Hyderabad →
AI Industry

AI Industry

Agency job
via Peak Hire Solutions by Dharati Thakkar

Mumbai, Bengaluru (Bangalore), Hyderabad, Gurugram · 6 - 10 years · ₹32L - ₹42L / yr · Posted 26 Mar 2026

ETL
SQL
Google Cloud Platform (GCP)
Data engineering
ELT
+17 more

Role & Responsibilities:

We are looking for a strong Data Engineer to join our growing team. The ideal candidate brings solid ETL fundamentals, hands-on pipeline experience, and cloud platform proficiency — with a preference for GCP / BigQuery expertise.


Responsibilities:

  • Design, build, and maintain scalable data pipelines and ETL/ELT workflows
  • Work with Dataform or DBT to implement transformation logic and data models
  • Develop and optimize data solutions on GCP (BigQuery, GCS) or AWS/Azure
  • Support data migration initiatives and data mesh architecture patterns
  • Collaborate with analysts, scientists, and business stakeholders to deliver reliable data products
  • Apply data governance and quality best practices across the data lifecycle
  • Troubleshoot pipeline issues and drive proactive monitoring and resolution


Ideal Candidate:

  • Strong Data Engineer Profile
  • Must have 6+ years of hands-on experience in Data Engineering, with strong ownership of end-to-end data pipeline development.
  • Must have strong experience in ETL/ELT pipeline design, transformation logic, and data workflow orchestration.
  • Must have hands-on experience with any one of the following: Dataform, dbt, or BigQuery, with practical exposure to data transformation, modeling, or cloud data warehousing.
  • Must have working experience on any cloud platform: GCP (preferred), AWS, or Azure, including object storage (GCS, S3, ADLS).
  • Must have strong SQL skills with experience in writing complex queries and optimizing performance.
  • Must have programming experience in Python and/or SQL for data processing.
  • Must have experience in building and maintaining scalable data pipelines and troubleshooting data issues.
  • Exposure to data migration projects and/or data mesh architecture concepts.
  • Experience with Spark / PySpark or large-scale data processing frameworks.
  • Experience working in product-based companies or data-driven environments.
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.


NOTE:

  • There will be an interview drive scheduled on 28th and 29th March 2026, and if shortlisted, they will be expected to be available on these Interview dates. Only Immediate joiners are considered.
Read more
Browse more ETL Jobs in Mumbai | ETL Job openings in Mumbai →
Global Digital Transformation Solutions Provider

Global Digital Transformation Solutions Provider

Agency job
via Peak Hire Solutions by Dharati Thakkar

Hyderabad · 5 - 7 years · ₹15L - ₹21L / yr · Posted 21 Feb 2026

skill iconPython
Terraform
PySpark
skill iconAmazon Web Services (AWS)

Job Details

- Job Title: Lead I - Data Engineering (Python, AWS Glue, Pyspark, Terraform)

- Industry: Global digital transformation solutions provider

- Domain - Information technology (IT)

- Experience Required: 5-7 years

- Employment Type: Full Time

- Job Location: Hyderabad

- CTC Range: Best in Industry

 

Job Description

Data Engineer with AWS, Python, Glue, Terraform, Step function and Spark

 

Skills: Python, AWS Glue, Pyspark, Terraform - All are mandatory

 

******

Notice period - 0 to 15 days only

Job stability is mandatory

Location: Hyderabad 

Read more
Browse more Python Jobs in Hyderabad | Python Job openings in Hyderabad →
Global digital transformation solutions provider.

Global digital transformation solutions provider.

Agency job
via Peak Hire Solutions by Dharati Thakkar

Hyderabad · 5 - 8 years · ₹11L - ₹20L / yr · Posted 13 Feb 2026

PySpark
Apache Kafka
Data architecture
skill iconAmazon Web Services (AWS)
EMR
+32 more

JOB DETAILS:

* Job Title: Lead II - Software Engineering - AWS, Apache Spark (PySpark/Scala), Apache Kafka

* Industry: Global digital transformation solutions provider

* Salary: Best in Industry

* Experience: 5-8 years

* Location: Hyderabad

 

Job Summary

We are seeking a skilled Data Engineer to design, build, and optimize scalable data pipelines and cloud-based data platforms. The role involves working with large-scale batch and real-time data processing systems, collaborating with cross-functional teams, and ensuring data reliability, security, and performance across the data lifecycle.


Key Responsibilities

ETL Pipeline Development & Optimization

  • Design, develop, and maintain complex end-to-end ETL pipelines for large-scale data ingestion and processing.
  • Optimize data pipelines for performance, scalability, fault tolerance, and reliability.

Big Data Processing

  • Develop and optimize batch and real-time data processing solutions using Apache Spark (PySpark/Scala) and Apache Kafka.
  • Ensure fault-tolerant, scalable, and high-performance data processing systems.

Cloud Infrastructure Development

  • Build and manage scalable, cloud-native data infrastructure on AWS.
  • Design resilient and cost-efficient data pipelines adaptable to varying data volume and formats.

Real-Time & Batch Data Integration

  • Enable seamless ingestion and processing of real-time streaming and batch data sources (e.g., AWS MSK).
  • Ensure consistency, data quality, and a unified view across multiple data sources and formats.

Data Analysis & Insights

  • Partner with business teams and data scientists to understand data requirements.
  • Perform in-depth data analysis to identify trends, patterns, and anomalies.
  • Deliver high-quality datasets and present actionable insights to stakeholders.

CI/CD & Automation

  • Implement and maintain CI/CD pipelines using Jenkins or similar tools.
  • Automate testing, deployment, and monitoring to ensure smooth production releases.

Data Security & Compliance

  • Collaborate with security teams to ensure compliance with organizational and regulatory standards (e.g., GDPR, HIPAA).
  • Implement data governance practices ensuring data integrity, security, and traceability.

Troubleshooting & Performance Tuning

  • Identify and resolve performance bottlenecks in data pipelines.
  • Apply best practices for monitoring, tuning, and optimizing data ingestion and storage.

Collaboration & Cross-Functional Work

  • Work closely with engineers, data scientists, product managers, and business stakeholders.
  • Participate in agile ceremonies, sprint planning, and architectural discussions.


Skills & Qualifications

Mandatory (Must-Have) Skills

  1. AWS Expertise
  • Hands-on experience with AWS Big Data services such as EMR, Managed Apache Airflow, Glue, S3, DMS, MSK, and EC2.
  • Strong understanding of cloud-native data architectures.
  1. Big Data Technologies
  • Proficiency in PySpark or Scala Spark and SQL for large-scale data transformation and analysis.
  • Experience with Apache Spark and Apache Kafka in production environments.
  1. Data Frameworks
  • Strong knowledge of Spark DataFrames and Datasets.
  1. ETL Pipeline Development
  • Proven experience in building scalable and reliable ETL pipelines for both batch and real-time data processing.
  1. Database Modeling & Data Warehousing
  • Expertise in designing scalable data models for OLAP and OLTP systems.
  1. Data Analysis & Insights
  • Ability to perform complex data analysis and extract actionable business insights.
  • Strong analytical and problem-solving skills with a data-driven mindset.
  1. CI/CD & Automation
  • Basic to intermediate experience with CI/CD pipelines using Jenkins or similar tools.
  • Familiarity with automated testing and deployment workflows.

 

Good-to-Have (Preferred) Skills

  • Knowledge of Java for data processing applications.
  • Experience with NoSQL databases (e.g., DynamoDB, Cassandra, MongoDB).
  • Familiarity with data governance frameworks and compliance tooling.
  • Experience with monitoring and observability tools such as AWS CloudWatch, Splunk, or Dynatrace.
  • Exposure to cost optimization strategies for large-scale cloud data platforms.

 

Skills: big data, scala spark, apache spark, ETL pipeline development

 

******

Notice period - 0 to 15 days only

Job stability is mandatory

Location: Hyderabad

Note: If a candidate is a short joiner, based in Hyderabad, and fits within the approved budget, we will proceed with an offer

F2F Interview: 14th Feb 2026

3 days in office, Hybrid model.

 


Read more
Browse more PySpark Jobs in India →
AI-First Company

AI-First Company

Agency job
via Peak Hire Solutions by Dharati Thakkar

Bengaluru (Bangalore), Mumbai, Hyderabad, Gurugram · 5 - 17 years · ₹30L - ₹45L / yr · Posted 20 Dec 2025

Data engineering
Data architecture
SQL
Data modeling
GCS
+47 more

ROLES AND RESPONSIBILITIES:

You will be responsible for architecting, implementing, and optimizing Dremio-based data Lakehouse environments integrated with cloud storage, BI, and data engineering ecosystems. The role requires a strong balance of architecture design, data modeling, query optimization, and governance enablement in large-scale analytical environments.


  • Design and implement Dremio lakehouse architecture on cloud (AWS/Azure/Snowflake/Databricks ecosystem).
  • Define data ingestion, curation, and semantic modeling strategies to support analytics and AI workloads.
  • Optimize Dremio reflections, caching, and query performance for diverse data consumption patterns.
  • Collaborate with data engineering teams to integrate data sources via APIs, JDBC, Delta/Parquet, and object storage layers (S3/ADLS).
  • Establish best practices for data security, lineage, and access control aligned with enterprise governance policies.
  • Support self-service analytics by enabling governed data products and semantic layers.
  • Develop reusable design patterns, documentation, and standards for Dremio deployment, monitoring, and scaling.
  • Work closely with BI and data science teams to ensure fast, reliable, and well-modeled access to enterprise data.


IDEAL CANDIDATE:

  • Bachelor’s or Master’s in Computer Science, Information Systems, or related field.
  • 5+ years in data architecture and engineering, with 3+ years in Dremio or modern lakehouse platforms.
  • Strong expertise in SQL optimization, data modeling, and performance tuning within Dremio or similar query engines (Presto, Trino, Athena).
  • Hands-on experience with cloud storage (S3, ADLS, GCS), Parquet/Delta/Iceberg formats, and distributed query planning.
  • Knowledge of data integration tools and pipelines (Airflow, DBT, Kafka, Spark, etc.).
  • Familiarity with enterprise data governance, metadata management, and role-based access control (RBAC).
  • Excellent problem-solving, documentation, and stakeholder communication skills.


PREFERRED:

  • Experience integrating Dremio with BI tools (Tableau, Power BI, Looker) and data catalogs (Collibra, Alation, Purview).
  • Exposure to Snowflake, Databricks, or BigQuery environments.
  • Experience in high-tech, manufacturing, or enterprise data modernization programs.
Read more
Browse more Data architecture Jobs in Mumbai | Data architecture Job openings in Mumbai →
One of the reputed Client in India

One of the reputed Client in India

Bengaluru (Bangalore), Mumbai, Delhi, Gurugram, Noida, Hyderabad, Pune · 6 - 8 years · ₹12L - ₹13L / yr · Posted 15 Oct 2025

skill iconAmazon Web Services (AWS)
skill iconPython
PySpark

Our Client is looking to hire Databricks Amin immediatly.


This is PAN-INDIA Bulk hiring


Minimum of 6-8+ years with Databricks, Pyspark/Python and AWS.

Must have AWS


Notice 15-30 days is preferred.


Share profiles at hr at etpspl dot com

Please refer/share our email to your friends/colleagues who are looking for job.

Read more
Browse more Python Jobs in Mumbai | Python Job openings in Mumbai →
Tata Consultancy Services

Chennai, Hyderabad, Kolkata, Delhi, Pune, Bengaluru (Bangalore) · 4 - 10 years · ₹6L - ₹30L / yr · Profitable · Posted 30 Sep 2025

Scala
PySpark
Spark
skill iconAmazon Web Services (AWS)

Job Title: PySpark/Scala Developer

 

Functional Skills: Experience in Credit Risk/Regulatory risk domain

Technical Skills: Spark ,PySpark, Python, Hive, Scala, MapReduce, Unix shell scripting

Good to Have Skills: Exposure to Machine Learning Techniques

Job Description:

5+ Years of experience with Developing/Fine tuning and implementing programs/applications

Using Python/PySpark/Scala on Big Data/Hadoop Platform.

Roles and Responsibilities:

a)     Work with a Leading Bank’s Risk Management team on specific projects/requirements pertaining to risk Models in

 consumer and wholesale banking

b)     Enhance Machine Learning Models using PySpark or Scala

c)     Work with Data Scientists to Build ML Models based on Business Requirements and Follow ML Cycle to Deploy them all

the way to Production Environment

d)     Participate Feature Engineering, Training Models, Scoring and retraining

e)     Architect Data Pipeline and Automate Data Ingestion and Model Jobs

 

Skills and competencies:

Required:

·       Strong analytical skills in conducting sophisticated statistical analysis using bureau/vendor data, customer performance

Data and macro-economic data to solve business problems.

·       Working experience in languages PySpark & Scala to develop code to validate and implement models and codes in

Credit Risk/Banking

·       Experience with distributed systems such as Hadoop/MapReduce, Spark, streaming data processing, cloud architecture.

  • Familiarity with machine learning frameworks and libraries (like scikit-learn, SparkML, tensorflow, pytorch etc.
  • Experience in systems integration, web services, batch processing
  • Experience in migrating codes to PySpark/Scala is big Plus
  • The ability to act as liaison conveying information needs of the business to IT and data constraints to the business

applies equal conveyance regarding business strategy and IT strategy, business processes and work flow

·       Flexibility in approach and thought process

·       Attitude to learn and comprehend the periodical changes in the regulatory requirement as per FED

 

 

Read more
Browse more PySpark Jobs in Chennai | PySpark Job openings in Chennai →
Tata Consultancy Services

Bengaluru (Bangalore), Hyderabad, Pune, Delhi, Kolkata, Chennai · 5 - 8 years · ₹7L - ₹30L / yr · Profitable · Posted 30 Sep 2025

skill iconScala
skill iconPython
PySpark
Apache Hive
Spark
+3 more

Skills and competencies:

Required:

·        Strong analytical skills in conducting sophisticated statistical analysis using bureau/vendor data, customer performance

Data and macro-economic data to solve business problems.

·        Working experience in languages PySpark & Scala to develop code to validate and implement models and codes in

Credit Risk/Banking

·        Experience with distributed systems such as Hadoop/MapReduce, Spark, streaming data processing, cloud architecture.

  • Familiarity with machine learning frameworks and libraries (like scikit-learn, SparkML, tensorflow, pytorch etc.
  • Experience in systems integration, web services, batch processing
  • Experience in migrating codes to PySpark/Scala is big Plus
  • The ability to act as liaison conveying information needs of the business to IT and data constraints to the business

applies equal conveyance regarding business strategy and IT strategy, business processes and work flow

·        Flexibility in approach and thought process

·        Attitude to learn and comprehend the periodical changes in the regulatory requirement as per FED

Read more
Browse more Scala Jobs in Hyderabad | Scala Job openings in Hyderabad →
VyTCDC
Gobinath Sundaram
Posted by Gobinath Sundaram

Chennai, Bengaluru (Bangalore), Hyderabad, Mumbai, Pune, Noida · 4 - 6 years · ₹3L - ₹21L / yr · Profitable · Posted 13 Aug 2025

AWS Data Engineer
skill iconAmazon Web Services (AWS)
skill iconPython
PySpark
databricks
+1 more

 Key Responsibilities

  • Design and implement ETL/ELT pipelines using Databricks, PySpark, and AWS Glue
  • Develop and maintain scalable data architectures on AWS (S3, EMR, Lambda, Redshift, RDS)
  • Perform data wrangling, cleansing, and transformation using Python and SQL
  • Collaborate with data scientists to integrate Generative AI models into analytics workflows
  • Build dashboards and reports to visualize insights using tools like Power BI or Tableau
  • Ensure data quality, governance, and security across all data assets
  • Optimize performance of data pipelines and troubleshoot bottlenecks
  • Work closely with stakeholders to understand data requirements and deliver actionable insights

🧪 Required Skills

Skill AreaTools & TechnologiesCloud PlatformsAWS (S3, Lambda, Glue, EMR, Redshift)Big DataDatabricks, Apache Spark, PySparkProgrammingPython, SQLData EngineeringETL/ELT, Data Lakes, Data WarehousingAnalyticsData Modeling, Visualization, BI ReportingGen AI IntegrationOpenAI, Hugging Face, LangChain (preferred)DevOps (Bonus)Git, Jenkins, Terraform, Docker

📚 Qualifications

  • Bachelor's or Master’s degree in Computer Science, Data Science, or related field
  • 3+ years of experience in data engineering or data analytics
  • Hands-on experience with Databricks, PySpark, and AWS
  • Familiarity with Generative AI tools and frameworks is a strong plus
  • Strong problem-solving and communication skills

🌟 Preferred Traits

  • Analytical mindset with attention to detail
  • Passion for data and emerging technologies
  • Ability to work independently and in cross-functional teams
  • Eagerness to learn and adapt in a fast-paced environment


Read more
Browse more AWS (Amazon Web Services) Jobs in Chennai | AWS (Amazon Web Services) Job openings in Chennai →
Tekit Software solution Pvt Ltd
himanshi Tripathi
Posted by himanshi Tripathi

Hyderabad, Bengaluru (Bangalore) · 8 - 10 years · ₹15L - ₹27L / yr · Raised funding · Posted 24 Jul 2025

skill iconAmazon Web Services (AWS)
skill iconPython
PySpark
SQL

🔍 Job Description:

We are looking for an experienced and highly skilled Technical Lead to guide the development and enhancement of a large-scale Data Observability solution built on AWS. This platform is pivotal in delivering monitoring, reporting, and actionable insights across the client's data landscape.

The Technical Lead will drive end-to-end feature delivery, mentor junior engineers, and uphold engineering best practices. The position reports to the Programme Technical Lead / Architect and involves close collaboration to align on platform vision, technical priorities, and success KPIs.

🎯 Key Responsibilities:

  • Lead the design, development, and delivery of features for the data observability solution.
  • Mentor and guide junior engineers, promoting technical growth and engineering excellence.
  • Collaborate with the architect to align on platform roadmap, vision, and success metrics.
  • Ensure high quality, scalability, and performance in data engineering solutions.
  • Contribute to code reviews, architecture discussions, and operational readiness.


🔧 Primary Must-Have Skills (Non-Negotiable):

  • 5+ years in Data Engineering or Software Engineering roles.
  • 3+ years in a technical team or squad leadership capacity.
  • Deep expertise in AWS Data Services: Glue, EMR, Kinesis, Lambda, Athena, S3.
  • Advanced programming experience with PySpark, Python, and SQL.
  • Proven experience in building scalable, production-grade data pipelines on cloud platforms.


Read more
Browse more AWS (Amazon Web Services) Jobs in Hyderabad | AWS (Amazon Web Services) Job openings in Hyderabad →
ZeMoSo Technologies

at ZeMoSo Technologies

11 recruiters
Agency job
via TIGI HR Solution Pvt. Ltd. by Vaidehi Sarkar

Mumbai, Bengaluru (Bangalore), Hyderabad, Chennai, Pune · 4 - 8 years · ₹10L - ₹15L / yr · Profitable · Posted 25 Apr 2025

Data engineering
skill iconPython
SQL
Data Warehouse (DWH)
skill iconAmazon Web Services (AWS)
+3 more

Work Mode: Hybrid


Need B.Tech, BE, M.Tech, ME candidates - Mandatory



Must-Have Skills:

● Educational Qualification :- B.Tech, BE, M.Tech, ME in any field.

● Minimum of 3 years of proven experience as a Data Engineer.

● Strong proficiency in Python programming language and SQL.

● Experience in DataBricks and setting up and managing data pipelines, data warehouses/lakes.

● Good comprehension and critical thinking skills.


● Kindly note Salary bracket will vary according to the exp. of the candidate - 

- Experience from 4 yrs to 6 yrs - Salary upto 22 LPA

- Experience from 5 yrs to 8 yrs - Salary upto 30 LPA

- Experience more than 8 yrs - Salary upto 40 LPA

Read more
Browse more Data engineering Jobs in Mumbai | Data engineering Job openings in Mumbai →
Deqode

at Deqode

1 recruiter
Alisha Das
Posted by Alisha Das

Bengaluru (Bangalore), Delhi, Gurugram, Noida, Ghaziabad, Faridabad, Mumbai, Pune, Hyderabad, Indore, Jaipur, Kolkata · 4 - 5 years · ₹2L - ₹18L / yr · Bootstrapped · Posted 22 Apr 2025

skill iconPython
PySpark

We are looking for a skilled and passionate Data Engineers with a strong foundation in Python programming and hands-on experience working with APIs, AWS cloud, and modern development practices. The ideal candidate will have a keen interest in building scalable backend systems and working with big data tools like PySpark.

Key Responsibilities:

  • Write clean, scalable, and efficient Python code.
  • Work with Python frameworks such as PySpark for data processing.
  • Design, develop, update, and maintain APIs (RESTful).
  • Deploy and manage code using GitHub CI/CD pipelines.
  • Collaborate with cross-functional teams to define, design, and ship new features.
  • Work on AWS cloud services for application deployment and infrastructure.
  • Basic database design and interaction with MySQL or DynamoDB.
  • Debugging and troubleshooting application issues and performance bottlenecks.

Required Skills & Qualifications:

  • 4+ years of hands-on experience with Python development.
  • Proficient in Python basics with a strong problem-solving approach.
  • Experience with AWS Cloud services (EC2, Lambda, S3, etc.).
  • Good understanding of API development and integration.
  • Knowledge of GitHub and CI/CD workflows.
  • Experience in working with PySpark or similar big data frameworks.
  • Basic knowledge of MySQL or DynamoDB.
  • Excellent communication skills and a team-oriented mindset.

Nice to Have:

  • Experience in containerization (Docker/Kubernetes).
  • Familiarity with Agile/Scrum methodologies.


Read more
Browse more Python Jobs in Mumbai | Python Job openings in Mumbai →
Xebia IT Architects

at Xebia IT Architects

2 recruiters
Vijay S
Posted by Vijay S

Bengaluru (Bangalore), Gurugram, Pune, Hyderabad, Chennai, Bhopal, Jaipur · 10 - 15 years · ₹30L - ₹40L / yr · Profitable · Posted 4 Apr 2025

Spark
Google Cloud Platform (GCP)
skill iconPython
Apache Airflow
PySpark
+1 more

We are looking for a Senior Data Engineer with strong expertise in GCP, Databricks, and Airflow to design and implement a GCP Cloud Native Data Processing Framework. The ideal candidate will work on building scalable data pipelines and help migrate existing workloads to a modern framework.


  • Shift: 2 PM 11 PM
  • Work Mode: Hybrid (3 days a week) across Xebia locations
  • Notice Period: Immediate joiners or those with a notice period of up to 30 days


Key Responsibilities:

  • Design and implement a GCP Native Data Processing Framework leveraging Spark and GCP Cloud Services.
  • Develop and maintain data pipelines using Databricks and Airflow for transforming Raw → Silver → Gold data layers.
  • Ensure data integrity, consistency, and availability across all systems.
  • Collaborate with data engineers, analysts, and stakeholders to optimize performance.
  • Document standards and best practices for data engineering workflows.

Required Experience:


  • 7-8 years of experience in data engineering, architecture, and pipeline development.
  • Strong knowledge of GCP, Databricks, PySpark, and BigQuery.
  • Experience with Orchestration tools like Airflow, Dagster, or GCP equivalents.
  • Understanding of Data Lake table formats (Delta, Iceberg, etc.).
  • Proficiency in Python for scripting and automation.
  • Strong problem-solving skills and collaborative mindset.


⚠️ Please apply only if you have not applied recently or are not currently in the interview process for any open roles at Xebia.


Looking forward to your response!


Best regards,

Vijay S

Assistant Manager - TAG

https://www.linkedin.com/in/vijay-selvarajan/

Read more
Browse more Spark Jobs in Pune | Spark Job openings in Pune →
Indigrators solutions
Afzal Mohammed
Posted by Afzal Mohammed

Hyderabad · 5 - 8 years · ₹18L - ₹24L / yr · Profitable · Posted 3 Jan 2025

skill iconPython
PySpark
Palantir Foundry
Palantir
Foundry

Job Description


Job Title: Data Engineer

Location: Hyderabad, India

Job Type: Full Time

Experience: 5 – 8 Years

Working Model: On-Site (No remote or work-from-home options available)

Work Schedule: Mountain Time Zone (3:00 PM to 11:00 PM IST)

Role Overview

The Data Engineer will be responsible for designing and implementing scalable backend systems, leveraging Python and PySpark to build high-performance solutions. The role requires a proactive and detail-orientated individual who can solve complex data engineering challenges while collaborating with cross-functional teams to deliver quality results.

Key Responsibilities

  • Develop and maintain backend systems using Python and PySpark.
  • Optimise and enhance system performance for large-scale data processing.
  • Collaborate with cross-functional teams to define requirements and deliver solutions.
  • Debug, troubleshoot, and resolve system issues and bottlenecks.
  • Follow coding best practices to ensure code quality and maintainability.
  • Utilise tools like Palantir Foundry for data management workflows (good to have).

Qualifications

  • Strong proficiency in Python backend development.
  • Hands-on experience with PySpark for data engineering.
  • Excellent problem-solving skills and attention to detail.
  • Good communication skills for effective team collaboration.
  • Experience with Palantir Foundry or similar platforms is a plus.

Preferred Skills

  • Experience with large-scale data processing and pipeline development.
  • Familiarity with agile methodologies and development tools.
  • Ability to optimise and streamline backend processes effectively.


Read more
Browse more Backend Developer Jobs in Hyderabad | Backend Developer Job openings in Hyderabad →
Frisco Analytics Pvt Ltd
Cedrick Mariadas
Posted by Cedrick Mariadas

Bengaluru (Bangalore), Hyderabad · 5 - 8 years · ₹15L - ₹20L / yr · Raised funding · Posted 10 May 2024

databricks
Apache Spark
skill iconPython
SQL
MySQL
+3 more

We are actively seeking a self-motivated Data Engineer with expertise in Azure cloud and Databricks, with a thorough understanding of Delta Lake and Lake-house Architecture. The ideal candidate should excel in developing scalable data solutions, crafting platform tools, and integrating systems, while demonstrating proficiency in cloud-native database solutions and distributed data processing.


Key Responsibilities:

  • Contribute to the development and upkeep of a scalable data platform, incorporating tools and frameworks that leverage Azure and Databricks capabilities.
  • Exhibit proficiency in various RDBMS databases such as MySQL and SQL-Server, emphasizing their integration in applications and pipeline development.
  • Design and maintain high-caliber code, including data pipelines and applications, utilizing Python, Scala, and PHP.
  • Implement effective data processing solutions via Apache Spark, optimizing Spark applications for large-scale data handling.
  • Optimize data storage using formats like Parquet and Delta Lake to ensure efficient data accessibility and reliable performance.
  • Demonstrate understanding of Hive Metastore, Unity Catalog Metastore, and the operational dynamics of external tables.
  • Collaborate with diverse teams to convert business requirements into precise technical specifications.

Requirements:

  • Bachelor’s degree in Computer Science, Engineering, or a related discipline.
  • Demonstrated hands-on experience with Azure cloud services and Databricks.
  • Proficient programming skills in Python, Scala, and PHP.
  • In-depth knowledge of SQL, NoSQL databases, and data warehousing principles.
  • Familiarity with distributed data processing and external table management.
  • Insight into enterprise data solutions for PIM, CDP, MDM, and ERP applications.
  • Exceptional problem-solving acumen and meticulous attention to detail.

Additional Qualifications :

  • Acquaintance with data security and privacy standards.
  • Experience in CI/CD pipelines and version control systems, notably Git.
  • Familiarity with Agile methodologies and DevOps practices.
  • Competence in technical writing for comprehensive documentation.


Read more

Bengaluru (Bangalore), Mumbai, Delhi, Gurugram, Pune, Hyderabad, Ahmedabad, Chennai · 3 - 7 years · ₹8L - ₹15L / yr · Bootstrapped · Posted 30 Apr 2024

AWS Lambda
Amazon S3
Amazon VPC
Amazon EC2
Amazon Redshift
+3 more

Technical Skills:


  • Ability to understand and translate business requirements into design.
  • Proficient in AWS infrastructure components such as S3, IAM, VPC, EC2, and Redshift.
  • Experience in creating ETL jobs using Python/PySpark.
  • Proficiency in creating AWS Lambda functions for event-based jobs.
  • Knowledge of automating ETL processes using AWS Step Functions.
  • Competence in building data warehouses and loading data into them.


Responsibilities:


  • Understand business requirements and translate them into design.
  • Assess AWS infrastructure needs for development work.
  • Develop ETL jobs using Python/PySpark to meet requirements.
  • Implement AWS Lambda for event-based tasks.
  • Automate ETL processes using AWS Step Functions.
  • Build data warehouses and manage data loading.
  • Engage with customers and stakeholders to articulate the benefits of proposed solutions and frameworks.
Read more
Browse more AWS Lambda Jobs in Mumbai | AWS Lambda Job openings in Mumbai →
Publicis Sapient

at Publicis Sapient

10 recruiters
Mohit Singh
Posted by Mohit Singh

Bengaluru (Bangalore), Pune, Hyderabad, Gurugram, Noida · 5 - 11 years · ₹20L - ₹36L / yr · Profitable · Posted 12 Apr 2024

PySpark
Data engineering
Big Data
Hadoop
Spark
+7 more

Publicis Sapient Overview:

The Senior Associate People Senior Associate L1 in Data Engineering, you will translate client requirements into technical design, and implement components for data engineering solution. Utilize deep understanding of data integration and big data design principles in creating custom solutions or implementing package solutions. You will independently drive design discussions to insure the necessary health of the overall solution 

.

Job Summary:

As Senior Associate L2 in Data Engineering, you will translate client requirements into technical design, and implement components for data engineering solution. Utilize deep understanding of data integration and big data design principles in creating custom solutions or implementing package solutions. You will independently drive design discussions to insure the necessary health of the overall solution

The role requires a hands-on technologist who has strong programming background like Java / Scala / Python, should have experience in Data Ingestion, Integration and data Wrangling, Computation, Analytics pipelines and exposure to Hadoop ecosystem components. You are also required to have hands-on knowledge on at least one of AWS, GCP, Azure cloud platforms.


Role & Responsibilities:

Your role is focused on Design, Development and delivery of solutions involving:

• Data Integration, Processing & Governance

• Data Storage and Computation Frameworks, Performance Optimizations

• Analytics & Visualizations

• Infrastructure & Cloud Computing

• Data Management Platforms

• Implement scalable architectural models for data processing and storage

• Build functionality for data ingestion from multiple heterogeneous sources in batch & real-time mode

• Build functionality for data analytics, search and aggregation

Experience Guidelines:

Mandatory Experience and Competencies:

# Competency

1.Overall 5+ years of IT experience with 3+ years in Data related technologies

2.Minimum 2.5 years of experience in Big Data technologies and working exposure in at least one cloud platform on related data services (AWS / Azure / GCP)

3.Hands-on experience with the Hadoop stack – HDFS, sqoop, kafka, Pulsar, NiFi, Spark, Spark Streaming, Flink, Storm, hive, oozie, airflow and other components required in building end to end data pipeline.

4.Strong experience in at least of the programming language Java, Scala, Python. Java preferable

5.Hands-on working knowledge of NoSQL and MPP data platforms like Hbase, MongoDb, Cassandra, AWS Redshift, Azure SQLDW, GCP BigQuery etc

6.Well-versed and working knowledge with data platform related services on at least 1 cloud platform, IAM and data security


Preferred Experience and Knowledge (Good to Have):

# Competency

1.Good knowledge of traditional ETL tools (Informatica, Talend, etc) and database technologies (Oracle, MySQL, SQL Server, Postgres) with hands on experience

2.Knowledge on data governance processes (security, lineage, catalog) and tools like Collibra, Alation etc

3.Knowledge on distributed messaging frameworks like ActiveMQ / RabbiMQ / Solace, search & indexing and Micro services architectures

4.Performance tuning and optimization of data pipelines

5.CI/CD – Infra provisioning on cloud, auto build & deployment pipelines, code quality

6.Cloud data specialty and other related Big data technology certifications


Personal Attributes:

• Strong written and verbal communication skills

• Articulation skills

• Good team player

• Self-starter who requires minimal oversight

• Ability to prioritize and manage multiple tasks

• Process orientation and the ability to define and set up processes


Read more
Browse more Spark Jobs in Pune | Spark Job openings in Pune →
Kanerika Software

at Kanerika Software

3 candid answers
2 recruiters
Meenakshi Ramagiri
Posted by Meenakshi Ramagiri

RIYADH (Saudi Arabia), Hyderabad · 6 - 12 years · ₹10L - ₹15L / yr · Profitable · Posted 21 Mar 2024

skill iconData Science
skill iconMachine Learning (ML)
Natural Language Processing (NLP)
Computer Vision
recommendation algorithm
+2 more

Job Description


Responsibilities:

- Collaborate with stakeholders to understand business objectives and requirements for AI/ML projects.

- Conduct research and stay up-to-date with the latest AI/ML algorithms, techniques, and frameworks.

- Design and develop machine learning models, algorithms, and data pipelines.

- Collect, preprocess, and clean large datasets to ensure data quality and reliability.

- Train, evaluate, and optimize machine learning models using appropriate evaluation metrics.

- Implement and deploy AI/ML models into production environments.

- Monitor model performance and propose enhancements or updates as needed.

- Collaborate with software engineers to integrate AI/ML capabilities into existing software systems.

- Perform data analysis and visualization to derive actionable insights.

- Stay informed about emerging trends and advancements in the field of AI/ML and apply them to improve existing solutions.

Strong experience in Apache pyspark is must

 

Requirements:

- Bachelor's or Master's degree in Computer Science, Engineering, or a related field.

- Proven experience of 3-5 years as an AI/ML Engineer or a similar role.

- Strong knowledge of machine learning algorithms, deep learning frameworks, and data science concepts.

- Proficiency in programming languages such as Python, Java, or C++.

- Experience with popular AI/ML libraries and frameworks, such as TensorFlow, Keras, PyTorch, or scikit-learn.

- Familiarity with cloud platforms, such as AWS, Azure, or GCP, and their AI/ML services.

- Solid understanding of data preprocessing, feature engineering, and model evaluation techniques.

- Experience in deploying and scaling machine learning models in production environments.

- Strong problem-solving skills and ability to work on multiple projects simultaneously.

- Excellent communication and teamwork skills.

 

Preferred Skills:

- Experience with natural language processing (NLP) techniques and tools.

- Familiarity with big data technologies, such as Hadoop, Spark, or Hive.

- Knowledge of containerization technologies like Docker and orchestration tools like Kubernetes.

- Understanding of DevOps practices for AI/ML model deployment

-Apache ,Pyspark



Read more
Browse more Natural Language Processing (NLP) Jobs in Hyderabad | Natural Language Processing (NLP) Job openings in Hyderabad →
Publicis Sapient

at Publicis Sapient

10 recruiters
Mohit Singh
Posted by Mohit Singh

Bengaluru (Bangalore), Gurugram, Pune, Hyderabad, Noida · 4 - 10 years · Profitable · Posted 25 Feb 2024

PySpark
Data engineering
Big Data
Hadoop
Spark
+6 more

Publicis Sapient Overview:

The Senior Associate People Senior Associate L1 in Data Engineering, you will translate client requirements into technical design, and implement components for data engineering solution. Utilize deep understanding of data integration and big data design principles in creating custom solutions or implementing package solutions. You will independently drive design discussions to insure the necessary health of the overall solution 

.

Job Summary:

As Senior Associate L1 in Data Engineering, you will do technical design, and implement components for data engineering solution. Utilize deep understanding of data integration and big data design principles in creating custom solutions or implementing package solutions. You will independently drive design discussions to insure the necessary health of the overall solution

The role requires a hands-on technologist who has strong programming background like Java / Scala / Python, should have experience in Data Ingestion, Integration and data Wrangling, Computation, Analytics pipelines and exposure to Hadoop ecosystem components. Having hands-on knowledge on at least one of AWS, GCP, Azure cloud platforms will be preferable.


Role & Responsibilities:

Job Title: Senior Associate L1 – Data Engineering

Your role is focused on Design, Development and delivery of solutions involving:

• Data Ingestion, Integration and Transformation

• Data Storage and Computation Frameworks, Performance Optimizations

• Analytics & Visualizations

• Infrastructure & Cloud Computing

• Data Management Platforms

• Build functionality for data ingestion from multiple heterogeneous sources in batch & real-time

• Build functionality for data analytics, search and aggregation


Experience Guidelines:

Mandatory Experience and Competencies:

# Competency

1.Overall 3.5+ years of IT experience with 1.5+ years in Data related technologies

2.Minimum 1.5 years of experience in Big Data technologies

3.Hands-on experience with the Hadoop stack – HDFS, sqoop, kafka, Pulsar, NiFi, Spark, Spark Streaming, Flink, Storm, hive, oozie, airflow and other components required in building end to end data pipeline. Working knowledge on real-time data pipelines is added advantage.

4.Strong experience in at least of the programming language Java, Scala, Python. Java preferable

5.Hands-on working knowledge of NoSQL and MPP data platforms like Hbase, MongoDb, Cassandra, AWS Redshift, Azure SQLDW, GCP BigQuery etc


Preferred Experience and Knowledge (Good to Have):

# Competency

1.Good knowledge of traditional ETL tools (Informatica, Talend, etc) and database technologies (Oracle, MySQL, SQL Server, Postgres) with hands on experience

2.Knowledge on data governance processes (security, lineage, catalog) and tools like Collibra, Alation etc

3.Knowledge on distributed messaging frameworks like ActiveMQ / RabbiMQ / Solace, search & indexing and Micro services architectures

4.Performance tuning and optimization of data pipelines

5.CI/CD – Infra provisioning on cloud, auto build & deployment pipelines, code quality

6.Working knowledge with data platform related services on at least 1 cloud platform, IAM and data security

7.Cloud data specialty and other related Big data technology certifications


Job Title: Senior Associate L1 – Data Engineering

Personal Attributes:

• Strong written and verbal communication skills

• Articulation skills

• Good team player

• Self-starter who requires minimal oversight

• Ability to prioritize and manage multiple tasks

• Process orientation and the ability to define and set up processes

Read more
Browse more Spark Jobs in Pune | Spark Job openings in Pune →
A LEADING US BASED MNC

A LEADING US BASED MNC

Agency job
via Zeal Consultants by Zeal Consultants

Bengaluru (Bangalore), Hyderabad, Delhi, Gurugram · 5 - 10 years · ₹14L - ₹15L / yr · Posted 10 Jan 2024

Google Cloud Platform (GCP)
Spark
PySpark
Apache Spark
"DATA STREAMING"

Data Engineering : Senior Engineer / Manager


As Senior Engineer/ Manager in Data Engineering, you will translate client requirements into technical design, and implement components for a data engineering solutions. Utilize a deep understanding of data integration and big data design principles in creating custom solutions or implementing package solutions. You will independently drive design discussions to insure the necessary health of the overall solution.


Must Have skills :


1. GCP


2. Spark streaming : Live data streaming experience is desired.


3. Any 1 coding language: Java/Pyhton /Scala



Skills & Experience :


- Overall experience of MINIMUM 5+ years with Minimum 4 years of relevant experience in Big Data technologies


- Hands-on experience with the Hadoop stack - HDFS, sqoop, kafka, Pulsar, NiFi, Spark, Spark Streaming, Flink, Storm, hive, oozie, airflow and other components required in building end to end data pipeline. Working knowledge on real-time data pipelines is added advantage.


- Strong experience in at least of the programming language Java, Scala, Python. Java preferable


- Hands-on working knowledge of NoSQL and MPP data platforms like Hbase, MongoDb, Cassandra, AWS Redshift, Azure SQLDW, GCP BigQuery etc.


- Well-versed and working knowledge with data platform related services on GCP


- Bachelor's degree and year of work experience of 6 to 12 years or any combination of education, training and/or experience that demonstrates the ability to perform the duties of the position


Your Impact :


- Data Ingestion, Integration and Transformation


- Data Storage and Computation Frameworks, Performance Optimizations


- Analytics & Visualizations


- Infrastructure & Cloud Computing


- Data Management Platforms


- Build functionality for data ingestion from multiple heterogeneous sources in batch & real-time


- Build functionality for data analytics, search and aggregation

Read more
Browse more Google Cloud Platform (GCP) Jobs in Hyderabad | Google Cloud Platform (GCP) Job openings in Hyderabad →
A fast growing Big Data company

A fast growing Big Data company

Agency job
via Careerconnects by Kumar Narayanan

Noida, Bengaluru (Bangalore), Chennai, Hyderabad · 6 - 8 years · ₹10L - ₹15L / yr · Posted 27 Oct 2023

AWS Glue
SQL
skill iconPython
PySpark
Data engineering
+6 more

AWS Glue Developer 

Work Experience: 6 to 8 Years

Work Location:  Noida, Bangalore, Chennai & Hyderabad

Must Have Skills: AWS Glue, DMS, SQL, Python, PySpark, Data integrations and Data Ops, 

Job Reference ID:BT/F21/IND


Job Description:

Design, build and configure applications to meet business process and application requirements.


Responsibilities:

7 years of work experience with ETL, Data Modelling, and Data Architecture Proficient in ETL optimization, designing, coding, and tuning big data processes using Pyspark Extensive experience to build data platforms on AWS using core AWS services Step function, EMR, Lambda, Glue and Athena, Redshift, Postgres, RDS etc and design/develop data engineering solutions. Orchestrate using Airflow.


Technical Experience:

Hands-on experience on developing Data platform and its components Data Lake, cloud Datawarehouse, APIs, Batch and streaming data pipeline Experience with building data pipelines and applications to stream and process large datasets at low latencies.


➢ Enhancements, new development, defect resolution and production support of Big data ETL development using AWS native services.

➢ Create data pipeline architecture by designing and implementing data ingestion solutions.

➢ Integrate data sets using AWS services such as Glue, Lambda functions/ Airflow.

➢ Design and optimize data models on AWS Cloud using AWS data stores such as Redshift, RDS, S3, Athena.

➢ Author ETL processes using Python, Pyspark.

➢ Build Redshift Spectrum direct transformations and data modelling using data in S3.

➢ ETL process monitoring using CloudWatch events.

➢ You will be working in collaboration with other teams. Good communication must.

➢ Must have experience in using AWS services API, AWS CLI and SDK


Professional Attributes:

➢ Experience operating very large data warehouses or data lakes Expert-level skills in writing and optimizing SQL Extensive, real-world experience designing technology components for enterprise solutions and defining solution architectures and reference architectures with a focus on cloud technology.

➢ Must have 6+ years of big data ETL experience using Python, S3, Lambda, Dynamo DB, Athena, Glue in AWS environment.

➢ Expertise in S3, RDS, Redshift, Kinesis, EC2 clusters highly desired.


Qualification:

➢ Degree in Computer Science, Computer Engineering or equivalent.


Salary: Commensurate with experience and demonstrated competence

Read more
Browse more DMS Jobs in Chennai | DMS Job openings in Chennai →
RandomTrees

at RandomTrees

1 recruiter
Amareswarreddt yaddula
Posted by Amareswarreddt yaddula

Hyderabad · 5 - 16 years · ₹1L - ₹30L / yr · Profitable · Posted 28 Jun 2023

ETL
Informatica
Data Warehouse (DWH)
skill iconAmazon Web Services (AWS)
SQL
+3 more

We are #hiring for AWS Data Engineer expert to join our team


Job Title: AWS Data Engineer

Experience: 5 Yrs to 10Yrs

Location: Remote

Notice: Immediate or Max 20 Days

Role: Permanent Role


Skillset: AWS, ETL, SQL, Python, Pyspark, Postgres DB, Dremio.


Job Description:

 Able to develop ETL jobs.

Able to help with data curation/cleanup, data transformation, and building ETL pipelines.

Strong Postgres DB exp and knowledge of Dremio data visualization/semantic layer between DB and the application is a plus.

Sql, Python, and Pyspark is a must.

Communication should be good





Read more
Browse more AWS (Amazon Web Services) Jobs in Hyderabad | AWS (Amazon Web Services) Job openings in Hyderabad →

Chennai, Hyderabad · 5 - 10 years · ₹10L - ₹25L / yr · Profitable · Posted 19 Nov 2022

PySpark
Data engineering
Big Data
Hadoop
Spark
+2 more

Bigdata with cloud:

 

Experience : 5-10 years

 

Location : Hyderabad/Chennai

 

Notice period : 15-20 days Max

 

1.  Expertise in building AWS Data Engineering pipelines with AWS Glue -> Athena -> Quick sight

2.  Experience in developing lambda functions with AWS Lambda

3.  Expertise with Spark/PySpark – Candidate should be hands on with PySpark code and should be able to do transformations with Spark

4.  Should be able to code in Python and Scala.

5.  Snowflake experience will be a plus

Read more
Browse more Spark Jobs in Chennai | Spark Job openings in Chennai →

Hyderabad · 5 - 15 years · ₹4L - ₹14L / yr · Profitable · Posted 8 Nov 2022

Spark
Hadoop
Big Data
Data engineering
PySpark
+4 more
Big Data Engineer:-


-Expertise in building AWS Data Engineering pipelines with AWS Glue -> Athena -> Quick sight.

-Experience in developing lambda functions with AWS Lambda.

-
Expertise with Spark/PySpark

– Candidate should be hands on with PySpark code and should be able to do transformations with Spark

-Should be able to code in Python and Scala.

-
Snowflake experience will be a plus
Read more
Browse more AWS Lambda Jobs in Hyderabad | AWS Lambda Job openings in Hyderabad →

Hyderabad · 4 - 8 years · ₹5L - ₹14L / yr · Profitable · Posted 25 Oct 2022

Spark
Hadoop
Big Data
Data engineering
PySpark
+4 more
Expertise in building AWS Data Engineering pipelines with AWS Glue -> Athena -> Quick sight
Experience in developing lambda functions with AWS Lambda
Expertise with Spark/PySpark – Candidate should be hands on with PySpark code and should be able to do transformations with Spark
Should be able to code in Python and Scala.
Snowflake experience will be a plus
Read more
Browse more PySpark Jobs in India →

Hyderabad · 4 - 8 years · ₹6L - ₹25L / yr · Profitable · Posted 19 Oct 2022

PySpark
Data engineering
Big Data
Hadoop
Spark
+4 more
  1. Expertise in building AWS Data Engineering pipelines with AWS Glue -> Athena -> Quick sight
  2. Experience in developing lambda functions with AWS Lambda
  3. Expertise with Spark/PySpark – Candidate should be hands on with PySpark code and should be able to do transformations with Spark
  4. Should be able to code in Python and Scala.
  5. Snowflake experience will be a plus

 

Read more
Browse more AWS Lambda Jobs in Hyderabad | AWS Lambda Job openings in Hyderabad →

Hyderabad · 3 - 7 years · ₹1L - ₹15L / yr · Posted 18 Oct 2022

Big Data
Spark
Hadoop
PySpark
skill iconAmazon Web Services (AWS)
+3 more

Big data Developer

Exp: 3yrs to 7 yrs.
Job Location: Hyderabad
Notice: Immediate / within 30 days

1. Expertise in building AWS Data Engineering pipelines with AWS Glue -> Athena -> Quick sight
2. Experience in developing lambda functions with AWS Lambda
3. Expertise with Spark/PySpark Candidate should be hands on with PySpark code and should be able to do transformations with Spark
4. Should be able to code in Python and Scala.
5. Snowflake experience will be a plus

We can start keeping Hadoop and Hive requirements as good to have or understanding of is enough rather than keeping it as a desirable requirement.

Read more
Browse more PySpark Jobs in India →
Aureus Tech Systems

at Aureus Tech Systems

3 recruiters
Naveen Yelleti
Posted by Naveen Yelleti

Kolkata, Hyderabad, Chennai, Bengaluru (Bangalore), Bhubaneswar, Visakhapatnam, Vijayawada, Trichur, Thiruvananthapuram, Mysore, Delhi, Noida, Gurugram, Nagpur · 1 - 7 years · ₹4L - ₹15L / yr · Profitable · Posted 27 Sep 2022

PySpark
Data engineering
Big Data
Hadoop
Spark
+2 more

Skills and requirements

  • Experience analyzing complex and varied data in a commercial or academic setting.
  • Desire to solve new and complex problems every day.
  • Excellent ability to communicate scientific results to both technical and non-technical team members.


Desirable

  • A degree in a numerically focused discipline such as, Maths, Physics, Chemistry, Engineering or Biological Sciences..
  • Hands on experience on Python, Pyspark, SQL
  • Hands on experience on building End to End Data Pipelines.
  • Hands on Experience on Azure Data Factory, Azure Data Bricks, Data Lake - added advantage
  • Hands on Experience in building data pipelines.
  • Experience with Bigdata Tools, Hadoop, Hive, Sqoop, Spark, SparkSQL
  • Experience with SQL or NoSQL databases for the purposes of data retrieval and management.
  • Experience in data warehousing and business intelligence tools, techniques and technology, as well as experience in diving deep on data analysis or technical issues to come up with effective solutions.
  • BS degree in math, statistics, computer science or equivalent technical field.
  • Experience in data mining structured and unstructured data (SQL, ETL, data warehouse, Machine Learning etc.) in a business environment with large-scale, complex data sets.
  • Proven ability to look at solutions in unconventional ways. Sees opportunities to innovate and can lead the way.
  • Willing to learn and work on Data Science, ML, AI.
Read more
Browse more PySpark Jobs in Kolkata | PySpark Job openings in Kolkata →
consulting & implementation services in the area of Oil & Gas, Mining and Manufacturing Industry

consulting & implementation services in the area of Oil & Gas, Mining and Manufacturing Industry

Agency job
via Jobdost by Sathish Kumar

Ahmedabad, Hyderabad, Pune, Delhi · 5 - 7 years · ₹18L - ₹25L / yr · Posted 21 Sep 2022

AWS Lambda
AWS Simple Notification Service (SNS)
AWS Simple Queuing Service (SQS)
skill iconPython
PySpark
+9 more
  1. Data Engineer

 Required skill set: AWS GLUE, AWS LAMBDA, AWS SNS/SQS, AWS ATHENA, SPARK, SNOWFLAKE, PYTHON

Mandatory Requirements  

  • Experience in AWS Glue
  • Experience in Apache Parquet 
  • Proficient in AWS S3 and data lake 
  • Knowledge of Snowflake
  • Understanding of file-based ingestion best practices.
  • Scripting language - Python & pyspark 

CORE RESPONSIBILITIES 

  • Create and manage cloud resources in AWS 
  • Data ingestion from different data sources which exposes data using different technologies, such as: RDBMS, REST HTTP API, flat files, Streams, and Time series data based on various proprietary systems. Implement data ingestion and processing with the help of Big Data technologies 
  • Data processing/transformation using various technologies such as Spark and Cloud Services. You will need to understand your part of business logic and implement it using the language supported by the base data platform 
  • Develop automated data quality check to make sure right data enters the platform and verifying the results of the calculations 
  • Develop an infrastructure to collect, transform, combine and publish/distribute customer data.
  • Define process improvement opportunities to optimize data collection, insights and displays.
  • Ensure data and results are accessible, scalable, efficient, accurate, complete and flexible 
  • Identify and interpret trends and patterns from complex data sets 
  • Construct a framework utilizing data visualization tools and techniques to present consolidated analytical and actionable results to relevant stakeholders. 
  • Key participant in regular Scrum ceremonies with the agile teams  
  • Proficient at developing queries, writing reports and presenting findings 
  • Mentor junior members and bring best industry practices 

 QUALIFICATIONS 

  • 5-7+ years’ experience as data engineer in consumer finance or equivalent industry (consumer loans, collections, servicing, optional product, and insurance sales) 
  • Strong background in math, statistics, computer science, data science or related discipline
  • Advanced knowledge one of language: Java, Scala, Python, C# 
  • Production experience with: HDFS, YARN, Hive, Spark, Kafka, Oozie / Airflow, Amazon Web Services (AWS), Docker / Kubernetes, Snowflake  
  • Proficient with
  • Data mining/programming tools (e.g. SAS, SQL, R, Python)
  • Database technologies (e.g. PostgreSQL, Redshift, Snowflake. and Greenplum)
  • Data visualization (e.g. Tableau, Looker, MicroStrategy)
  • Comfortable learning about and deploying new technologies and tools. 
  • Organizational skills and the ability to handle multiple projects and priorities simultaneously and meet established deadlines. 
  • Good written and oral communication skills and ability to present results to non-technical audiences 
  • Knowledge of business intelligence and analytical tools, technologies and techniques.

  

Familiarity and experience in the following is a plus:  

  • AWS certification
  • Spark Streaming 
  • Kafka Streaming / Kafka Connect 
  • ELK Stack 
  • Cassandra / MongoDB 
  • CI/CD: Jenkins, GitLab, Jira, Confluence other related tools
Read more
Browse more AWS Lambda Jobs in Ahmedabad | AWS Lambda Job openings in Ahmedabad →
SenecaGlobal

at SenecaGlobal

6 recruiters
Shiva V
Posted by Shiva V

Remote, Hyderabad · 4 - 6 years · ₹15L - ₹20L / yr · Profitable · Remote friendly · Posted 17 May 2022

skill iconPython
PySpark
Spark
skill iconScala
Microsoft Azure Data factory
Should have good experience with Python or Scala/PySpark/Spark/
• Experience with Advanced SQL
• Experience with Azure data factory, data bricks,
• Experience with Azure IOT, Cosmos DB, BLOB Storage
• API management, FHIR API development,
• Proficient with Git and CI/CD best practices
• Experience working with Snowflake is a plus
Read more
Browse more Python Jobs in Hyderabad | Python Job openings in Hyderabad →
Indium Software

at Indium Software

16 recruiters
Karunya P
Posted by Karunya P

Bengaluru (Bangalore), Hyderabad · 1 - 9 years · ₹1L - ₹15L / yr · Profitable · Posted 6 Apr 2022

SQL
skill iconPython
Hadoop
HiveQL
Spark
+1 more

Responsibilities:

 

* 3+ years of Data Engineering Experience - Design, develop, deliver and maintain data infrastructures.

* SQL Specialist – Strong knowledge and Seasoned experience with SQL Queries

* Languages: Python

* Good communicator, shows initiative, works well with stakeholders.

* Experience working closely with Data Analysts and provide the data they need and guide them on the issues.

* Solid ETL experience and Hadoop/Hive/Pyspark/Presto/ SparkSQL

* Solid communication and articulation skills

* Able to handle stakeholders independently with less interventions of reporting manager.

* Develop strategies to solve problems in logical yet creative ways.

* Create custom reports and presentations accompanied by strong data visualization and storytelling

 

We would be excited if you have:

 

* Excellent communication and interpersonal skills

* Ability to meet deadlines and manage project delivery

* Excellent report-writing and presentation skills

* Critical thinking and problem-solving capabilities

Read more
Browse more SQL Jobs in Hyderabad | SQL Job openings in Hyderabad →
Consulting and Services company

Consulting and Services company

Agency job
via Jobdost by Sathish Kumar

Hyderabad, Ahmedabad · 5 - 10 years · ₹5L - ₹30L / yr · Posted 31 Mar 2022

skill iconAmazon Web Services (AWS)
Apache
skill iconPython
PySpark

Data Engineer 

  

Mandatory Requirements  

  • Experience in AWS Glue 
  • Experience in Apache Parquet  
  • Proficient in AWS S3 and data lake  
  • Knowledge of Snowflake 
  • Understanding of file-based ingestion best practices. 
  • Scripting language - Python & pyspark 

  

CORE RESPONSIBILITIES 

  • Create and manage cloud resources in AWS  
  • Data ingestion from different data sources which exposes data using different technologies, such as: RDBMS, REST HTTP API, flat files, Streams, and Time series data based on various proprietary systems. Implement data ingestion and processing with the help of Big Data technologies  
  • Data processing/transformation using various technologies such as Spark and Cloud Services. You will need to understand your part of business logic and implement it using the language supported by the base data platform  
  • Develop automated data quality check to make sure right data enters the platform and verifying the results of the calculations  
  • Develop an infrastructure to collect, transform, combine and publish/distribute customer data. 
  • Define process improvement opportunities to optimize data collection, insights and displays. 
  • Ensure data and results are accessible, scalable, efficient, accurate, complete and flexible  
  • Identify and interpret trends and patterns from complex data sets  
  • Construct a framework utilizing data visualization tools and techniques to present consolidated analytical and actionable results to relevant stakeholders.  
  • Key participant in regular Scrum ceremonies with the agile teams   
  • Proficient at developing queries, writing reports and presenting findings  
  • Mentor junior members and bring best industry practices  

  

QUALIFICATIONS 

  • 5-7+ years’ experience as data engineer in consumer finance or equivalent industry (consumer loans, collections, servicing, optional product, and insurance sales)  
  • Strong background in math, statistics, computer science, data science or related discipline 
  • Advanced knowledge one of language: Java, Scala, Python, C#  
  • Production experience with: HDFS, YARN, Hive, Spark, Kafka, Oozie / Airflow, Amazon Web Services (AWS), Docker / Kubernetes, Snowflake   
  • Proficient with 
  • Data mining/programming tools (e.g. SAS, SQL, R, Python) 
  • Database technologies (e.g. PostgreSQL, Redshift, Snowflake. and Greenplum) 
  • Data visualization (e.g. Tableau, Looker, MicroStrategy) 
  • Comfortable learning about and deploying new technologies and tools.  
  • Organizational skills and the ability to handle multiple projects and priorities simultaneously and meet established deadlines.  
  • Good written and oral communication skills and ability to present results to non-technical audiences  
  • Knowledge of business intelligence and analytical tools, technologies and techniques. 

  

Familiarity and experience in the following is a plus:  

  • AWS certification 
  • Spark Streaming  
  • Kafka Streaming / Kafka Connect  
  • ELK Stack  
  • Cassandra / MongoDB  
  • CI/CD: Jenkins, GitLab, Jira, Confluence other related tools 
Read more
Browse more AWS (Amazon Web Services) Jobs in Hyderabad | AWS (Amazon Web Services) Job openings in Hyderabad →
Persistent System Ltd

Persistent System Ltd

Agency job
via Milestone Hr Consultancy by Haina khan

Pune, Bengaluru (Bangalore), Hyderabad · 4 - 9 years · ₹8L - ₹27L / yr · Posted 30 Mar 2022

skill iconPython
PySpark
skill iconAmazon Web Services (AWS)
Spark
skill iconScala
Greetings..

We have urgent requirement of Data Engineer/Sr Data Engineer for reputed MNC company.

Exp: 4-9yrs

Location: Pune/Bangalore/Hyderabad

Skills: We need candidate either Python AWS or Pyspark AWS or Spark Scala
Read more
Browse more Backend Developer Jobs in Pune | Backend Developer Job openings in Pune →
Persistent Systems

at Persistent Systems

1 video
1 recruiter
Agency job
via Milestone Hr Consultancy by Haina khan

Pune, Bengaluru (Bangalore), Hyderabad, Nagpur · 4 - 9 years · ₹4L - ₹15L / yr · Profitable · Posted 21 Mar 2022

Spark
Hadoop
Big Data
Data engineering
PySpark
+3 more
Greetings..

We have an urgent requirements of Big Data Developer profiles in our reputed MNC company.

Location: Pune/Bangalore/Hyderabad/Nagpur
Experience: 4-9yrs

Skills: Pyspark,AWS
or Spark,Scala,AWS
or Python Aws
Read more
Browse more PySpark Jobs in Pune | PySpark Job openings in Pune →
Picture the future

Picture the future

Agency job
via Jobdost by Sathish Kumar

Hyderabad · 4 - 7 years · ₹5L - ₹15L / yr · Posted 19 Mar 2022

PySpark
Data engineering
Big Data
Hadoop
Spark
+7 more

CORE RESPONSIBILITIES

  • Create and manage cloud resources in AWS 
  • Data ingestion from different data sources which exposes data using different technologies, such as: RDBMS, REST HTTP API, flat files, Streams, and Time series data based on various proprietary systems. Implement data ingestion and processing with the help of Big Data technologies 
  • Data processing/transformation using various technologies such as Spark and Cloud Services. You will need to understand your part of business logic and implement it using the language supported by the base data platform 
  • Develop automated data quality check to make sure right data enters the platform and verifying the results of the calculations 
  • Develop an infrastructure to collect, transform, combine and publish/distribute customer data.
  • Define process improvement opportunities to optimize data collection, insights and displays.
  • Ensure data and results are accessible, scalable, efficient, accurate, complete and flexible 
  • Identify and interpret trends and patterns from complex data sets 
  • Construct a framework utilizing data visualization tools and techniques to present consolidated analytical and actionable results to relevant stakeholders. 
  • Key participant in regular Scrum ceremonies with the agile teams  
  • Proficient at developing queries, writing reports and presenting findings 
  • Mentor junior members and bring best industry practices 

 

QUALIFICATIONS

  • 5-7+ years’ experience as data engineer in consumer finance or equivalent industry (consumer loans, collections, servicing, optional product, and insurance sales) 
  • Strong background in math, statistics, computer science, data science or related discipline
  • Advanced knowledge one of language: Java, Scala, Python, C# 
  • Production experience with: HDFS, YARN, Hive, Spark, Kafka, Oozie / Airflow, Amazon Web Services (AWS), Docker / Kubernetes, Snowflake  
  • Proficient with
  • Data mining/programming tools (e.g. SAS, SQL, R, Python)
  • Database technologies (e.g. PostgreSQL, Redshift, Snowflake. and Greenplum)
  • Data visualization (e.g. Tableau, Looker, MicroStrategy)
  • Comfortable learning about and deploying new technologies and tools. 
  • Organizational skills and the ability to handle multiple projects and priorities simultaneously and meet established deadlines. 
  • Good written and oral communication skills and ability to present results to non-technical audiences 
  • Knowledge of business intelligence and analytical tools, technologies and techniques.


Mandatory Requirements 

  • Experience in AWS Glue
  • Experience in Apache Parquet 
  • Proficient in AWS S3 and data lake 
  • Knowledge of Snowflake
  • Understanding of file-based ingestion best practices.
  • Scripting language - Python & pyspark

 

Read more
Browse more AWS (Amazon Web Services) Jobs in Hyderabad | AWS (Amazon Web Services) Job openings in Hyderabad →
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Why apply via Cutshort?
Connect with actual hiring teams and get their fast response. No spam.
Find more jobs
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort