Data Engineer at VyTCDC · Bengaluru (Bangalore) · 5 - 8 years · ₹4L - ₹25L / yr · Profitable · Posted 3 Jul 2025

🛠️ Key Responsibilities
- Design, build, and maintain scalable data pipelines using Python and Apache Spark (PySpark or Scala APIs)
- Develop and optimize ETL processes for batch and real-time data ingestion
- Collaborate with data scientists, analysts, and DevOps teams to support data-driven solutions
- Ensure data quality, integrity, and governance across all stages of the data lifecycle
- Implement data validation, monitoring, and alerting mechanisms for production pipelines
- Work with cloud platforms (AWS, GCP, or Azure) and tools like Airflow, Kafka, and Delta Lake
- Participate in code reviews, performance tuning, and documentation
🎓 Qualifications
- Bachelor’s or Master’s degree in Computer Science, Engineering, or related field
- 3–6 years of experience in data engineering with a focus on Python and Spark
- Experience with distributed computing and handling large-scale datasets (10TB+)
- Familiarity with data security, PII handling, and compliance standards is a plus

Similar jobs (10)
Description
We are looking for Senior Data Engineers to join our Data Platform team and build scalable, high-performance data platforms that power data processing, analytics, and downstream applications.
The ideal candidate will have strong experience in distributed data processing, ETL pipelines, and Big Data technologies, with hands-on expertise in Apache Spark and Python Scala.
You will be responsible for designing, developing, and optimizing large-scale data pipelines while collaborating closely with cross-functional engineering teams to build reliable, production-grade data solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL and data processing pipelines for large-scale datasets.
- Build and optimize distributed data applications using Apache Spark and Python Scala.
- Develop reliable, high-performance data pipelines for batch and streaming workloads.
- Design and manage data workflows using Apache Airflow.
- Build and operate data workloads on AWS, with strong usage of Amazon S3 for large-scale data storage.
- Work with large datasets to ensure data quality, consistency, reliability, and performance.
- Collaborate with engineering, product, analytics, and other platform teams to deliver robust data solutions.
- Optimize data workflows for scalability, reliability, performance, and cost efficiency.
- Troubleshoot production issues, identify bottlenecks, and continuously improve platform performance.
Requirements
Candidates who demonstrate:
- 5+ years of experience in Data Engineering, Big Data Engineering, or a similar role.
- Strong hands-on experience with Apache Spark and Scala.
- Experience designing, building, and maintaining large-scale ETL pipelines.
- Strong hands-on experience with AWS, particularly Amazon S3.
- Hands-on experience with Apache Airflow for workflow orchestration and scheduling.
- Strong SQL skills and a solid understanding of distributed data processing concepts.
- Experience working with batch and/or streaming data pipelines.
- Excellent debugging, problem-solving, and performance optimization skills.
- Strong communication and collaboration skills.
Good to Have
- Experience with Databricks and the broader Databricks data platform.
- Familiarity with streaming technologies such as Apache Kafka.
- Experience working on large-scale data platforms handling high-volume data workloads.
- Exposure to additional AWS data services and cloud-native data architectures.
Data Engineer Short Hiring Post
🚨 Hiring: Data Engineer
🔹 Experience: 5–9 Years
🔹 Location: Bangalore / Hyderabad
🔹 Skills: PySpark, Python, SQL, ETL, CI/CD, Data Modeling
🔹 Process: L1 Virtual → L2 F2F Karat Test
🔹 F2F: Bangalore / Hyderabad Location
🔹 Positions: Immediate requirement
⚠️ Note: Candidates must be available for F2F Karat immediately after L1.
#Hiring #DataEngineer #PySpark #Python #SQL #BangaloreJobs #HyderabadJobs #Mphasis #ImmediateJoiners
Skills Referential (Required knowledge, skills and abilities)
Technical Skills:
Python
Pyspark
SQL
ETL Aws, Azure, gcp
Data Engineer Hiring Post
🚨 Hiring: Data Engineer | PySpark + Python + SQL
We are looking for experienced Data Engineers to join our team!
🔹 Experience: 5 to 9 Years
🔹 Locations: Bangalore / Hyderabad
🔹 Interview Process:
• 1st Round – Virtual
• 2nd Round – Face-to-Face (Karat Test)
🔑 Key Skills:
✅ PySpark
✅ SQL
✅ Python
✅ ETL
📩 Interested candidates can share their updated resume.
#Hiring #DataEngineer #PySpark #Python #SQL #ETL #BangaloreJobs #HyderabadJobs #TechHiring #ImmediateHiring
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Design, develop, and maintain ETL pipelines involving large-scale data.
Develop data processing and analytics applications primarily using PySpark and Python.
Build scalable and distributed data processing solutions using Apache Spark.
Develop and deploy data applications on AWS cloud.
Work with AWS services related to storage, compute, ETL, data warehousing, analytics, and streaming.
Implement distributed storage and processing solutions capable of handling high-volume datasets.
Design data processing applications with a focus on performance, scalability, reliability, and optimization.
Work with both SQL and NoSQL databases for data storage, processing, and analytics.
Write, optimize, and analyze SQL, HQL, and NoSQL queries.
Troubleshoot data pipeline and processing issues and ensure data quality and reliability.
Collaborate with data engineers, analysts, architects, and other technical teams to deliver data-driven solutions.
About the Role
We are looking for a Senior Data Engineer with strong hands-on expertise in Databricks, Python, PySpark, and SQL to build scalable, high-performance data engineering solutions. You’ll architect and develop large scale, high-performance data pipelines capable of handling massive real-time and batch data volumes across multiple business systems. Databricks is the core enterprise data and processing platform for this role. You will also use Apache Airflow for workflow orchestration and dbt for ELT transformations, and will contribute to designing reliable, secure, and governed data platforms that enable analytics, reporting, and AI-driven use cases.
Key Responsibilities
- Design and implement large-scale data pipelines using Python/PySpark, Databricks, and Microsoft Fabric.
- Develop and optimize data processing workloads in Databricks using PySpark and Spark SQL, with a strong focus on scalability, reliability, performance, and maintainability.
- Develop and maintain dbt models including layered architecture, incremental models, snapshots, macros, testing, and documentation.
- Design, develop, and maintain Apache Airflow DAGs for orchestrating reliable, scalable, and observable data pipelines.
- Design and implement data quality, observability, and governance frameworks, including automated testing, monitoring, lineage, access control, and data privacy standards.
- Partner with analytics, product, and business stakeholders to turn requirements into trustworthy datasets, and raise the engineering bar through design discussions, code reviews, and mentoring junior engineers.
Required Skills
- Strong expertise in Python for developing scalable, modular, and production-ready data engineering applications.
- Strong expertise in PySpark, including DataFrame API, Spark SQL, Structured Streaming, partitioning strategies, joins, caching, handling data skew, and Spark performance optimization.
- Strong hands-on experience with Databricks for data ingestion, transformation, processing, and optimization, including Delta Lake, Unity Catalog, Databricks Workflows, notebooks, jobs, and Databricks-native data engineering capabilities.
- Strong experience in Databricks/Spark performance tuning, including query and job optimization, partitioning, file sizing, caching, join optimization, handling data skew, and efficient use of compute resources.
- Hands-on experience with Delta Lake, including transactional data processing, schema management, incremental data processing, and reliable batch and streaming data pipelines.
- Hands-on experience in developing dbt projects using layered architecture, incremental models, snapshots, macros/Jinja, testing, documentation, and deployment best practices.
- Expertise in advanced SQL and data modelling — dimensional modeling, slowly changing dimensions, schema evolution, and query optimization.
- Hands-on experience in developing and managing Apache Airflow DAGs, scheduling workflows, dependency management, retries, backfills, and operational monitoring.
- Hands-on experience with at least one major cloud platform (AWS, Azure or GCP).
- Strong problem-solving skills and the ability to work independently with business and analytics stakeholders.
Nice to Have
- Hands-on exposure to Microsoft Fabric for data integration and analytics.
- Experience using AI coding assistants (e.g. Claude Code, GitHub Copilot) as part of a development workflow.
- Familiarity with modern DevOps practices, including CI/CD pipelines, Infrastructure as Code (IaC), and containerization (Docker/Kubernetes).
- Domain expertise in financial services.
Job Title : Senior Data Engineer – Databricks
Experience : 14 to 20 Years
Location : HSR Layout, Bangalore
Work Mode : Hybrid – 3 Days WFO
Shift : 11:30 AM – 07:30 PM IST
Positions : 2
Notice Period : Immediate Joiners Only
Interview : 1 Technical Round + 2 Client Rounds
Role Overview :
We are looking for a Senior Data Engineer to build and lead enterprise-scale data platforms for a Switzerland-based commodity client.
The role requires a strong hands-on Data Engineering professional with expertise in Databricks, PySpark, Python, SQL, and AWS, along with technical leadership and stakeholder management experience.
Must-Have Skills :
- 14 to 20 years of Data Engineering experience
- Databricks & Apache Spark / PySpark
- Python & SQL
- AWS Cloud
- Lakehouse Architecture
- ETL / ELT & Distributed Data Processing
- Batch & Streaming Pipelines
- Data Pipeline Optimization & Data Modeling
- CDC & Incremental Processing
- Git, CI/CD & Testing
- Data Quality, Monitoring & Observability
- Technical Leadership & Stakeholder Management
Key Responsibilities :
- Design and build scalable data pipelines using Databricks, PySpark, Python, SQL, and AWS.
- Own data products from design through production.
- Develop batch / streaming pipelines and reusable ETL / ELT frameworks.
- Optimize pipelines for performance, scalability, reliability, and cost.
- Design scalable data architectures and data models.
- Implement data quality, monitoring, lineage, and CI/CD practices.
- Lead technical discussions and mentor engineering teams.
- Collaborate with business stakeholders, architects, product owners, and engineering teams.
- Remain hands-on while providing technical leadership.
Ideal Candidate :
A 14 to 20 years experienced, hands-on Data Engineering leader with strong Databricks + PySpark + AWS expertise, excellent communication, stakeholder management, and experience delivering enterprise-scale data platforms.
🔴 Super Urgent : Only Bangalore-based immediate joiners.
We are looking for a Data Engineer with at least 1 year of hands-on experience building solutions on Snowflake. The candidate should be comfortable designing, building, and managing reliable data pipelines that move data from multiple sources into a central data platform.
Responsibilities
- Build and maintain data pipelines for ingesting, transforming, and loading data into Snowflake
- Design scalable data models, schemas, tables, and views in Snowflake
- Develop ETL/ELT workflows using SQL, Python, or data orchestration tools
- Integrate data from APIs, databases, files, and third-party platforms
- Monitor pipeline performance, failures, data quality, and freshness
- Optimize Snowflake queries, warehouses, storage, and compute usage
- Implement incremental loads, change data capture, and scheduled workflows
- Work with engineering and business teams to understand data requirements
- Maintain documentation for pipelines, datasets, and data transformations
Requirements
- 1+ year of hands-on experience working with Snowflake
- Strong SQL skills and experience writing complex queries
- Experience building and managing ETL or ELT data pipelines
- Knowledge of data warehousing concepts, dimensional modelling, and data quality
- Experience with Python or another scripting language
- Familiarity with orchestration tools such as Airflow, Dagster, Prefect, dbt, or similar
- Understanding of APIs, relational databases, file formats, and cloud storage
- Ability to troubleshoot pipeline failures and performance issues
- Strong analytical, problem-solving, and communication skills
Good to Have
- Experience with dbt and Snowflake Tasks, Streams, Snowpipe, or Dynamic Tables
- Knowledge of AWS, Azure, or Google Cloud
- Experience with Kafka or other streaming platforms
- Familiarity with CI/CD, Git, monitoring, and data governance practices
- Experience integrating ERP, finance, or operational systems
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Title: Data Engineer – PySpark | Oracle | GCP
Experience: 5–7 Years
Location: Hyderabad
Notice Period: Immediate Joiners Preferred
Job Summary
We are seeking an experienced Data Engineer with strong expertise in PySpark, Oracle, and Google Cloud Platform (GCP) to design, develop, and optimize scalable data pipelines. The ideal candidate should have hands-on experience in ETL development, data integration, and cloud-based data engineering solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL/data pipelines using PySpark.
- Extract, transform, and load data from Oracle databases into GCP environments.
- Build and optimize batch data processing workflows for high performance and reliability.
- Develop data engineering solutions using GCP services.
- Ensure data quality through validation, monitoring, and troubleshooting.
- Optimize SQL queries and ETL jobs for performance and scalability.
Required Skills
- 5–7 years of experience as a Data Engineer.
- Strong hands-on experience with PySpark.
- Solid experience with Oracle Database and advanced SQL.
- Hands-on experience with Google Cloud Platform (GCP).
- Strong understanding of ETL processes and data warehousing concepts.
Work Location: Hyderabad
Notice Period: Immediate Joiners Preferred
Key Responsibilities
Build and maintain data transformation pipelines using java Spark
Develop and optimize large-scale/CPU intensive data processing using Apache Spark
Orchestrate workflows using Airflow
Implement data quality checks, testing, and monitoring for pipeline. Good to have exposer into managing metadata, cataloguing, and lineage
Support schema evolution, backfills, and incremental processing
Ensure pipelines meet SLAs for freshness, reliability, and performance
Expertise/working knowledge in Spark and HBase(semantic layer, virtual datasets, Reflections)
Required Skills & Qualifications
Strong hands-on experience with Apache Spark
Experience with HBase/SQL or similar lakehouse query engines
Airflow
Understanding of data catalogs and lineage (e.g., OpenLineage, DataHub, Apache Polaris , openlineage)
Proficiency in Java
Experience with Git-based development and CI/CD






