Data Engineer at Planet Spark · Gurugram · 2 - 5 years · ₹7L - ₹18L / yr · Posted 19 Oct 2021

- Create and maintain optimal data pipeline architecture
- Assemble large, complex data sets that meet business requirements
- Identifying, designing, and implementing internal process improvements including redesigning infrastructure for greater scalability, optimizing data delivery, and automating manual processes
- Work with Data, Analytics & Tech team to extract, arrange and analyze data
- Build the infrastructure required for optimal extraction, transformation, and loading of data from a wide variety of data sources using SQL and AWS technologies
- Building analytical tools to utilize the data pipeline, providing actionable insight into key business performance metrics including operational efficiency and customer acquisition
- Works closely with all business units and engineering teams to develop a strategy for long-term data platform architecture.
- Working with stakeholders including data, design, product, and executive teams, and assisting them with data-related technical issues
- Working with stakeholders including the Executive, Product, Data, and Design teams to support their data infrastructure needs while assisting with data-related technical issues.
- SQL
- Ruby or Python(Ruby preferred)
- Apache-Hadoop based analytics
- Data warehousing
- Data architecture
- Schema design
- ML
- Prior experience of 2 to 5 years as a Data Engineer.
- Ability in managing and communicating data warehouse plans to internal teams.
- Experience designing, building, and maintaining data processing systems.
- Ability to perform root cause analysis on external and internal processes and data to identify opportunities for improvement and answer questions.
- Excellent analytic skills associated with working on unstructured datasets.
- Ability to build processes that support data transformation, workload management, data structures, dependency, and metadata.

About Planet Spark
About
PlanetSpark is on its journey to becoming the global leader in the large and untapped communication skills segment. We are Series A funded by some top VCs and are on a 30% month-on-month growth curve. We have our footprint in India, the Middle East, North America, and Australia. Come join a passionate team of over 500 young and energetic members and 400+ expert and handpicked teachers on this roller coaster ride to build the most loved brand for kids who will move the world!
Connect with the team
Company social profiles
Similar jobs (10)
Data Engineer Short Hiring Post
🚨 Hiring: Data Engineer
🔹 Experience: 5–9 Years
🔹 Location: Bangalore / Hyderabad
🔹 Skills: PySpark, Python, SQL, ETL, CI/CD, Data Modeling
🔹 Process: L1 Virtual → L2 F2F Karat Test
🔹 F2F: Bangalore / Hyderabad Location
🔹 Positions: Immediate requirement
⚠️ Note: Candidates must be available for F2F Karat immediately after L1.
#Hiring #DataEngineer #PySpark #Python #SQL #BangaloreJobs #HyderabadJobs #Mphasis #ImmediateJoiners
About Us:
The QX Impact was launched with a mission to make A.I accessible and affordable and deliver AI Products/Solutions at scale for the enterprises by bringing the power of Data, AI, and Engineering to drive digital transformation. We believe without insights; businesses will continue to face challenges to better understand their customers and even lose them. Secondly, without insights businesses won't’ be able to deliver differentiated products/services; and finally, without insights, businesses can’t achieve a new level of “Operational Excellence” is crucial to remain competitive, meeting rising customer expectations, expanding markets, and digitalization.
Job Summary:
We are looking for a Senior Data Engineer who is creative, collaborative, and adaptable to join our agile team of data scientists, engineers, and UX developers. The role focuses on building and maintaining robust data pipelines to support advanced analytics, data science, and BI solutions.
As a Senior Data Engineer, you will work with internal and external data, collaborate with data scientists, and contribute to the design, development, and deployment of innovative solutions.
Key Responsibilities:
- Design, develop, test, and maintain optimal data pipeline and ETL architectures.
- Map out data systems and define/design required integrations, ETL, BI, and AI systems/processes.
- Prepare and optimize data for predictive and prescriptive modeling.
- Collaborate with teams to integrate ERP data into the enterprise data lake, ensuring seamless flow and quality.
- Enhance cloud data infrastructure on AWS or Azure for scalability and performance.
- Utilize big data tools and frameworks to optimize data acquisition and preparation.
- Build architectures to move data to/from data lakes and data warehouses for advanced analytics.
- Develop and curate data models for analytics, dashboards, and reports.
- Conduct code reviews, maintain production-level code, and implement testing approaches.
- Monitor, troubleshoot, and resolve data ingestion workflows to maintain reliability and uptime.
- Drive innovation and implement efficient new approaches to data engineering tasks.
Must-Have Skills:
- Bachelor’s degree in Computer Science, Mathematics, Engineering, or a related field.
- 5+ years of experience working with enterprise data platforms, including building and managing data lakes.
- 3–5 years of experience designing and implementing data warehouse solutions.
- Expertise in SQL, including developing stored procedures (SP) and applying advanced data design concepts.
- Proficiency in Spark (Python/Scala) and Spark Streaming for real-time data pipelines.
- Experience with AWS or Azure services (e.g., AWS Glue, Azure Data Factory, Redshift, Snowflake).
- Familiarity with big data tools such as Apache Kafka, Apache Spark, or Flink.
- Hands-on experience with orchestration tools (e.g., Apache Airflow, Prefect).
- Knowledge of CI/CD processes, version control (e.g., Git, Jenkins), and deployment automation.
- Strong problem-solving, communication, and collaboration skills.
Good-to-Have Skills:
- Experience in integrating ERP data into data lakes.
- Experience with traditional ETL tools (e.g., Talend, Pentaho).
Competencies:
- Tech Savvy - Anticipating and adopting innovations in business-building digital and technology applications.
- Self-Development - Actively seeking new ways to grow and be challenged using both formal and informal development channels.
- Action Oriented - Taking on new opportunities and tough challenges with a sense of urgency, high energy, and enthusiasm.
- Customer Focus - Building strong customer relationships and delivering customer-centric solutions.
- Optimize Work Processes - Knowing the most effective and efficient processes to get things done, with a focus on continuous improvement.
Why Join Us?
- Be part of a collaborative and agile team driving cutting-edge AI and data engineering solutions.
- Work on impactful projects that make a difference across industries.
- Opportunities for professional growth and continuous learning.
- Competitive salary and benefits package.
Application Details
Ready to make an impact? Apply today and become part of the QX Impact team!
Description
We are looking for Senior Data Engineers to join our Data Platform team and build scalable, high-performance data platforms that power data processing, analytics, and downstream applications.
The ideal candidate will have strong experience in distributed data processing, ETL pipelines, and Big Data technologies, with hands-on expertise in Apache Spark and Python Scala.
You will be responsible for designing, developing, and optimizing large-scale data pipelines while collaborating closely with cross-functional engineering teams to build reliable, production-grade data solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL and data processing pipelines for large-scale datasets.
- Build and optimize distributed data applications using Apache Spark and Python Scala.
- Develop reliable, high-performance data pipelines for batch and streaming workloads.
- Design and manage data workflows using Apache Airflow.
- Build and operate data workloads on AWS, with strong usage of Amazon S3 for large-scale data storage.
- Work with large datasets to ensure data quality, consistency, reliability, and performance.
- Collaborate with engineering, product, analytics, and other platform teams to deliver robust data solutions.
- Optimize data workflows for scalability, reliability, performance, and cost efficiency.
- Troubleshoot production issues, identify bottlenecks, and continuously improve platform performance.
Requirements
Candidates who demonstrate:
- 5+ years of experience in Data Engineering, Big Data Engineering, or a similar role.
- Strong hands-on experience with Apache Spark and Scala.
- Experience designing, building, and maintaining large-scale ETL pipelines.
- Strong hands-on experience with AWS, particularly Amazon S3.
- Hands-on experience with Apache Airflow for workflow orchestration and scheduling.
- Strong SQL skills and a solid understanding of distributed data processing concepts.
- Experience working with batch and/or streaming data pipelines.
- Excellent debugging, problem-solving, and performance optimization skills.
- Strong communication and collaboration skills.
Good to Have
- Experience with Databricks and the broader Databricks data platform.
- Familiarity with streaming technologies such as Apache Kafka.
- Experience working on large-scale data platforms handling high-volume data workloads.
- Exposure to additional AWS data services and cloud-native data architectures.
Job Summary
We are seeking a skilled Data Engineer to design, build, and maintain scalable data pipelines and infrastructure. The ideal candidate should have strong expertise in SQL, Python, Linux, and modern data engineering practices to support data integration, transformation, and analytics.
Key Responsibilities
- Design, develop, and maintain ETL/ELT data pipelines.
- Write efficient and optimized SQL queries for data extraction, transformation, and reporting.
- Develop automation scripts using Python for data processing and workflow optimization.
- Work with Linux environments for deployment, monitoring, and troubleshooting.
- Ensure data quality, integrity, and reliability across data platforms.
- Collaborate with data analysts, software engineers, and business stakeholders to deliver data solutions.
- Monitor, troubleshoot, and optimize data pipelines for performance and scalability.
- Implement best practices for data security, governance, and documentation.
Required Skills
- Strong experience in Data Engineering concepts and ETL/ELT processes.
- Proficiency in SQL, including query optimization and database design.
- Strong programming skills in Python.
- Hands-on experience with Linux commands, shell scripting, and system administration basics.
- Experience with relational databases such as PostgreSQL, MySQL, SQL Server, or Oracle.
- Familiarity with Git/version control.
- Strong analytical and problem-solving skills.
Preferred Skills
- Experience with cloud platforms (AWS, Azure, or GCP).
- Knowledge of Apache Spark, Airflow, Kafka, or similar data engineering tools.
- Experience with data warehousing solutions and big data technologies.
- Understanding of CI/CD pipelines and containerization (Docker/Kubernetes).
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Relevant certifications in cloud or data engineering are an added advantage.
Job Summary
We are looking for a skilled and experienced Data Engineer to join our growing data team. The ideal candidate will have strong expertise in Python, PySpark, Data Modeling, and Power BI, with hands-on experience in designing, developing, and optimizing scalable data solutions. The role requires working closely with business stakeholders, data architects, and analytics teams to build robust data pipelines and semantic models that enable data-driven decision-making.
Technical Skills
- Strong hands-on experience in Python and PySpark development.
- Expertise in building and optimizing Data Engineering solutions and ETL pipelines.
- Strong understanding of Data Modeling concepts (Star Schema, Snowflake Schema, Dimensional Modeling).
- Experience with Power BI Data Modeling and Semantic Layer development.
- Proficiency in DAX (Data Analysis Expressions).
- Experience designing and managing Semantic Models in Power BI.
- Strong SQL skills and experience working with large datasets.
- Knowledge of data warehousing concepts and best practices.
Preferred Skills
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Exposure to modern data platforms like Databricks.
- Understanding of data governance and data quality frameworks.
Job Description
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, SQL, CI/CD, and DevOps to design, develop, and maintain scalable data pipelines and data processing solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines.
- Develop data processing solutions using Python.
- Write complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain data ingestion and integration workflows.
- Implement data quality, validation, monitoring, and error-handling processes.
- Develop and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps tools and practices for automated build, deployment, and infrastructure management.
- Collaborate with data analysts, data scientists, software engineers, and business teams.
- Optimize data pipelines for performance, reliability, and scalability.
- Troubleshoot production data issues and ensure timely resolution.
- Follow best practices for version control, code quality, testing, and deployment.
Mandatory Skills
- Python
- ETL
- SQL
- CI/CD
- DevOps
- Git / Version Control
- Strong problem-solving and debugging skills
Job Title : Senior Data Engineer – Databricks
Experience : 14 to 20 Years
Location : HSR Layout, Bangalore
Work Mode : Hybrid – 3 Days WFO
Shift : 11:30 AM – 07:30 PM IST
Positions : 2
Notice Period : Immediate Joiners Only
Interview : 1 Technical Round + 2 Client Rounds
Role Overview :
We are looking for a Senior Data Engineer to build and lead enterprise-scale data platforms for a Switzerland-based commodity client.
The role requires a strong hands-on Data Engineering professional with expertise in Databricks, PySpark, Python, SQL, and AWS, along with technical leadership and stakeholder management experience.
Must-Have Skills :
- 14 to 20 years of Data Engineering experience
- Databricks & Apache Spark / PySpark
- Python & SQL
- AWS Cloud
- Lakehouse Architecture
- ETL / ELT & Distributed Data Processing
- Batch & Streaming Pipelines
- Data Pipeline Optimization & Data Modeling
- CDC & Incremental Processing
- Git, CI/CD & Testing
- Data Quality, Monitoring & Observability
- Technical Leadership & Stakeholder Management
Key Responsibilities :
- Design and build scalable data pipelines using Databricks, PySpark, Python, SQL, and AWS.
- Own data products from design through production.
- Develop batch / streaming pipelines and reusable ETL / ELT frameworks.
- Optimize pipelines for performance, scalability, reliability, and cost.
- Design scalable data architectures and data models.
- Implement data quality, monitoring, lineage, and CI/CD practices.
- Lead technical discussions and mentor engineering teams.
- Collaborate with business stakeholders, architects, product owners, and engineering teams.
- Remain hands-on while providing technical leadership.
Ideal Candidate :
A 14 to 20 years experienced, hands-on Data Engineering leader with strong Databricks + PySpark + AWS expertise, excellent communication, stakeholder management, and experience delivering enterprise-scale data platforms.
🔴 Super Urgent : Only Bangalore-based immediate joiners.
Data Engineer Hiring Post
🚨 Hiring: Data Engineer | PySpark + Python + SQL
We are looking for experienced Data Engineers to join our team!
🔹 Experience: 5 to 9 Years
🔹 Locations: Bangalore / Hyderabad
🔹 Interview Process:
• 1st Round – Virtual
• 2nd Round – Face-to-Face (Karat Test)
🔑 Key Skills:
✅ PySpark
✅ SQL
✅ Python
✅ ETL
📩 Interested candidates can share their updated resume.
#Hiring #DataEngineer #PySpark #Python #SQL #ETL #BangaloreJobs #HyderabadJobs #TechHiring #ImmediateHiring
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Roles & Responsibilities
- Design, develop, and deliver scalable end-to-end data pipelines using Azure Data Factory, ensuring robust integration
of enterprise-wide data from diverse sources
• Build and optimize data engineering workflows using Databricks and PySpark
• Write efficient, high-performance SQL for data transformation and analysis
• Work with the Azure Cloud platform and associated services, applying strong understanding of data warehousing,
data models, and pipelines
• Provide technical leadership to a team of developers, including code reviews and enforcing best practices across the
development lifecycle
• Oversee CI/CD implementation using Azure DevOps, managing deployments across development, QA, and production
environments with proper change control processes
• Collaborate with cross-functional teams to translate business requirements into scalable data solutions
• Ensure data quality, reliability, and performance across all pipelines and platforms
Ideal Candidate
1Strong Azure Databricks Engineer / Senior Data Engineer Profile
2Mandatory (Experience 1) – Must have minimum 8+ years of overall experience in Data Engineering, Data Development, or related data technology roles, with strong hands-on experience in enterprise data pipeline development.
3Mandatory (Experience 2) – Must have strong hands-on experience with Azure Databricks, including development and optimization of scalable data engineering workflows using Databricks and PySpark.
4Mandatory (Experience 3) – Must have strong hands-on proficiency in PySpark/Python and SQL, with proven experience developing complex data transformations, processing workflows, and performance-optimized queries.
5Mandatory (Experience 4) – Must have hands-on experience with Azure Data Factory (ADF) for designing, developing, and orchestrating end-to-end data pipelines and integrating data from multiple sources.
6Mandatory (Experience 5) – Must have strong experience working on the Azure Cloud platform and associated data services, with solid understanding of data warehousing, data modeling, pipeline architecture, and enterprise data solutions.
7Mandatory (Experience 6) – Must have hands-on experience implementing CI/CD using Azure DevOps, including deployment and release management across development, QA, and production environments.
8Mandatory (Experience 7) – Must have proven technical leadership experience, including code reviews, enforcing development best practices, mentoring developers, and providing technical guidance to a data engineering team.
9Mandatory (Notice Period) – Immediate joiners or candidates who can join within 15 days.
10Mandatory (Note) - The position is open across all Cognizant offices pan India. Candidates must be willing to attend the F2F interview at the nearest Cognizant office location.
Job Summary
Role Overview
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, Advanced SQL, CI/CD, DevOps, and Data Analytics. The ideal candidate should have hands-on experience designing and developing scalable data pipelines, transforming large datasets, and supporting data-driven applications.
Experience with Google Cloud Platform (GCP) will be an added advantage.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Python and SQL.
- Develop complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain reliable data integration workflows across multiple data sources.
- Perform data cleansing, validation, transformation, and quality checks.
- Analyze data and provide insights to support business and technical requirements.
- Implement and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps practices and tools to automate deployments, monitoring, and infrastructure processes.
- Troubleshoot data pipeline failures, performance issues, and production incidents.
- Optimize data processing workflows for performance, scalability, and reliability.
- Collaborate with Data Analysts, Data Scientists, Developers, and other stakeholders.
- Follow best practices for version control, testing, documentation, and deployment.
- Contribute to cloud-based data engineering initiatives, preferably on GCP.
Required Skills
- 5–7 years of hands-on experience in Data Engineering.
- Strong programming skills in Python.
- Strong expertise in Advanced SQL and database concepts.
- Hands-on experience with ETL/ELT processes and data pipelines.
- Good understanding of Data Warehousing and Data Modeling concepts.
- Experience with CI/CD practices and tools.
- Strong understanding of DevOps principles, automation, and deployment processes.
- Strong data analytics and problem-solving skills.
- Experience working with large datasets and performance optimization.
- Good understanding of Git/version control and software development best practices.
Good to Have
- Hands-on experience with Google Cloud Platform (GCP).
- Exposure to GCP data services such as BigQuery, Cloud Storage, Dataflow, Composer, or Pub/Sub.
- Experience with containerization/orchestration technologies such as Docker/Kubernetes.
- Experience with workflow orchestration tools such as Airflow.
- Knowledge of cloud-based data architecture and distributed data processing.
Preferred Candidate Profile
- Strong analytical and problem-solving abilities.
- Good communication and stakeholder management skills.
- Ability to work independently as well as in a collaborative team environment.
- Strong ownership of data pipelines and production systems.
- Candidates who can join at short notice are preferred.
Mandatory Skills
Data Engineer, Python , ETL, GCP, Advanced SQL, Strong Data Analytics skills, CICD, Devops






