Data Engineer at Propellor.ai · Remote only · 2 - 5 years · ₹5L - ₹15L / yr · Raised funding · Remote only · Posted 1 Nov 2022

Job Description - Data Engineer
About us
Propellor is aimed at bringing Marketing Analytics and other Business Workflows to the Cloud ecosystem. We work with International Clients to make their Analytics ambitions come true, by deploying the latest tech stack and data science and engineering methods, making their business data insightful and actionable.
What is the role?
This team is responsible for building a Data Platform for many different units. This platform will be built on Cloud and therefore in this role, the individual will be organizing and orchestrating different data sources, and
giving recommendations on the services that fulfil goals based on the type of data
Qualifications:
• Experience with Python, SQL, Spark
• Knowledge/notions of JavaScript
• Knowledge of data processing, data modeling, and algorithms
• Strong in data, software, and system design patterns and architecture
• API building and maintaining
• Strong soft skills, communication
Nice to have:
• Experience with cloud: Google Cloud Platform, AWS, Azure
• Knowledge of Google Analytics 360 and/or GA4.
Key Responsibilities
• Work on the core backend and ensure it meets the performance benchmarks.
• Designing and developing APIs for the front end to consume.
• Constantly improve the architecture of the application by clearing the technical backlog.
• Meeting both technical and consumer needs.
• Staying abreast of developments in web applications and programming languages.
Key Responsibilities
• Design and develop platform based on microservices architecture.
• Work on the core backend and ensure it meets the performance benchmarks.
• Work on the front end with ReactJS.
• Designing and developing APIs for the front end to consume.
• Constantly improve the architecture of the application by clearing the technical backlog.
• Meeting both technical and consumer needs.
• Staying abreast of developments in web applications and programming languages.
What are we looking for?
An enthusiastic individual with the following skills. Please do not hesitate to apply if you do not match all of it. We are open to promising candidates who are passionate about their work and are team players.
• Education - BE/MCA or equivalent.
• Agnostic/Polyglot with multiple tech stacks.
• Worked on open-source technologies – NodeJS, ReactJS, MySQL, NoSQL, MongoDB, DynamoDB.
• Good experience with Front-end technologies like ReactJS.
• Backend exposure – good knowledge of building API.
• Worked on serverless technologies.
• Efficient in building microservices in combining server & front-end.
• Knowledge of cloud architecture.
• Should have sound working experience with relational and columnar DB.
• Should be innovative and communicative in approach.
• Will be responsible for the functional/technical track of a project.
Whom will you work with?
You will closely work with the engineering team and support the Product Team.
Hiring Process includes :
a. Written Test on Python and SQL
b. 2 - 3 rounds of Interviews
Immediate Joiners will be preferred

About Propellor.ai
About
Who we are
At Propellor, we are passionate about solving Data Unification challenges faced by our clients. We build solutions using the latest tech stack. We believe all solutions lie in the congruence of Business, Technology, and Data Science. Combining the 3, our team of young Data professionals solves some real-world problems. Here's what we live by:
Skin in the game
We believe that Individual and Collective success orientations both propel us ahead.
Cross Fertility
Borrowing from and building on one another’s varied perspectives means we are always viewing business problems with a fresh lens.
Sub 25's
A bunch of young turks, who keep our explorer mindset alive and kicking.
Future-proofing
Keeping an eye ahead, we are upskilling constantly, staying relevant at any given point in time.
Tech Agile
Tech changes quickly. Whatever your stack, we adapt speedily and easily.
If you are evaluating us to be your next employer, we urge you to read more about our team and culture here: https://bit.ly/3ExSNA2. We assure you, it's worth a read!
Tech stack
Company video


Candid answers by the company


Photos
Connect with the team
Similar jobs (10)
Job Summary
We are looking for a skilled and experienced Data Engineer to join our growing data team. The ideal candidate will have strong expertise in Python, PySpark, Data Modeling, and Power BI, with hands-on experience in designing, developing, and optimizing scalable data solutions. The role requires working closely with business stakeholders, data architects, and analytics teams to build robust data pipelines and semantic models that enable data-driven decision-making.
Technical Skills
- Strong hands-on experience in Python and PySpark development.
- Expertise in building and optimizing Data Engineering solutions and ETL pipelines.
- Strong understanding of Data Modeling concepts (Star Schema, Snowflake Schema, Dimensional Modeling).
- Experience with Power BI Data Modeling and Semantic Layer development.
- Proficiency in DAX (Data Analysis Expressions).
- Experience designing and managing Semantic Models in Power BI.
- Strong SQL skills and experience working with large datasets.
- Knowledge of data warehousing concepts and best practices.
Preferred Skills
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Exposure to modern data platforms like Databricks.
- Understanding of data governance and data quality frameworks.
Job Description
• Design and Implement Data Solutions: Lead the design, development, and implementation of scalable and secure Azure-
based data platforms, ensuring integration with various data sources and business systems. Deliver at least 2 major
projects every year with a focus on data engineering best practices.
• Optimize Data Pipelines: Build and optimize end-to-end data pipelines using Azure Data Factory, Azure Databricks, and
Azure Synapse, with an emphasis on automating data workflows. Achieve a 20% reduction in pipeline execution times
within the first 6 months.
• Cloud Infrastructure Management: Manage and maintain the Azure data environment, ensuring high availability, disaster
recovery, and cost optimization. Track and improve system uptime to exceed 99.9% reliability.
• Collaborate with Cross-Functional Teams: Partner with data scientists, data analysts, and business stakeholders to translate
business requirements into efficient data solutions. Facilitate at least 3 collaborative sessions per quarter to address key
business use cases.
• Ensure Data Security & Compliance: Implement data security measures, ensuring compliance with industry regulations
(GDPR, HIPAA, etc.) and company policies. Achieve and maintain full compliance in all data environments within the first
quarter of onboarding.
• Continuous Learning & Knowledge Sharing: Stay up-to-date with emerging Azure technologies, and mentor junior
engineers to promote knowledge sharing. Complete 1 Azure certification annually and conduct at least 2 internal
knowledge-sharing sessions per year.
• Sound knowledge of data governance practices, data quality management, and data security principles.
• Play a pivotal role in shaping our organization's data-driven journey, driving innovation through data analytics and insights.
• Optimize data storage, processing and retrieval mechanisms for performance, cost, and scalability using data storage
services (such as Azure Data Lake Storage, Azure SQL Database, etc.), data processing services (such as Azure Data Bricks,
Azure Synapse, etc.) and data visualization (PowerBI, Qlik, etc.) & integration services (Data API builder, logic apps, etc.)
• Monitor and troubleshoot data platform performance, identify and resolve issues, and provide recommendations for
continuous improvement.
• Collaborate with DevOps teams to automate deployment, configuration, and monitoring processes using Azure DevOps,
PowerShell, or other relevant tools.
• Stay up to date with the latest trends and advancements in cloud data services and provide recommendations on adopting
new technologies or features to enhance the data platform.
• Document technical designs, procedures, and guidelines for data platform engineering and operations
Knowledge, Skills & Experience
Job Experience • Bachelor's degree in Computer Science, Engineering, or a related field. Advanced
degree preferred.
• Proven 6-10 years experience in playing platform engineer or admin role
• Experience with big data technologies such as Apache Spark, Hadoop, or similar
frameworks.
• Solid understanding of cloud computing concepts and experience with cloud
infrastructure management and provisioning.
• Solid understanding of network security concepts and technologies (such as
firewalls, VPNs, intrusion detection/prevention systems, etc.) and data security
concepts and technologies (such as access controls, encryption, observability,
privacy laws/regulations, etc.)
• Experience in a Retail setup is preferred.
Required Skills The position will require someone with the following:
• Strategic Planning
Public
• Communication and Collaboration
• Problem Solving Skills A/B testing & experimentation
• SQL, BI tools, and storytelling with data
4 - 10 years of experience in designing and buildingarchitecting highly resilient data platforms
∙Strong knowledge of data engineering, architecture and data modeling
∙Experience in platforms like Databricks and Snowflake
∙Experience on building applications on cloud (AWS or Azure or Google Cloud)
∙Strong analytical and problem-solving skills
∙Prior experience in developing data or computation intensive (e.g. grid based) backend applications is an
advantage
∙OOP design skills with an understanding or at least personal interest towards the concepts of Functional
Programming
∙Willingness to understand and enhance other people’s code, being able to work in an environment where
developers will oversee and work on wider components also dealing with older “legacy” code
∙Strong programming skills (Java/ Scala / Python) skills with the willingness to pick up the other language if not
already mastered at a sufficient level is important
∙Spring knowledge is an advantage, but in general willingness to learn, work with and even enhance in-house
developed frameworks is a must
∙Prior experience in working with Git, Bitbucket, Jenkins, working with PR-s, using JIRA, following the Scrum Agile
methodology is an advantage
∙Prior knowledge of financial products is an advantage
∙Bachelors or Masters in any relevant field of IT/Engineering area is an advantage
Job Title : Senior Data Engineer – Databricks
Experience : 14 to 20 Years
Location : HSR Layout, Bangalore
Work Mode : Hybrid – 3 Days WFO
Shift : 11:30 AM – 07:30 PM IST
Positions : 2
Notice Period : Immediate Joiners Only
Interview : 1 Technical Round + 2 Client Rounds
Role Overview :
We are looking for a Senior Data Engineer to build and lead enterprise-scale data platforms for a Switzerland-based commodity client.
The role requires a strong hands-on Data Engineering professional with expertise in Databricks, PySpark, Python, SQL, and AWS, along with technical leadership and stakeholder management experience.
Must-Have Skills :
- 14 to 20 years of Data Engineering experience
- Databricks & Apache Spark / PySpark
- Python & SQL
- AWS Cloud
- Lakehouse Architecture
- ETL / ELT & Distributed Data Processing
- Batch & Streaming Pipelines
- Data Pipeline Optimization & Data Modeling
- CDC & Incremental Processing
- Git, CI/CD & Testing
- Data Quality, Monitoring & Observability
- Technical Leadership & Stakeholder Management
Key Responsibilities :
- Design and build scalable data pipelines using Databricks, PySpark, Python, SQL, and AWS.
- Own data products from design through production.
- Develop batch / streaming pipelines and reusable ETL / ELT frameworks.
- Optimize pipelines for performance, scalability, reliability, and cost.
- Design scalable data architectures and data models.
- Implement data quality, monitoring, lineage, and CI/CD practices.
- Lead technical discussions and mentor engineering teams.
- Collaborate with business stakeholders, architects, product owners, and engineering teams.
- Remain hands-on while providing technical leadership.
Ideal Candidate :
A 14 to 20 years experienced, hands-on Data Engineering leader with strong Databricks + PySpark + AWS expertise, excellent communication, stakeholder management, and experience delivering enterprise-scale data platforms.
🔴 Super Urgent : Only Bangalore-based immediate joiners.
Senior Data Engineer – PySpark & Oracle
Experience: 7+ Years
Location: Bangalore
Notice Period: Immediate to 10 Days
Key Skills:
- Strong expertise in Data Modeling, Data Design & Modernization
- Primary skills: PySpark, Oracle SQL/PLSQL
- Secondary skills: Python, ETL & Data Pipelines
- Experience with Kafka and Hadoop
- Exposure to AWS / Azure / GCP
- Good knowledge of Git and JIRA
Roles & Responsibilities:
- Design, develop, and modernize scalable data models and data architecture.
- Develop and optimize data processing solutions using PySpark and Oracle SQL/PLSQL.
- Build and maintain robust ETL workflows and data pipelines.
- Work with Kafka, Hadoop, and cloud platforms for data processing and integration.
- Perform data transformation, optimization, and performance tuning.
- Collaborate with technical teams on data design, development, testing, and deployment.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Title : Data Engineer – Databricks
Experience : 6+ Years
Location : Noida / Hyderabad / Chennai / Pune / Bengaluru (Hybrid)
Shift : IST (Normal Shift)
Job Summary :
We are seeking an experienced Data Engineer with strong expertise in Databricks, Snowflake, Python, and Spark to build and optimize scalable data pipelines and support AI/ML model deployments. The ideal candidate should have experience working with cloud-based data platforms and preferably possess exposure to the Healthcare domain.
Required Skills :
- Databricks (Preferred)
- Snowflake
- Python
- Apache Spark
- SQL
- Azure Cloud
- Kubernetes
- Apache Airflow
- GitHub & CI/CD Pipelines
- AI/ML Model Deployment
- Data Analytics
Preferred :
- Experience in the Healthcare domain.
- Strong understanding of scalable data engineering architectures and best practices.
Role Overview
We are looking for a GCP Data Engineer with 10+ years of experience to design, develop, and optimize scalable cloud-based data solutions. The ideal candidate will have strong hands-on expertise in GCP, BigQuery, and advanced SQL, with experience building data pipelines and working with large-scale datasets.
Key Responsibilities
- Design and develop scalable data pipelines and ETL/ELT processes on GCP.
- Build, optimize, and maintain data solutions using Google BigQuery.
- Develop complex SQL queries for data transformation, aggregation, and analysis.
- Design efficient data models and optimize pipelines for performance, scalability, and cost.
- Integrate data from multiple sources and ensure data quality, reliability, and availability.
- Troubleshoot pipeline and data issues and drive continuous improvement.
- Collaborate with data architects, analysts, application teams, and business stakeholders.
- Follow best practices for cloud security, data governance, testing, and documentation.
Required Skills
- 8+ years of Data Engineering experience
- Strong hands-on experience with GCP, Django, and MongoDB
- Extensive experience with BigQuery
- Advanced SQL skills
- Strong understanding of ETL/ELT and data pipeline development
- Data modeling and data warehousing experience
- Experience handling large-scale datasets and performance optimization
- Strong problem-solving and communication skills
Good to Have
- GCP services such as Cloud Storage, Dataflow, Pub/Sub, Cloud Composer, or Cloud Functions
- Python or other data engineering languages
- Experience with data governance and security
- Agile development experience
Description
We are looking for Senior Data Engineers to join our Data Platform team and build scalable, high-performance data platforms that power data processing, analytics, and downstream applications.
The ideal candidate will have strong experience in distributed data processing, ETL pipelines, and Big Data technologies, with hands-on expertise in Apache Spark and Python Scala.
You will be responsible for designing, developing, and optimizing large-scale data pipelines while collaborating closely with cross-functional engineering teams to build reliable, production-grade data solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL and data processing pipelines for large-scale datasets.
- Build and optimize distributed data applications using Apache Spark and Python Scala.
- Develop reliable, high-performance data pipelines for batch and streaming workloads.
- Design and manage data workflows using Apache Airflow.
- Build and operate data workloads on AWS, with strong usage of Amazon S3 for large-scale data storage.
- Work with large datasets to ensure data quality, consistency, reliability, and performance.
- Collaborate with engineering, product, analytics, and other platform teams to deliver robust data solutions.
- Optimize data workflows for scalability, reliability, performance, and cost efficiency.
- Troubleshoot production issues, identify bottlenecks, and continuously improve platform performance.
Requirements
Candidates who demonstrate:
- 5+ years of experience in Data Engineering, Big Data Engineering, or a similar role.
- Strong hands-on experience with Apache Spark and Scala.
- Experience designing, building, and maintaining large-scale ETL pipelines.
- Strong hands-on experience with AWS, particularly Amazon S3.
- Hands-on experience with Apache Airflow for workflow orchestration and scheduling.
- Strong SQL skills and a solid understanding of distributed data processing concepts.
- Experience working with batch and/or streaming data pipelines.
- Excellent debugging, problem-solving, and performance optimization skills.
- Strong communication and collaboration skills.
Good to Have
- Experience with Databricks and the broader Databricks data platform.
- Familiarity with streaming technologies such as Apache Kafka.
- Experience working on large-scale data platforms handling high-volume data workloads.
- Exposure to additional AWS data services and cloud-native data architectures.
About Us
We believe the future of software development is AI-native — where engineers operate at a higher level of abstraction and quality remains non-negotiable.
Incubyte is a software craft consultancy where the “how” of building software matters as much as the “what”.
We partner with companies of all sizes, from helping enterprises build, scale, and modernize to early-stage founders bring their ideas to life.
Our engineers operate in an AI-native development model, using AI as a collaborator across the SDLC to accelerate development while upholding the discipline of software craftsmanship. Guided by Software Craftsmanship and Extreme Programming practices, we build reliable, maintainable, and scalable systems with speed, without compromising quality. If this way of building software resonates with you, we’d like to talk.
Our Guiding Principles
These principles define how we work at Incubyte. They are non-negotiable.
Relentless Pursuit of Quality with Pragmatism
We build high-quality systems without losing sight of delivery.
Extreme Ownership
We take responsibility end-to-end for decisions, execution, and outcomes.
Proactive Collaboration
We collaborate closely, challenge each other, and solve problems together.
Active Pursuit of Mastery
We continuously improve our craft and raise our bar.
Invite, Give, and Act on Feedback
We seek, give, and act on feedback to get better every day.
Ensuring Client Success
We act as trusted partners and focus on real outcomes, not just output.
Job Description
This is a remote position.
Experience Level
2+ years of experience in SQL, Python, and Snowflake (or equivalent cloud data warehouse), Azure Cloud services.
Role Overview
If you're a Data Craftsperson who takes pride in clean, well-tested data solutions and believes in the principles of Extreme Programming, we'd love to meet you. At Incubyte, we're a DevOps organization where developers own the entire release cycle — you'll get hands-on experience across data engineering, analytics, cloud infrastructure, and direct client communication. This role sits primarily in data engineering (80%) with a meaningful analytics component (20%), supporting our client's data systems end-to-end.
What You'll Do
- Design, build, and maintain data pipelines and infrastructure using SQL and Python
- Work within Snowflake to build and optimize data models supporting business use cases
- Parse and process structured and semi-structured data (JSON, XML) from varied sources
- Diagnose issues across raw, intermediate, and summary tables
- Build SQL queries to support repeatable analytics use cases based on stakeholder requirements
- Investigate and resolve data quality issues, including time-sensitive or urgent ones
- Identify opportunities to consolidate models and maintain a single source of truth (SSOT)
Requirements
What We're Looking For
- 2+ years of experience with SQL and relational databases, with the ability to understand complex data relationships and transformations (required)
- 2+ years of experience with Python for data engineering tasks (required)
- Experience with Snowflake or an equivalent cloud data warehouse (required)
- Experience working with Snowflake Coco or any other AI tools(required)
- Experience parsing JSON and XML data (a plus)
- A strong eye for data quality and attention to detail
- Knowledge of Git (required)
- Knowledge of Azure cloud services such as Azure Data Factory, Azure Blob Storage, and Azure SQL Database (required)
- Knowledge of data infrastructure/modeling tools like DBT, Fivetran (a plus)
- Experience with BI tools like Power BI(a plus, not core to this role)
- Knowledge of Docker, Linux, Shell/Bash, and virtualization technologies (a plus)
- Knowledge of SSIS packages (a plus)
- Familiarity with CI/CD methodologies
Benefits
Life at Incubyte
We are a remote-first company with structured flexibility. Teams commit to shared rhythms during core hours, ensuring smooth collaboration while maintaining autonomy. Twice a year, we come together in person for a co-working sprint and once a year for a retreat - with all travel expenses covered.
Our environment is built for crafters: pairing, refactoring, experimenting with AI, and pushing the boundaries of software excellence. We are all lifelong learners, and our work is our passion.
Perks
- Dedicated learning & development budget.
- Sponsorship for conference talks.
- Comprehensive medical & term insurance.
- Employee-friendly leave policies.
- Home Office fund
- Medical Insurance
About the Role
We are looking for a Senior Data Engineer with strong hands-on expertise in Databricks, Python, PySpark, and SQL to build scalable, high-performance data engineering solutions. You’ll architect and develop large scale, high-performance data pipelines capable of handling massive real-time and batch data volumes across multiple business systems. Databricks is the core enterprise data and processing platform for this role. You will also use Apache Airflow for workflow orchestration and dbt for ELT transformations, and will contribute to designing reliable, secure, and governed data platforms that enable analytics, reporting, and AI-driven use cases.
Key Responsibilities
- Design and implement large-scale data pipelines using Python/PySpark, Databricks, and Microsoft Fabric.
- Develop and optimize data processing workloads in Databricks using PySpark and Spark SQL, with a strong focus on scalability, reliability, performance, and maintainability.
- Develop and maintain dbt models including layered architecture, incremental models, snapshots, macros, testing, and documentation.
- Design, develop, and maintain Apache Airflow DAGs for orchestrating reliable, scalable, and observable data pipelines.
- Design and implement data quality, observability, and governance frameworks, including automated testing, monitoring, lineage, access control, and data privacy standards.
- Partner with analytics, product, and business stakeholders to turn requirements into trustworthy datasets, and raise the engineering bar through design discussions, code reviews, and mentoring junior engineers.
Required Skills
- Strong expertise in Python for developing scalable, modular, and production-ready data engineering applications.
- Strong expertise in PySpark, including DataFrame API, Spark SQL, Structured Streaming, partitioning strategies, joins, caching, handling data skew, and Spark performance optimization.
- Strong hands-on experience with Databricks for data ingestion, transformation, processing, and optimization, including Delta Lake, Unity Catalog, Databricks Workflows, notebooks, jobs, and Databricks-native data engineering capabilities.
- Strong experience in Databricks/Spark performance tuning, including query and job optimization, partitioning, file sizing, caching, join optimization, handling data skew, and efficient use of compute resources.
- Hands-on experience with Delta Lake, including transactional data processing, schema management, incremental data processing, and reliable batch and streaming data pipelines.
- Hands-on experience in developing dbt projects using layered architecture, incremental models, snapshots, macros/Jinja, testing, documentation, and deployment best practices.
- Expertise in advanced SQL and data modelling — dimensional modeling, slowly changing dimensions, schema evolution, and query optimization.
- Hands-on experience in developing and managing Apache Airflow DAGs, scheduling workflows, dependency management, retries, backfills, and operational monitoring.
- Hands-on experience with at least one major cloud platform (AWS, Azure or GCP).
- Strong problem-solving skills and the ability to work independently with business and analytics stakeholders.
Nice to Have
- Hands-on exposure to Microsoft Fabric for data integration and analytics.
- Experience using AI coding assistants (e.g. Claude Code, GitHub Copilot) as part of a development workflow.
- Familiarity with modern DevOps practices, including CI/CD pipelines, Infrastructure as Code (IaC), and containerization (Docker/Kubernetes).
- Domain expertise in financial services.












