Data Engineer at world’s fastest growing consumer internet company · Bengaluru (Bangalore) · 5 - 8 years · ₹20L - ₹35L / yr (ESOP available) · Posted 26 Oct 2021

Data Engineer
at world’s fastest growing consumer internet company
Data Engineer JD:
- Designing, developing, constructing, installing, testing and maintaining the complete data management & processing systems.
- Building highly scalable, robust, fault-tolerant, & secure user data platform adhering to data protection laws.
- Taking care of the complete ETL (Extract, Transform & Load) process.
- Ensuring architecture is planned in such a way that it meets all the business requirements.
- Exploring new ways of using existing data, to provide more insights out of it.
- Proposing ways to improve data quality, reliability & efficiency of the whole system.
- Creating data models to reduce system complexity and hence increase efficiency & reduce cost.
- Introducing new data management tools & technologies into the existing system to make it more efficient.
- Setting up monitoring and alarming on data pipeline jobs to detect failures and anomalies
What do we expect from you?
- BS/MS in Computer Science or equivalent experience
- 5 years of recent experience in Big Data Engineering.
- Good experience in working with Hadoop and Big Data technologies like HDFS, Pig, Hive, Zookeeper, Storm, Spark, Airflow and NoSQL systems
- Excellent programming and debugging skills in Java or Python.
- Apache spark, python, hands on experience in deploying ML models
- Has worked on streaming and realtime pipelines
- Experience with Apache Kafka or has worked with any of Spark Streaming, Flume or Storm
Focus Area:
|
R1 |
Data structure & Algorithms |
|
R2 |
Problem solving + Coding |
|
R3 |
Design (LLD) |

Similar jobs (10)
Job Summary
Role Overview
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, Advanced SQL, CI/CD, DevOps, and Data Analytics. The ideal candidate should have hands-on experience designing and developing scalable data pipelines, transforming large datasets, and supporting data-driven applications.
Experience with Google Cloud Platform (GCP) will be an added advantage.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Python and SQL.
- Develop complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain reliable data integration workflows across multiple data sources.
- Perform data cleansing, validation, transformation, and quality checks.
- Analyze data and provide insights to support business and technical requirements.
- Implement and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps practices and tools to automate deployments, monitoring, and infrastructure processes.
- Troubleshoot data pipeline failures, performance issues, and production incidents.
- Optimize data processing workflows for performance, scalability, and reliability.
- Collaborate with Data Analysts, Data Scientists, Developers, and other stakeholders.
- Follow best practices for version control, testing, documentation, and deployment.
- Contribute to cloud-based data engineering initiatives, preferably on GCP.
Required Skills
- 5–7 years of hands-on experience in Data Engineering.
- Strong programming skills in Python.
- Strong expertise in Advanced SQL and database concepts.
- Hands-on experience with ETL/ELT processes and data pipelines.
- Good understanding of Data Warehousing and Data Modeling concepts.
- Experience with CI/CD practices and tools.
- Strong understanding of DevOps principles, automation, and deployment processes.
- Strong data analytics and problem-solving skills.
- Experience working with large datasets and performance optimization.
- Good understanding of Git/version control and software development best practices.
Good to Have
- Hands-on experience with Google Cloud Platform (GCP).
- Exposure to GCP data services such as BigQuery, Cloud Storage, Dataflow, Composer, or Pub/Sub.
- Experience with containerization/orchestration technologies such as Docker/Kubernetes.
- Experience with workflow orchestration tools such as Airflow.
- Knowledge of cloud-based data architecture and distributed data processing.
Preferred Candidate Profile
- Strong analytical and problem-solving abilities.
- Good communication and stakeholder management skills.
- Ability to work independently as well as in a collaborative team environment.
- Strong ownership of data pipelines and production systems.
- Candidates who can join at short notice are preferred.
Mandatory Skills
Data Engineer, Python , ETL, GCP, Advanced SQL, Strong Data Analytics skills, CICD, Devops
Job Summary
We are looking for a skilled and experienced Data Engineer to join our growing data team. The ideal candidate will have strong expertise in Python, PySpark, Data Modeling, and Power BI, with hands-on experience in designing, developing, and optimizing scalable data solutions. The role requires working closely with business stakeholders, data architects, and analytics teams to build robust data pipelines and semantic models that enable data-driven decision-making.
Technical Skills
- Strong hands-on experience in Python and PySpark development.
- Expertise in building and optimizing Data Engineering solutions and ETL pipelines.
- Strong understanding of Data Modeling concepts (Star Schema, Snowflake Schema, Dimensional Modeling).
- Experience with Power BI Data Modeling and Semantic Layer development.
- Proficiency in DAX (Data Analysis Expressions).
- Experience designing and managing Semantic Models in Power BI.
- Strong SQL skills and experience working with large datasets.
- Knowledge of data warehousing concepts and best practices.
Preferred Skills
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Exposure to modern data platforms like Databricks.
- Understanding of data governance and data quality frameworks.
Job Description
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, SQL, CI/CD, and DevOps to design, develop, and maintain scalable data pipelines and data processing solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines.
- Develop data processing solutions using Python.
- Write complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain data ingestion and integration workflows.
- Implement data quality, validation, monitoring, and error-handling processes.
- Develop and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps tools and practices for automated build, deployment, and infrastructure management.
- Collaborate with data analysts, data scientists, software engineers, and business teams.
- Optimize data pipelines for performance, reliability, and scalability.
- Troubleshoot production data issues and ensure timely resolution.
- Follow best practices for version control, code quality, testing, and deployment.
Mandatory Skills
- Python
- ETL
- SQL
- CI/CD
- DevOps
- Git / Version Control
- Strong problem-solving and debugging skills
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
About AuxoAI:
AuxoAI is a global platform-based services firm. We help companies—turn their strategies into practical digital and AI solutions. By understanding how our clients make decisions, we use digital and Artificial Intelligence (AI) technologies to drive growth, enhance their operations, improve customer experiences, and provide clear, actionable insights from their data. What We Do We work across various industries such as healthcare, high-tech, consumer packaged goods (CPG), finance etc., and in sales, marketing, and customer support functions.
We help our clients with accelerating their digital and AI journeys through:
• AI Application Development
• Data, Digital and Cloud acceleration using AI
• AI Native Product Engineering
We are seeking a skilled and experienced Data Engineer to join our dynamic team. The ideal candidate will have 6+ years of prior experience in data engineering, with a strong background in AWS (Amazon Web Services) technologies. This role offers an exciting opportunity to work on diverse projects, collaborating with cross-functional teams to design, build, and optimize data pipelines and infrastructure.
Responsibilities:
* Design, develop, and maintain scalable data pipelines and ETL processes leveraging AWS services such as S3, Glue, EMR, Lambda, and Redshift.
* Collaborate with data scientists and analysts to understand data requirements and implement solutions that support analytics and machine learning initiatives.
* Optimize data storage and retrieval mechanisms to ensure performance, reliability, and cost-effectiveness.
* Implement data governance and security best practices to ensure compliance and data integrity.
* Troubleshoot and debug data pipeline issues, providing timely resolution and proactive monitoring.
* Stay abreast of emerging technologies and industry trends, recommending innovative solutions to enhance data engineering capabilities.
Requirements :
* Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
* 6+ years of prior experience in data engineering, with a focus on designing and building data pipelines.
* Proficiency in AWS services, particularly S3, Glue, EMR, Lambda, and Redshift.
* Strong programming skills in languages such as Python, Java, or Scala.
* Experience with SQL and NoSQL databases, data warehousing concepts, and big data technologies.
* Familiarity with containerization technologies (e.g., Docker, Kubernetes) and orchestration tools (e.g., Apache Airflow) is a plus.
Description
We are looking for Senior Data Engineers to join our Data Platform team and build scalable, high-performance data platforms that power data processing, analytics, and downstream applications.
The ideal candidate will have strong experience in distributed data processing, ETL pipelines, and Big Data technologies, with hands-on expertise in Apache Spark and Python Scala.
You will be responsible for designing, developing, and optimizing large-scale data pipelines while collaborating closely with cross-functional engineering teams to build reliable, production-grade data solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL and data processing pipelines for large-scale datasets.
- Build and optimize distributed data applications using Apache Spark and Python Scala.
- Develop reliable, high-performance data pipelines for batch and streaming workloads.
- Design and manage data workflows using Apache Airflow.
- Build and operate data workloads on AWS, with strong usage of Amazon S3 for large-scale data storage.
- Work with large datasets to ensure data quality, consistency, reliability, and performance.
- Collaborate with engineering, product, analytics, and other platform teams to deliver robust data solutions.
- Optimize data workflows for scalability, reliability, performance, and cost efficiency.
- Troubleshoot production issues, identify bottlenecks, and continuously improve platform performance.
Requirements
Candidates who demonstrate:
- 5+ years of experience in Data Engineering, Big Data Engineering, or a similar role.
- Strong hands-on experience with Apache Spark and Scala.
- Experience designing, building, and maintaining large-scale ETL pipelines.
- Strong hands-on experience with AWS, particularly Amazon S3.
- Hands-on experience with Apache Airflow for workflow orchestration and scheduling.
- Strong SQL skills and a solid understanding of distributed data processing concepts.
- Experience working with batch and/or streaming data pipelines.
- Excellent debugging, problem-solving, and performance optimization skills.
- Strong communication and collaboration skills.
Good to Have
- Experience with Databricks and the broader Databricks data platform.
- Familiarity with streaming technologies such as Apache Kafka.
- Experience working on large-scale data platforms handling high-volume data workloads.
- Exposure to additional AWS data services and cloud-native data architectures.
Job Summary
We are seeking a skilled Data Engineer to design, build, and maintain scalable data pipelines and infrastructure. The ideal candidate should have strong expertise in SQL, Python, Linux, and modern data engineering practices to support data integration, transformation, and analytics.
Key Responsibilities
- Design, develop, and maintain ETL/ELT data pipelines.
- Write efficient and optimized SQL queries for data extraction, transformation, and reporting.
- Develop automation scripts using Python for data processing and workflow optimization.
- Work with Linux environments for deployment, monitoring, and troubleshooting.
- Ensure data quality, integrity, and reliability across data platforms.
- Collaborate with data analysts, software engineers, and business stakeholders to deliver data solutions.
- Monitor, troubleshoot, and optimize data pipelines for performance and scalability.
- Implement best practices for data security, governance, and documentation.
Required Skills
- Strong experience in Data Engineering concepts and ETL/ELT processes.
- Proficiency in SQL, including query optimization and database design.
- Strong programming skills in Python.
- Hands-on experience with Linux commands, shell scripting, and system administration basics.
- Experience with relational databases such as PostgreSQL, MySQL, SQL Server, or Oracle.
- Familiarity with Git/version control.
- Strong analytical and problem-solving skills.
Preferred Skills
- Experience with cloud platforms (AWS, Azure, or GCP).
- Knowledge of Apache Spark, Airflow, Kafka, or similar data engineering tools.
- Experience with data warehousing solutions and big data technologies.
- Understanding of CI/CD pipelines and containerization (Docker/Kubernetes).
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Relevant certifications in cloud or data engineering are an added advantage.
Role Summary
We are hiring a Data Engineer / ML Data Pipeline Engineer to build and operate the data backbone of the Enterprise AI platform:
What You'll Own
- Ingestion & ETL/ELT pipelines for heterogeneous project folders (PDF drawings, SVG files, IFC models, BBS.json bar-bending-schedule data, Excel exports, and AI agent output JSON).
- AWS-based data architecture: S3 raw/staging/curated/outputs structuring, partitioning, versioning, and lifecycle management; querying via Athena/Glue and warehousing via Redshift or Snowflake as needed.
- Data validation frameworks: GUID cross-referencing between SVG and BBS data, schema enforcement, duplicate/orphan detection, reference integrity checks, and structured validation reporting.
- Agent run logging & observability: designing the database schema and pipelines that track every AI agent run (inputs, outputs, status, errors, cost, retries, reviewer feedback).
- AI Factory monitoring dashboards: operational dashboards (failure rates, retries, latency, data quality) and business dashboards (throughput, cost per run, rework rate) for Power BI/QuickSight or equivalent.
- ML data pipeline support: dataset preparation, labeling/annotation workflows, human-in-the-loop review tooling, and dataset versioning for models that classify or QC drawing issues.
- APIs: designing and building FastAPI/Flask endpoints to trigger validation runs and expose agent processing status to internal tools.
- Data quality & testing discipline: idempotent pipelines, quarantine/reject handling, regression and reconciliation testing, and root-cause debugging when pipelines or query performance degrade in production.
Key Skills — Non-Negotiable (Must-Have, Strong Level)
- Python — production-grade scripting: file/folder handling, JSON/schema processing, clean error handling, not just notebook-level scripting.
- SQL — strong hands-on ability, including GROUP BY/HAVING for duplicate detection, window functions, and daily aggregate/rate calculations (e.g., success-rate queries).
- AWS S3 data handling — practical experience structuring buckets for raw/staging/curated data, versioning, and avoiding overwrite issues at scale.
- Data validation — demonstrable experience building validation logic (set comparisons, duplicate/missing detection, structured pass/fail reporting), not just "I write assertions."
- ETL/ELT pipeline design — end-to-end ownership of at least one pipeline: source → transform → storage → validation → monitoring → business outcome, with clear articulation of what they personally built.
- Query/warehouse engine judgment — working knowledge of when to use Athena vs. Redshift vs. Snowflake (or equivalent), partitioning, clustering, sort/distribution keys, and storage format trade-offs (Parquet vs. JSON vs. CSV).
Key Skills — Good to Have
- Dashboarding — Power BI / QuickSight (or equivalent) fact/dimension table design, KPI cards, drill-downs; medium-to-strong level is a plus but trainable.
- FastAPI / Flask — building real endpoints with request/response schemas and basic error handling; especially valuable for validation-trigger and agent-status APIs.
- ML data pipeline experience — dataset labeling, annotation platform design, train/test/validation splitting, dataset versioning; strong on the pipeline/data side rather than model training itself.
- Human-in-the-loop / review tooling — experience building or contributing to browser-based labeling/review platforms (session persistence, label schema, export formats).
- Large-scale metadata querying — experience making file discovery fast across large volumes (1,000+ projects, thousands of files each) via metadata index tables, event-based ingestion, or catalog tools like AWS Glue.
Design, develop, and maintain ETL pipelines involving large-scale data.
Develop data processing and analytics applications primarily using PySpark and Python.
Build scalable and distributed data processing solutions using Apache Spark.
Develop and deploy data applications on AWS cloud.
Work with AWS services related to storage, compute, ETL, data warehousing, analytics, and streaming.
Implement distributed storage and processing solutions capable of handling high-volume datasets.
Design data processing applications with a focus on performance, scalability, reliability, and optimization.
Work with both SQL and NoSQL databases for data storage, processing, and analytics.
Write, optimize, and analyze SQL, HQL, and NoSQL queries.
Troubleshoot data pipeline and processing issues and ensure data quality and reliability.
Collaborate with data engineers, analysts, architects, and other technical teams to deliver data-driven solutions.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Summary
We are seeking a motivated Data Engineer with strong skills in SQL, Python, and Linux to design, build, and maintain scalable data pipelines and support data-driven decision-making. The ideal candidate should have experience working with large datasets, ETL processes, and relational databases while ensuring data quality and performance.
Key Responsibilities
- Design, develop, and maintain ETL/ELT data pipelines.
- Write optimized SQL queries, stored procedures, and database objects.
- Develop Python scripts for data extraction, transformation, and automation.
- Work in Linux environments to manage scripts, cron jobs, and system processes.
- Monitor and troubleshoot data pipeline failures.
- Ensure data integrity, consistency, and quality across systems.
- Collaborate with data analysts, software engineers, and business stakeholders.
- Optimize database performance and query execution.
- Participate in code reviews and follow best engineering practices.
Required Skills
- Strong proficiency in SQL (joins, subqueries, window functions, CTEs, indexing, query optimization).
- Good programming experience in Python.
- Hands-on experience with Linux commands and shell scripting.
- Understanding of ETL/ELT concepts and data warehousing.
- Knowledge of relational databases such as PostgreSQL, MySQL, Oracle, or SQL Server.
- Familiarity with Git for version control.
- Strong problem-solving and analytical skills.
Data Engineer – Microsoft Fabric
Location: Pune, India
Work Mode: Hybrid
Experience: 6+ Years
Employment Type: Full-time contactor
Compensation: As per market standards, commensurate with experience and expertise
Shift Timings: 2:00 PM – 11:00 PM IST
Notice Period: 0 – 15 days
About the Role
Jade Business Services (JBS) is seeking a Data Engineer – Microsoft Fabric to join our Pune team and work on enterprise-scale data transformation and analytics initiatives.
We are looking for a hands-on Data Engineer with strong experience in Microsoft Fabric, SQL, Python/PySpark and modern data engineering practices. The candidate will be responsible for building scalable data pipelines, implementing Lakehouse and Warehouse solutions, developing data models and supporting governed, reliable and AI-ready data platforms.
The ideal candidate should be comfortable working with architects, engineering teams and client stakeholders to translate business requirements into scalable and production-ready data solutions.
Roles and Responsibilities
- Design and develop data solutions using Microsoft Fabric, including OneLake, Lakehouse, Warehouse and Data Factory pipelines.
- Build and maintain scalable ETL/ELT pipelines for batch and incremental data processing.
- Develop data ingestion and transformation pipelines using Fabric Data Factory, SQL, Python and/or PySpark.
- Implement Medallion Architecture using Bronze, Silver and Gold layers.
- Work with Lakehouse and Fabric Warehouse for enterprise data processing and analytics.
- Develop and maintain data models, tables, views and optimized SQL queries.
- Build and support semantic models for Power BI and analytical workloads.
- Implement data quality, validation, monitoring and error-handling mechanisms.
- Work with metadata, lineage and governance requirements using Microsoft Purview.
- Implement data security, access controls and role-based permissions across data platforms.
- Support Data Product and domain-oriented data architecture principles.
- Follow DataOps practices including CI/CD, deployment, monitoring and production support.
- Troubleshoot pipeline failures, performance issues and data quality problems.
- Optimize data pipelines, queries and storage for performance and cost efficiency.
- Work closely with Data Architects and business stakeholders to understand requirements and implement technical solutions.
- Participate in technical design discussions, code reviews and architecture reviews.
- Maintain technical documentation, data flow diagrams and pipeline documentation.
- Support production deployments, incident resolution and SLA-driven data platform operations.
- Identify opportunities for automation and AI-assisted improvements across data engineering processes.
Qualifications and Skills
- 6+ years of experience in Data Engineering, Data Integration or Data Platform development.
- Strong hands-on experience with Microsoft Fabric.
- Experience with:
- Microsoft Fabric Lakehouse
- Fabric Warehouse
- OneLake
- Fabric Data Factory / Pipelines
- Semantic Models
- Strong understanding of Lakehouse and Medallion Architecture.
- Strong SQL development and query optimization skills.
- Hands-on experience with Python and/or PySpark.
- Experience developing enterprise ETL/ELT and data integration pipelines.
- Experience with batch and incremental data processing.
- Understanding of data modelling concepts including dimensional modelling.
- Knowledge of data quality, metadata, lineage and data governance.
- Working knowledge of Microsoft Purview.
- Understanding of Data Mesh and Data Product concepts.
- Experience with CI/CD, version control, monitoring and DataOps practices.
- Understanding of cloud security, access controls and data privacy.
- Good troubleshooting and problem-solving skills.
- Strong communication skills and ability to work with distributed and client-facing teams.
Preferred Skills
- Microsoft Fabric or Azure Data certifications.
- Experience migrating workloads from Azure Synapse, SQL Server, Databricks or other data platforms to Microsoft Fabric.
- Experience implementing Medallion Architecture on Microsoft Fabric.
- Experience with Power BI and semantic modelling.
- Exposure to AI/ML, Generative AI or Agentic AI use cases on enterprise data platforms.
- Experience working with Data Products or domain-oriented data solutions.
- Experience in Energy & Utilities, Healthcare, Financial Services or Insurance.
- Experience working with US or international enterprise clients.
What We Expect
The ideal candidate should be hands-on first and capable of independently building, troubleshooting and optimizing Fabric data solutions. You should be able to explain the technical decisions behind your implementation and work effectively with architects and engineering teams to deliver production-ready solutions.







