Sr. Data Engineer DataBricks at Exponentia.ai · Mumbai · 4 - 6 years · ₹12L - ₹19L / yr · Bootstrapped · Posted 18 Jul 2023

Job DescriptionPosition: Sr Data Engineer – Databricks & AWS
Experience: 4 - 5 Years
Company Profile:
Exponentia.ai is an AI tech organization with a presence across India, Singapore, the Middle East, and the UK. We are an innovative and disruptive organization, working on cutting-edge technology to help our clients transform into the enterprises of the future. We provide artificial intelligence-based products/platforms capable of automated cognitive decision-making to improve productivity, quality, and economics of the underlying business processes. Currently, we are transforming ourselves and rapidly expanding our business.
Exponentia.ai has developed long-term relationships with world-class clients such as PayPal, PayU, SBI Group, HDFC Life, Kotak Securities, Wockhardt and Adani Group amongst others.
One of the top partners of Cloudera (leading analytics player) and Qlik (leader in BI technologies), Exponentia.ai has recently been awarded the ‘Innovation Partner Award’ by Qlik in 2017.
Get to know more about us on our website: http://www.exponentia.ai/ and Life @Exponentia.
Role Overview:
· A Data Engineer understands the client requirements and develops and delivers the data engineering solutions as per the scope.
· The role requires good skills in the development of solutions using various services required for data architecture on Databricks Delta Lake, streaming, AWS, ETL Development, and data modeling.
Job Responsibilities
• Design of data solutions on Databricks including delta lake, data warehouse, data marts and other data solutions to support the analytics needs of the organization.
• Apply best practices during design in data modeling (logical, physical) and ETL pipelines (streaming and batch) using cloud-based services.
• Design, develop and manage the pipelining (collection, storage, access), data engineering (data quality, ETL, Data Modelling) and understanding (documentation, exploration) of the data.
• Interact with stakeholders regarding data landscape understanding, conducting discovery exercises, developing proof of concepts and demonstrating it to stakeholders.
Technical Skills
• Has more than 2 Years of experience in developing data lakes, and datamarts on the Databricks platform.
• Proven skill sets in AWS Data Lake services such as - AWS Glue, S3, Lambda, SNS, IAM, and skills in Spark, Python, and SQL.
• Experience in Pentaho
• Good understanding of developing a data warehouse, data marts etc.
• Has a good understanding of system architectures, and design patterns and should be able to design and develop applications using these principles.
Personality Traits
• Good collaboration and communication skills
• Excellent problem-solving skills to be able to structure the right analytical solutions.
• Strong sense of teamwork, ownership, and accountability
• Analytical and conceptual thinking
• Ability to work in a fast-paced environment with tight schedules.
• Good presentation skills with the ability to convey complex ideas to peers and management.
Education:
BE / ME / MS/MCA.

About Exponentia.ai
About
Data is the new Oil and AI is the most powerful value accelerator today - this is one of the most important belief we live by. We are an award-winning AI-Tech MNC transforming businesses through AI, ML and Data Analytics Solutions.
We help organizations solve complex business challenges by combining industry experience, data and AI first engineering practices, and our propriety technology solutions to achieve business outcomes at scale.
Our propriety solutions include - OneTAP for Voice Analytics, Intelligent Nudges, Customer Intelligence Platform and Engagely.ai and we specialise in Data (ML, BI, Cloud, DWH & Data Lakes) and AI Solutions (NLP, Conversational Analytics).
Exponentia.ai was founded in the year 2014 and is headquartered in Mumbai, India. We have expanded our services to the UK, Singapore and US. We’ve been honored with the Innovation Award by Qlik & Excellence in Business Process Automation by Automation Anywhere.
Visit our website- www.exponentia.ai to learn more about our products and services.
Tech stack
Product showcase
Connect with the team
Similar jobs (10)
Job Title : Senior Data Engineer – Databricks
Experience : 14 to 20 Years
Location : HSR Layout, Bangalore
Work Mode : Hybrid – 3 Days WFO
Shift : 11:30 AM – 07:30 PM IST
Positions : 2
Notice Period : Immediate Joiners Only
Interview : 1 Technical Round + 2 Client Rounds
Role Overview :
We are looking for a Senior Data Engineer to build and lead enterprise-scale data platforms for a Switzerland-based commodity client.
The role requires a strong hands-on Data Engineering professional with expertise in Databricks, PySpark, Python, SQL, and AWS, along with technical leadership and stakeholder management experience.
Must-Have Skills :
- 14 to 20 years of Data Engineering experience
- Databricks & Apache Spark / PySpark
- Python & SQL
- AWS Cloud
- Lakehouse Architecture
- ETL / ELT & Distributed Data Processing
- Batch & Streaming Pipelines
- Data Pipeline Optimization & Data Modeling
- CDC & Incremental Processing
- Git, CI/CD & Testing
- Data Quality, Monitoring & Observability
- Technical Leadership & Stakeholder Management
Key Responsibilities :
- Design and build scalable data pipelines using Databricks, PySpark, Python, SQL, and AWS.
- Own data products from design through production.
- Develop batch / streaming pipelines and reusable ETL / ELT frameworks.
- Optimize pipelines for performance, scalability, reliability, and cost.
- Design scalable data architectures and data models.
- Implement data quality, monitoring, lineage, and CI/CD practices.
- Lead technical discussions and mentor engineering teams.
- Collaborate with business stakeholders, architects, product owners, and engineering teams.
- Remain hands-on while providing technical leadership.
Ideal Candidate :
A 14 to 20 years experienced, hands-on Data Engineering leader with strong Databricks + PySpark + AWS expertise, excellent communication, stakeholder management, and experience delivering enterprise-scale data platforms.
🔴 Super Urgent : Only Bangalore-based immediate joiners.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Title : Data Engineer – Databricks
Experience : 6+ Years
Location : Noida / Hyderabad / Chennai / Pune / Bengaluru (Hybrid)
Shift : IST (Normal Shift)
Job Summary :
We are seeking an experienced Data Engineer with strong expertise in Databricks, Snowflake, Python, and Spark to build and optimize scalable data pipelines and support AI/ML model deployments. The ideal candidate should have experience working with cloud-based data platforms and preferably possess exposure to the Healthcare domain.
Required Skills :
- Databricks (Preferred)
- Snowflake
- Python
- Apache Spark
- SQL
- Azure Cloud
- Kubernetes
- Apache Airflow
- GitHub & CI/CD Pipelines
- AI/ML Model Deployment
- Data Analytics
Preferred :
- Experience in the Healthcare domain.
- Strong understanding of scalable data engineering architectures and best practices.
About the Role
We are looking for a Senior Data Engineer with strong hands-on expertise in Databricks, Python, PySpark, and SQL to build scalable, high-performance data engineering solutions. You’ll architect and develop large scale, high-performance data pipelines capable of handling massive real-time and batch data volumes across multiple business systems. Databricks is the core enterprise data and processing platform for this role. You will also use Apache Airflow for workflow orchestration and dbt for ELT transformations, and will contribute to designing reliable, secure, and governed data platforms that enable analytics, reporting, and AI-driven use cases.
Key Responsibilities
- Design and implement large-scale data pipelines using Python/PySpark, Databricks, and Microsoft Fabric.
- Develop and optimize data processing workloads in Databricks using PySpark and Spark SQL, with a strong focus on scalability, reliability, performance, and maintainability.
- Develop and maintain dbt models including layered architecture, incremental models, snapshots, macros, testing, and documentation.
- Design, develop, and maintain Apache Airflow DAGs for orchestrating reliable, scalable, and observable data pipelines.
- Design and implement data quality, observability, and governance frameworks, including automated testing, monitoring, lineage, access control, and data privacy standards.
- Partner with analytics, product, and business stakeholders to turn requirements into trustworthy datasets, and raise the engineering bar through design discussions, code reviews, and mentoring junior engineers.
Required Skills
- Strong expertise in Python for developing scalable, modular, and production-ready data engineering applications.
- Strong expertise in PySpark, including DataFrame API, Spark SQL, Structured Streaming, partitioning strategies, joins, caching, handling data skew, and Spark performance optimization.
- Strong hands-on experience with Databricks for data ingestion, transformation, processing, and optimization, including Delta Lake, Unity Catalog, Databricks Workflows, notebooks, jobs, and Databricks-native data engineering capabilities.
- Strong experience in Databricks/Spark performance tuning, including query and job optimization, partitioning, file sizing, caching, join optimization, handling data skew, and efficient use of compute resources.
- Hands-on experience with Delta Lake, including transactional data processing, schema management, incremental data processing, and reliable batch and streaming data pipelines.
- Hands-on experience in developing dbt projects using layered architecture, incremental models, snapshots, macros/Jinja, testing, documentation, and deployment best practices.
- Expertise in advanced SQL and data modelling — dimensional modeling, slowly changing dimensions, schema evolution, and query optimization.
- Hands-on experience in developing and managing Apache Airflow DAGs, scheduling workflows, dependency management, retries, backfills, and operational monitoring.
- Hands-on experience with at least one major cloud platform (AWS, Azure or GCP).
- Strong problem-solving skills and the ability to work independently with business and analytics stakeholders.
Nice to Have
- Hands-on exposure to Microsoft Fabric for data integration and analytics.
- Experience using AI coding assistants (e.g. Claude Code, GitHub Copilot) as part of a development workflow.
- Familiarity with modern DevOps practices, including CI/CD pipelines, Infrastructure as Code (IaC), and containerization (Docker/Kubernetes).
- Domain expertise in financial services.
Responsibilities and JD
Job Description: We are looking for a Senior Developer with strong expertise in PySpark, Databricks, and Snowflake to build scalable data engineering solutions and enterprise data platforms.
Key Responsibilities:
- Design, develop, and maintain ETL/ELT pipelines using PySpark, Databricks, and Snowflake.
- Develop batch and real-time data processing solutions for structured and semi-structured data.
- Build and optimize Databricks notebooks, workflows, and Delta Lake solutions.
- Design and implement Snowflake databases, schemas, views, stored procedures, tasks, and streams.
- Develop scalable data models, data marts, and data warehouse solutions.
- Optimize PySpark jobs, Databricks workloads, and Snowflake queries for performance and cost efficiency.
- Implement data quality, validation, governance, and security controls.
- Collaborate with business stakeholders, architects, and cross-functional teams to deliver data solutions.
- Manage source control and CI/CD deployments using Git and Azure DevOps.
- Troubleshoot production issues, perform root cause analysis, and ensure pipeline reliability.
- Mentor junior team members and participate in code reviews and technical design discussions.
Required Skills: PySpark, Databricks, Snowflake, Python, SQL.
Experience: 5+ years of Data Engineering experience with strong hands-on expertise in PySpark, Databricks, and Snowflake.
4 - 10 years of experience in designing and buildingarchitecting highly resilient data platforms
∙Strong knowledge of data engineering, architecture and data modeling
∙Experience in platforms like Databricks and Snowflake
∙Experience on building applications on cloud (AWS or Azure or Google Cloud)
∙Strong analytical and problem-solving skills
∙Prior experience in developing data or computation intensive (e.g. grid based) backend applications is an
advantage
∙OOP design skills with an understanding or at least personal interest towards the concepts of Functional
Programming
∙Willingness to understand and enhance other people’s code, being able to work in an environment where
developers will oversee and work on wider components also dealing with older “legacy” code
∙Strong programming skills (Java/ Scala / Python) skills with the willingness to pick up the other language if not
already mastered at a sufficient level is important
∙Spring knowledge is an advantage, but in general willingness to learn, work with and even enhance in-house
developed frameworks is a must
∙Prior experience in working with Git, Bitbucket, Jenkins, working with PR-s, using JIRA, following the Scrum Agile
methodology is an advantage
∙Prior knowledge of financial products is an advantage
∙Bachelors or Masters in any relevant field of IT/Engineering area is an advantage
Hiring for Data Engineer - Delivery Manager
Exp : 10 - 15 yrs
Edu : BE/B.Tech/MCA
Work Loation : Hyderabad
.Roles & Responsibilitie:
Own end-to-end delivery of data engineering programs ensuring alignment with business goals, timelines, and quality standards.
Drive execution across multiple data initiatives within Azure and Databricks environments.
Provide technical leadership in designing and implementing scalable data pipelines using Python, PySpark, and Spark.
Required Skills:
Strong experience with Databricks, PySpark, Python, and Spark.
Expertise in Azure Data Services including ADF, ADLS, and Synapse.
Proven experience in delivery management, stakeholder management, and Agile execution
Job Description
• Design and Implement Data Solutions: Lead the design, development, and implementation of scalable and secure Azure-
based data platforms, ensuring integration with various data sources and business systems. Deliver at least 2 major
projects every year with a focus on data engineering best practices.
• Optimize Data Pipelines: Build and optimize end-to-end data pipelines using Azure Data Factory, Azure Databricks, and
Azure Synapse, with an emphasis on automating data workflows. Achieve a 20% reduction in pipeline execution times
within the first 6 months.
• Cloud Infrastructure Management: Manage and maintain the Azure data environment, ensuring high availability, disaster
recovery, and cost optimization. Track and improve system uptime to exceed 99.9% reliability.
• Collaborate with Cross-Functional Teams: Partner with data scientists, data analysts, and business stakeholders to translate
business requirements into efficient data solutions. Facilitate at least 3 collaborative sessions per quarter to address key
business use cases.
• Ensure Data Security & Compliance: Implement data security measures, ensuring compliance with industry regulations
(GDPR, HIPAA, etc.) and company policies. Achieve and maintain full compliance in all data environments within the first
quarter of onboarding.
• Continuous Learning & Knowledge Sharing: Stay up-to-date with emerging Azure technologies, and mentor junior
engineers to promote knowledge sharing. Complete 1 Azure certification annually and conduct at least 2 internal
knowledge-sharing sessions per year.
• Sound knowledge of data governance practices, data quality management, and data security principles.
• Play a pivotal role in shaping our organization's data-driven journey, driving innovation through data analytics and insights.
• Optimize data storage, processing and retrieval mechanisms for performance, cost, and scalability using data storage
services (such as Azure Data Lake Storage, Azure SQL Database, etc.), data processing services (such as Azure Data Bricks,
Azure Synapse, etc.) and data visualization (PowerBI, Qlik, etc.) & integration services (Data API builder, logic apps, etc.)
• Monitor and troubleshoot data platform performance, identify and resolve issues, and provide recommendations for
continuous improvement.
• Collaborate with DevOps teams to automate deployment, configuration, and monitoring processes using Azure DevOps,
PowerShell, or other relevant tools.
• Stay up to date with the latest trends and advancements in cloud data services and provide recommendations on adopting
new technologies or features to enhance the data platform.
• Document technical designs, procedures, and guidelines for data platform engineering and operations
Knowledge, Skills & Experience
Job Experience • Bachelor's degree in Computer Science, Engineering, or a related field. Advanced
degree preferred.
• Proven 6-10 years experience in playing platform engineer or admin role
• Experience with big data technologies such as Apache Spark, Hadoop, or similar
frameworks.
• Solid understanding of cloud computing concepts and experience with cloud
infrastructure management and provisioning.
• Solid understanding of network security concepts and technologies (such as
firewalls, VPNs, intrusion detection/prevention systems, etc.) and data security
concepts and technologies (such as access controls, encryption, observability,
privacy laws/regulations, etc.)
• Experience in a Retail setup is preferred.
Required Skills The position will require someone with the following:
• Strategic Planning
Public
• Communication and Collaboration
• Problem Solving Skills A/B testing & experimentation
• SQL, BI tools, and storytelling with data
We are looking for an AWS Data Engineer to build reliable, scalable data pipelines on AWS.
Responsibilities
- Build ETL and ELT pipelines with AWS Glue and dbt
- Model and optimise data warehouses in Redshift and Snowflake
- Orchestrate workflows with Apache Airflow
- Build streaming ingestion with Amazon Kinesis
- Ensure data quality, monitoring and cost efficiency
Requirements
- 2+ years of data engineering on AWS
- Hands-on with Glue, Redshift or Snowflake, and Airflow
- Strong data modelling and warehousing skills
Job Summary
We are looking for a skilled and experienced Data Engineer to join our growing data team. The ideal candidate will have strong expertise in Python, PySpark, Data Modeling, and Power BI, with hands-on experience in designing, developing, and optimizing scalable data solutions. The role requires working closely with business stakeholders, data architects, and analytics teams to build robust data pipelines and semantic models that enable data-driven decision-making.
Technical Skills
- Strong hands-on experience in Python and PySpark development.
- Expertise in building and optimizing Data Engineering solutions and ETL pipelines.
- Strong understanding of Data Modeling concepts (Star Schema, Snowflake Schema, Dimensional Modeling).
- Experience with Power BI Data Modeling and Semantic Layer development.
- Proficiency in DAX (Data Analysis Expressions).
- Experience designing and managing Semantic Models in Power BI.
- Strong SQL skills and experience working with large datasets.
- Knowledge of data warehousing concepts and best practices.
Preferred Skills
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Exposure to modern data platforms like Databricks.
- Understanding of data governance and data quality frameworks.
About the Role
We are looking for a Forward Deployed Engineer with strong hands-on experience in Databricks and AI-assisted software development using Cursor. The role is suited for an engineer who can work directly with clients and internal teams to understand business problems, rapidly build solutions, and take them from prototype to production.
The ideal candidate should have strong expertise in Python, SQL, Databricks, PySpark, data engineering, APIs, and modern AI-assisted development workflows, along with excellent problem-solving and client-facing skills.
Key Responsibilities
Forward Deployed Engineering
- Work directly with clients and stakeholders to understand business and technical requirements.
- Translate business problems into scalable technical and data solutions.
- Rapidly prototype, test, iterate, and productionize solutions.
- Collaborate with engineering, data, AI, and delivery teams to implement customer solutions.
- Troubleshoot production issues and continuously improve deployed solutions.
- Act as a technical bridge between clients and internal engineering teams.
Databricks & Data Engineering
- Design and develop scalable data solutions using Databricks, PySpark, Python, and SQL.
- Build and optimize ETL/ELT and data processing pipelines.
- Work with Databricks Lakehouse, Delta Lake, Unity Catalog, and Databricks Workflows.
- Develop data ingestion and transformation pipelines for structured and semi-structured data.
- Optimize Databricks workloads for performance, scalability, reliability, and cost.
- Integrate Databricks with databases, APIs, cloud platforms, and enterprise applications.
Cursor & AI-Assisted Development
- Use Cursor and AI-assisted development workflows to accelerate software development, debugging, refactoring, and documentation.
- Effectively use AI coding assistants to understand existing codebases and develop new features.
- Apply appropriate engineering judgment to review, validate, test, and secure AI-generated code.
- Use AI-assisted development for rapid prototyping and proof-of-concept development.
- Work with modern AI/LLM APIs and tools where required for customer solutions.
- Stay current with emerging AI-assisted software engineering practices.
Production & Deployment
- Develop production-ready applications, APIs, and data pipelines.
- Work with Git, CI/CD, APIs, containers, and cloud environments.
- Monitor application and pipeline performance and resolve production issues.
- Ensure solutions meet requirements for scalability, security, reliability, and maintainability.
- Collaborate with Data Scientists and ML Engineers to integrate AI/ML capabilities into production systems.
Required Skills & Experience
- 4+ years of experience in Software Engineering, Data Engineering, AI Engineering, or a related field.
- Strong hands-on experience with Databricks.
- Strong proficiency in Python, PySpark, and SQL.
- Experience with Delta Lake and Lakehouse architecture.
- Experience building production-grade data pipelines.
- Hands-on experience with Cursor or similar AI-powered coding assistants.
- Strong understanding of REST APIs and system integrations.
- Experience working with cloud platforms such as AWS, Azure, or GCP.
- Strong debugging, problem-solving, and analytical skills.
- Excellent communication and client-facing abilities.
Preferred Skills
- Experience with Databricks Unity Catalog, Workflows, and MLflow.
- Experience with Generative AI / LLM applications.
- Knowledge of Claude, OpenAI, Azure OpenAI, or other LLM platforms.
- Experience with RAG, vector databases, embeddings, or AI agents.
- Experience with Docker, Kubernetes, and CI/CD.
- Experience in a consulting, customer-facing engineering, or professional services environment.
- Exposure to Agile/Scrum methodologies.
Key Competencies
- Strong problem-solving and ownership mindset
- Ability to work in ambiguous and fast-paced environments.
- Strong client/stakeholder management skills.
- Ability to understand business requirements and convert them into technical solutions.
- Strong communication and presentation skills.
- Ability to rapidly learn new technologies and tools.
- Comfortable working with AI-assisted development while maintaining high engineering standards.
Education
- Bachelor's or master’s degree in computer science, Information Technology, Engineering, Data Science, or a related discipline.








