Data Engineer at EASEBUZZ · Pune · 2 - 4 years · ₹2L - ₹20L / yr · Raised funding · Posted 17 Mar 2022

Company Profile:
Easebuzz is a payment solutions (fintech organisation) company which enables online merchants to accept, process and disburse payments through developer friendly APIs. We are focusing on building plug n play products including the payment infrastructure to solve complete business problems. Definitely a wonderful place where all the actions related to payments, lending, subscription, eKYC is happening at the same time.
We have been consistently profitable and are constantly developing new innovative products, as a result, we are able to grow 4x over the past year alone. We are well capitalised and have recently closed a fundraise of $4M in March, 2021 from prominent VC firms and angel investors. The company is based out of Pune and has a total strength of 180 employees. Easebuzz’s corporate culture is tied into the vision of building a workplace which breeds open communication and minimal bureaucracy. An equal opportunity employer, we welcome and encourage diversity in the workplace. One thing you can be sure of is that you will be surrounded by colleagues who are committed to helping each other grow.
Easebuzz Pvt. Ltd. has its presence in Pune, Bangalore, Gurugram.
Salary: As per company standards.
Designation: Data Engineering
Location: Pune
Experience with ETL, Data Modeling, and Data Architecture
Design, build and operationalize large scale enterprise data solutions and applications using one or more of AWS data and analytics services in combination with 3rd parties
- Spark, EMR, DynamoDB, RedShift, Kinesis, Lambda, Glue.
Experience with AWS cloud data lake for development of real-time or near real-time use cases
Experience with messaging systems such as Kafka/Kinesis for real time data ingestion and processing
Build data pipeline frameworks to automate high-volume and real-time data delivery
Create prototypes and proof-of-concepts for iterative development.
Experience with NoSQL databases, such as DynamoDB, MongoDB etc
Create and maintain optimal data pipeline architecture,
Identify, design, and implement internal process improvements: automating manual processes, optimizing data delivery, re-designing infrastructure for greater scalability, etc.
Build the infrastructure required for optimal extraction, transformation, and loading of data from a wide variety of data sources using SQL and AWS ‘big data’ technologies.
Build analytics tools that utilize the data pipeline to provide actionable insights into customer acquisition, operational efficiency and other key business performance metrics.
Work with stakeholders including the Executive, Product, Data and Design teams to assist with data-related technical issues and support their data infrastructure needs.
Keep our data separated and secure across national boundaries through multiple data centers and AWS regions.
Create data tools for analytics and data scientist team members that assist them in building and optimizing our product into an innovative industry leader.
Evangelize a very high standard of quality, reliability and performance for data models and algorithms that can be streamlined into the engineering and sciences workflow
Build and enhance data pipeline architecture by designing and implementing data ingestion solutions.
Employment Type
Full-time

Similar jobs (10)
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
About AuxoAI:
AuxoAI is a global platform-based services firm. We help companies—turn their strategies into practical digital and AI solutions. By understanding how our clients make decisions, we use digital and Artificial Intelligence (AI) technologies to drive growth, enhance their operations, improve customer experiences, and provide clear, actionable insights from their data. What We Do We work across various industries such as healthcare, high-tech, consumer packaged goods (CPG), finance etc., and in sales, marketing, and customer support functions.
We help our clients with accelerating their digital and AI journeys through:
• AI Application Development
• Data, Digital and Cloud acceleration using AI
• AI Native Product Engineering
We are seeking a skilled and experienced Data Engineer to join our dynamic team. The ideal candidate will have 6+ years of prior experience in data engineering, with a strong background in AWS (Amazon Web Services) technologies. This role offers an exciting opportunity to work on diverse projects, collaborating with cross-functional teams to design, build, and optimize data pipelines and infrastructure.
Responsibilities:
* Design, develop, and maintain scalable data pipelines and ETL processes leveraging AWS services such as S3, Glue, EMR, Lambda, and Redshift.
* Collaborate with data scientists and analysts to understand data requirements and implement solutions that support analytics and machine learning initiatives.
* Optimize data storage and retrieval mechanisms to ensure performance, reliability, and cost-effectiveness.
* Implement data governance and security best practices to ensure compliance and data integrity.
* Troubleshoot and debug data pipeline issues, providing timely resolution and proactive monitoring.
* Stay abreast of emerging technologies and industry trends, recommending innovative solutions to enhance data engineering capabilities.
Requirements :
* Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
* 6+ years of prior experience in data engineering, with a focus on designing and building data pipelines.
* Proficiency in AWS services, particularly S3, Glue, EMR, Lambda, and Redshift.
* Strong programming skills in languages such as Python, Java, or Scala.
* Experience with SQL and NoSQL databases, data warehousing concepts, and big data technologies.
* Familiarity with containerization technologies (e.g., Docker, Kubernetes) and orchestration tools (e.g., Apache Airflow) is a plus.
Data Engineer
Location: Bengaluru, India (Hybrid)
Employment Type: Full-time
Experience: 3-5 years
Role Overview
What We’re Looking For:
- Bachelor’s degree in Computer Science/Engineering or equivalent experience required.
- Experience designing and shipping cloud services products.
- Experience driving and managing technical and architectural dependencies on AWS Cloud.
- A firm understanding of system architecture, cloud computing, PaaS/SaaS design principles, S3, DynamoDB, RDS mandatory.
- Experience in building or maintaining ETL processes and tools, i.e., AWS Glue or any open-source tool.
- Proven system-level design contribution to a current “Live” (in production / under daily high load) multi-region SaaS or PaaS offering.
- Proven experience with S3, DynamoDB, SQL, and AWS RDS services.
- Proficiency in programming languages such as Python.
- Strong analytical and problem-solving skills.
Required Skills & Experience
- Experience with Python, SQL, and data visualization/exploration tools.
- Familiarity with the AWS ecosystem, specifically S3, DynamoDB, and RDS.
- Communication skills, especially for explaining technical concepts to nontechnical business leaders.
- Ability to work on a dynamic, research-oriented team that has concurrent projects.
- Experience in AWS cost optimization (Savings Plans, Reserved Instances, Spot Instances) and governance frameworks.
- Experience developing solutions using infrastructure orchestration tools (SSM, automation account, Ansible, etc.).
- Excellent leadership, stakeholder management, and communication skills.
What We Offer
- Work with some of the brightest minds in the emerging EV industry.
- Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
- Freedom to suggest, implement, and innovate on systems, processes, and technologies.
- Daily ownership in a high-growth, challenging environment.
- Flexible work environment with hybrid schedules and virtualization options.
- Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Role Summary
We are hiring a Data Engineer / ML Data Pipeline Engineer to build and operate the data backbone of the Enterprise AI platform:
What You'll Own
- Ingestion & ETL/ELT pipelines for heterogeneous project folders (PDF drawings, SVG files, IFC models, BBS.json bar-bending-schedule data, Excel exports, and AI agent output JSON).
- AWS-based data architecture: S3 raw/staging/curated/outputs structuring, partitioning, versioning, and lifecycle management; querying via Athena/Glue and warehousing via Redshift or Snowflake as needed.
- Data validation frameworks: GUID cross-referencing between SVG and BBS data, schema enforcement, duplicate/orphan detection, reference integrity checks, and structured validation reporting.
- Agent run logging & observability: designing the database schema and pipelines that track every AI agent run (inputs, outputs, status, errors, cost, retries, reviewer feedback).
- AI Factory monitoring dashboards: operational dashboards (failure rates, retries, latency, data quality) and business dashboards (throughput, cost per run, rework rate) for Power BI/QuickSight or equivalent.
- ML data pipeline support: dataset preparation, labeling/annotation workflows, human-in-the-loop review tooling, and dataset versioning for models that classify or QC drawing issues.
- APIs: designing and building FastAPI/Flask endpoints to trigger validation runs and expose agent processing status to internal tools.
- Data quality & testing discipline: idempotent pipelines, quarantine/reject handling, regression and reconciliation testing, and root-cause debugging when pipelines or query performance degrade in production.
Key Skills — Non-Negotiable (Must-Have, Strong Level)
- Python — production-grade scripting: file/folder handling, JSON/schema processing, clean error handling, not just notebook-level scripting.
- SQL — strong hands-on ability, including GROUP BY/HAVING for duplicate detection, window functions, and daily aggregate/rate calculations (e.g., success-rate queries).
- AWS S3 data handling — practical experience structuring buckets for raw/staging/curated data, versioning, and avoiding overwrite issues at scale.
- Data validation — demonstrable experience building validation logic (set comparisons, duplicate/missing detection, structured pass/fail reporting), not just "I write assertions."
- ETL/ELT pipeline design — end-to-end ownership of at least one pipeline: source → transform → storage → validation → monitoring → business outcome, with clear articulation of what they personally built.
- Query/warehouse engine judgment — working knowledge of when to use Athena vs. Redshift vs. Snowflake (or equivalent), partitioning, clustering, sort/distribution keys, and storage format trade-offs (Parquet vs. JSON vs. CSV).
Key Skills — Good to Have
- Dashboarding — Power BI / QuickSight (or equivalent) fact/dimension table design, KPI cards, drill-downs; medium-to-strong level is a plus but trainable.
- FastAPI / Flask — building real endpoints with request/response schemas and basic error handling; especially valuable for validation-trigger and agent-status APIs.
- ML data pipeline experience — dataset labeling, annotation platform design, train/test/validation splitting, dataset versioning; strong on the pipeline/data side rather than model training itself.
- Human-in-the-loop / review tooling — experience building or contributing to browser-based labeling/review platforms (session persistence, label schema, export formats).
- Large-scale metadata querying — experience making file discovery fast across large volumes (1,000+ projects, thousands of files each) via metadata index tables, event-based ingestion, or catalog tools like AWS Glue.
Description
We are looking for Senior Data Engineers to join our Data Platform team and build scalable, high-performance data platforms that power data processing, analytics, and downstream applications.
The ideal candidate will have strong experience in distributed data processing, ETL pipelines, and Big Data technologies, with hands-on expertise in Apache Spark and Python Scala.
You will be responsible for designing, developing, and optimizing large-scale data pipelines while collaborating closely with cross-functional engineering teams to build reliable, production-grade data solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL and data processing pipelines for large-scale datasets.
- Build and optimize distributed data applications using Apache Spark and Python Scala.
- Develop reliable, high-performance data pipelines for batch and streaming workloads.
- Design and manage data workflows using Apache Airflow.
- Build and operate data workloads on AWS, with strong usage of Amazon S3 for large-scale data storage.
- Work with large datasets to ensure data quality, consistency, reliability, and performance.
- Collaborate with engineering, product, analytics, and other platform teams to deliver robust data solutions.
- Optimize data workflows for scalability, reliability, performance, and cost efficiency.
- Troubleshoot production issues, identify bottlenecks, and continuously improve platform performance.
Requirements
Candidates who demonstrate:
- 5+ years of experience in Data Engineering, Big Data Engineering, or a similar role.
- Strong hands-on experience with Apache Spark and Scala.
- Experience designing, building, and maintaining large-scale ETL pipelines.
- Strong hands-on experience with AWS, particularly Amazon S3.
- Hands-on experience with Apache Airflow for workflow orchestration and scheduling.
- Strong SQL skills and a solid understanding of distributed data processing concepts.
- Experience working with batch and/or streaming data pipelines.
- Excellent debugging, problem-solving, and performance optimization skills.
- Strong communication and collaboration skills.
Good to Have
- Experience with Databricks and the broader Databricks data platform.
- Familiarity with streaming technologies such as Apache Kafka.
- Experience working on large-scale data platforms handling high-volume data workloads.
- Exposure to additional AWS data services and cloud-native data architectures.
Example:
We are looking for an experienced Data Architect to design, develop, and manage the organization's enterprise data architecture. The candidate will be responsible for building scalable data platforms, ensuring data quality and governance, and supporting business analytics through modern data solutions.
Experience Required
Mention the minimum years of experience.
Example:
- Minimum 12 years of experience in Data Architecture, Data Engineering, or related fields.
- 5+ years of experience in the Banking/Financial Services domain is preferred.
Educational Qualification
Mention the required degree.
Example:
- BE/BTech in Computer Science, Information Technology, Software Engineering, Electronics & Communication Engineering, or equivalent.
- OR MCA/MTech/MSc in Computer Science, IT, or related disciplines.
- MBA is preferred.
Technical Skills
List the skills the candidate must have.
Example:
- AWS, Azure, or GCP
- Data Warehousing (DWH)
- ETL/ELT
- Database Management
- Data Modeling
- Data Analytics
- Data Lakes
- Data Governance
Key Responsibilities
Convert the points you received into simple action statements.
Example:
- Design and maintain enterprise data architecture.
- Develop data warehouses and data lakes.
- Define data standards and governance policies.
- Ensure data quality and security.
- Design ETL/ELT processes.
- Integrate data from multiple systems.
- Plan and execute data migration projects.
- Review existing data architecture and recommend improvements.
- Provide technical guidance to project teams.
- Evaluate new data technologies and tools.
Preferred Skills
These are not mandatory but are an advantage.
Example:
- Banking domain experience
- Strong analytical and problem-solving skills
- Good communication skills
- Leadership and stakeholder management
- Experience mentoring technical teams
About Us:
The QX Impact was launched with a mission to make A.I accessible and affordable and deliver AI Products/Solutions at scale for the enterprises by bringing the power of Data, AI, and Engineering to drive digital transformation. We believe without insights; businesses will continue to face challenges to better understand their customers and even lose them. Secondly, without insights businesses won't’ be able to deliver differentiated products/services; and finally, without insights, businesses can’t achieve a new level of “Operational Excellence” is crucial to remain competitive, meeting rising customer expectations, expanding markets, and digitalization.
Job Summary:
We are looking for a Senior Data Engineer who is creative, collaborative, and adaptable to join our agile team of data scientists, engineers, and UX developers. The role focuses on building and maintaining robust data pipelines to support advanced analytics, data science, and BI solutions.
As a Senior Data Engineer, you will work with internal and external data, collaborate with data scientists, and contribute to the design, development, and deployment of innovative solutions.
Key Responsibilities:
- Design, develop, test, and maintain optimal data pipeline and ETL architectures.
- Map out data systems and define/design required integrations, ETL, BI, and AI systems/processes.
- Prepare and optimize data for predictive and prescriptive modeling.
- Collaborate with teams to integrate ERP data into the enterprise data lake, ensuring seamless flow and quality.
- Enhance cloud data infrastructure on AWS or Azure for scalability and performance.
- Utilize big data tools and frameworks to optimize data acquisition and preparation.
- Build architectures to move data to/from data lakes and data warehouses for advanced analytics.
- Develop and curate data models for analytics, dashboards, and reports.
- Conduct code reviews, maintain production-level code, and implement testing approaches.
- Monitor, troubleshoot, and resolve data ingestion workflows to maintain reliability and uptime.
- Drive innovation and implement efficient new approaches to data engineering tasks.
Must-Have Skills:
- Bachelor’s degree in Computer Science, Mathematics, Engineering, or a related field.
- 5+ years of experience working with enterprise data platforms, including building and managing data lakes.
- 3–5 years of experience designing and implementing data warehouse solutions.
- Expertise in SQL, including developing stored procedures (SP) and applying advanced data design concepts.
- Proficiency in Spark (Python/Scala) and Spark Streaming for real-time data pipelines.
- Experience with AWS or Azure services (e.g., AWS Glue, Azure Data Factory, Redshift, Snowflake).
- Familiarity with big data tools such as Apache Kafka, Apache Spark, or Flink.
- Hands-on experience with orchestration tools (e.g., Apache Airflow, Prefect).
- Knowledge of CI/CD processes, version control (e.g., Git, Jenkins), and deployment automation.
- Strong problem-solving, communication, and collaboration skills.
Good-to-Have Skills:
- Experience in integrating ERP data into data lakes.
- Experience with traditional ETL tools (e.g., Talend, Pentaho).
Competencies:
- Tech Savvy - Anticipating and adopting innovations in business-building digital and technology applications.
- Self-Development - Actively seeking new ways to grow and be challenged using both formal and informal development channels.
- Action Oriented - Taking on new opportunities and tough challenges with a sense of urgency, high energy, and enthusiasm.
- Customer Focus - Building strong customer relationships and delivering customer-centric solutions.
- Optimize Work Processes - Knowing the most effective and efficient processes to get things done, with a focus on continuous improvement.
Why Join Us?
- Be part of a collaborative and agile team driving cutting-edge AI and data engineering solutions.
- Work on impactful projects that make a difference across industries.
- Opportunities for professional growth and continuous learning.
- Competitive salary and benefits package.
Application Details
Ready to make an impact? Apply today and become part of the QX Impact team!
Job Title : Senior Data Engineer – Databricks
Experience : 14 to 20 Years
Location : HSR Layout, Bangalore
Work Mode : Hybrid – 3 Days WFO
Shift : 11:30 AM – 07:30 PM IST
Positions : 2
Notice Period : Immediate Joiners Only
Interview : 1 Technical Round + 2 Client Rounds
Role Overview :
We are looking for a Senior Data Engineer to build and lead enterprise-scale data platforms for a Switzerland-based commodity client.
The role requires a strong hands-on Data Engineering professional with expertise in Databricks, PySpark, Python, SQL, and AWS, along with technical leadership and stakeholder management experience.
Must-Have Skills :
- 14 to 20 years of Data Engineering experience
- Databricks & Apache Spark / PySpark
- Python & SQL
- AWS Cloud
- Lakehouse Architecture
- ETL / ELT & Distributed Data Processing
- Batch & Streaming Pipelines
- Data Pipeline Optimization & Data Modeling
- CDC & Incremental Processing
- Git, CI/CD & Testing
- Data Quality, Monitoring & Observability
- Technical Leadership & Stakeholder Management
Key Responsibilities :
- Design and build scalable data pipelines using Databricks, PySpark, Python, SQL, and AWS.
- Own data products from design through production.
- Develop batch / streaming pipelines and reusable ETL / ELT frameworks.
- Optimize pipelines for performance, scalability, reliability, and cost.
- Design scalable data architectures and data models.
- Implement data quality, monitoring, lineage, and CI/CD practices.
- Lead technical discussions and mentor engineering teams.
- Collaborate with business stakeholders, architects, product owners, and engineering teams.
- Remain hands-on while providing technical leadership.
Ideal Candidate :
A 14 to 20 years experienced, hands-on Data Engineering leader with strong Databricks + PySpark + AWS expertise, excellent communication, stakeholder management, and experience delivering enterprise-scale data platforms.
🔴 Super Urgent : Only Bangalore-based immediate joiners.
Design, develop, and maintain ETL pipelines involving large-scale data.
Develop data processing and analytics applications primarily using PySpark and Python.
Build scalable and distributed data processing solutions using Apache Spark.
Develop and deploy data applications on AWS cloud.
Work with AWS services related to storage, compute, ETL, data warehousing, analytics, and streaming.
Implement distributed storage and processing solutions capable of handling high-volume datasets.
Design data processing applications with a focus on performance, scalability, reliability, and optimization.
Work with both SQL and NoSQL databases for data storage, processing, and analytics.
Write, optimize, and analyze SQL, HQL, and NoSQL queries.
Troubleshoot data pipeline and processing issues and ensure data quality and reliability.
Collaborate with data engineers, analysts, architects, and other technical teams to deliver data-driven solutions.

About the Role
You'll be at the forefront of designing and implementing robust data platform solutions that power advanced analytics, AI, and machine learning. Working with modern cloud technologies, you'll build scalable data foundations that enable clients to make smarter, data-driven decisions.
Key Responsibilities
- Build scalable data pipelines using Snowflake, AWS, GCP, and Databricks.
- Design and optimize data models for AI and machine learning workloads.
- Develop reliable data foundations for MLOps, governance, and data lineage.
- Integrate data from multiple sources into modern data platforms.
- Leverage Snowpark ML and Snowflake's native AI capabilities.
- Ensure data platforms are secure, scalable, and high-performing.
What We're Looking For
- 5+ years of hands-on experience with Snowflake.
- Strong proficiency in SQL and Python.
- Experience with AWS, Azure, or GCP.
- Knowledge of cloud storage services such as S3, ADLS, or GCS.
- Strong understanding of Dimensional Modeling and Data Vault.
- Experience with Scala or Java is a plus.
Tech Stack
- Data Warehouse: Snowflake
- Programming: SQL, Python, Scala (Good to Have), Java (Good to Have)
- Cloud: AWS, Azure, GCP
- Storage: S3, ADLS, GCS
- AI/ML: Snowpark ML, MLOps
Perks & Benefits
- Public Speaking & Communication Program
- Mentoring Program with Senior Support Leads
- 360° Progress Reviews
- Weekly Learning Sessions & Guilds
- Paid Certifications
- Hackathons & Innovation Days
- Recognition & Rewards Programs
- Team Socials & Annual Offsites
- Employee Assistance Program (24/7 Wellbeing Support)
The Data People Shaping Tomorrow
Our client helps organizations unlock the power of data through modern cloud, analytics, and AI solutions. We believe in creating an environment where talented technologists can learn, innovate, and make a real impact while building cutting-edge data platforms for global clients. If you're passionate about data engineering and want to work with the latest technologies in AI, cloud, and analytics, we'd love to hear from you.
Job Summary
We are looking for a skilled and experienced Data Engineer to join our growing data team. The ideal candidate will have strong expertise in Python, PySpark, Data Modeling, and Power BI, with hands-on experience in designing, developing, and optimizing scalable data solutions. The role requires working closely with business stakeholders, data architects, and analytics teams to build robust data pipelines and semantic models that enable data-driven decision-making.
Technical Skills
- Strong hands-on experience in Python and PySpark development.
- Expertise in building and optimizing Data Engineering solutions and ETL pipelines.
- Strong understanding of Data Modeling concepts (Star Schema, Snowflake Schema, Dimensional Modeling).
- Experience with Power BI Data Modeling and Semantic Layer development.
- Proficiency in DAX (Data Analysis Expressions).
- Experience designing and managing Semantic Models in Power BI.
- Strong SQL skills and experience working with large datasets.
- Knowledge of data warehousing concepts and best practices.
Preferred Skills
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Exposure to modern data platforms like Databricks.
- Understanding of data governance and data quality frameworks.





