Data Engineer at Credilio Financial Technologies Pvt. Ltd. · Mumbai · 3 - 8 years · Raised funding · Posted 30 Jul 2026
We are looking for a hands-on Data Engineer to help build and manage our data platform for reporting, analytics, and future data science use cases.
The role will involve working with SQL, Python, PySpark, AWS, ETL/ELT pipelines, data warehouses, and BI/reporting tools. Our architecture may use technologies such as ClickHouse, Redshift, AWS Glue, Airflow, Step Functions, Lambda, S3, and CDC-based replication tools based on scale, cost, and operational needs.
Key Responsibilities
· Design, build, and maintain scalable ETL/ELT data pipelines from multiple databases, applications, and external systems.
· Build raw, cleaned, and business-ready data layers to support reporting, analytics, and future data science use cases.
· Write efficient SQL, Python, and PySpark jobs for data ingestion, transformation, validation, and processing.
· Implement workflow orchestration using Apache Airflow, AWS Glue, Step Functions, or similar tools.
· Work with data warehouses such as ClickHouse, Redshift, Snowflake, BigQuery, or similar.
· Support cross-service reporting as the architecture moves towards independent microservice databases.
· Build reusable reporting tables, aggregates, summaries, and basic data marts.
· Monitor, troubleshoot, and optimize data pipelines, warehouse queries, and processing jobs.
· Implement data quality checks for freshness, completeness, consistency, duplicates, and reconciliation.
· Support BI/reporting needs through tools such as Power BI, Metabase, Superset, Redash, or similar.
· Apply data governance, access control, security, and PII-handling best practices.
· Collaborate with engineering, DevOps, product, business, finance, risk, and support teams.
Required Skills
· Strong expertise in SQL for joins, aggregations, window functions, query optimization, and analytical reporting.
· Hands-on experience with Python for data processing, automation, validation, and scripting.
· Working experience with PySpark / Apache Spark for processing large datasets.
· Good understanding of ETL/ELT pipelines, data warehousing, and data modelling concepts.
· Experience with workflow orchestration using Apache Airflow, AWS Step Functions, AWS Glue, or similar tools.
· Experience with AWS data services such as S3, Glue, Lambda, Step Functions, Redshift, DMS, CloudWatch, or similar.
· Experience with any data warehouse such as ClickHouse, Redshift, Snowflake, BigQuery, or similar.
· Understanding of relational databases, preferably PostgreSQL.
· Ability to debug data mismatches, failed pipelines, slow queries, and data quality issues.
· Exposure to BI tools such as Power BI, Metabase, Superset, Redash, or similar.
Good ownership, problem-solving, communication, and collaboration skills.

About Credilio Financial Technologies Pvt. Ltd.
About
Credilio is founded with the idea of digitizing the distribution value chain of personal finance products. The mission of Credilio is to create a nationwide army of multi-branded trained advisors who can assist in the digital onboarding of loans and cards. Credilio enables a small financial advisor to educate customers and recommend lending products that are best suited to meet their unique requirements and further enables them to subscribe for credit cards or loans digitally.
The company does this by providing a large range of product offerings and a customized recommendation tool through a simple mobile app that empowers advisors to be truly digital and future-ready. In addition, the platform allows for instant online approval by leveraging open banking APIs of Banks and NBFCs, thus making the purchase process transparent and paperless. It has reduced the average prospecting-to-onboarding time to less than 10 minutes in a single session as compared to the industry average of an interrupted process over 2-3 days.
A challenger start-up, Credilio is the brainchild of second-time entrepreneurial duo Mr. Aditya Gupta and Mr. Anand Kapadia. It’s backed by a clutch of Venture Capital funds namely Cornerstone Venture Capital Partners, Exfinity Venture Partners, and Param Capital founder Mukul Agarwal.
The platform currently partners with 20+ leading Banks and NBFCs and has over 10,000 Financial Advisors who have assisted more than 0.5 million customer applications getting processed fully digitally in a span of less than a year. The capital raised will power Credilio’s ambitious plans to meet an ARR of 100 cr. by March 2023 and serve 25 million customers through a million advisors in the next 3 years.
Tech stack
Photos
Connect with the team
Similar jobs (10)
Data Engineer Short Hiring Post
🚨 Hiring: Data Engineer
🔹 Experience: 5–9 Years
🔹 Location: Bangalore / Hyderabad
🔹 Skills: PySpark, Python, SQL, ETL, CI/CD, Data Modeling
🔹 Process: L1 Virtual → L2 F2F Karat Test
🔹 F2F: Bangalore / Hyderabad Location
🔹 Positions: Immediate requirement
⚠️ Note: Candidates must be available for F2F Karat immediately after L1.
#Hiring #DataEngineer #PySpark #Python #SQL #BangaloreJobs #HyderabadJobs #Mphasis #ImmediateJoiners
Skills Referential (Required knowledge, skills and abilities)
Technical Skills:
Python
Pyspark
SQL
ETL Aws, Azure, gcp
Job Summary
We are seeking a skilled Data Engineer to design, build, and maintain scalable data pipelines and infrastructure. The ideal candidate should have strong expertise in SQL, Python, Linux, and modern data engineering practices to support data integration, transformation, and analytics.
Key Responsibilities
- Design, develop, and maintain ETL/ELT data pipelines.
- Write efficient and optimized SQL queries for data extraction, transformation, and reporting.
- Develop automation scripts using Python for data processing and workflow optimization.
- Work with Linux environments for deployment, monitoring, and troubleshooting.
- Ensure data quality, integrity, and reliability across data platforms.
- Collaborate with data analysts, software engineers, and business stakeholders to deliver data solutions.
- Monitor, troubleshoot, and optimize data pipelines for performance and scalability.
- Implement best practices for data security, governance, and documentation.
Required Skills
- Strong experience in Data Engineering concepts and ETL/ELT processes.
- Proficiency in SQL, including query optimization and database design.
- Strong programming skills in Python.
- Hands-on experience with Linux commands, shell scripting, and system administration basics.
- Experience with relational databases such as PostgreSQL, MySQL, SQL Server, or Oracle.
- Familiarity with Git/version control.
- Strong analytical and problem-solving skills.
Preferred Skills
- Experience with cloud platforms (AWS, Azure, or GCP).
- Knowledge of Apache Spark, Airflow, Kafka, or similar data engineering tools.
- Experience with data warehousing solutions and big data technologies.
- Understanding of CI/CD pipelines and containerization (Docker/Kubernetes).
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Relevant certifications in cloud or data engineering are an added advantage.
We are looking for a Data Engineer with at least 1 year of hands-on experience building solutions on Snowflake. The candidate should be comfortable designing, building, and managing reliable data pipelines that move data from multiple sources into a central data platform.
Responsibilities
- Build and maintain data pipelines for ingesting, transforming, and loading data into Snowflake
- Design scalable data models, schemas, tables, and views in Snowflake
- Develop ETL/ELT workflows using SQL, Python, or data orchestration tools
- Integrate data from APIs, databases, files, and third-party platforms
- Monitor pipeline performance, failures, data quality, and freshness
- Optimize Snowflake queries, warehouses, storage, and compute usage
- Implement incremental loads, change data capture, and scheduled workflows
- Work with engineering and business teams to understand data requirements
- Maintain documentation for pipelines, datasets, and data transformations
Requirements
- 1+ year of hands-on experience working with Snowflake
- Strong SQL skills and experience writing complex queries
- Experience building and managing ETL or ELT data pipelines
- Knowledge of data warehousing concepts, dimensional modelling, and data quality
- Experience with Python or another scripting language
- Familiarity with orchestration tools such as Airflow, Dagster, Prefect, dbt, or similar
- Understanding of APIs, relational databases, file formats, and cloud storage
- Ability to troubleshoot pipeline failures and performance issues
- Strong analytical, problem-solving, and communication skills
Good to Have
- Experience with dbt and Snowflake Tasks, Streams, Snowpipe, or Dynamic Tables
- Knowledge of AWS, Azure, or Google Cloud
- Experience with Kafka or other streaming platforms
- Familiarity with CI/CD, Git, monitoring, and data governance practices
- Experience integrating ERP, finance, or operational systems
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Summary
We are seeking a motivated Data Engineer with strong skills in SQL, Python, and Linux to design, build, and maintain scalable data pipelines and support data-driven decision-making. The ideal candidate should have experience working with large datasets, ETL processes, and relational databases while ensuring data quality and performance.
Key Responsibilities
- Design, develop, and maintain ETL/ELT data pipelines.
- Write optimized SQL queries, stored procedures, and database objects.
- Develop Python scripts for data extraction, transformation, and automation.
- Work in Linux environments to manage scripts, cron jobs, and system processes.
- Monitor and troubleshoot data pipeline failures.
- Ensure data integrity, consistency, and quality across systems.
- Collaborate with data analysts, software engineers, and business stakeholders.
- Optimize database performance and query execution.
- Participate in code reviews and follow best engineering practices.
Required Skills
- Strong proficiency in SQL (joins, subqueries, window functions, CTEs, indexing, query optimization).
- Good programming experience in Python.
- Hands-on experience with Linux commands and shell scripting.
- Understanding of ETL/ELT concepts and data warehousing.
- Knowledge of relational databases such as PostgreSQL, MySQL, Oracle, or SQL Server.
- Familiarity with Git for version control.
- Strong problem-solving and analytical skills.
Job Description
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, SQL, CI/CD, and DevOps to design, develop, and maintain scalable data pipelines and data processing solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines.
- Develop data processing solutions using Python.
- Write complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain data ingestion and integration workflows.
- Implement data quality, validation, monitoring, and error-handling processes.
- Develop and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps tools and practices for automated build, deployment, and infrastructure management.
- Collaborate with data analysts, data scientists, software engineers, and business teams.
- Optimize data pipelines for performance, reliability, and scalability.
- Troubleshoot production data issues and ensure timely resolution.
- Follow best practices for version control, code quality, testing, and deployment.
Mandatory Skills
- Python
- ETL
- SQL
- CI/CD
- DevOps
- Git / Version Control
- Strong problem-solving and debugging skills
Description
We are looking for Senior Data Engineers to join our Data Platform team and build scalable, high-performance data platforms that power data processing, analytics, and downstream applications.
The ideal candidate will have strong experience in distributed data processing, ETL pipelines, and Big Data technologies, with hands-on expertise in Apache Spark and Python Scala.
You will be responsible for designing, developing, and optimizing large-scale data pipelines while collaborating closely with cross-functional engineering teams to build reliable, production-grade data solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL and data processing pipelines for large-scale datasets.
- Build and optimize distributed data applications using Apache Spark and Python Scala.
- Develop reliable, high-performance data pipelines for batch and streaming workloads.
- Design and manage data workflows using Apache Airflow.
- Build and operate data workloads on AWS, with strong usage of Amazon S3 for large-scale data storage.
- Work with large datasets to ensure data quality, consistency, reliability, and performance.
- Collaborate with engineering, product, analytics, and other platform teams to deliver robust data solutions.
- Optimize data workflows for scalability, reliability, performance, and cost efficiency.
- Troubleshoot production issues, identify bottlenecks, and continuously improve platform performance.
Requirements
Candidates who demonstrate:
- 5+ years of experience in Data Engineering, Big Data Engineering, or a similar role.
- Strong hands-on experience with Apache Spark and Scala.
- Experience designing, building, and maintaining large-scale ETL pipelines.
- Strong hands-on experience with AWS, particularly Amazon S3.
- Hands-on experience with Apache Airflow for workflow orchestration and scheduling.
- Strong SQL skills and a solid understanding of distributed data processing concepts.
- Experience working with batch and/or streaming data pipelines.
- Excellent debugging, problem-solving, and performance optimization skills.
- Strong communication and collaboration skills.
Good to Have
- Experience with Databricks and the broader Databricks data platform.
- Familiarity with streaming technologies such as Apache Kafka.
- Experience working on large-scale data platforms handling high-volume data workloads.
- Exposure to additional AWS data services and cloud-native data architectures.
About Us:
The QX Impact was launched with a mission to make A.I accessible and affordable and deliver AI Products/Solutions at scale for the enterprises by bringing the power of Data, AI, and Engineering to drive digital transformation. We believe without insights; businesses will continue to face challenges to better understand their customers and even lose them. Secondly, without insights businesses won't’ be able to deliver differentiated products/services; and finally, without insights, businesses can’t achieve a new level of “Operational Excellence” is crucial to remain competitive, meeting rising customer expectations, expanding markets, and digitalization.
Job Summary:
We are looking for a Senior Data Engineer who is creative, collaborative, and adaptable to join our agile team of data scientists, engineers, and UX developers. The role focuses on building and maintaining robust data pipelines to support advanced analytics, data science, and BI solutions.
As a Senior Data Engineer, you will work with internal and external data, collaborate with data scientists, and contribute to the design, development, and deployment of innovative solutions.
Key Responsibilities:
- Design, develop, test, and maintain optimal data pipeline and ETL architectures.
- Map out data systems and define/design required integrations, ETL, BI, and AI systems/processes.
- Prepare and optimize data for predictive and prescriptive modeling.
- Collaborate with teams to integrate ERP data into the enterprise data lake, ensuring seamless flow and quality.
- Enhance cloud data infrastructure on AWS or Azure for scalability and performance.
- Utilize big data tools and frameworks to optimize data acquisition and preparation.
- Build architectures to move data to/from data lakes and data warehouses for advanced analytics.
- Develop and curate data models for analytics, dashboards, and reports.
- Conduct code reviews, maintain production-level code, and implement testing approaches.
- Monitor, troubleshoot, and resolve data ingestion workflows to maintain reliability and uptime.
- Drive innovation and implement efficient new approaches to data engineering tasks.
Must-Have Skills:
- Bachelor’s degree in Computer Science, Mathematics, Engineering, or a related field.
- 5+ years of experience working with enterprise data platforms, including building and managing data lakes.
- 3–5 years of experience designing and implementing data warehouse solutions.
- Expertise in SQL, including developing stored procedures (SP) and applying advanced data design concepts.
- Proficiency in Spark (Python/Scala) and Spark Streaming for real-time data pipelines.
- Experience with AWS or Azure services (e.g., AWS Glue, Azure Data Factory, Redshift, Snowflake).
- Familiarity with big data tools such as Apache Kafka, Apache Spark, or Flink.
- Hands-on experience with orchestration tools (e.g., Apache Airflow, Prefect).
- Knowledge of CI/CD processes, version control (e.g., Git, Jenkins), and deployment automation.
- Strong problem-solving, communication, and collaboration skills.
Good-to-Have Skills:
- Experience in integrating ERP data into data lakes.
- Experience with traditional ETL tools (e.g., Talend, Pentaho).
Competencies:
- Tech Savvy - Anticipating and adopting innovations in business-building digital and technology applications.
- Self-Development - Actively seeking new ways to grow and be challenged using both formal and informal development channels.
- Action Oriented - Taking on new opportunities and tough challenges with a sense of urgency, high energy, and enthusiasm.
- Customer Focus - Building strong customer relationships and delivering customer-centric solutions.
- Optimize Work Processes - Knowing the most effective and efficient processes to get things done, with a focus on continuous improvement.
Why Join Us?
- Be part of a collaborative and agile team driving cutting-edge AI and data engineering solutions.
- Work on impactful projects that make a difference across industries.
- Opportunities for professional growth and continuous learning.
- Competitive salary and benefits package.
Application Details
Ready to make an impact? Apply today and become part of the QX Impact team!
About Us
We believe the future of software development is AI-native — where engineers operate at a higher level of abstraction and quality remains non-negotiable.
Incubyte is a software craft consultancy where the “how” of building software matters as much as the “what”.
We partner with companies of all sizes, from helping enterprises build, scale, and modernize to early-stage founders bring their ideas to life.
Our engineers operate in an AI-native development model, using AI as a collaborator across the SDLC to accelerate development while upholding the discipline of software craftsmanship. Guided by Software Craftsmanship and Extreme Programming practices, we build reliable, maintainable, and scalable systems with speed, without compromising quality. If this way of building software resonates with you, we’d like to talk.
Our Guiding Principles
These principles define how we work at Incubyte. They are non-negotiable.
Relentless Pursuit of Quality with Pragmatism
We build high-quality systems without losing sight of delivery.
Extreme Ownership
We take responsibility end-to-end for decisions, execution, and outcomes.
Proactive Collaboration
We collaborate closely, challenge each other, and solve problems together.
Active Pursuit of Mastery
We continuously improve our craft and raise our bar.
Invite, Give, and Act on Feedback
We seek, give, and act on feedback to get better every day.
Ensuring Client Success
We act as trusted partners and focus on real outcomes, not just output.
Job Description
This is a remote position.
Experience Level
2+ years of experience in SQL, Python, and Snowflake (or equivalent cloud data warehouse), Azure Cloud services.
Role Overview
If you're a Data Craftsperson who takes pride in clean, well-tested data solutions and believes in the principles of Extreme Programming, we'd love to meet you. At Incubyte, we're a DevOps organization where developers own the entire release cycle — you'll get hands-on experience across data engineering, analytics, cloud infrastructure, and direct client communication. This role sits primarily in data engineering (80%) with a meaningful analytics component (20%), supporting our client's data systems end-to-end.
What You'll Do
- Design, build, and maintain data pipelines and infrastructure using SQL and Python
- Work within Snowflake to build and optimize data models supporting business use cases
- Parse and process structured and semi-structured data (JSON, XML) from varied sources
- Diagnose issues across raw, intermediate, and summary tables
- Build SQL queries to support repeatable analytics use cases based on stakeholder requirements
- Investigate and resolve data quality issues, including time-sensitive or urgent ones
- Identify opportunities to consolidate models and maintain a single source of truth (SSOT)
Requirements
What We're Looking For
- 2+ years of experience with SQL and relational databases, with the ability to understand complex data relationships and transformations (required)
- 2+ years of experience with Python for data engineering tasks (required)
- Experience with Snowflake or an equivalent cloud data warehouse (required)
- Experience parsing JSON and XML data (a plus)
- A strong eye for data quality and attention to detail
- Knowledge of Git (required)
- Knowledge of Azure cloud services such as Azure Data Factory, Azure Blob Storage, and Azure SQL Database (required)
- Knowledge of data infrastructure/modeling tools like DBT, Fivetran (a plus)
- Experience with BI tools like Power BI(a plus, not core to this role)
- Knowledge of Docker, Linux, Shell/Bash, and virtualization technologies (a plus)
- Knowledge of SSIS packages (a plus)
- Familiarity with CI/CD methodologies
Benefits
Life at Incubyte
We are a remote-first company with structured flexibility. Teams commit to shared rhythms during core hours, ensuring smooth collaboration while maintaining autonomy. Twice a year, we come together in person for a co-working sprint and once a year for a retreat - with all travel expenses covered.
Our environment is built for crafters: pairing, refactoring, experimenting with AI, and pushing the boundaries of software excellence. We are all lifelong learners, and our work is our passion.
Perks
- Dedicated learning & development budget.
- Sponsorship for conference talks.
- Comprehensive medical & term insurance.
- Employee-friendly leave policies.
- Home Office fund
- Medical Insurance
- Design, build, and maintain scalable ETL/ELT pipelines for batch and real-time data ingestion and transformation.
- Develop and optimize data lake and data warehouse architectures (e.g., Snowflake, BigQuery, Redshift).
- Work with cloud platforms GCP, Azure to manage data infrastructure.
- GCP as mandatory skills
- Collaborate with analytics and product teams to understand data needs and deliver solutions.
- Ensure data quality, reliability, security, and compliance across all data systems.
- Mentor junior data engineers and contribute to best practices and code reviews.
- Monitor and troubleshoot data pipeline performance and resolve data-related issues.
- Automate data validation, monitoring, and alerting processes.
- 8+ years of experience in data engineering or software engineering with a data focus.
- Proficient in SQL and at least one programming language (e.g., Python, Scala, Java).
- Experience with modern data warehousing tools (e.g., Snowflake, Redshift, BigQuery).
- Strong understanding of data modeling, data lakes, and ETL/ELT design.
- Hands-on experience with orchestration tools like Airflow, dbt, or similar.
- Solid experience with cloud data platforms (AWS/GCP/Azure).
- Familiarity with CI/CD pipelines, containerization (Docker/Kubernetes), and version control (Git).
- Experience working in a DevOps or DataOps environment.
- Knowledge of data governance, lineage, and cataloging tools (e.g., Collibra, Alation).
- Familiarity with streaming technologies (Kafka, Spark Streaming, Flink).
- Experience supporting machine learning workflows and data science initiatives.







