Data Engineer at Big revolution in the e-gaming industry. (GK1) · Bengaluru (Bangalore) · 2 - 3 years · ₹15L - ₹20L / yr · Posted 10 Jan 2022

- We are looking for a Data Engineer to build the next-generation mobile applications for our world-class fintech product.
- The candidate will be responsible for expanding and optimising our data and data pipeline architecture, as well as optimising data flow and collection for cross-functional teams.
- The ideal candidate is an experienced data pipeline builder and data wrangler who enjoys optimising data systems and building them from the ground up.
- Looking for a person with a strong ability to analyse and provide valuable insights to the product and business team to solve daily business problems.
- You should be able to work in a high-volume environment, have outstanding planning and organisational skills.
Qualifications for Data Engineer
- Working SQL knowledge and experience working with relational databases, query authoring (SQL) as well as working familiarity with a variety of databases.
- Experience building and optimising ‘big data’ data pipelines, architectures, and data sets.
- Experience performing root cause analysis on internal and external data and processes to answer specific business questions and identify opportunities for improvement.
- Strong analytic skills related to working with unstructured datasets. Build processes supporting data transformation, data structures, metadata, dependency and workload management.
- Experience supporting and working with cross-functional teams in a dynamic environment.
- Looking for a candidate with 2-3 years of experience in a Data Engineer role, who is a CS graduate or has an equivalent experience.
What we're looking for?
- Experience with big data tools: Hadoop, Spark, Kafka and other alternate tools.
- Experience with relational SQL and NoSQL databases, including MySql/Postgres and Mongodb.
- Experience with data pipeline and workflow management tools: Luigi, Airflow.
- Experience with AWS cloud services: EC2, EMR, RDS, Redshift.
- Experience with stream-processing systems: Storm, Spark-Streaming.
- Experience with object-oriented/object function scripting languages: Python, Java, Scala.

Similar jobs (10)
Job Summary
Role Overview
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, Advanced SQL, CI/CD, DevOps, and Data Analytics. The ideal candidate should have hands-on experience designing and developing scalable data pipelines, transforming large datasets, and supporting data-driven applications.
Experience with Google Cloud Platform (GCP) will be an added advantage.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Python and SQL.
- Develop complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain reliable data integration workflows across multiple data sources.
- Perform data cleansing, validation, transformation, and quality checks.
- Analyze data and provide insights to support business and technical requirements.
- Implement and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps practices and tools to automate deployments, monitoring, and infrastructure processes.
- Troubleshoot data pipeline failures, performance issues, and production incidents.
- Optimize data processing workflows for performance, scalability, and reliability.
- Collaborate with Data Analysts, Data Scientists, Developers, and other stakeholders.
- Follow best practices for version control, testing, documentation, and deployment.
- Contribute to cloud-based data engineering initiatives, preferably on GCP.
Required Skills
- 5–7 years of hands-on experience in Data Engineering.
- Strong programming skills in Python.
- Strong expertise in Advanced SQL and database concepts.
- Hands-on experience with ETL/ELT processes and data pipelines.
- Good understanding of Data Warehousing and Data Modeling concepts.
- Experience with CI/CD practices and tools.
- Strong understanding of DevOps principles, automation, and deployment processes.
- Strong data analytics and problem-solving skills.
- Experience working with large datasets and performance optimization.
- Good understanding of Git/version control and software development best practices.
Good to Have
- Hands-on experience with Google Cloud Platform (GCP).
- Exposure to GCP data services such as BigQuery, Cloud Storage, Dataflow, Composer, or Pub/Sub.
- Experience with containerization/orchestration technologies such as Docker/Kubernetes.
- Experience with workflow orchestration tools such as Airflow.
- Knowledge of cloud-based data architecture and distributed data processing.
Preferred Candidate Profile
- Strong analytical and problem-solving abilities.
- Good communication and stakeholder management skills.
- Ability to work independently as well as in a collaborative team environment.
- Strong ownership of data pipelines and production systems.
- Candidates who can join at short notice are preferred.
Mandatory Skills
Data Engineer, Python , ETL, GCP, Advanced SQL, Strong Data Analytics skills, CICD, Devops
Must-Have Skills
- Minimum 3 years of experience in Data Engineering / Analytics Engineering / Fintech Data roles
- Must have worked on SMS Parsing, intelligent platform, converting RAW customer SMS data into structured actionable financial signals and enabling downstream usage of SMS derived variables
- Must have established a continuous learning cycle to expand parser coverage
- Experience in Lending / NBFC / Fintech domain
- Experience working with Bureau, SMS, Device, or Banking data
- Strong Python and SQL (production level)
- Experience handling unstructured data (SMS, logs, JSON, APIs)
- Experience building data pipelines, schedulers, and cron jobs
- Strong database design and data modelling skills
- Ability to work in a startup environment with high ownership
- Familiarity with modern platforms like AWS, Snowflake, Google BigQuery, Redshift
Good to Have
- Experience in STPL, especially less than 25K ticket size
- Experience with streaming (Kafka/Kinesis) and orchestration (Airflow or Step Functions)
- Experience with feature stores and risk analytics datasets
- Knowledge of regex, NLP basics for SMS parsing
- Experience supporting real-time decision engines/underwriting systems
Role Summary
This role will be responsible for owning the end-to-end data-structuring layer across the organisation. The individual will transform large volumes of raw, unstructured, and semi-structured data (such as SMS, device, bureau, and app data) into clean, standardised, and analysis-ready datasets. These structured datasets will directly power risk analytics, fraud detection, marketing insights, collections strategy, and policy decisioning.
Key Objective of the Role
Ensure all raw lending data (SMS, Bureau, Device, AA, App logs) is captured, parsed, structured, and stored in a clean analytics-ready format inside databases (PostgreSQL, DynamoDB, AWS stack) so that the Risk and Data Science team can directly use it for feature creation, policy building, and portfolio monitoring.
Core Responsibilities
- End-to-End Data Ownership
- Design, build, and maintain end-to-end data pipelines (batch + streaming) using AWS native services (Glue, Lambda, Step Functions, Kinesis, S3, Athena, Redshift, EMR/Spark, etc.): ingestion
→ parsing → structuring → storage
- Work closely with Tech, Product, and Data Science to define what data should be captured
- Maintain data documentation, data dictionaries, and schema governance
- Ensure data quality, consistency, and version control
- Unstructured Data Processing (Highest Priority)
- Parse raw SMS dumps and categorise into salary, EMI, loan apps, collections, credits, debits, OTP, etc.
- Process device fingerprint, behavioural logs, and vendor data (FinBox, AA, Bureau APIs)
- Convert JSON, logs, and raw API responses into structured feature tables
- Build regex/keyword-based parsers for financial SMS classification
- Feature Implementation (From Risk & Data Science Team)
- Implement feature creation logic provided by Risk/Data Science team
- Translate business and policy logic into SQL/Python pipelines
- Create reusable feature layers for underwriting, fraud, collections, and monitoring
- Maintain a feature store for consistent model and policy usage
- Lending Data Understanding (Domain-Specific Requirement)
- Work with Bureau data
- Structure SMS-derived financial variables (income, stress, EMI signals)
- Work with Account Aggregator and bank transaction datasets
- Understand fintech alternate data used in underwriting and fraud detection
- Data Pipelines & Automation
- Build and maintain ETL/ELT pipelines using Python & SQL
- Create cron jobs for automated data ingestion and feature refresh
- Automate vendor data pulls (Bureau, SMS SDK, AA, device data)
- Ensure low-latency pipelines for real-time underwriting use cases
- Database Structuring & Storage Architecture
- Structure clean datasets in PostgreSQL (analytics layer)
- Manage raw data storage in DynamoDB / S3 data lake
- Design normalized and denormalised tables for risk analytics
- Optimise database performance for large-scale query workloads
- Dashboards & Readable Data Layer
- Create analytics-ready datasets, implement & write Metabase queries and convert into dashboards (Metabase / Power BI)
- Enable self-serve data access for Risk, Business, and Founders
- Support ad-hoc analysis requirements from leadership
- Cross-Functional Collaboration (Very Important)
- The role requires close collaboration with data science, tech, product, and business teams to ensure reliable data pipelines, well-defined schemas, API integrations, logging architecture and high data quality, enabling faster and more accurate decision-making across lending workflows.
Tech Stack (Current Environment)
- AWS Services
- PostgreSQL (Primary analytics DB)
- DynamoDB (Raw/NoSQL storage)
- Python (Pandas, NumPy, ETL frameworks)
- Advanced SQL
- APIs, JSON, and Log Data Handling
Role Summary
We are hiring a Data Engineer / ML Data Pipeline Engineer to build and operate the data backbone of the Enterprise AI platform:
What You'll Own
- Ingestion & ETL/ELT pipelines for heterogeneous project folders (PDF drawings, SVG files, IFC models, BBS.json bar-bending-schedule data, Excel exports, and AI agent output JSON).
- AWS-based data architecture: S3 raw/staging/curated/outputs structuring, partitioning, versioning, and lifecycle management; querying via Athena/Glue and warehousing via Redshift or Snowflake as needed.
- Data validation frameworks: GUID cross-referencing between SVG and BBS data, schema enforcement, duplicate/orphan detection, reference integrity checks, and structured validation reporting.
- Agent run logging & observability: designing the database schema and pipelines that track every AI agent run (inputs, outputs, status, errors, cost, retries, reviewer feedback).
- AI Factory monitoring dashboards: operational dashboards (failure rates, retries, latency, data quality) and business dashboards (throughput, cost per run, rework rate) for Power BI/QuickSight or equivalent.
- ML data pipeline support: dataset preparation, labeling/annotation workflows, human-in-the-loop review tooling, and dataset versioning for models that classify or QC drawing issues.
- APIs: designing and building FastAPI/Flask endpoints to trigger validation runs and expose agent processing status to internal tools.
- Data quality & testing discipline: idempotent pipelines, quarantine/reject handling, regression and reconciliation testing, and root-cause debugging when pipelines or query performance degrade in production.
Key Skills — Non-Negotiable (Must-Have, Strong Level)
- Python — production-grade scripting: file/folder handling, JSON/schema processing, clean error handling, not just notebook-level scripting.
- SQL — strong hands-on ability, including GROUP BY/HAVING for duplicate detection, window functions, and daily aggregate/rate calculations (e.g., success-rate queries).
- AWS S3 data handling — practical experience structuring buckets for raw/staging/curated data, versioning, and avoiding overwrite issues at scale.
- Data validation — demonstrable experience building validation logic (set comparisons, duplicate/missing detection, structured pass/fail reporting), not just "I write assertions."
- ETL/ELT pipeline design — end-to-end ownership of at least one pipeline: source → transform → storage → validation → monitoring → business outcome, with clear articulation of what they personally built.
- Query/warehouse engine judgment — working knowledge of when to use Athena vs. Redshift vs. Snowflake (or equivalent), partitioning, clustering, sort/distribution keys, and storage format trade-offs (Parquet vs. JSON vs. CSV).
Key Skills — Good to Have
- Dashboarding — Power BI / QuickSight (or equivalent) fact/dimension table design, KPI cards, drill-downs; medium-to-strong level is a plus but trainable.
- FastAPI / Flask — building real endpoints with request/response schemas and basic error handling; especially valuable for validation-trigger and agent-status APIs.
- ML data pipeline experience — dataset labeling, annotation platform design, train/test/validation splitting, dataset versioning; strong on the pipeline/data side rather than model training itself.
- Human-in-the-loop / review tooling — experience building or contributing to browser-based labeling/review platforms (session persistence, label schema, export formats).
- Large-scale metadata querying — experience making file discovery fast across large volumes (1,000+ projects, thousands of files each) via metadata index tables, event-based ingestion, or catalog tools like AWS Glue.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
About AuxoAI:
AuxoAI is a global platform-based services firm. We help companies—turn their strategies into practical digital and AI solutions. By understanding how our clients make decisions, we use digital and Artificial Intelligence (AI) technologies to drive growth, enhance their operations, improve customer experiences, and provide clear, actionable insights from their data. What We Do We work across various industries such as healthcare, high-tech, consumer packaged goods (CPG), finance etc., and in sales, marketing, and customer support functions.
We help our clients with accelerating their digital and AI journeys through:
• AI Application Development
• Data, Digital and Cloud acceleration using AI
• AI Native Product Engineering
We are seeking a skilled and experienced Data Engineer to join our dynamic team. The ideal candidate will have 6+ years of prior experience in data engineering, with a strong background in AWS (Amazon Web Services) technologies. This role offers an exciting opportunity to work on diverse projects, collaborating with cross-functional teams to design, build, and optimize data pipelines and infrastructure.
Responsibilities:
* Design, develop, and maintain scalable data pipelines and ETL processes leveraging AWS services such as S3, Glue, EMR, Lambda, and Redshift.
* Collaborate with data scientists and analysts to understand data requirements and implement solutions that support analytics and machine learning initiatives.
* Optimize data storage and retrieval mechanisms to ensure performance, reliability, and cost-effectiveness.
* Implement data governance and security best practices to ensure compliance and data integrity.
* Troubleshoot and debug data pipeline issues, providing timely resolution and proactive monitoring.
* Stay abreast of emerging technologies and industry trends, recommending innovative solutions to enhance data engineering capabilities.
Requirements :
* Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
* 6+ years of prior experience in data engineering, with a focus on designing and building data pipelines.
* Proficiency in AWS services, particularly S3, Glue, EMR, Lambda, and Redshift.
* Strong programming skills in languages such as Python, Java, or Scala.
* Experience with SQL and NoSQL databases, data warehousing concepts, and big data technologies.
* Familiarity with containerization technologies (e.g., Docker, Kubernetes) and orchestration tools (e.g., Apache Airflow) is a plus.
We are looking for a dynamic Data Engineer to join our team of technology enthusiasts. You will leverage data to drive strategic decision-making and pioneering solutions, working with complex datasets, collaborating closely with stakeholders, and transforming data into actionable insights to drive innovation.
Qualifications and Skills:
- Minimum 4 years of experience as a Data Engineer
- Hands-on experience with Azure cloud-based data solutions
- Fabric experience is a must – designing, implementing, and managing data workflows and pipelines
- Expertise in database design and management, including SQL databases such as SQL Server
- Proficient in ETL (Extract, Transform, Load) design for data integration and processing
- Strong knowledge of data modeling principles and techniques
- Experience with Azure Data Factory (ADF) for orchestrating data workflows
- Ability to analyze and translate data into actionable insights, reports, and visualizations
- Proficiency in Power BI for reporting and data visualization
Desirable Skills:
- Experience with Power BI Report Builder / Reporting Services
- Knowledge of statistical analysis or Data Science
- Experience within the UK Insurance industry is a plus
- Python or R coding skills
Responsibilities:
- Implement efficient data exchange between internal and external systems to increase efficiency and reduce re-keying and translation errors
- Support the Broking business by developing high-quality information resources, ensuring data availability and accessibility for decision-making
- Engineer data inputs and outputs from core applications and semi-structured remote service data through data syncs between data lake, ODS (SQL database), and leveraging Fabric and ADF
- Perform data engineering tasks including ingestion, cleansing, and collation from a wide range of internal and external sources
- Implement different methods of streaming data and create reconciliations for datasets
- Build analytical models to support reporting and analytics
- Collaborate with an agile delivery team to work on the backlog of specified work
Job Title : Data Engineer – Databricks
Experience : 6+ Years
Location : Noida / Hyderabad / Chennai / Pune / Bengaluru (Hybrid)
Shift : IST (Normal Shift)
Job Summary :
We are seeking an experienced Data Engineer with strong expertise in Databricks, Snowflake, Python, and Spark to build and optimize scalable data pipelines and support AI/ML model deployments. The ideal candidate should have experience working with cloud-based data platforms and preferably possess exposure to the Healthcare domain.
Required Skills :
- Databricks (Preferred)
- Snowflake
- Python
- Apache Spark
- SQL
- Azure Cloud
- Kubernetes
- Apache Airflow
- GitHub & CI/CD Pipelines
- AI/ML Model Deployment
- Data Analytics
Preferred :
- Experience in the Healthcare domain.
- Strong understanding of scalable data engineering architectures and best practices.
Job Description
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, SQL, CI/CD, and DevOps to design, develop, and maintain scalable data pipelines and data processing solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines.
- Develop data processing solutions using Python.
- Write complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain data ingestion and integration workflows.
- Implement data quality, validation, monitoring, and error-handling processes.
- Develop and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps tools and practices for automated build, deployment, and infrastructure management.
- Collaborate with data analysts, data scientists, software engineers, and business teams.
- Optimize data pipelines for performance, reliability, and scalability.
- Troubleshoot production data issues and ensure timely resolution.
- Follow best practices for version control, code quality, testing, and deployment.
Mandatory Skills
- Python
- ETL
- SQL
- CI/CD
- DevOps
- Git / Version Control
- Strong problem-solving and debugging skills
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Summary
We are seeking a motivated Data Engineer with strong skills in SQL, Python, and Linux to design, build, and maintain scalable data pipelines and support data-driven decision-making. The ideal candidate should have experience working with large datasets, ETL processes, and relational databases while ensuring data quality and performance.
Key Responsibilities
- Design, develop, and maintain ETL/ELT data pipelines.
- Write optimized SQL queries, stored procedures, and database objects.
- Develop Python scripts for data extraction, transformation, and automation.
- Work in Linux environments to manage scripts, cron jobs, and system processes.
- Monitor and troubleshoot data pipeline failures.
- Ensure data integrity, consistency, and quality across systems.
- Collaborate with data analysts, software engineers, and business stakeholders.
- Optimize database performance and query execution.
- Participate in code reviews and follow best engineering practices.
Required Skills
- Strong proficiency in SQL (joins, subqueries, window functions, CTEs, indexing, query optimization).
- Good programming experience in Python.
- Hands-on experience with Linux commands and shell scripting.
- Understanding of ETL/ELT concepts and data warehousing.
- Knowledge of relational databases such as PostgreSQL, MySQL, Oracle, or SQL Server.
- Familiarity with Git for version control.
- Strong problem-solving and analytical skills.
Job Summary
We are looking for a skilled and experienced Data Engineer to join our growing data team. The ideal candidate will have strong expertise in Python, PySpark, Data Modeling, and Power BI, with hands-on experience in designing, developing, and optimizing scalable data solutions. The role requires working closely with business stakeholders, data architects, and analytics teams to build robust data pipelines and semantic models that enable data-driven decision-making.
Technical Skills
- Strong hands-on experience in Python and PySpark development.
- Expertise in building and optimizing Data Engineering solutions and ETL pipelines.
- Strong understanding of Data Modeling concepts (Star Schema, Snowflake Schema, Dimensional Modeling).
- Experience with Power BI Data Modeling and Semantic Layer development.
- Proficiency in DAX (Data Analysis Expressions).
- Experience designing and managing Semantic Models in Power BI.
- Strong SQL skills and experience working with large datasets.
- Knowledge of data warehousing concepts and best practices.
Preferred Skills
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Exposure to modern data platforms like Databricks.
- Understanding of data governance and data quality frameworks.
About the Role
We are seeking motivated Data Engineering Interns to join our team remotely for a 3-month internship. This role is designed for students or recent graduates interested in working with data pipelines, ETL processes, and big data tools. You will gain practical experience in building scalable data solutions. While this is an unpaid internship, interns who successfully complete the program will receive a Completion Certificate and a Letter of Recommendation.
Responsibilities
- Assist in designing and building data pipelines for structured and unstructured data.
- Support ETL (Extract, Transform, Load) processes to prepare data for analytics.
- Work with databases (SQL/NoSQL) for data storage and retrieval.
- Help optimize data workflows for performance and scalability.
- Collaborate with data scientists and analysts to ensure data quality and consistency.
- Document workflows, schemas, and technical processes.
Requirements
- Strong interest in data engineering, databases, and big data systems.
- Basic knowledge of SQL and relational database concepts.
- Familiarity with Python, Java, or Scala for data processing.
- Understanding of ETL concepts and data pipelines.
- Exposure to cloud platforms (AWS, Azure, or GCP) is a plus.
- Familiarity with big data frameworks (Hadoop, Spark, Kafka) is an advantage.
- Good problem-solving skills and ability to work independently in a remote setup.
What You’ll Gain
- Hands-on experience in data engineering and ETL pipelines.
- Exposure to real-world data workflows.
- Mentorship and guidance from experienced engineers.
- Completion Certificate upon successful completion.
- Letter of Recommendation based on performance.
Internship Details
- Duration: 3 months
- Location: Remote (Work from Home)
- Stipend: Unpaid
- Perks: Completion Certificate + Letter of Recommendation







