Data Scientist at One of our Premium Client · Chennai · 3 - 8 years · ₹3L - ₹17L / yr · Posted 28 Oct 2022

Job Description – Data Science
Basic Qualification:
- ME/MS from premier institute with a background in Mechanical/Industrial/Chemical/Materials engineering.
- Strong Analytical skills and application of Statistical techniques to problem solving
- Expertise in algorithms, data structures and performance optimization techniques
- Proven track record of demonstrating end to end ownership involving taking an idea from incubator to market
- Minimum years of experience in data analysis (2+), statistical analysis, data mining, algorithms for optimization.
Responsibilities
The Data Engineer/Analyst will
- Work with stakeholders throughout the organization to identify opportunities for leveraging company data to drive business solutions.
- Clear interaction with Business teams including product planning, sales, marketing, finance for defining the projects, objectives.
- Mine and analyze data from company databases to drive optimization and improvement of product and process development, marketing techniques and business strategies
- Coordinate with different R&D and Business teams to implement models and monitor outcomes.
- Mentor team members towards developing quick solutions for business impact.
- Skilled at all stages of the analysis process including defining key business questions, recommending measures, data sources, methodology and study design, dataset creation, analysis execution, interpretation and presentation and publication of results.
- 4+ years’ experience in MNC environment with projects involving ML, DL and/or DS
- Experience in Machine Learning, Data Mining or Machine Intelligence (Artificial Intelligence)
- Knowledge on Microsoft Azure will be desired.
- Expertise in machine learning such as Classification, Data/Text Mining, NLP, Image Processing, Decision Trees, Random Forest, Neural Networks, Deep Learning Algorithms
- Proficient in Python and its various libraries such as Numpy, MatPlotLib, Pandas
- Superior verbal and written communication skills, ability to convey rigorous mathematical concepts and considerations to Business Teams.
- Experience in infra development / building platforms is highly desired.
- A drive to learn and master new technologies and techniques.

Similar jobs (10)
Job Summary
We are looking for a skilled and experienced Data Engineer to join our growing data team. The ideal candidate will have strong expertise in Python, PySpark, Data Modeling, and Power BI, with hands-on experience in designing, developing, and optimizing scalable data solutions. The role requires working closely with business stakeholders, data architects, and analytics teams to build robust data pipelines and semantic models that enable data-driven decision-making.
Technical Skills
- Strong hands-on experience in Python and PySpark development.
- Expertise in building and optimizing Data Engineering solutions and ETL pipelines.
- Strong understanding of Data Modeling concepts (Star Schema, Snowflake Schema, Dimensional Modeling).
- Experience with Power BI Data Modeling and Semantic Layer development.
- Proficiency in DAX (Data Analysis Expressions).
- Experience designing and managing Semantic Models in Power BI.
- Strong SQL skills and experience working with large datasets.
- Knowledge of data warehousing concepts and best practices.
Preferred Skills
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Exposure to modern data platforms like Databricks.
- Understanding of data governance and data quality frameworks.
About Us
We believe the future of software development is AI-native — where engineers operate at a higher level of abstraction and quality remains non-negotiable.
Incubyte is a software craft consultancy where the “how” of building software matters as much as the “what”.
We partner with companies of all sizes, from helping enterprises build, scale, and modernize to early-stage founders bring their ideas to life.
Our engineers operate in an AI-native development model, using AI as a collaborator across the SDLC to accelerate development while upholding the discipline of software craftsmanship. Guided by Software Craftsmanship and Extreme Programming practices, we build reliable, maintainable, and scalable systems with speed, without compromising quality. If this way of building software resonates with you, we’d like to talk.
Our Guiding Principles
These principles define how we work at Incubyte. They are non-negotiable.
Relentless Pursuit of Quality with Pragmatism
We build high-quality systems without losing sight of delivery.
Extreme Ownership
We take responsibility end-to-end for decisions, execution, and outcomes.
Proactive Collaboration
We collaborate closely, challenge each other, and solve problems together.
Active Pursuit of Mastery
We continuously improve our craft and raise our bar.
Invite, Give, and Act on Feedback
We seek, give, and act on feedback to get better every day.
Ensuring Client Success
We act as trusted partners and focus on real outcomes, not just output.
Job Description
This is a remote position.
Experience Level
2+ years of experience in SQL, Python, and Snowflake (or equivalent cloud data warehouse), Azure Cloud services.
Role Overview
If you're a Data Craftsperson who takes pride in clean, well-tested data solutions and believes in the principles of Extreme Programming, we'd love to meet you. At Incubyte, we're a DevOps organization where developers own the entire release cycle — you'll get hands-on experience across data engineering, analytics, cloud infrastructure, and direct client communication. This role sits primarily in data engineering (80%) with a meaningful analytics component (20%), supporting our client's data systems end-to-end.
What You'll Do
- Design, build, and maintain data pipelines and infrastructure using SQL and Python
- Work within Snowflake to build and optimize data models supporting business use cases
- Parse and process structured and semi-structured data (JSON, XML) from varied sources
- Diagnose issues across raw, intermediate, and summary tables
- Build SQL queries to support repeatable analytics use cases based on stakeholder requirements
- Investigate and resolve data quality issues, including time-sensitive or urgent ones
- Identify opportunities to consolidate models and maintain a single source of truth (SSOT)
Requirements
What We're Looking For
- 2+ years of experience with SQL and relational databases, with the ability to understand complex data relationships and transformations (required)
- 2+ years of experience with Python for data engineering tasks (required)
- Experience with Snowflake or an equivalent cloud data warehouse (required)
- Experience working with Snowflake Coco or any other AI tools(required)
- Experience parsing JSON and XML data (a plus)
- A strong eye for data quality and attention to detail
- Knowledge of Git (required)
- Knowledge of Azure cloud services such as Azure Data Factory, Azure Blob Storage, and Azure SQL Database (required)
- Knowledge of data infrastructure/modeling tools like DBT, Fivetran (a plus)
- Experience with BI tools like Power BI(a plus, not core to this role)
- Knowledge of Docker, Linux, Shell/Bash, and virtualization technologies (a plus)
- Knowledge of SSIS packages (a plus)
- Familiarity with CI/CD methodologies
Benefits
Life at Incubyte
We are a remote-first company with structured flexibility. Teams commit to shared rhythms during core hours, ensuring smooth collaboration while maintaining autonomy. Twice a year, we come together in person for a co-working sprint and once a year for a retreat - with all travel expenses covered.
Our environment is built for crafters: pairing, refactoring, experimenting with AI, and pushing the boundaries of software excellence. We are all lifelong learners, and our work is our passion.
Perks
- Dedicated learning & development budget.
- Sponsorship for conference talks.
- Comprehensive medical & term insurance.
- Employee-friendly leave policies.
- Home Office fund
- Medical Insurance
Job Summary
Role Overview
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, Advanced SQL, CI/CD, DevOps, and Data Analytics. The ideal candidate should have hands-on experience designing and developing scalable data pipelines, transforming large datasets, and supporting data-driven applications.
Experience with Google Cloud Platform (GCP) will be an added advantage.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Python and SQL.
- Develop complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain reliable data integration workflows across multiple data sources.
- Perform data cleansing, validation, transformation, and quality checks.
- Analyze data and provide insights to support business and technical requirements.
- Implement and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps practices and tools to automate deployments, monitoring, and infrastructure processes.
- Troubleshoot data pipeline failures, performance issues, and production incidents.
- Optimize data processing workflows for performance, scalability, and reliability.
- Collaborate with Data Analysts, Data Scientists, Developers, and other stakeholders.
- Follow best practices for version control, testing, documentation, and deployment.
- Contribute to cloud-based data engineering initiatives, preferably on GCP.
Required Skills
- 5–7 years of hands-on experience in Data Engineering.
- Strong programming skills in Python.
- Strong expertise in Advanced SQL and database concepts.
- Hands-on experience with ETL/ELT processes and data pipelines.
- Good understanding of Data Warehousing and Data Modeling concepts.
- Experience with CI/CD practices and tools.
- Strong understanding of DevOps principles, automation, and deployment processes.
- Strong data analytics and problem-solving skills.
- Experience working with large datasets and performance optimization.
- Good understanding of Git/version control and software development best practices.
Good to Have
- Hands-on experience with Google Cloud Platform (GCP).
- Exposure to GCP data services such as BigQuery, Cloud Storage, Dataflow, Composer, or Pub/Sub.
- Experience with containerization/orchestration technologies such as Docker/Kubernetes.
- Experience with workflow orchestration tools such as Airflow.
- Knowledge of cloud-based data architecture and distributed data processing.
Preferred Candidate Profile
- Strong analytical and problem-solving abilities.
- Good communication and stakeholder management skills.
- Ability to work independently as well as in a collaborative team environment.
- Strong ownership of data pipelines and production systems.
- Candidates who can join at short notice are preferred.
Mandatory Skills
Data Engineer, Python , ETL, GCP, Advanced SQL, Strong Data Analytics skills, CICD, Devops
This role will be permanent with NAM info and deploy to client location Hyderabad & Pune.
Work Mode: WORK FROM OFFICE
Role Descriptions:
- Perform detailed data analysis and support business decision-making
- Gather and document business requirements and translate them into technical specifications
- Work closely with stakeholders to define data needs and reporting requirements
- Create user stories, functional specifications, and support UAT activities
- Ensure alignment between business objectives and data solutions
Required Skills:
- Strong expertise in SQL and data querying
- Proven experience in data analysis, requirement gathering, and stakeholder management
- Ability to translate business requirements into technical solutions and user stories
- Good understanding of data models, reporting, and analytics concepts
Skills: Business Analysis~ORACLE SQL
Locations: ~HYDERABAD~PUNE~
Desire candidate
- Candidate should have valid PF.
Must-Have Skills
- Minimum 3 years of experience in Data Engineering / Analytics Engineering / Fintech Data roles
- Must have worked on SMS Parsing, intelligent platform, converting RAW customer SMS data into structured actionable financial signals and enabling downstream usage of SMS derived variables
- Must have established a continuous learning cycle to expand parser coverage
- Experience in Lending / NBFC / Fintech domain
- Experience working with Bureau, SMS, Device, or Banking data
- Strong Python and SQL (production level)
- Experience handling unstructured data (SMS, logs, JSON, APIs)
- Experience building data pipelines, schedulers, and cron jobs
- Strong database design and data modelling skills
- Ability to work in a startup environment with high ownership
- Familiarity with modern platforms like AWS, Snowflake, Google BigQuery, Redshift
Good to Have
- Experience in STPL, especially less than 25K ticket size
- Experience with streaming (Kafka/Kinesis) and orchestration (Airflow or Step Functions)
- Experience with feature stores and risk analytics datasets
- Knowledge of regex, NLP basics for SMS parsing
- Experience supporting real-time decision engines/underwriting systems
Role Summary
This role will be responsible for owning the end-to-end data-structuring layer across the organisation. The individual will transform large volumes of raw, unstructured, and semi-structured data (such as SMS, device, bureau, and app data) into clean, standardised, and analysis-ready datasets. These structured datasets will directly power risk analytics, fraud detection, marketing insights, collections strategy, and policy decisioning.
Key Objective of the Role
Ensure all raw lending data (SMS, Bureau, Device, AA, App logs) is captured, parsed, structured, and stored in a clean analytics-ready format inside databases (PostgreSQL, DynamoDB, AWS stack) so that the Risk and Data Science team can directly use it for feature creation, policy building, and portfolio monitoring.
Core Responsibilities
- End-to-End Data Ownership
- Design, build, and maintain end-to-end data pipelines (batch + streaming) using AWS native services (Glue, Lambda, Step Functions, Kinesis, S3, Athena, Redshift, EMR/Spark, etc.): ingestion
→ parsing → structuring → storage
- Work closely with Tech, Product, and Data Science to define what data should be captured
- Maintain data documentation, data dictionaries, and schema governance
- Ensure data quality, consistency, and version control
- Unstructured Data Processing (Highest Priority)
- Parse raw SMS dumps and categorise into salary, EMI, loan apps, collections, credits, debits, OTP, etc.
- Process device fingerprint, behavioural logs, and vendor data (FinBox, AA, Bureau APIs)
- Convert JSON, logs, and raw API responses into structured feature tables
- Build regex/keyword-based parsers for financial SMS classification
- Feature Implementation (From Risk & Data Science Team)
- Implement feature creation logic provided by Risk/Data Science team
- Translate business and policy logic into SQL/Python pipelines
- Create reusable feature layers for underwriting, fraud, collections, and monitoring
- Maintain a feature store for consistent model and policy usage
- Lending Data Understanding (Domain-Specific Requirement)
- Work with Bureau data
- Structure SMS-derived financial variables (income, stress, EMI signals)
- Work with Account Aggregator and bank transaction datasets
- Understand fintech alternate data used in underwriting and fraud detection
- Data Pipelines & Automation
- Build and maintain ETL/ELT pipelines using Python & SQL
- Create cron jobs for automated data ingestion and feature refresh
- Automate vendor data pulls (Bureau, SMS SDK, AA, device data)
- Ensure low-latency pipelines for real-time underwriting use cases
- Database Structuring & Storage Architecture
- Structure clean datasets in PostgreSQL (analytics layer)
- Manage raw data storage in DynamoDB / S3 data lake
- Design normalized and denormalised tables for risk analytics
- Optimise database performance for large-scale query workloads
- Dashboards & Readable Data Layer
- Create analytics-ready datasets, implement & write Metabase queries and convert into dashboards (Metabase / Power BI)
- Enable self-serve data access for Risk, Business, and Founders
- Support ad-hoc analysis requirements from leadership
- Cross-Functional Collaboration (Very Important)
- The role requires close collaboration with data science, tech, product, and business teams to ensure reliable data pipelines, well-defined schemas, API integrations, logging architecture and high data quality, enabling faster and more accurate decision-making across lending workflows.
Tech Stack (Current Environment)
- AWS Services
- PostgreSQL (Primary analytics DB)
- DynamoDB (Raw/NoSQL storage)
- Python (Pandas, NumPy, ETL frameworks)
- Advanced SQL
- APIs, JSON, and Log Data Handling
We are looking for a dynamic Data Engineer to join our team of technology enthusiasts. You will leverage data to drive strategic decision-making and pioneering solutions, working with complex datasets, collaborating closely with stakeholders, and transforming data into actionable insights to drive innovation.
Qualifications and Skills:
- Minimum 5 years of experience as a Data Engineer
- Hands-on experience with Azure cloud-based data solutions
- Fabric experience is a must – designing, implementing, and managing data workflows and pipelines
- Expertise in database design and management, including SQL databases such as SQL Server
- Proficient in ETL (Extract, Transform, Load) design for data integration and processing
- Strong knowledge of data modeling principles and techniques
- Experience with Azure Data Factory (ADF) for orchestrating data workflows
- Ability to analyze and translate data into actionable insights, reports, and visualizations
- Proficiency in Power BI for reporting and data visualization
Desirable Skills:
- Experience with Power BI Report Builder / Reporting Services
- Knowledge of statistical analysis or Data Science
- Experience within the UK Insurance industry is a plus
- Python or R coding skills
Responsibilities:
- Implement efficient data exchange between internal and external systems to increase efficiency and reduce re-keying and translation errors
- Support the Broking business by developing high-quality information resources, ensuring data availability and accessibility for decision-making
- Engineer data inputs and outputs from core applications and semi-structured remote service data through data syncs between data lake, ODS (SQL database), and leveraging Fabric and ADF
- Perform data engineering tasks including ingestion, cleansing, and collation from a wide range of internal and external sources
- Implement different methods of streaming data and create reconciliations for datasets
- Build analytical models to support reporting and analytics
- Collaborate with an agile delivery team to work on the backlog of specified work
About VerbaFlo.ai:
VerbaFlo.ai is a fast-growing AI SaaS startup revolutionizing how businesses leverage AI-powered solutions. As a part of our dynamic team, you’ll work alongside industry leaders and visionaries to drive innovation and execution across multiple functions.
Role Overview:
We are seeking a Senior Business Data Analyst with strong experience in analytics, SQL, and dashboarding (preferably Metabase) who can independently lead complex analytical initiatives, translate business problems into scalable data solutions, and mentor junior analysts. This role demands someone who can think strategically, operate with high ownership, and ensure the company runs on accurate, timely, and actionable insights.
Responsibilities:
Strategic & Cross-Functional Ownership
- Partner with leadership (Product, Ops, Growth, Finance) to translate business goals into analytical frameworks, KPIs, and measurable outcomes.
- Influence strategic decisions by providing data-driven recommendations, forecasting, and scenario modeling.
- Drive adoption of data-first practices across teams and proactively identify high-impact opportunity areas.
Analytics & Dashboarding
- Own end-to-end development of dashboards and analytics systems in Metabase (or similar BI tools).
- Build scalable KPI frameworks, business reports, and automated insights to support day-to-day and long-term decision-making.
- Ensure data availability, accuracy, and reliability across reporting layers.
Advanced Data Analysis
- Write, optimize, and review complex SQL queries for deep-dives, cohort analysis, funnel performance, and product/operations diagnostics.
- Conduct root-cause analysis, hypothesis testing, and generate actionable insights with clear recommendations.
Data Infrastructure Collaboration
- Work closely with engineering/data teams to define data requirements, improve data models, and support robust pipelines.
- Identify data quality issues, define fixes, and ensure consistency across systems and sources.
Leadership & Process Excellence
- Mentor junior analysts, review their work, and establish best practices across analytics.
- Standardize reporting processes, create documentation, and improve analytical efficiency.
Core Requirements:
- 5–8 years of experience as a Business Analyst, Data Analyst, Product Analyst, or similar role in a fast-paced environment.
- Strong proficiency in SQL, relational databases, and building scalable dashboards (Metabase preferred)
- Demonstrated experience in converting raw data into structured analysis, insights, and business recommendations.
- Strong understanding of product funnels, operational metrics, and business workflows.
- Ability to communicate complex analytical findings to both technical and non-technical stakeholders.
- Proven track record of independently driving cross-functional initiatives from problem definition to execution.
- Experience with Python, Git, or BI tools like Looker/Power BI/Tableau.
- Hands-on with data warehouses (BigQuery, Redshift, Snowflake).
- Familiarity with product analytics tools (Mixpanel, GA4, Amplitude).
- Exposure to forecasting, financial modeling, or experimentation (A/B testing).
Why Join Us?
- Work directly with top leadership in a high-impact role.
- Be part of an innovative and fast-growing AI startup.
- Opportunity to take ownership of key projects and drive efficiency.
- A collaborative, ambitious, and fast-paced work environment.
- Perks & Benefits: gym membership benefit, workation policy, and company-sponsored lunch.
If you’re looking for an exciting role that combines strategy, execution, and leadership exposure, we’d love to hear from you!
Apply now to join VerbaFlo.AI on this journey.
Job Summary
We are seeking a skilled Data Engineer to design, build, and maintain scalable data pipelines and infrastructure. The ideal candidate should have strong expertise in SQL, Python, Linux, and modern data engineering practices to support data integration, transformation, and analytics.
Key Responsibilities
- Design, develop, and maintain ETL/ELT data pipelines.
- Write efficient and optimized SQL queries for data extraction, transformation, and reporting.
- Develop automation scripts using Python for data processing and workflow optimization.
- Work with Linux environments for deployment, monitoring, and troubleshooting.
- Ensure data quality, integrity, and reliability across data platforms.
- Collaborate with data analysts, software engineers, and business stakeholders to deliver data solutions.
- Monitor, troubleshoot, and optimize data pipelines for performance and scalability.
- Implement best practices for data security, governance, and documentation.
Required Skills
- Strong experience in Data Engineering concepts and ETL/ELT processes.
- Proficiency in SQL, including query optimization and database design.
- Strong programming skills in Python.
- Hands-on experience with Linux commands, shell scripting, and system administration basics.
- Experience with relational databases such as PostgreSQL, MySQL, SQL Server, or Oracle.
- Familiarity with Git/version control.
- Strong analytical and problem-solving skills.
Preferred Skills
- Experience with cloud platforms (AWS, Azure, or GCP).
- Knowledge of Apache Spark, Airflow, Kafka, or similar data engineering tools.
- Experience with data warehousing solutions and big data technologies.
- Understanding of CI/CD pipelines and containerization (Docker/Kubernetes).
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Relevant certifications in cloud or data engineering are an added advantage.
1st virtual , 2nd round F2F
Python pyspark, SQL, data engineer
5+yrs
Bang/hyderabad
immediate to 15days.
Job Description – Azure Data Engineer
Role: Azure Data Engineer
Experience: 9+ Years
Location: Bangalore / Hyderabad
Notice Period: Immediate to 15 Days
Interview Process: 1st Round – Virtual | 2nd Round – F2F
Mandatory Skills
- Python
- PySpark
- SQL
- Azure Data Engineering
Job Description
We are looking for an experienced Azure Data Engineer with 9+ years of experience and strong hands-on expertise in Python, PySpark, SQL, and Azure Data Engineering.
Key Responsibilities
- Develop and maintain scalable data engineering solutions using Azure.
- Build and optimize data processing pipelines using PySpark and Python.
- Write complex SQL queries for data extraction and transformation.
- Work with Azure data services and cloud-based data platforms.
- Perform data processing, transformation, and integration.
- Troubleshoot data pipeline and production issues.
- Collaborate with technical and business teams to deliver data solutions.
Preferred: Immediate to 15 Days joiners.







