Data Scientist at Perfios · Bengaluru (Bangalore) · 4 - 6 years · ₹4L - ₹15L / yr · Posted 26 Sep 2022

1. ROLE AND RESPONSIBILITIES
1.1. Implement next generation intelligent data platform solutions that help build high performance distributed systems.
1.2. Proactively diagnose problems and envisage long term life of the product focusing on reusable, extensible components.
1.3. Ensure agile delivery processes.
1.4. Work collaboratively with stake holders including product and engineering teams.
1.5. Build best-practices in the engineering team.
2. PRIMARY SKILL REQUIRED
2.1. Having a 2-6 years of core software product development experience.
2.2. Experience of working with data-intensive projects, with a variety of technology stacks including different programming languages (Java,
Python, Scala)
2.3. Experience in building infrastructure required for optimal extraction, transformation, and loading of data from a wide variety of data
sources to support other teams to run pipelines/jobs/reports etc.
2.4. Experience in Open-source stack
2.5. Experiences of working with RDBMS databases, NoSQL Databases
2.6. Knowledge of enterprise data lakes, data analytics, reporting, in-memory data handling, etc.
2.7. Have core computer science academic background
2.8. Aspire to continue to pursue career in technical stream
3. Optional Skill Required:
3.1. Understanding of Big Data technologies and Machine learning/Deep learning
3.2. Understanding of diverse set of databases like MongoDB, Cassandra, Redshift, Postgres, etc.
3.3. Understanding of Cloud Platform: AWS, Azure, GCP, etc.
3.4. Experience in BFSI domain is a plus.
4. PREFERRED SKILLS
4.1. A Startup mentality: comfort with ambiguity, a willingness to test, learn and improve rapidl

Similar jobs (10)
Job Description
• Design and Implement Data Solutions: Lead the design, development, and implementation of scalable and secure Azure-
based data platforms, ensuring integration with various data sources and business systems. Deliver at least 2 major
projects every year with a focus on data engineering best practices.
• Optimize Data Pipelines: Build and optimize end-to-end data pipelines using Azure Data Factory, Azure Databricks, and
Azure Synapse, with an emphasis on automating data workflows. Achieve a 20% reduction in pipeline execution times
within the first 6 months.
• Cloud Infrastructure Management: Manage and maintain the Azure data environment, ensuring high availability, disaster
recovery, and cost optimization. Track and improve system uptime to exceed 99.9% reliability.
• Collaborate with Cross-Functional Teams: Partner with data scientists, data analysts, and business stakeholders to translate
business requirements into efficient data solutions. Facilitate at least 3 collaborative sessions per quarter to address key
business use cases.
• Ensure Data Security & Compliance: Implement data security measures, ensuring compliance with industry regulations
(GDPR, HIPAA, etc.) and company policies. Achieve and maintain full compliance in all data environments within the first
quarter of onboarding.
• Continuous Learning & Knowledge Sharing: Stay up-to-date with emerging Azure technologies, and mentor junior
engineers to promote knowledge sharing. Complete 1 Azure certification annually and conduct at least 2 internal
knowledge-sharing sessions per year.
• Sound knowledge of data governance practices, data quality management, and data security principles.
• Play a pivotal role in shaping our organization's data-driven journey, driving innovation through data analytics and insights.
• Optimize data storage, processing and retrieval mechanisms for performance, cost, and scalability using data storage
services (such as Azure Data Lake Storage, Azure SQL Database, etc.), data processing services (such as Azure Data Bricks,
Azure Synapse, etc.) and data visualization (PowerBI, Qlik, etc.) & integration services (Data API builder, logic apps, etc.)
• Monitor and troubleshoot data platform performance, identify and resolve issues, and provide recommendations for
continuous improvement.
• Collaborate with DevOps teams to automate deployment, configuration, and monitoring processes using Azure DevOps,
PowerShell, or other relevant tools.
• Stay up to date with the latest trends and advancements in cloud data services and provide recommendations on adopting
new technologies or features to enhance the data platform.
• Document technical designs, procedures, and guidelines for data platform engineering and operations
Knowledge, Skills & Experience
Job Experience • Bachelor's degree in Computer Science, Engineering, or a related field. Advanced
degree preferred.
• Proven 6-10 years experience in playing platform engineer or admin role
• Experience with big data technologies such as Apache Spark, Hadoop, or similar
frameworks.
• Solid understanding of cloud computing concepts and experience with cloud
infrastructure management and provisioning.
• Solid understanding of network security concepts and technologies (such as
firewalls, VPNs, intrusion detection/prevention systems, etc.) and data security
concepts and technologies (such as access controls, encryption, observability,
privacy laws/regulations, etc.)
• Experience in a Retail setup is preferred.
Required Skills The position will require someone with the following:
• Strategic Planning
Public
• Communication and Collaboration
• Problem Solving Skills A/B testing & experimentation
• SQL, BI tools, and storytelling with data
Roles & Responsibilities
- Design, develop, and deliver scalable end-to-end data pipelines using Azure Data Factory, ensuring robust integration
of enterprise-wide data from diverse sources
• Build and optimize data engineering workflows using Databricks and PySpark
• Write efficient, high-performance SQL for data transformation and analysis
• Work with the Azure Cloud platform and associated services, applying strong understanding of data warehousing,
data models, and pipelines
• Provide technical leadership to a team of developers, including code reviews and enforcing best practices across the
development lifecycle
• Oversee CI/CD implementation using Azure DevOps, managing deployments across development, QA, and production
environments with proper change control processes
• Collaborate with cross-functional teams to translate business requirements into scalable data solutions
• Ensure data quality, reliability, and performance across all pipelines and platforms
Ideal Candidate
1Strong Azure Databricks Engineer / Senior Data Engineer Profile
2Mandatory (Experience 1) – Must have minimum 8+ years of overall experience in Data Engineering, Data Development, or related data technology roles, with strong hands-on experience in enterprise data pipeline development.
3Mandatory (Experience 2) – Must have strong hands-on experience with Azure Databricks, including development and optimization of scalable data engineering workflows using Databricks and PySpark.
4Mandatory (Experience 3) – Must have strong hands-on proficiency in PySpark/Python and SQL, with proven experience developing complex data transformations, processing workflows, and performance-optimized queries.
5Mandatory (Experience 4) – Must have hands-on experience with Azure Data Factory (ADF) for designing, developing, and orchestrating end-to-end data pipelines and integrating data from multiple sources.
6Mandatory (Experience 5) – Must have strong experience working on the Azure Cloud platform and associated data services, with solid understanding of data warehousing, data modeling, pipeline architecture, and enterprise data solutions.
7Mandatory (Experience 6) – Must have hands-on experience implementing CI/CD using Azure DevOps, including deployment and release management across development, QA, and production environments.
8Mandatory (Experience 7) – Must have proven technical leadership experience, including code reviews, enforcing development best practices, mentoring developers, and providing technical guidance to a data engineering team.
9Mandatory (Notice Period) – Immediate joiners or candidates who can join within 15 days.
10Mandatory (Note) - The position is open across all Cognizant offices pan India. Candidates must be willing to attend the F2F interview at the nearest Cognizant office location.
Job Summary
The Technical Lead will be responsible for overseeing and leading projects related to Azure Data Factory (ADF), Azure Databricks, SQL, Oracle PL/SQL, and Python. The role involves designing, developing, and implementing data solutions while ensuring they meet the business requirements and align with best practices. (1.) Key Responsibilities
1. Lead and manage end-to-end data engineering projects using azure data factory, azure databricks, sql, oracle pl/sql, and python.
2. Collaborate with stakeholders to gather and understand requirements for data pipelines and analytics solutions.
3. Design and develop etl processes, data models, and data integration solutions.
4. Provide technical guidance and mentorship to the team members.
5. Ensure data quality, data governance, and data security standards are maintained throughout the project lifecycle.
6. Troubleshoot and optimize data pipelines and processes for performance and efficiency.
7. Stay updated on the latest trends and technologies in data engineering and contribute to continuous improvement efforts.
Skill Requirements
1. Proficiency in azure data factory (adf) and azure databricks for building and managing data pipelines.
2. Strong experience with sql and oracle pl/sql for data querying and manipulation.
3. Advanced programming skills in python for scripting and data processing tasks.
4. Knowledge of data modeling, data warehousing concepts, and database design principles.
5. Ability to work in a collaborative team environment and communicate effectively with stakeholders.
6. Strong analytical and problem-solving skills with attention to detail.
7. Experience in data visualization tools and techniques is a plus.
Certifications: Relevant certifications in Azure Data Factory, Azure Databricks, SQL, Oracle PL/SQL, or Python are advantageous.
Skill (Primary)
Data Fabric-Azure-Azure Data Factory (ADF)
Hiring for Data Engineer - Delivery Manager
Exp : 10 - 15 yrs
Edu : BE/B.Tech/MCA
Work Loation : Hyderabad
.Roles & Responsibilitie:
Own end-to-end delivery of data engineering programs ensuring alignment with business goals, timelines, and quality standards.
Drive execution across multiple data initiatives within Azure and Databricks environments.
Provide technical leadership in designing and implementing scalable data pipelines using Python, PySpark, and Spark.
Required Skills:
Strong experience with Databricks, PySpark, Python, and Spark.
Expertise in Azure Data Services including ADF, ADLS, and Synapse.
Proven experience in delivery management, stakeholder management, and Agile execution
Data Engineer – Microsoft Fabric
Location: Pune, India
Work Mode: Hybrid
Experience: 6+ Years
Employment Type: Full-time contactor
Compensation: As per market standards, commensurate with experience and expertise
Shift Timings: 2:00 PM – 11:00 PM IST
Notice Period: 0 – 15 days
About the Role
Jade Business Services (JBS) is seeking a Data Engineer – Microsoft Fabric to join our Pune team and work on enterprise-scale data transformation and analytics initiatives.
We are looking for a hands-on Data Engineer with strong experience in Microsoft Fabric, SQL, Python/PySpark and modern data engineering practices. The candidate will be responsible for building scalable data pipelines, implementing Lakehouse and Warehouse solutions, developing data models and supporting governed, reliable and AI-ready data platforms.
The ideal candidate should be comfortable working with architects, engineering teams and client stakeholders to translate business requirements into scalable and production-ready data solutions.
Roles and Responsibilities
- Design and develop data solutions using Microsoft Fabric, including OneLake, Lakehouse, Warehouse and Data Factory pipelines.
- Build and maintain scalable ETL/ELT pipelines for batch and incremental data processing.
- Develop data ingestion and transformation pipelines using Fabric Data Factory, SQL, Python and/or PySpark.
- Implement Medallion Architecture using Bronze, Silver and Gold layers.
- Work with Lakehouse and Fabric Warehouse for enterprise data processing and analytics.
- Develop and maintain data models, tables, views and optimized SQL queries.
- Build and support semantic models for Power BI and analytical workloads.
- Implement data quality, validation, monitoring and error-handling mechanisms.
- Work with metadata, lineage and governance requirements using Microsoft Purview.
- Implement data security, access controls and role-based permissions across data platforms.
- Support Data Product and domain-oriented data architecture principles.
- Follow DataOps practices including CI/CD, deployment, monitoring and production support.
- Troubleshoot pipeline failures, performance issues and data quality problems.
- Optimize data pipelines, queries and storage for performance and cost efficiency.
- Work closely with Data Architects and business stakeholders to understand requirements and implement technical solutions.
- Participate in technical design discussions, code reviews and architecture reviews.
- Maintain technical documentation, data flow diagrams and pipeline documentation.
- Support production deployments, incident resolution and SLA-driven data platform operations.
- Identify opportunities for automation and AI-assisted improvements across data engineering processes.
Qualifications and Skills
- 6+ years of experience in Data Engineering, Data Integration or Data Platform development.
- Strong hands-on experience with Microsoft Fabric.
- Experience with:
- Microsoft Fabric Lakehouse
- Fabric Warehouse
- OneLake
- Fabric Data Factory / Pipelines
- Semantic Models
- Strong understanding of Lakehouse and Medallion Architecture.
- Strong SQL development and query optimization skills.
- Hands-on experience with Python and/or PySpark.
- Experience developing enterprise ETL/ELT and data integration pipelines.
- Experience with batch and incremental data processing.
- Understanding of data modelling concepts including dimensional modelling.
- Knowledge of data quality, metadata, lineage and data governance.
- Working knowledge of Microsoft Purview.
- Understanding of Data Mesh and Data Product concepts.
- Experience with CI/CD, version control, monitoring and DataOps practices.
- Understanding of cloud security, access controls and data privacy.
- Good troubleshooting and problem-solving skills.
- Strong communication skills and ability to work with distributed and client-facing teams.
Preferred Skills
- Microsoft Fabric or Azure Data certifications.
- Experience migrating workloads from Azure Synapse, SQL Server, Databricks or other data platforms to Microsoft Fabric.
- Experience implementing Medallion Architecture on Microsoft Fabric.
- Experience with Power BI and semantic modelling.
- Exposure to AI/ML, Generative AI or Agentic AI use cases on enterprise data platforms.
- Experience working with Data Products or domain-oriented data solutions.
- Experience in Energy & Utilities, Healthcare, Financial Services or Insurance.
- Experience working with US or international enterprise clients.
What We Expect
The ideal candidate should be hands-on first and capable of independently building, troubleshooting and optimizing Fabric data solutions. You should be able to explain the technical decisions behind your implementation and work effectively with architects and engineering teams to deliver production-ready solutions.
4 - 10 years of experience in designing and buildingarchitecting highly resilient data platforms
∙Strong knowledge of data engineering, architecture and data modeling
∙Experience in platforms like Databricks and Snowflake
∙Experience on building applications on cloud (AWS or Azure or Google Cloud)
∙Strong analytical and problem-solving skills
∙Prior experience in developing data or computation intensive (e.g. grid based) backend applications is an
advantage
∙OOP design skills with an understanding or at least personal interest towards the concepts of Functional
Programming
∙Willingness to understand and enhance other people’s code, being able to work in an environment where
developers will oversee and work on wider components also dealing with older “legacy” code
∙Strong programming skills (Java/ Scala / Python) skills with the willingness to pick up the other language if not
already mastered at a sufficient level is important
∙Spring knowledge is an advantage, but in general willingness to learn, work with and even enhance in-house
developed frameworks is a must
∙Prior experience in working with Git, Bitbucket, Jenkins, working with PR-s, using JIRA, following the Scrum Agile
methodology is an advantage
∙Prior knowledge of financial products is an advantage
∙Bachelors or Masters in any relevant field of IT/Engineering area is an advantage
Azure Data Factory and Azure Databricks, processing 5 million+ records/week from 4+ source systems into a governed Lakehouse.
Job Summary
We are looking for a skilled and experienced Data Engineer to join our growing data team. The ideal candidate will have strong expertise in Python, PySpark, Data Modeling, and Power BI, with hands-on experience in designing, developing, and optimizing scalable data solutions. The role requires working closely with business stakeholders, data architects, and analytics teams to build robust data pipelines and semantic models that enable data-driven decision-making.
Technical Skills
- Strong hands-on experience in Python and PySpark development.
- Expertise in building and optimizing Data Engineering solutions and ETL pipelines.
- Strong understanding of Data Modeling concepts (Star Schema, Snowflake Schema, Dimensional Modeling).
- Experience with Power BI Data Modeling and Semantic Layer development.
- Proficiency in DAX (Data Analysis Expressions).
- Experience designing and managing Semantic Models in Power BI.
- Strong SQL skills and experience working with large datasets.
- Knowledge of data warehousing concepts and best practices.
Preferred Skills
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Exposure to modern data platforms like Databricks.
- Understanding of data governance and data quality frameworks.
Job Summary
We are seeking a skilled Data Engineer to design, build, and maintain scalable data pipelines and infrastructure. The ideal candidate should have strong expertise in SQL, Python, Linux, and modern data engineering practices to support data integration, transformation, and analytics.
Key Responsibilities
- Design, develop, and maintain ETL/ELT data pipelines.
- Write efficient and optimized SQL queries for data extraction, transformation, and reporting.
- Develop automation scripts using Python for data processing and workflow optimization.
- Work with Linux environments for deployment, monitoring, and troubleshooting.
- Ensure data quality, integrity, and reliability across data platforms.
- Collaborate with data analysts, software engineers, and business stakeholders to deliver data solutions.
- Monitor, troubleshoot, and optimize data pipelines for performance and scalability.
- Implement best practices for data security, governance, and documentation.
Required Skills
- Strong experience in Data Engineering concepts and ETL/ELT processes.
- Proficiency in SQL, including query optimization and database design.
- Strong programming skills in Python.
- Hands-on experience with Linux commands, shell scripting, and system administration basics.
- Experience with relational databases such as PostgreSQL, MySQL, SQL Server, or Oracle.
- Familiarity with Git/version control.
- Strong analytical and problem-solving skills.
Preferred Skills
- Experience with cloud platforms (AWS, Azure, or GCP).
- Knowledge of Apache Spark, Airflow, Kafka, or similar data engineering tools.
- Experience with data warehousing solutions and big data technologies.
- Understanding of CI/CD pipelines and containerization (Docker/Kubernetes).
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Relevant certifications in cloud or data engineering are an added advantage.
About Us
We believe the future of software development is AI-native — where engineers operate at a higher level of abstraction and quality remains non-negotiable.
Incubyte is a software craft consultancy where the “how” of building software matters as much as the “what”.
We partner with companies of all sizes, from helping enterprises build, scale, and modernize to early-stage founders bring their ideas to life.
Our engineers operate in an AI-native development model, using AI as a collaborator across the SDLC to accelerate development while upholding the discipline of software craftsmanship. Guided by Software Craftsmanship and Extreme Programming practices, we build reliable, maintainable, and scalable systems with speed, without compromising quality. If this way of building software resonates with you, we’d like to talk.
Our Guiding Principles
These principles define how we work at Incubyte. They are non-negotiable.
Relentless Pursuit of Quality with Pragmatism
We build high-quality systems without losing sight of delivery.
Extreme Ownership
We take responsibility end-to-end for decisions, execution, and outcomes.
Proactive Collaboration
We collaborate closely, challenge each other, and solve problems together.
Active Pursuit of Mastery
We continuously improve our craft and raise our bar.
Invite, Give, and Act on Feedback
We seek, give, and act on feedback to get better every day.
Ensuring Client Success
We act as trusted partners and focus on real outcomes, not just output.
Job Description
This is a remote position.
Experience Level
2+ years of experience in SQL, Python, and Snowflake (or equivalent cloud data warehouse), Azure Cloud services.
Role Overview
If you're a Data Craftsperson who takes pride in clean, well-tested data solutions and believes in the principles of Extreme Programming, we'd love to meet you. At Incubyte, we're a DevOps organization where developers own the entire release cycle — you'll get hands-on experience across data engineering, analytics, cloud infrastructure, and direct client communication. This role sits primarily in data engineering (80%) with a meaningful analytics component (20%), supporting our client's data systems end-to-end.
What You'll Do
- Design, build, and maintain data pipelines and infrastructure using SQL and Python
- Work within Snowflake to build and optimize data models supporting business use cases
- Parse and process structured and semi-structured data (JSON, XML) from varied sources
- Diagnose issues across raw, intermediate, and summary tables
- Build SQL queries to support repeatable analytics use cases based on stakeholder requirements
- Investigate and resolve data quality issues, including time-sensitive or urgent ones
- Identify opportunities to consolidate models and maintain a single source of truth (SSOT)
Requirements
What We're Looking For
- 2+ years of experience with SQL and relational databases, with the ability to understand complex data relationships and transformations (required)
- 2+ years of experience with Python for data engineering tasks (required)
- Experience with Snowflake or an equivalent cloud data warehouse (required)
- Experience parsing JSON and XML data (a plus)
- A strong eye for data quality and attention to detail
- Knowledge of Git (required)
- Knowledge of Azure cloud services such as Azure Data Factory, Azure Blob Storage, and Azure SQL Database (required)
- Knowledge of data infrastructure/modeling tools like DBT, Fivetran (a plus)
- Experience with BI tools like Power BI(a plus, not core to this role)
- Knowledge of Docker, Linux, Shell/Bash, and virtualization technologies (a plus)
- Knowledge of SSIS packages (a plus)
- Familiarity with CI/CD methodologies
Benefits
Life at Incubyte
We are a remote-first company with structured flexibility. Teams commit to shared rhythms during core hours, ensuring smooth collaboration while maintaining autonomy. Twice a year, we come together in person for a co-working sprint and once a year for a retreat - with all travel expenses covered.
Our environment is built for crafters: pairing, refactoring, experimenting with AI, and pushing the boundaries of software excellence. We are all lifelong learners, and our work is our passion.
Perks
- Dedicated learning & development budget.
- Sponsorship for conference talks.
- Comprehensive medical & term insurance.
- Employee-friendly leave policies.
- Home Office fund
- Medical Insurance






