Cutshort logo
For Employers
Public Listed - Product Based company logo
Data Scientist
Public Listed - Product Based company

Data Scientist at Public Listed - Product Based company · Bengaluru (Bangalore) · 4 - 8 years · ₹25L - ₹70L / yr · Posted 6 Apr 2026

Recruiting Bond's logo

Data Scientist

at Public Listed - Product Based company

Agency job
4 - 8 yrs
₹25L - ₹70L / yr
Bengaluru (Bangalore)
Skills
skill iconData Science
data platforms
Data-flow analysis
Data pipelines
AI Infrastructure
Large Language Models (LLM)
Large Language Models (LLM) tuning
Agentic Systems
AI Agents
skill iconMachine Learning (ML)
skill iconDeep Learning
Natural Language Processing (NLP)
Statistical Modeling
Gen AI
Retrieval Augmented Generation (RAG)
Data Warehouse (DWH)
Data modeling
skill iconPython
PyTorch
AI/ML
Multi-modal AI
SQL
skill iconPostgreSQL
skill iconMongoDB
Query planning
Prompt engineering
ETL
ELT
AI copilots
Generative AI
Artificial Intelligence (AI)
Google Vertex AI
Vector database

🤖 Data Scientist – Frontier AI for Data Platforms & Distributed Systems (4–8 Years)

Experience: 4–8 Years

Location: Bengaluru (On-site / Hybrid)

Company: Publicly Listed, Global Product Platform


🧠 About the Mission

We are building a Top 1% AI-Native Engineering & Data Organization — from first principles.

This is not incremental improvement.

This is a full-stack transformation of a large-scale enterprise into an AI-native data platform company.

We are re-architecting:

  • Legacy systems → AI-native architectures
  • Static pipelines → autonomous, self-healing systems
  • Data platforms → intelligent, learning systems
  • Software workflows → agentic execution layers

This is the kind of shift you would expect from companies like Google or Microsoft —

Except here, you will build it from day zero and scale it globally.


🧠 The Opportunity: This role sits at the intersection of three high-impact domains:

1. Frontier AI Systems: Large Language Models (LLMs), Small Language Models (SLMs), and Agentic AI

2. Data Platforms: Warehouses, Lakehouses, Streaming Systems, Query Engines

3. Distributed Systems: High-throughput, low-latency, multi-region infrastructure


We are building systems where:

  • Data platforms optimize themselves using ML/LLMs
  • Pipelines are autonomous, self-healing, and adaptive
  • Queries are generated, optimized, and executed intelligently
  • Infrastructure learns from usage and evolves continuously

This is: AI as the control plane for data infrastructure


🧩 What You’ll Work On

You will design and build AI-native systems deeply embedded inside data infrastructure.

1. AI-Native Data Platforms

  • Build LLM-powered interfaces:
  • Natural language → SQL / pipelines / transformations
  • Design semantic data layers:
  • Embeddings, vector search, knowledge graphs
  • Develop AI copilots:
  • For data engineers, analysts, and platform users

2. Autonomous Data Pipelines

  • Build self-healing ETL/ELT systems using AI agents
  • Create pipelines that:
  • Detect anomalies in real time
  • Automatically debug failures
  • Dynamically optimize transformations

3. Intelligent Query & Compute Optimization

  • Apply ML/LLMs to:
  • Query planning and execution
  • Cost-based optimization using learned models
  • Workload prediction and scheduling
  • Build systems that:
  • Learn from query patterns
  • Continuously improve performance and cost efficiency

4. Distributed Data + AI Infrastructure

  • Architect systems operating at:
  • Billions of events per day
  • Petabyte-scale data
  • Work with:
  • Distributed compute engines (Spark / Flink / Ray class systems)
  • Streaming systems (Kafka-class infra)
  • Vector databases and hybrid retrieval systems

5. Learning Systems & Feedback Loops

  • Build closed-loop AI systems:
  • Execution → feedback → model updates
  • Develop:
  • Continual learning pipelines
  • Online learning systems for infra optimization
  • Experimentation frameworks (A/B, bandits, eval pipelines)

6. LLM & Agentic Systems (Infra-Aware)

  • Build agents that understand data systems
  • Enable:
  • Autonomous pipeline debugging
  • Root cause analysis for infra failures
  • Intelligent orchestration of data workflows


🧠 What We’re Looking For

Core Foundations

  • Strong grounding in:
  • Machine Learning, Deep Learning, NLP
  • Statistics, optimization, probabilistic systems
  • Distributed systems fundamentals
  • Deep understanding of:
  • Transformer architectures
  • Modern LLM ecosystems

Hands-On Expertise

  • Experience building:
  • LLM / GenAI systems (RAG, fine-tuning, embeddings)
  • Data platforms (warehouse, lake, lakehouse architectures)
  • Distributed pipelines and compute systems
  • Strong programming skills:
  • Python (ML/AI stack)
  • SQL (deep understanding — query planning, optimization mindset)


Systems Thinking (Critical)

You think in systems, not components.

  • Built or worked on:
  • Large-scale data pipelines
  • High-throughput distributed systems
  • Low-latency, high-concurrency architectures
  • Understand:
  • Query optimization and execution
  • Data partitioning, indexing, caching
  • Trade-offs in distributed systems


🔥 What Sets You Apart (Top 1%)

  • Built AI-powered data platforms or infra systems in production
  • Designed or contributed to:
  • Query engines / optimizers
  • Data observability / lineage systems
  • AI-driven infra or AIOps platforms
  • Experience with:
  • Multi-modal AI (logs, metrics, traces, text)
  • Agentic AI systems
  • Autonomous infrastructure
  • Worked on systems at scale comparable to:
  • Google (BigQuery-like systems)
  • Meta (real-time analytics infra)
  • Snowflake / Databricks (lakehouse architectures)


🧬 Ideal Background (Not Mandatory)

We often see strong candidates from:

  • Data infrastructure or platform engineering teams
  • AI-first startups or research-driven environments
  • High-scale product companies

Experience building:

  • Internal platforms used by 1000s of engineers
  • Systems serving millions of users / high throughput workloads
  • Multi-region, distributed cloud systems


🧠 The Kind of Problems You’ll Solve

  • Can LLMs replace traditional query optimizers?
  • How do we build self-healing data pipelines at scale?
  • Can data systems learn from every query and improve automatically?
  • How do we embed reasoning and planning into infrastructure layers?
  • What does a fully autonomous data platform look like?


Background: We Commonly See (But Not Limited To)

Our team often includes engineers from top-tier institutions and strong research or product backgrounds, including:

  • Leading engineering schools in India and globally
  • Engineers with experience in top product companies, AI startups, or research-driven environments
  • That said, we care far more about demonstrated ability, depth, and impact than pedigree alone.


Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

Similar jobs (10)

company logo
Remote only
5 - 10 yrs
Best in industry
skill iconPython
SQL
skill iconMachine Learning (ML)
databricks
Apache Airflow
+1 more

Description

We’re seeking a highly skilled, execution-focused Senior Data Scientist with a minimum of 5 years of experience. This role demands hands-on expertise in building, deploying, and optimizing machine learning models at scale, while working with big data technologies and modern cloud platforms. You will be responsible for driving data-driven solutions from experimentation to production, leveraging advanced tools and frameworks across Python, SQL, Spark, and AWS. The role requires strong technical depth, problem-solving ability, and ownership in delivering business impact through data science.


Responsibilities

  • Design, build, and deploy scalable machine learning models into production systems.
  • Develop advanced analytics and predictive models using Python, SQL, and popular ML/DL frameworks (Pandas, Scikit-learn, TensorFlow, PyTorch).
  • Leverage Databricks, Apache Spark, and Hadoop for large-scale data processing and model training.
  • Implement workflows and pipelines using Airflow and AWS EMR for automation and orchestration.
  • Collaborate with engineering teams to integrate models into cloud-based applications on AWS.
  • Optimize query performance, storage usage, and data pipelines for efficiency.
  • Conduct end-to-end experiments, including data preprocessing, feature engineering, model training, validation, and deployment.
  • Drive initiatives independently with high ownership and accountability.
  • Stay up to date with industry best practices in machine learning, big data, and cloud-native deployments.


Requirements

  • Minimum 5 years of experience in Data Science or Applied Machine Learning.
  • Strong proficiency in Python, SQL, and ML libraries (Pandas, Scikit-learn, TensorFlow, PyTorch).
  • Proven expertise in deploying ML models into production systems.
  • Experience with big data platforms (Hadoop, Spark) and distributed data processing.
  • Hands-on experience with Databricks, Airflow, and AWS EMR.
  • Strong knowledge of AWS cloud services (S3, Lambda, SageMaker, EC2, etc.).
  • Solid understanding of query optimization, storage systems, and data pipelines.
  • Excellent problem-solving skills, with the ability to design scalable solutions.
  • Strong communication and collaboration skills to work in cross-functional teams.


Benefits

  • Best-in-class salary: We hire strong talent and compensate accordingly.
  • Proximity Talks: Meet and learn from designers, engineers, product leaders, and AI practitioners.
  • Continuous learning: Work with a world-class team and stay close to the latest in AI, engineering, and product development.
  • High-impact work: Build AI-first systems and products used at scale by global clients.



About Us

Proximity is the trusted technology, design, and consulting partner for some of the biggest Sports, Media, and Entertainment companies in the world. We’re headquartered in San Francisco and have offices in Palo Alto, Dubai, Mumbai, and Bangalore.

Since 2019, Proximity has built high-impact, scalable products used by millions of users every day. Today, we are a global team of engineers, designers, product managers, and experts solving complex problems and building cutting-edge technology at scale.


Read more
Service Co
Service Co
Agency job
via by Rishika Teja
Pune
5 - 12 yrs
₹15L - ₹34L / yr
SQL
skill iconPython
skill iconData Science
Spark

Hiring for Data Scientist / Senior Data Scientist


Exp : 4 - 12 yrs

Edu : BE/B.tech/MCA

Work Location : Pune

Notice Period : Immediate - 15 days


Skills :


4+ years of experience in data engineering, data science, or related domains.


Hands-on experience with SQL, Python, and distributed data systems.


Knowledge of machine learning techniques and statistical analysis.


Experience with cloud data platforms (Azure Data Factory, AWS Glue, GCP BigQuery).


Familiarity with DevOps practices and CI/CD for data pipelines.


Platforms & Operations Experience (Preferred)

- Experience working with Azure, AWS, or Google Cloud data tools.


Operational experience with data orchestration tools (Airflow, ADF, Glue).


Understanding of Kubernetes, Docker, or containerized environments.


Hands-on experience with data warehousing platforms (Snowflake, Redshift, BigQuery).


Experience in monitoring, logging, and alerting operations for data workflows.

Read more
company logo
Mayank Choudhary
Posted by Mayank Choudhary
Pune
3 - 5 yrs
₹21L - ₹25L / yr
skill iconData Science
Artificial Intelligence (AI)

Strong Data Scientist / AI Engineer / Generative AI Engineer profile.

2

Mandatory (Experience 1) - Must have 3+ years of hands-on experience in Data Science, Artificial Intelligence, Machine Learning, Deep Learning, NLP, or Generative AI application development.

3

Mandatory (Experience 2) - Must have strong hands-on experience in Python programming, backend development, API development, and production-grade application support.

4

Mandatory (Experience 3) - Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, or Scikit-learn.

5

Mandatory (Experience 4) - Must have hands-on experience in NLP use cases such as text classification, sentiment analysis, entity recognition (NER), semantic search, embeddings, or document understanding.

6

Mandatory (Experience 5) - Must have experience working with Large Language Models (LLMs) such as GPT, LLaMA, Mistral, Phi, Claude, Gemini, or similar models.

7

Mandatory (Experience 6) - Must have hands-on experience building or implementing Retrieval Augmented Generation (RAG) solutions, vector search, semantic search, or knowledge-based AI applications.

8

Mandatory (Experience 7) - Must have experience with Prompt Engineering and Generative AI frameworks such as LangChain, LangGraph, AI Agents, Azure OpenAI, or similar technologies.

9

Mandatory (Experience 8) - Must have experience developing, consuming, or integrating APIs using Python frameworks such as FastAPI, Flask, or similar technologies.

10

Mandatory (CTC) - The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.

11

Preferred (Experience 1) - Experience with LLMOps/MLOps tools for monitoring, evaluation, experimentation, and versioning of AI models.

12

Preferred (Experience 2) - Exposure to Azure OpenAI, Azure Kubernetes Service (AKS), Kubernetes, cloud-native AI deployments, or distributed systems.

13

Preferred (Experience 3) - Experience working with PostgreSQL, MongoDB, Redis, Kafka, or large-scale data platforms.

14

Preferred (Experience 4) - Familiarity with Docker, Kubernetes, cloud platforms, and scalable deployment architecture.

15

Preferred (Company) - Candidates from AI-first startups, product companies, SaaS organizations, fintech, or data-driven technology companies.

16

Mandatory ( Age ) - Candidate Should be Below 28 Years.

17

Mandatory ( Pedigree) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered.

Read more
Pune
3 - 5 yrs
₹21L - ₹25L / yr
skill iconData Science
Artificial Intelligence (AI)

Strong Data Scientist / AI Engineer / Generative AI Engineer profile.

2

Mandatory (Experience 1) - Must have 3+ years of hands-on experience in Data Science, Artificial Intelligence, Machine Learning, Deep Learning, NLP, or Generative AI application development.

3

Mandatory (Experience 2) - Must have strong hands-on experience in Python programming, backend development, API development, and production-grade application support.

4

Mandatory (Experience 3) - Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, or Scikit-learn.

5

Mandatory (Experience 4) - Must have hands-on experience in NLP use cases such as text classification, sentiment analysis, entity recognition (NER), semantic search, embeddings, or document understanding.

6

Mandatory (Experience 5) - Must have experience working with Large Language Models (LLMs) such as GPT, LLaMA, Mistral, Phi, Claude, Gemini, or similar models.

7

Mandatory (Experience 6) - Must have hands-on experience building or implementing Retrieval Augmented Generation (RAG) solutions, vector search, semantic search, or knowledge-based AI applications.

8

Mandatory (Experience 7) - Must have experience with Prompt Engineering and Generative AI frameworks such as LangChain, LangGraph, AI Agents, Azure OpenAI, or similar technologies.

9

Mandatory (Experience 8) - Must have experience developing, consuming, or integrating APIs using Python frameworks such as FastAPI, Flask, or similar technologies.

10

Mandatory (CTC) - The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.

11

Preferred (Experience 1) - Experience with LLMOps/MLOps tools for monitoring, evaluation, experimentation, and versioning of AI models.

12

Preferred (Experience 2) - Exposure to Azure OpenAI, Azure Kubernetes Service (AKS), Kubernetes, cloud-native AI deployments, or distributed systems.

13

Preferred (Experience 3) - Experience working with PostgreSQL, MongoDB, Redis, Kafka, or large-scale data platforms.

14

Preferred (Experience 4) - Familiarity with Docker, Kubernetes, cloud platforms, and scalable deployment architecture.

15

Preferred (Company) - Candidates from AI-first startups, product companies, SaaS organizations, fintech, or data-driven technology companies.

16

Mandatory ( Age ) - Candidate Should be Below 28 Years.

17

Mandatory ( Pedigree) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered.

Read more
company logo
Pune
3 - 5 yrs
₹15L - ₹20L / yr
Data Scientist
Retrieval Augmented Generation (RAG)
Artificial Intelligence (AI)
skill iconMachine Learning (ML)
Natural Language Processing (NLP)

Roles & Responsibilities

  • Design and develop intelligent AI-based applications using advanced NLP and LLM techniques to solve real-world business challenges in financial services.
  • Build and optimize Retrieval-Augmented Generation (RAG) pipelines leveraging structured and unstructured financial data.
  • Integrate and orchestrate LLMs/SLMs for question-answering, summarization, semantic search, and document understanding.
  • Develop and maintain RESTful APIs (sync and async) to serve NLP models and chatbot interfaces using frameworks like FastAPI, Flask, etc.
  • Should have knowledge of advanced prompting techniques.
  • Implement semantic search, hybrid search, and text retrieval systems using Elasticsearch and vector databases (e.g., FAISS, Pinecone, Weaviate).
  • Perform NLP tasks such as entity recognition, text classification, intent detection, embedding generation, and sentiment analysis where required.
  • Monitor and fine-tune LLM/SLM performance with real-world user data to improve relevance, latency, and accuracy.
  • Exposure to LLMOps tools for monitoring, evaluation, and versioning of AI models in production.
  • Build, train, and evaluate deep learning models for NLP tasks including classification, NER, summarization, and embedding generation.
  • Develop traditional machine learning models (e.g., regression, decision trees, clustering) for structured data analysis and prediction tasks.
  • Interact with cross-functional teams to understand system issues and follow up with respective teams to get them fixed.
  • Understand and identify areas of improvement across businesses and participate in solution identification and implementation.
  • Should be able to work as an Individual Contributor on new and existing projects.
  • Positive and problem-solving attitude, must work as an independent contributor.

Ideal Candidate

1.Strong Data Scientist / AI Engineer / Generative AI Engineer profile.

2.Mandatory (Experience 1) - Must have 3+ years of hands-on experience in Data Science, Artificial Intelligence, Machine Learning, Deep Learning, NLP, or Generative AI application development.

3.Mandatory (Experience 2) - Must have strong hands-on experience in Python programming, backend development, API development, and production-grade application support

4.Mandatory (Experience 3) - Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, or Scikit-learn.

5.Mandatory (Experience 4) - Must have hands-on experience in NLP use cases such as text classification, sentiment analysis, entity recognition (NER), semantic search, embeddings, or document understanding.

6.Mandatory (Experience 5) - Must have experience working with Large Language Models (LLMs) such as GPT, LLaMA, Mistral, Phi, Claude, Gemini, or similar models

.7.Mandatory (Experience 6) - Must have hands-on experience building or implementing Retrieval Augmented Generation (RAG) solutions, vector search, semantic search, or knowledge-based AI applications.

8.Mandatory (Experience 7) - Must have experience with Prompt Engineering and Generative AI frameworks such as LangChain, LangGraph, AI Agents, Azure OpenAI, or similar technologies.

9.Mandatory (Experience 8) - Must have experience developing, consuming, or integrating APIs using Python frameworks such as FastAPI, Flask, or similar technologies.

10.Mandatory (CTC) - The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.

11.Preferred (Experience 1) - Experience with LLMOps/MLOps tools for monitoring, evaluation, experimentation, and versioning of AI models.

12.Preferred (Experience 2) - Exposure to Azure OpenAI, Azure Kubernetes Service (AKS), Kubernetes, cloud-native AI deployments, or distributed systems.

13.Preferred (Experience 3) - Experience working with PostgreSQL, MongoDB, Redis, Kafka, or large-scale data platforms.

14.Preferred (Experience 4) - Familiarity with Docker, Kubernetes, cloud platforms, and scalable deployment architecture.

15.Preferred (Company) - Candidates from AI-first startups, product companies, SaaS organizations, fintech, or data-driven technology companies.

16.Mandatory ( Age ) - Candidate Should be Below 28 Years.

17.Mandatory ( Pedigree) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered.

Read more
Hiring for IT Product based (MNC)
Hiring for IT Product based (MNC)
Agency job
via by Sneha k
Pune
2 - 4 yrs
₹15L - ₹20L / yr
skill iconData Science
Large Language Models (LLM) tuning
skill iconPython
Large Language Models (LLM)
Generative AI
+2 more

Role Overview 

As a Data Scientist, you will work with business stakeholders, AI engineers, and domain experts to transform data into actionable insights and intelligent solutions. You will develop machine learning models, perform statistical analysis, and contribute to AI-driven products that create measurable business impact. 



Key Responsibilities 

Data Science & Machine Learning 

  • Analyze structured and unstructured data to identify patterns, trends, and business opportunities.  
  • Perform exploratory data analysis (EDA), feature engineering, and data preparation.  
  • Develop, evaluate, and optimize machine learning models for prediction, classification, clustering, and forecasting.  
  • Apply statistical techniques to solve business problems and validate model performance.  
  • Design and execute experiments to improve model accuracy and business outcomes.  

 AI Solution Development 

  • Collaborate with AI Engineers, Data Engineers, and domain experts to build AI-powered solutions.  
  • Translate business requirements into scalable data science approaches.  
  • Contribute to Generative AI and advanced analytics initiatives where applicable.  
  • Document methodologies, model performance, and key findings.  


 Required Technical Skills 

  • Strong programming skills in Python and SQL for data analysis, feature engineering, and machine learning.  
  • Strong understanding of Statistics, Probability, Linear Algebra, and Calculus as applied to machine learning and data science.  
  • Experience with Exploratory Data Analysis (EDA), data preprocessing, feature engineering, feature selection, and handling missing or imbalanced data.  
  • Good understanding of Supervised, Unsupervised, and Ensemble Machine Learning algorithms, including their assumptions, strengths, limitations, and appropriate use cases.  
  • Strong knowledge of Regression, Classification, Clustering, Time Series Forecasting, Dimensionality Reduction, Recommendation Systems, and Anomaly Detection techniques.  
  • Experience with Model Evaluation, Cross-Validation, Hyperparameter Optimization, Bias-Variance Trade-off, Feature Importance, Explainable AI (XAI), and Performance Metrics.  
  • Understanding of Statistical Inference, Hypothesis Testing, Probability Distributions, Sampling Techniques, Confidence Intervals, and A/B Testing.  
  • Experience translating business problems into analytical approaches and developing scalable, data-driven solutions.  
  • Working knowledge of Generative AI, Large Language Models (LLMs), Prompt Engineering, and Retrieval-Augmented Generation (RAG) is preferred. 

Preferred Qualifications 

  • Bachelor's or master's degree in computer science, Artificial Intelligence, Data Science, Statistics, Mathematics, Engineering, or a related field.  
  • 2–4 years of experience developing machine learning or data science solutions.  
  • Experience working on end-to-end data science projects in a business environment.  


 Nice to Have 

  • Exposure to Generative AI, LLMs, RAG, or Agentic AI.  
  • Experience with Computer Vision or Natural Language Processing (NLP).  
  • Familiarity with cloud-based AI platforms.  
  • Knowledge of construction, engineering, manufacturing, or industrial domains.  
  • Participation in hackathons, research, Kaggle competitions, or open-source projects.  


 Soft Skills 

Strong analytical and problem-solving skills, effective communication and collaboration, ownership mindset, adaptability, continuous learning, and a passion for innovation.

Read more
company logo
Bengaluru (Bangalore), Delhi, Gurugram, Noida, Ghaziabad, Faridabad, Chennai
4 - 15 yrs
₹30L - ₹40L / yr
skill iconMachine Learning (ML)
Natural Language Processing (NLP)
Generative AI
skill iconPython
Scikit-Learn
+4 more

About the Role

 

We are looking for a highly skilled Data Scientist with strong expertise in Machine Learning, MLOps, and Generative AI. The ideal candidate will have hands-on experience in building scalable ML models, deploying them in production, and working with modern AI frameworks, including GenAI technologies.

 

 

 

Key Responsibilities

 

·      Design, develop, and deploy machine learning models for real-world business problems

·      Work on end-to-end ML lifecycle: data preprocessing, model building, evaluation, deployment, and monitoring

·      Implement and manage MLOps pipelines for scalable and reproducible workflows

·      Utilize tools like MLflow for experiment tracking, model versioning, and lifecycle management

·      Develop and integrate Generative AI (GenAI) solutions such as LLM-based applications

·      Collaborate with cross-functional teams (engineering, product, business) to translate requirements into AI solutions

·      Optimize model performance and ensure production stability

·      Stay updated with the latest advancements in AI/ML and GenAI ecosystems

 

 

 

Required Skills & Qualifications

 

·      4+ years of experience in Data Science / Machine Learning

·      Strong programming skills in Python

·      Hands-on experience with ML modeling techniques (supervised, unsupervised, NLP, etc.)

·      Solid understanding of MLOps practices and tools

·      Experience with MLflow or similar model lifecycle tools 

·      Practical experience in Generative AI (GenAI), including working with LLMs

·      Experience with libraries/frameworks like Scikit-learn, TensorFlow, PyTorch

·      Strong understanding of data structures, algorithms, and statistics

·      Experience with cloud platforms (AWS/GCP/Azure) is a plus


Good to Have

 

·      Experience with LLM fine-tuning, prompt engineering, or RAG pipelines

·      Exposure to Docker, Kubernetes, and CI/CD pipelines

·      Knowledge of data engineering workflows 



Read more
company logo
Remote only
3 - 8 yrs
₹20L - ₹35L / yr
MLFlow
MLOps
Fine-tuning LLMs

We are looking for a talented and driven Data Scientist to join our growing Analytics team in India. In this role, you will work at the intersection of advanced machine learning, scalable MLOps infrastructure, and domain-specific healthcare analytics. You will collaborate closely with cross-functional teams to build, deploy, and maintain production-grade ML models that drive real-world impact in clinical trials and healthcare operations.


KEY RESPONSIBILITIES

End-to-End ML Development

•     Design, build, and optimize predictive models across the full ML lifecycle—from data ingestion to model serving.

•     Conduct rigorous Exploratory Data Analysis (EDA) to surface insights and drive feature engineering decisions.

•     Validate model performance using appropriate statistical techniques and domain knowledge.


MLOps & Production Deployment

•     Deploy, monitor, and maintain production-grade ML models using Databricks MLFlow endpoints and Unity Catalog.

•     Implement CI/CD pipelines for model versioning, experiment tracking, and automated retraining.

•     Ensure model reliability, observability, and performance in live production environments.


Language Models & LLM Applications

•     Apply transformer-based models (BERT, ClinicalBERT, Trial2Vec) for NLP tasks including classification, NER, and information extraction.

•     Build and maintain vector similarity search pipelines for semantic retrieval and recommendation use cases.

•     Fine-tune pre-trained models for domain-specific applications in clinical and healthcare contexts.

•     Support exploratory work around LLM integration and prompt engineering for internal tooling.


Domain-Driven Analytics

•     Apply advanced analytics within complex healthcare and clinical trial datasets—including patient records, trial protocols, and adverse event data.

•     Translate ambiguous business problems into structured analytical frameworks with measurable outcomes.

•     Partner with domain experts, product managers, and engineering teams to deliver data-driven solutions.


REQUIRED QUALIFICATIONS

Education

•     Bachelor’s or Master’s degree in Computer Science, Statistics, Mathematics, Bioinformatics, or a closely related field.


Experience

•     2–4 years of hands-on experience in a data science or machine learning role.

•     Demonstrable experience deploying ML models in production environments (not just prototyping).


Technical Skills

•     Strong proficiency in Python (pandas, NumPy, scikit-learn, PyTorch / TensorFlow).

•     Experience with Databricks, MLFlow (experiment tracking, model registry, endpoints), and Unity Catalog.

•     Hands-on experience with BERT-family models and Hugging Face Transformers library.

•     Familiarity with vector databases (e.g., FAISS, Pinecone, Weaviate) and embedding-based retrieval.

•     Solid understanding of SQL and working with large structured/unstructured datasets.

•     Exposure to cloud platforms (AWS / GCP / Azure) and distributed computing frameworks (Spark).


GOOD TO HAVE

•     Prior experience with clinical trial data standards (CDISC, CDASH, SDTM) or healthcare ontologies (SNOMED, ICD-10).

•     Familiarity with Trial2Vec or similar trial-to-vector embedding approaches.

•     Experience with LLM fine-tuning, RAG pipelines, or prompt engineering in a production setting.

•     Knowledge of regulatory and compliance considerations in healthcare AI (e.g., FDA guidelines, HIPAA).

•     Contributions to open-source ML projects or published research.


THIS ROLE IS NOT FOR YOU IF…

•     You have strong SQL/BI skills but limited hands-on ML modelling experience — or you’ve built models only in notebooks without ever deploying them to production.

•     Your LLM exposure is limited to API calls and prompt engineering — with no experience fine-tuning models, working with embeddings, or building vector search pipelines.

Read more
Remote, Delhi, Gurugram, Noida, Ghaziabad, Faridabad, Pune
2 - 4 yrs
₹30L - ₹40L / yr
databricks
MLFlow
skill iconPython
BERT
Large Language Models (LLM) tuning
+1 more

We are looking for a talented and driven Data Scientist to join our growing Analytics team in India. In this role, you will work at the intersection of advanced machine learning, scalable MLOps infrastructure, and domain-specific healthcare analytics. You will collaborate closely with cross-functional teams to build, deploy, and maintain production-grade ML models that drive real-world impact in clinical trials and healthcare operations.


KEY RESPONSIBILITIES

End-to-End ML Development

•     Design, build, and optimize predictive models across the full ML lifecycle—from data ingestion to model serving.

•     Conduct rigorous Exploratory Data Analysis (EDA) to surface insights and drive feature engineering decisions.

•     Validate model performance using appropriate statistical techniques and domain knowledge.


MLOps & Production Deployment

•     Deploy, monitor, and maintain production-grade ML models using Databricks MLFlow endpoints and Unity Catalog.

•     Implement CI/CD pipelines for model versioning, experiment tracking, and automated retraining.

•     Ensure model reliability, observability, and performance in live production environments.


Language Models & LLM Applications

•     Apply transformer-based models (BERT, ClinicalBERT, Trial2Vec) for NLP tasks including classification, NER, and information extraction.

•     Build and maintain vector similarity search pipelines for semantic retrieval and recommendation use cases.

•     Fine-tune pre-trained models for domain-specific applications in clinical and healthcare contexts.

•     Support exploratory work around LLM integration and prompt engineering for internal tooling.


Domain-Driven Analytics

•     Apply advanced analytics within complex healthcare and clinical trial datasets—including patient records, trial protocols, and adverse event data.

•     Translate ambiguous business problems into structured analytical frameworks with measurable outcomes.

•     Partner with domain experts, product managers, and engineering teams to deliver data-driven solutions.


REQUIRED QUALIFICATIONS

Education

•     Bachelor’s or Master’s degree in Computer Science, Statistics, Mathematics, Bioinformatics, or a closely related field.


Experience

•     2–4 years of hands-on experience in a data science or machine learning role.

•     Demonstrable experience deploying ML models in production environments (not just prototyping).


Technical Skills

•     Strong proficiency in Python (pandas, NumPy, scikit-learn, PyTorch / TensorFlow).

•     Experience with Databricks, MLFlow (experiment tracking, model registry, endpoints), and Unity Catalog.

•     Hands-on experience with BERT-family models and Hugging Face Transformers library.

•     Familiarity with vector databases (e.g., FAISS, Pinecone, Weaviate) and embedding-based retrieval.

•     Solid understanding of SQL and working with large structured/unstructured datasets.

•     Exposure to cloud platforms (AWS / GCP / Azure) and distributed computing frameworks (Spark).


GOOD TO HAVE

•     Prior experience with clinical trial data standards (CDISC, CDASH, SDTM) or healthcare ontologies (SNOMED, ICD-10).

•     Familiarity with Trial2Vec or similar trial-to-vector embedding approaches.

•     Experience with LLM fine-tuning, RAG pipelines, or prompt engineering in a production setting.

•     Knowledge of regulatory and compliance considerations in healthcare AI (e.g., FDA guidelines, HIPAA).

•     Contributions to open-source ML projects or published research.


THIS ROLE IS NOT FOR YOU IF…

•     You have strong SQL/BI skills but limited hands-on ML modelling experience — or you’ve built models only in notebooks without ever deploying them to production.

•     Your LLM exposure is limited to API calls and prompt engineering — with no experience fine-tuning models, working with embeddings, or building vector search pipelines.

Read more
company logo
Anupam Arya
Posted by Anupam Arya
Bengaluru (Bangalore), Mumbai, Delhi, Gurugram, Hyderabad
5 - 12 yrs
₹30L - ₹40L / yr
skill iconMachine Learning (ML)
skill iconDeep Learning
Generative AI (GenAI)

AuxoAI is hiring a Senior Data Scientist with strong expertise in AI, machine learning engineering (MLE), and generative AI. You will play a leading role in designing, deploying, and scaling production-grade ML systems — including large language model (LLM)-based pipelines, AI copilots, and agentic workflows. This role is ideal for someone who thrives on balancing cutting-edge research with production rigor and loves mentoring while building impact-first AI applications. 


Location - Mumbai/Bangalore/Hyderabad/Gurgaon (Hybrid - 3 Days a week in Office)​​


Responsibilities: 

  • Own the full ML lifecycle: model design, training, evaluation, deployment 
  • Design production-ready ML pipelines with CI/CD, testing, monitoring, and drift detection 
  • Fine-tune LLMs and implement retrieval-augmented generation (RAG) pipelines 
  • Build agentic workflows for reasoning, planning, and decision-making 
  • Develop both real-time and batch inference systems using Docker, Kubernetes, and Spark 
  • Leverage state-of-the-art architectures: transformers, diffusion models, RLHF, and multimodal pipelines 
  • Collaborate with product and engineering teams to integrate AI models into business applications 
  • Mentor junior team members and promote MLOps, scalable architecture, and responsible AI best practices 



Requirements

  • 5+ years of experience in designing, deploying, and scaling ML/DL systems in production 
  • Proficient in Python and deep learning frameworks such as PyTorch, TensorFlow, or JAX 
  • Experience with LLM fine-tuning, LoRA/QLoRA, vector search (Weaviate/PGVector), and RAG pipelines 
  • Familiarity with agent-based development (e.g., ReAct agents, function-calling, orchestration) 
  • Solid understanding of MLOps: Docker, Kubernetes, Spark, model registries, and deployment workflows 
  • Strong software engineering background with experience in testing, version control, and APIs 
  • Proven ability to balance innovation with scalable deployment 
  • B.S./M.S./Ph.D. in Computer Science, Data Science, or a related field 
  • Bonus: Open-source contributions, GenAI research, or applied systems at scale 


Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos