Machine learning/AI expert at CSIR Institute of Microbial Technology · Chandigarh · 3 - 3 years · ₹9.6L - ₹9.7L / yr · Raised funding · Posted 22 Jan 2025

Project titled: “Machine Learning Models to Predict MIC in Indian Priority Pathogens & Identification of Novel Antimicrobial Resistance Mechanisms”
Project Code: GAP-0242
Position: Project Research Scientist-II (Non-Medical) =01 Essential
Qualification: 1. Post Graduate Degree, including the integrated PG degrees, with three years post qualification Experience or Ph.D. 2. For Engineering/ IT/ CS - Graduate degree of Four years with three years’ post qualification experience.
Desirable: Experience in Machine learning and AI methods, previous experience in next generation sequence data analysis is expected. Hands-on experience in wet-lab (basic microbiology/molecular biology) is beneficial not mandatory.
Upper Age limit (years) - 40
Monthly Emoluments: Rs. 67,000/- +HRA, as admissible.

About CSIR Institute of Microbial Technology
About
Similar jobs (10)
🚨 Hiring – Data Scientist | Python + Agentic AI
💼 Experience: 5+ Years
Must Have:
• Strong Data Science experience
• Python
• Agentic AI / AI Agents
• Generative AI / LLMs
• RAG / Vector Databases
• LangChain / LangGraph or similar Agent Frameworks
• Machine Learning & NLP
Role Overview
As a Data Scientist, you will work with business stakeholders, AI engineers, and domain experts to transform data into actionable insights and intelligent solutions. You will develop machine learning models, perform statistical analysis, and contribute to AI-driven products that create measurable business impact.
Key Responsibilities
Data Science & Machine Learning
- Analyze structured and unstructured data to identify patterns, trends, and business opportunities.
- Perform exploratory data analysis (EDA), feature engineering, and data preparation.
- Develop, evaluate, and optimize machine learning models for prediction, classification, clustering, and forecasting.
- Apply statistical techniques to solve business problems and validate model performance.
- Design and execute experiments to improve model accuracy and business outcomes.
AI Solution Development
- Collaborate with AI Engineers, Data Engineers, and domain experts to build AI-powered solutions.
- Translate business requirements into scalable data science approaches.
- Contribute to Generative AI and advanced analytics initiatives where applicable.
- Document methodologies, model performance, and key findings.
Required Technical Skills
- Strong programming skills in Python and SQL for data analysis, feature engineering, and machine learning.
- Strong understanding of Statistics, Probability, Linear Algebra, and Calculus as applied to machine learning and data science.
- Experience with Exploratory Data Analysis (EDA), data preprocessing, feature engineering, feature selection, and handling missing or imbalanced data.
- Good understanding of Supervised, Unsupervised, and Ensemble Machine Learning algorithms, including their assumptions, strengths, limitations, and appropriate use cases.
- Strong knowledge of Regression, Classification, Clustering, Time Series Forecasting, Dimensionality Reduction, Recommendation Systems, and Anomaly Detection techniques.
- Experience with Model Evaluation, Cross-Validation, Hyperparameter Optimization, Bias-Variance Trade-off, Feature Importance, Explainable AI (XAI), and Performance Metrics.
- Understanding of Statistical Inference, Hypothesis Testing, Probability Distributions, Sampling Techniques, Confidence Intervals, and A/B Testing.
- Experience translating business problems into analytical approaches and developing scalable, data-driven solutions.
- Working knowledge of Generative AI, Large Language Models (LLMs), Prompt Engineering, and Retrieval-Augmented Generation (RAG) is preferred.
Preferred Qualifications
- Bachelor's or master's degree in computer science, Artificial Intelligence, Data Science, Statistics, Mathematics, Engineering, or a related field.
- 2–4 years of experience developing machine learning or data science solutions.
- Experience working on end-to-end data science projects in a business environment.
Nice to Have
- Exposure to Generative AI, LLMs, RAG, or Agentic AI.
- Experience with Computer Vision or Natural Language Processing (NLP).
- Familiarity with cloud-based AI platforms.
- Knowledge of construction, engineering, manufacturing, or industrial domains.
- Participation in hackathons, research, Kaggle competitions, or open-source projects.
Soft Skills
Strong analytical and problem-solving skills, effective communication and collaboration, ownership mindset, adaptability, continuous learning, and a passion for innovation.
Experience - 4 to 6 year
Location – Ahmedabad/Pune/Indore
- Additional Job Description
Additional Job Description
Required Skills and Experience:
- Strong proficiency in Python and experience with ML/AI libraries (scikit-learn, TensorFlow, PyTorch, Hugging Face ecosystem).
- Hands-on experience with LLMs, RAG, vector databases, and retrieval pipelines.
- Practical experience deploying agentic workflows and building multi-step, tool-enabled agents.
- Experience using Garak (or similar LLM red-teaming/vulnerability scanners) to identify model weaknesses and harden deployments.
- Demonstrated experience implementing content filtering / moderation systems.
- Solid skills working with structured and unstructured data and advanced feature engineering.
- Familiarity with cloud GenAI platforms and services (Azure AI Services preferred; AWS/GCP acceptable).
- Experience building APIs/microservices; containerization (Docker), orchestration (Kubernetes).
- Strong understanding of model evaluation, performance profiling, inference cost optimization, and observability.
- Good knowledge of security, data governance, and privacy best practices for AI systems.

🚀 We’re Hiring | Data Scientist 🧠📊
Ready to turn data into real-world intelligence? Join us and work on exciting AI/ML & data-driven solutions!
🔹 Experience: 8+ Years
🔹 Must-Have Skills:
🐍 Python | 🤖 Machine Learning | ☁️ Cloud | 🧠 NLP | 📊 Data Visualization
📍 Location: Pune
💼 Work Mode: Work from Office
If you're passionate about Data Science, AI & solving complex business problems, we’d love to hear from you!
📩 Interested? Kindly text
#Hiring #DataScientist #DataScience #MachineLearning #Python #NLP #AI #Cloud #DataVisualization #TechJobs #HiringNow

Are you looking to work in the cutting edge area of applying data-science to help global customers get a better insight into their health? If so, read on and apply.
Role Name: Senior Data Scientist
Science Team | Full-Time | In-Office | Bangalore
The Role
The Ultrahuman Science Team builds the algorithms behind the Ring, M1 CGM, blood and urine biomarkers, and Performance Lab assessments. We are hiring a Senior Data Scientist to own those algorithms end to end: from the raw sensor signal to a model that is shipped, monitored, and trusted in users' hands.
This is a build role with real scope. In a typical month you will improve a production algorithm, root-cause a metric users are complaining about, and stand up the data pipeline the next model needs. The common thread is ownership: you take a vague question and return a working answer, without waiting to be handed scope.
What You'll Do
· Own algorithms end to end: sleep staging, activity detection, sensor-derived metrics, and health scores. You frame the problem, build the features, train and evaluate the model, and see it live
· Ship models, not notebooks: you prove a change on our own cohort before it reaches users, and a model is done only when it runs in production and you can tell how it is behaving
· Validate against reference standards: design evaluations against gold standards, reference devices, and study ground truth, and know when a result is real and when it is an artifact
· Own the data layer: cohort extraction, feature pipelines, study data, and raw sensor data, so the next model starts from clean inputs
What This Looks Like in Practice
1. Improving production algorithms - Take an existing production model like sleep staging, root-cause the failure modes against reference data, and ship a fix you can defend with numbers.
2. Building new models - Train an activity classifier on raw sensor data, design the labeled data collection that expands it, and pick the operating point so false positives never erode trust.
3. Proving it before it ships - Run a new steps algorithm against reference-device cohorts, decide with data when it is ready, and monitor how it behaves after rollout.
Who You Are
The two things we can't coach
· High ownership, end to end: you take a problem from a vague question to a shipped model without waiting to be handed scope, and you can point to something you owned from raw data all the way to production
· Hungry for more scope: you have outgrown your current role and want problems biggerthan your title, with the technical depth to be trusted with them
Also important
· You've worked with human health data: wearables, physiological signals, or clinical data.
If your experience is close but not exact, show us why you will ramp fast
· You've built at a startup: or somewhere small enough that nobody handed you clean data, clear specs, or a mature ML platform
· You work like it's 2026: coding agents and AI tooling are part of how you build every day, and you can tell which new capabilities are worth adopting
· You communicate: you can explain a model and its limits to a product manager, an engineer, or a founder, and hold your own with our scientists Core Technical Skills
· Languages and data: Python and SQL daily, comfortable working in a real codebase
· Machine learning: PyTorch or TensorFlow, scikit-learn, and gradient boosting, with the judgment to know which the problem needs
· Advanced machine learning: time series and sequence models, deep learning on continuous physiological signals, and ensembles
· Statistics and evaluation: hypothesis testing, experiment and A/B design, model evaluation, and error analysis against a reference standard
· Scale and cloud: Spark or equivalent on large datasets, and AWS, GCP, or Azure
· Production ML and MLOps: training pipelines, model versioning, deployment, monitoring, and drift detection
· LLMs and agentic systems: fine-tuning and serving models, building agentic pipelines, and using coding agents to move faster
Experience:
- 4 to 5 years building and shipping machine learning systems. We index on what you have shipped and on trajectory, not the exact number of years; if you are a little earlier but have clearly outgrown your current scope, we want to hear from you.
- Bachelor's or higher in engineering, computer science, statistics, or a related field.
How We Work and Who Thrives Here
- The Science team is small and moves fast, and much of the work has no precedent to copy.
- People do their best work here when they are energized by ambiguity, low on ego, quick to adopt a better idea no matter where it comes from, and comfortable owning something before anyone has told them how. If you need a mature data org, clean labelled datasets, and clear guardrails to thrive, this particular role will not be the right fit, and that is worth knowing up front.
What You'll Gain
· Ownership of algorithms that hundreds of thousands of people see every morning
· A dataset most scientists never get to touch: 100M+ nights of sleep and continuous physiological signals at scale
· Direct collaboration with the engineering, product, and design teams building Ultrahuman
Strong Data Scientist / AI Engineer / Generative AI Engineer profile.
2
Mandatory (Experience 1) - Must have 3+ years of hands-on experience in Data Science, Artificial Intelligence, Machine Learning, Deep Learning, NLP, or Generative AI application development.
3
Mandatory (Experience 2) - Must have strong hands-on experience in Python programming, backend development, API development, and production-grade application support.
4
Mandatory (Experience 3) - Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, or Scikit-learn.
5
Mandatory (Experience 4) - Must have hands-on experience in NLP use cases such as text classification, sentiment analysis, entity recognition (NER), semantic search, embeddings, or document understanding.
6
Mandatory (Experience 5) - Must have experience working with Large Language Models (LLMs) such as GPT, LLaMA, Mistral, Phi, Claude, Gemini, or similar models.
7
Mandatory (Experience 6) - Must have hands-on experience building or implementing Retrieval Augmented Generation (RAG) solutions, vector search, semantic search, or knowledge-based AI applications.
8
Mandatory (Experience 7) - Must have experience with Prompt Engineering and Generative AI frameworks such as LangChain, LangGraph, AI Agents, Azure OpenAI, or similar technologies.
9
Mandatory (Experience 8) - Must have experience developing, consuming, or integrating APIs using Python frameworks such as FastAPI, Flask, or similar technologies.
10
Mandatory (CTC) - The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
11
Preferred (Experience 1) - Experience with LLMOps/MLOps tools for monitoring, evaluation, experimentation, and versioning of AI models.
12
Preferred (Experience 2) - Exposure to Azure OpenAI, Azure Kubernetes Service (AKS), Kubernetes, cloud-native AI deployments, or distributed systems.
Key Responsibilities:
- Develop and deploy machine learning, deep learning, and NLP models for various business use cases.
- Build end-to-end ML pipelines including data preprocessing, feature engineering, training, evaluation, and production deployment.
- Optimize model performance and ensure scalability in production environments.
- Work closely with data scientists, product teams, and engineers to translate business requirements into AI solutions.
- Conduct data analysis to identify trends and insights.
- Implement MLOps practices for versioning, monitoring, and automating ML workflows.
- Research and evaluate new AI/ML techniques, tools, and frameworks.
- Document system architecture, model design, and development processes.
Required Skills:
- Strong programming skills in Python (NumPy, Pandas, Scikit-learn, TensorFlow, PyTorch, Keras).
- Hands-on experience in building and deploying, finetuning ML/DL models in production.
- Good understanding of machine learning algorithms, neural networks, NLP, and computer vision.
- Experience with REST APIs, Docker, Kubernetes, and cloud platforms (AWS/GCP/Azure).
- Working knowledge of MLOps tools such as MLflow, Airflow, DVC, or Kubeflow.
- Familiarity with data pipelines and big data technologies (Spark, Hadoop) is a plus.
- Strong analytical skills and ability to work with large datasets.
- Excellent communication and problem-solving abilities.
- Experience in deploying models using cloud services (AWS Sagemaker, GCP Vertex AI, etc.).
- Experience in LLM fine-tuning or Generative AI, Voice AI, is an added advantage.
Educational Qualification:
- Bachelor’s or Master’s degree in Computer Science, Data Science, AI, Machine Learning, IT, from IIT/NIT colleges strongly preferred
We are looking for a Data Science & Machine Learning Senior Associate with 3–5 years of relevant experience in data science, machine learning, and analytics. The candidate will be responsible for developing predictive models, analyzing complex datasets, building scalable ML solutions, and supporting production-grade data science applications on cloud platforms.
Key Responsibilities
- Develop and implement machine learning and predictive analytics models.
- Perform data analysis, statistical modeling, and optimization to solve business problems.
- Build demand forecasting and predictive models using time-series and other advanced techniques.
- Work with large datasets using Python, SQL, PySpark, and cloud-based data platforms.
- Develop and maintain scalable data pipelines for ML model development and deployment.
- Implement MLOps practices including model deployment, monitoring, retraining, and data-drift detection.
- Collaborate with software engineers, product teams, and business stakeholders to convert business requirements into analytical solutions.
- Validate models for accuracy, robustness, bias, and production readiness.
- Create meaningful visualizations and communicate analytical insights to technical and non-technical stakeholders.
Required Skills
- Python
- SQL
- Data Science & Machine Learning
- Predictive Modeling
- Statistical Analysis
- Machine Learning Algorithms
- Optimization Techniques
- Google Cloud Platform (GCP)
- BigQuery
- Dataflow
- Dataproc
- Data Fusion
- Cloud SQL
- Airflow
- PySpark
- PostgreSQL
- Terraform
- Tekton
- APIs
- MLOps
Preferred Skill
- Java
Good to Have
- Demand Forecasting
- Time-Series Analysis
- Neural Networks
- Ensemble Methods
- Support Vector Machines (SVM)
- Regression and Cluster Analysis
- ML Model Testing
- Bias Detection and Data Drift Monitoring
- Production ML Deployment
- QlikSense
- Automotive or Supply Chain Analytics
Education
Bachelor's degree in Computer Science, Data Science, Engineering, Statistics, Mathematics, or a related technical field.Master's degree in a relevant quantitative or technical field is preferred.
Sr.Data Scientist,Python, AI ML
We are looking for a skilled Data Scientist to analyze complex datasets, develop predictive models, and generate actionable insights that support business decisions. The ideal candidate should have strong statistical, analytical, and programming skills, along with hands-on experience in machine learning.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
We are looking for a talented and driven Data Scientist to join our growing Analytics team in India. In this role, you will work at the intersection of advanced machine learning, scalable MLOps infrastructure, and domain-specific healthcare analytics. You will collaborate closely with cross-functional teams to build, deploy, and maintain production-grade ML models that drive real-world impact in clinical trials and healthcare operations.
KEY RESPONSIBILITIES
End-to-End ML Development
• Design, build, and optimize predictive models across the full ML lifecycle—from data ingestion to model serving.
• Conduct rigorous Exploratory Data Analysis (EDA) to surface insights and drive feature engineering decisions.
• Validate model performance using appropriate statistical techniques and domain knowledge.
MLOps & Production Deployment
• Deploy, monitor, and maintain production-grade ML models using Databricks MLFlow endpoints and Unity Catalog.
• Implement CI/CD pipelines for model versioning, experiment tracking, and automated retraining.
• Ensure model reliability, observability, and performance in live production environments.
Language Models & LLM Applications
• Apply transformer-based models (BERT, ClinicalBERT, Trial2Vec) for NLP tasks including classification, NER, and information extraction.
• Build and maintain vector similarity search pipelines for semantic retrieval and recommendation use cases.
• Fine-tune pre-trained models for domain-specific applications in clinical and healthcare contexts.
• Support exploratory work around LLM integration and prompt engineering for internal tooling.
Domain-Driven Analytics
• Apply advanced analytics within complex healthcare and clinical trial datasets—including patient records, trial protocols, and adverse event data.
• Translate ambiguous business problems into structured analytical frameworks with measurable outcomes.
• Partner with domain experts, product managers, and engineering teams to deliver data-driven solutions.
REQUIRED QUALIFICATIONS
Education
• Bachelor’s or Master’s degree in Computer Science, Statistics, Mathematics, Bioinformatics, or a closely related field.
Experience
• 2–4 years of hands-on experience in a data science or machine learning role.
• Demonstrable experience deploying ML models in production environments (not just prototyping).
Technical Skills
• Strong proficiency in Python (pandas, NumPy, scikit-learn, PyTorch / TensorFlow).
• Experience with Databricks, MLFlow (experiment tracking, model registry, endpoints), and Unity Catalog.
• Hands-on experience with BERT-family models and Hugging Face Transformers library.
• Familiarity with vector databases (e.g., FAISS, Pinecone, Weaviate) and embedding-based retrieval.
• Solid understanding of SQL and working with large structured/unstructured datasets.
• Exposure to cloud platforms (AWS / GCP / Azure) and distributed computing frameworks (Spark).
GOOD TO HAVE
• Prior experience with clinical trial data standards (CDISC, CDASH, SDTM) or healthcare ontologies (SNOMED, ICD-10).
• Familiarity with Trial2Vec or similar trial-to-vector embedding approaches.
• Experience with LLM fine-tuning, RAG pipelines, or prompt engineering in a production setting.
• Knowledge of regulatory and compliance considerations in healthcare AI (e.g., FDA guidelines, HIPAA).
• Contributions to open-source ML projects or published research.
THIS ROLE IS NOT FOR YOU IF…
• You have strong SQL/BI skills but limited hands-on ML modelling experience — or you’ve built models only in notebooks without ever deploying them to production.
• Your LLM exposure is limited to API calls and prompt engineering — with no experience fine-tuning models, working with embeddings, or building vector search pipelines.





