Senior Data Scientist pharmacy company only at Staffnixcom · Bengaluru (Bangalore) · 4 - 8 years · ₹21L - ₹27L / yr · Bootstrapped · Posted 14 May 2026

Strong Pharma Analytics Profile
Mandatory (Experience) : Must have 4+ years of experience as an analytics consultant with atleast 2 years in pharma domain
Mandatory (Skill 1) : Must have hands-on experience working with patient-level datasets (claims, EHR, lab, pharmacy data)
Mandatory (Skill 2) : Must have worked on patient journey analysis, treatment patterns, disease progression and advanced analytics
Mandatory (Skill 3) : Must have experience with SQL, Python and Predictive modelling (regression, classification, clustering)
Mandatory (Skill 4) : Must have experience combining multiple healthcare datasets and building longitudinal patient views
Mandatory (Skill 5) : Ability to translate complex analysis into actionable business/clinical insights
Mandatory (Skill 6): Must have experience with time-series analysis and/or survival analysis - specifically to study treatment duration, patient drop-off, or retention trends
Mandatory (Skill 7): Must have experience building risk stratification models using ML techniques to prioritize patients based on clinical or behavioural risk factors
Mandatory (Company) : PharmaTech/life sciences companies

Similar jobs (10)

Are you looking to work in the cutting edge area of applying data-science to help global customers get a better insight into their health? If so, read on and apply.
Role Name: Senior Data Scientist
Science Team | Full-Time | In-Office | Bangalore
The Role
The Ultrahuman Science Team builds the algorithms behind the Ring, M1 CGM, blood and urine biomarkers, and Performance Lab assessments. We are hiring a Senior Data Scientist to own those algorithms end to end: from the raw sensor signal to a model that is shipped, monitored, and trusted in users' hands.
This is a build role with real scope. In a typical month you will improve a production algorithm, root-cause a metric users are complaining about, and stand up the data pipeline the next model needs. The common thread is ownership: you take a vague question and return a working answer, without waiting to be handed scope.
What You'll Do
· Own algorithms end to end: sleep staging, activity detection, sensor-derived metrics, and health scores. You frame the problem, build the features, train and evaluate the model, and see it live
· Ship models, not notebooks: you prove a change on our own cohort before it reaches users, and a model is done only when it runs in production and you can tell how it is behaving
· Validate against reference standards: design evaluations against gold standards, reference devices, and study ground truth, and know when a result is real and when it is an artifact
· Own the data layer: cohort extraction, feature pipelines, study data, and raw sensor data, so the next model starts from clean inputs
What This Looks Like in Practice
1. Improving production algorithms - Take an existing production model like sleep staging, root-cause the failure modes against reference data, and ship a fix you can defend with numbers.
2. Building new models - Train an activity classifier on raw sensor data, design the labeled data collection that expands it, and pick the operating point so false positives never erode trust.
3. Proving it before it ships - Run a new steps algorithm against reference-device cohorts, decide with data when it is ready, and monitor how it behaves after rollout.
Who You Are
The two things we can't coach
· High ownership, end to end: you take a problem from a vague question to a shipped model without waiting to be handed scope, and you can point to something you owned from raw data all the way to production
· Hungry for more scope: you have outgrown your current role and want problems biggerthan your title, with the technical depth to be trusted with them
Also important
· You've worked with human health data: wearables, physiological signals, or clinical data.
If your experience is close but not exact, show us why you will ramp fast
· You've built at a startup: or somewhere small enough that nobody handed you clean data, clear specs, or a mature ML platform
· You work like it's 2026: coding agents and AI tooling are part of how you build every day, and you can tell which new capabilities are worth adopting
· You communicate: you can explain a model and its limits to a product manager, an engineer, or a founder, and hold your own with our scientists Core Technical Skills
· Languages and data: Python and SQL daily, comfortable working in a real codebase
· Machine learning: PyTorch or TensorFlow, scikit-learn, and gradient boosting, with the judgment to know which the problem needs
· Advanced machine learning: time series and sequence models, deep learning on continuous physiological signals, and ensembles
· Statistics and evaluation: hypothesis testing, experiment and A/B design, model evaluation, and error analysis against a reference standard
· Scale and cloud: Spark or equivalent on large datasets, and AWS, GCP, or Azure
· Production ML and MLOps: training pipelines, model versioning, deployment, monitoring, and drift detection
· LLMs and agentic systems: fine-tuning and serving models, building agentic pipelines, and using coding agents to move faster
Experience:
- 4 to 5 years building and shipping machine learning systems. We index on what you have shipped and on trajectory, not the exact number of years; if you are a little earlier but have clearly outgrown your current scope, we want to hear from you.
- Bachelor's or higher in engineering, computer science, statistics, or a related field.
How We Work and Who Thrives Here
- The Science team is small and moves fast, and much of the work has no precedent to copy.
- People do their best work here when they are energized by ambiguity, low on ego, quick to adopt a better idea no matter where it comes from, and comfortable owning something before anyone has told them how. If you need a mature data org, clean labelled datasets, and clear guardrails to thrive, this particular role will not be the right fit, and that is worth knowing up front.
What You'll Gain
· Ownership of algorithms that hundreds of thousands of people see every morning
· A dataset most scientists never get to touch: 100M+ nights of sleep and continuous physiological signals at scale
· Direct collaboration with the engineering, product, and design teams building Ultrahuman
We are looking for a talented and driven Data Scientist to join our growing Analytics team in India. In this role, you will work at the intersection of advanced machine learning, scalable MLOps infrastructure, and domain-specific healthcare analytics. You will collaborate closely with cross-functional teams to build, deploy, and maintain production-grade ML models that drive real-world impact in clinical trials and healthcare operations.
KEY RESPONSIBILITIES
End-to-End ML Development
• Design, build, and optimize predictive models across the full ML lifecycle—from data ingestion to model serving.
• Conduct rigorous Exploratory Data Analysis (EDA) to surface insights and drive feature engineering decisions.
• Validate model performance using appropriate statistical techniques and domain knowledge.
MLOps & Production Deployment
• Deploy, monitor, and maintain production-grade ML models using Databricks MLFlow endpoints and Unity Catalog.
• Implement CI/CD pipelines for model versioning, experiment tracking, and automated retraining.
• Ensure model reliability, observability, and performance in live production environments.
Language Models & LLM Applications
• Apply transformer-based models (BERT, ClinicalBERT, Trial2Vec) for NLP tasks including classification, NER, and information extraction.
• Build and maintain vector similarity search pipelines for semantic retrieval and recommendation use cases.
• Fine-tune pre-trained models for domain-specific applications in clinical and healthcare contexts.
• Support exploratory work around LLM integration and prompt engineering for internal tooling.
Domain-Driven Analytics
• Apply advanced analytics within complex healthcare and clinical trial datasets—including patient records, trial protocols, and adverse event data.
• Translate ambiguous business problems into structured analytical frameworks with measurable outcomes.
• Partner with domain experts, product managers, and engineering teams to deliver data-driven solutions.
REQUIRED QUALIFICATIONS
Education
• Bachelor’s or Master’s degree in Computer Science, Statistics, Mathematics, Bioinformatics, or a closely related field.
Experience
• 2–4 years of hands-on experience in a data science or machine learning role.
• Demonstrable experience deploying ML models in production environments (not just prototyping).
Technical Skills
• Strong proficiency in Python (pandas, NumPy, scikit-learn, PyTorch / TensorFlow).
• Experience with Databricks, MLFlow (experiment tracking, model registry, endpoints), and Unity Catalog.
• Hands-on experience with BERT-family models and Hugging Face Transformers library.
• Familiarity with vector databases (e.g., FAISS, Pinecone, Weaviate) and embedding-based retrieval.
• Solid understanding of SQL and working with large structured/unstructured datasets.
• Exposure to cloud platforms (AWS / GCP / Azure) and distributed computing frameworks (Spark).
GOOD TO HAVE
• Prior experience with clinical trial data standards (CDISC, CDASH, SDTM) or healthcare ontologies (SNOMED, ICD-10).
• Familiarity with Trial2Vec or similar trial-to-vector embedding approaches.
• Experience with LLM fine-tuning, RAG pipelines, or prompt engineering in a production setting.
• Knowledge of regulatory and compliance considerations in healthcare AI (e.g., FDA guidelines, HIPAA).
• Contributions to open-source ML projects or published research.
THIS ROLE IS NOT FOR YOU IF…
• You have strong SQL/BI skills but limited hands-on ML modelling experience — or you’ve built models only in notebooks without ever deploying them to production.
• Your LLM exposure is limited to API calls and prompt engineering — with no experience fine-tuning models, working with embeddings, or building vector search pipelines.
Sr.Data Scientist,Python, AI ML
We are looking for a skilled Data Scientist to analyze complex datasets, develop predictive models, and generate actionable insights that support business decisions. The ideal candidate should have strong statistical, analytical, and programming skills, along with hands-on experience in machine learning.

🚀 We’re Hiring | Data Scientist 🧠📊
Ready to turn data into real-world intelligence? Join us and work on exciting AI/ML & data-driven solutions!
🔹 Experience: 8+ Years
🔹 Must-Have Skills:
🐍 Python | 🤖 Machine Learning | ☁️ Cloud | 🧠 NLP | 📊 Data Visualization
📍 Location: Pune
💼 Work Mode: Work from Office
If you're passionate about Data Science, AI & solving complex business problems, we’d love to hear from you!
📩 Interested? Kindly text
#Hiring #DataScientist #DataScience #MachineLearning #Python #NLP #AI #Cloud #DataVisualization #TechJobs #HiringNow
Location – Hyderabad (Hybrid)
Work Experience – 5 to 7 years
CTC – upto 20 LPA
Roles & Responsibilities:
· We are looking for a Senior Data Engineering who will be majorly responsible for designing, building and maintaining ETL/ ELT pipelines.
· Integration of data from multiple sources or vendors to provide the holistic insights from data.
· You are expected to build and manage Data warehouse solutions, designing data models, creating ETL processes, implementing data quality mechanisms etc.
· Performs EDA (exploratory data analysis) required to troubleshoot data related issues and assist in the resolution of data issues.
· Should have experience in client interaction.
· Experience in mentoring juniors and providing required guidance.
Required Technical Skills
· Extensive hands on experience in Python, Pyspark, SQL, Dataiku.
· Strong experience in Data Warehouse, ETL, Data Modelling, building ETL Pipelines, Snowflake database.
· Working knowledge in Databricks, Redshift, ADF etc.
· Hands-on experience in cloud services like Azure, AWS- S3, Glue, Lambda, CloudWatch, Athena.
· Sound knowledge in end-to-end Data management, Data ops, quality and data governance.
· Familiar with SFDC, Waterfall/ Agile methodology.
· Strong domain knowledge in Pharma domain/ life sciences commercial data operations.
Qualifications
· Bachelor’s or master’s Engineering/ MCA or equivalent degree.
· 5-7 years of relevant industry experience as Data Engineer.
· Experience working on Pharma syndicated data such as IQVIA, Veeva, Symphony; Claims, CRM, Sales etc.
· High motivation, good work ethic, maturity, self-organized and personal initiative.
· Ability to work collaboratively and providing the support to the team.
· Excellent written and verbal communication skills.
· Strong analytical and problem-solving skills.
Role Overview
As a Data Scientist, you will work with business stakeholders, AI engineers, and domain experts to transform data into actionable insights and intelligent solutions. You will develop machine learning models, perform statistical analysis, and contribute to AI-driven products that create measurable business impact.
Key Responsibilities
Data Science & Machine Learning
- Analyze structured and unstructured data to identify patterns, trends, and business opportunities.
- Perform exploratory data analysis (EDA), feature engineering, and data preparation.
- Develop, evaluate, and optimize machine learning models for prediction, classification, clustering, and forecasting.
- Apply statistical techniques to solve business problems and validate model performance.
- Design and execute experiments to improve model accuracy and business outcomes.
AI Solution Development
- Collaborate with AI Engineers, Data Engineers, and domain experts to build AI-powered solutions.
- Translate business requirements into scalable data science approaches.
- Contribute to Generative AI and advanced analytics initiatives where applicable.
- Document methodologies, model performance, and key findings.
Required Technical Skills
- Strong programming skills in Python and SQL for data analysis, feature engineering, and machine learning.
- Strong understanding of Statistics, Probability, Linear Algebra, and Calculus as applied to machine learning and data science.
- Experience with Exploratory Data Analysis (EDA), data preprocessing, feature engineering, feature selection, and handling missing or imbalanced data.
- Good understanding of Supervised, Unsupervised, and Ensemble Machine Learning algorithms, including their assumptions, strengths, limitations, and appropriate use cases.
- Strong knowledge of Regression, Classification, Clustering, Time Series Forecasting, Dimensionality Reduction, Recommendation Systems, and Anomaly Detection techniques.
- Experience with Model Evaluation, Cross-Validation, Hyperparameter Optimization, Bias-Variance Trade-off, Feature Importance, Explainable AI (XAI), and Performance Metrics.
- Understanding of Statistical Inference, Hypothesis Testing, Probability Distributions, Sampling Techniques, Confidence Intervals, and A/B Testing.
- Experience translating business problems into analytical approaches and developing scalable, data-driven solutions.
- Working knowledge of Generative AI, Large Language Models (LLMs), Prompt Engineering, and Retrieval-Augmented Generation (RAG) is preferred.
Preferred Qualifications
- Bachelor's or master's degree in computer science, Artificial Intelligence, Data Science, Statistics, Mathematics, Engineering, or a related field.
- 2–4 years of experience developing machine learning or data science solutions.
- Experience working on end-to-end data science projects in a business environment.
Nice to Have
- Exposure to Generative AI, LLMs, RAG, or Agentic AI.
- Experience with Computer Vision or Natural Language Processing (NLP).
- Familiarity with cloud-based AI platforms.
- Knowledge of construction, engineering, manufacturing, or industrial domains.
- Participation in hackathons, research, Kaggle competitions, or open-source projects.
Soft Skills
Strong analytical and problem-solving skills, effective communication and collaboration, ownership mindset, adaptability, continuous learning, and a passion for innovation.
We’re looking for a dynamic and driven Data Analyst to join our team of technology enthusiasts. This role is crucial in transforming data into insights that support strategic decision-making and innovation within the insurance technology (InsurTech) space. If you’re passionate about working with data, understanding systems, and delivering value through analytics, we’d love to hear from you.
What We’re Looking For
- Proven experience working as a Data Analyst or in a similar analytical role
- 5+ Years of experience in the field
- Strong command of SQL for querying and manipulating relational databases
- Experience with Power BI for building impactful dashboards and reports
- Familiarity with QlikView and Qlik Sense is a plus
- Ability to communicate findings clearly to technical and non-technical stakeholders
- Knowledge of Python or R for data manipulation is nice to have
- Bachelor’s degree in Computer Science, Statistics, Mathematics, Economics, or a related field
- Understanding of the insurance industry or InsurTech is a strong advantage
What You’ll Be Doing:
- Delivering timely and insightful reports to support strategic decision-making
- Working extensively with Policy Administration System (PAS) data to uncover patterns and trends
- Ensuring data accuracy and consistency across reports and systems
- Collaborating with clients, underwriters, and brokers to translate business needs into data solutions
- Organizing and structuring datasets, contributing to data engineering workflows and pipelines
- Producing analytics to support business development and market strategy
We are looking for a talented and driven Data Scientist to join our growing Analytics team in India. In this role, you will work at the intersection of advanced machine learning, scalable MLOps infrastructure, and domain-specific healthcare analytics. You will collaborate closely with cross-functional teams to build, deploy, and maintain production-grade ML models that drive real-world impact in clinical trials and healthcare operations.
KEY RESPONSIBILITIES
End-to-End ML Development
• Design, build, and optimize predictive models across the full ML lifecycle—from data ingestion to model serving.
• Conduct rigorous Exploratory Data Analysis (EDA) to surface insights and drive feature engineering decisions.
• Validate model performance using appropriate statistical techniques and domain knowledge.
MLOps & Production Deployment
• Deploy, monitor, and maintain production-grade ML models using Databricks MLFlow endpoints and Unity Catalog.
• Implement CI/CD pipelines for model versioning, experiment tracking, and automated retraining.
• Ensure model reliability, observability, and performance in live production environments.
Language Models & LLM Applications
• Apply transformer-based models (BERT, ClinicalBERT, Trial2Vec) for NLP tasks including classification, NER, and information extraction.
• Build and maintain vector similarity search pipelines for semantic retrieval and recommendation use cases.
• Fine-tune pre-trained models for domain-specific applications in clinical and healthcare contexts.
• Support exploratory work around LLM integration and prompt engineering for internal tooling.
Domain-Driven Analytics
• Apply advanced analytics within complex healthcare and clinical trial datasets—including patient records, trial protocols, and adverse event data.
• Translate ambiguous business problems into structured analytical frameworks with measurable outcomes.
• Partner with domain experts, product managers, and engineering teams to deliver data-driven solutions.
REQUIRED QUALIFICATIONS
Education
• Bachelor’s or Master’s degree in Computer Science, Statistics, Mathematics, Bioinformatics, or a closely related field.
Experience
• 2–4 years of hands-on experience in a data science or machine learning role.
• Demonstrable experience deploying ML models in production environments (not just prototyping).
Technical Skills
• Strong proficiency in Python (pandas, NumPy, scikit-learn, PyTorch / TensorFlow).
• Experience with Databricks, MLFlow (experiment tracking, model registry, endpoints), and Unity Catalog.
• Hands-on experience with BERT-family models and Hugging Face Transformers library.
• Familiarity with vector databases (e.g., FAISS, Pinecone, Weaviate) and embedding-based retrieval.
• Solid understanding of SQL and working with large structured/unstructured datasets.
• Exposure to cloud platforms (AWS / GCP / Azure) and distributed computing frameworks (Spark).
GOOD TO HAVE
• Prior experience with clinical trial data standards (CDISC, CDASH, SDTM) or healthcare ontologies (SNOMED, ICD-10).
• Familiarity with Trial2Vec or similar trial-to-vector embedding approaches.
• Experience with LLM fine-tuning, RAG pipelines, or prompt engineering in a production setting.
• Knowledge of regulatory and compliance considerations in healthcare AI (e.g., FDA guidelines, HIPAA).
• Contributions to open-source ML projects or published research.
THIS ROLE IS NOT FOR YOU IF…
• You have strong SQL/BI skills but limited hands-on ML modelling experience — or you’ve built models only in notebooks without ever deploying them to production.
• Your LLM exposure is limited to API calls and prompt engineering — with no experience fine-tuning models, working with embeddings, or building vector search pipelines.
- 10+ years of software development experience
- 3+ years in a technical leadership role
- Strong expertise in Python and SQL
- Experience building scalable APIs and backend systems
- Solid understanding of database design and performance tuning
- Experience with Azure cloud services (AWS familiarity preferred)
- Working knowledge of ML/AI integration in enterprise systems
- Experience in client-facing or consulting environments preferred
- Experience with Databricks or modern data platforms
- Exposure to ETL tools such as Talend
- Experience with BI tools (e.g., Power BI)
- Exposure to regulated domains such as Pharma, Healthcare
We are seeking a Senior Data Science & ML Associate with 4+ years of applied ML experience to build and ship models end-to-end from data prep and feature engineering to training, evaluation, and deployment driving measurable business impact.
Key Responsibilities
• Build, train, and evaluate ML and deep-learning models
• Engineer features and prepare data at scale
• Deploy models and monitor production performance
• Partner with stakeholders to frame problems and metrics
• Communicate results and drive decisions
• Iterate on models from business feedback
Mandatory Skills
• 4+ years applied machine learning
• Strong Python (Pandas, NumPy, scikit-learn)
• Classical ML and deep learning (TensorFlow/PyTorch)
• Solid statistics and experiment design
• SQL and data wrangling at scale
• Model deployment / MLOps exposure
Nice to Have: NLP or computer vision; cloud ML (SageMaker, Azure ML)





