Data Scientist at Smartan.ai · Chennai · 4 - 8 years · ₹5L - ₹15L / yr · Profitable · Posted 12 Nov 2024

Role Overview:
We are seeking a highly skilled and motivated Data Scientist to join our growing team. The ideal candidate will be responsible for developing and deploying machine learning models from scratch to production level, focusing on building robust data-driven products. You will work closely with software engineers, product managers, and other stakeholders to ensure our AI-driven solutions meet the needs of our users and align with the company's strategic goals.
Key Responsibilities:
- Develop, implement, and optimize machine learning models and algorithms to support product development.
- Work on the end-to-end lifecycle of data science projects, including data collection, preprocessing, model training, evaluation, and deployment.
- Collaborate with cross-functional teams to define data requirements and product taxonomy.
- Design and build scalable data pipelines and systems to support real-time data processing and analysis.
- Ensure the accuracy and quality of data used for modeling and analytics.
- Monitor and evaluate the performance of deployed models, making necessary adjustments to maintain optimal results.
- Implement best practices for data governance, privacy, and security.
- Document processes, methodologies, and technical solutions to maintain transparency and reproducibility.
Qualifications:
- Bachelor's or Master's degree in Data Science, Computer Science, Engineering, or a related field.
- 5+ years of experience in data science, machine learning, or a related field, with a track record of developing and deploying products from scratch to production.
- Strong programming skills in Python and experience with data analysis and machine learning libraries (e.g., Pandas, NumPy, TensorFlow, PyTorch).
- Experience with cloud platforms (e.g., AWS, GCP, Azure) and containerization technologies (e.g., Docker).
- Proficiency in building and optimizing data pipelines, ETL processes, and data storage solutions.
- Hands-on experience with data visualization tools and techniques.
- Strong understanding of statistics, data analysis, and machine learning concepts.
- Excellent problem-solving skills and attention to detail.
- Ability to work collaboratively in a fast-paced, dynamic environment.
Preferred Qualifications:
- Knowledge of microservices architecture and RESTful APIs.
- Familiarity with Agile development methodologies.
- Experience in building taxonomy for data products.
- Strong communication skills and the ability to explain complex technical concepts to non-technical stakeholders.

About Smartan.ai
About
Smartan Fit is an innovative fitness tech company dedicated to transforming the gym experience for owners, trainers, and members. We leverage advanced technology to provide real-time insights and personalized recommendations, empowering the fitness community to achieve their goals more efficiently and effectively.
Candid answers by the company
At Smartan-Fit, we aim to empower gym owners, trainers, and members alike by revolutionizing gym management through cutting-edge technology. We believe that fitness should be personalized, data-driven, and seamlessly integrated into the gym experience.
Member-Centric Approach:
- Smartan-Fit prioritizes user experience and results.
- Members receive real-time feedback, celebrate milestones, and stay motivated.
- No more guesswork—just data-driven progress.
Privacy and Trust:
- We respect privacy. Smartan-Fit’s camera-based tracking is non-intrusive and GDPR-compliant.
- Members control their data, and transparency is our commitment.
Gym Efficiency:
- Smartan-Fit streamlines operations reduces paperwork and enhances staff productivity.
- Managers can focus on what matters—delivering exceptional fitness experiences.
Company social profiles
Similar jobs (10)
About the Role
We are looking for a highly skilled Data Scientist with strong expertise in Machine Learning, MLOps, and Generative AI. The ideal candidate will have hands-on experience in building scalable ML models, deploying them in production, and working with modern AI frameworks, including GenAI technologies.
Key Responsibilities
· Design, develop, and deploy machine learning models for real-world business problems
· Work on end-to-end ML lifecycle: data preprocessing, model building, evaluation, deployment, and monitoring
· Implement and manage MLOps pipelines for scalable and reproducible workflows
· Utilize tools like MLflow for experiment tracking, model versioning, and lifecycle management
· Develop and integrate Generative AI (GenAI) solutions such as LLM-based applications
· Collaborate with cross-functional teams (engineering, product, business) to translate requirements into AI solutions
· Optimize model performance and ensure production stability
· Stay updated with the latest advancements in AI/ML and GenAI ecosystems
Required Skills & Qualifications
· 4+ years of experience in Data Science / Machine Learning
· Strong programming skills in Python
· Hands-on experience with ML modeling techniques (supervised, unsupervised, NLP, etc.)
· Solid understanding of MLOps practices and tools
· Experience with MLflow or similar model lifecycle tools
· Practical experience in Generative AI (GenAI), including working with LLMs
· Experience with libraries/frameworks like Scikit-learn, TensorFlow, PyTorch
· Strong understanding of data structures, algorithms, and statistics
· Experience with cloud platforms (AWS/GCP/Azure) is a plus
Good to Have
· Experience with LLM fine-tuning, prompt engineering, or RAG pipelines
· Exposure to Docker, Kubernetes, and CI/CD pipelines
· Knowledge of data engineering workflows
AuxoAI is hiring a Senior Data Scientist with strong expertise in AI, machine learning engineering (MLE), and generative AI. You will play a leading role in designing, deploying, and scaling production-grade ML systems — including large language model (LLM)-based pipelines, AI copilots, and agentic workflows. This role is ideal for someone who thrives on balancing cutting-edge research with production rigor and loves mentoring while building impact-first AI applications.
Location - Mumbai/Bangalore/Hyderabad/Gurgaon (Hybrid - 3 Days a week in Office)
Responsibilities:
- Own the full ML lifecycle: model design, training, evaluation, deployment
- Design production-ready ML pipelines with CI/CD, testing, monitoring, and drift detection
- Fine-tune LLMs and implement retrieval-augmented generation (RAG) pipelines
- Build agentic workflows for reasoning, planning, and decision-making
- Develop both real-time and batch inference systems using Docker, Kubernetes, and Spark
- Leverage state-of-the-art architectures: transformers, diffusion models, RLHF, and multimodal pipelines
- Collaborate with product and engineering teams to integrate AI models into business applications
- Mentor junior team members and promote MLOps, scalable architecture, and responsible AI best practices
Requirements
- 5+ years of experience in designing, deploying, and scaling ML/DL systems in production
- Proficient in Python and deep learning frameworks such as PyTorch, TensorFlow, or JAX
- Experience with LLM fine-tuning, LoRA/QLoRA, vector search (Weaviate/PGVector), and RAG pipelines
- Familiarity with agent-based development (e.g., ReAct agents, function-calling, orchestration)
- Solid understanding of MLOps: Docker, Kubernetes, Spark, model registries, and deployment workflows
- Strong software engineering background with experience in testing, version control, and APIs
- Proven ability to balance innovation with scalable deployment
- B.S./M.S./Ph.D. in Computer Science, Data Science, or a related field
- Bonus: Open-source contributions, GenAI research, or applied systems at scale
Description
We’re seeking a highly skilled, execution-focused Senior Data Scientist with a minimum of 5 years of experience. This role demands hands-on expertise in building, deploying, and optimizing machine learning models at scale, while working with big data technologies and modern cloud platforms. You will be responsible for driving data-driven solutions from experimentation to production, leveraging advanced tools and frameworks across Python, SQL, Spark, and AWS. The role requires strong technical depth, problem-solving ability, and ownership in delivering business impact through data science.
Responsibilities
- Design, build, and deploy scalable machine learning models into production systems.
- Develop advanced analytics and predictive models using Python, SQL, and popular ML/DL frameworks (Pandas, Scikit-learn, TensorFlow, PyTorch).
- Leverage Databricks, Apache Spark, and Hadoop for large-scale data processing and model training.
- Implement workflows and pipelines using Airflow and AWS EMR for automation and orchestration.
- Collaborate with engineering teams to integrate models into cloud-based applications on AWS.
- Optimize query performance, storage usage, and data pipelines for efficiency.
- Conduct end-to-end experiments, including data preprocessing, feature engineering, model training, validation, and deployment.
- Drive initiatives independently with high ownership and accountability.
- Stay up to date with industry best practices in machine learning, big data, and cloud-native deployments.
Requirements
- Minimum 5 years of experience in Data Science or Applied Machine Learning.
- Strong proficiency in Python, SQL, and ML libraries (Pandas, Scikit-learn, TensorFlow, PyTorch).
- Proven expertise in deploying ML models into production systems.
- Experience with big data platforms (Hadoop, Spark) and distributed data processing.
- Hands-on experience with Databricks, Airflow, and AWS EMR.
- Strong knowledge of AWS cloud services (S3, Lambda, SageMaker, EC2, etc.).
- Solid understanding of query optimization, storage systems, and data pipelines.
- Excellent problem-solving skills, with the ability to design scalable solutions.
- Strong communication and collaboration skills to work in cross-functional teams.
Benefits
- Best-in-class salary: We hire strong talent and compensate accordingly.
- Proximity Talks: Meet and learn from designers, engineers, product leaders, and AI practitioners.
- Continuous learning: Work with a world-class team and stay close to the latest in AI, engineering, and product development.
- High-impact work: Build AI-first systems and products used at scale by global clients.
About Us
Proximity is the trusted technology, design, and consulting partner for some of the biggest Sports, Media, and Entertainment companies in the world. We’re headquartered in San Francisco and have offices in Palo Alto, Dubai, Mumbai, and Bangalore.
Since 2019, Proximity has built high-impact, scalable products used by millions of users every day. Today, we are a global team of engineers, designers, product managers, and experts solving complex problems and building cutting-edge technology at scale.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Strong Data Scientist / AI Engineer / Generative AI Engineer profile.
2
Mandatory (Experience 1) - Must have 3+ years of hands-on experience in Data Science, Artificial Intelligence, Machine Learning, Deep Learning, NLP, or Generative AI application development.
3
Mandatory (Experience 2) - Must have strong hands-on experience in Python programming, backend development, API development, and production-grade application support.
4
Mandatory (Experience 3) - Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, or Scikit-learn.
5
Mandatory (Experience 4) - Must have hands-on experience in NLP use cases such as text classification, sentiment analysis, entity recognition (NER), semantic search, embeddings, or document understanding.
6
Mandatory (Experience 5) - Must have experience working with Large Language Models (LLMs) such as GPT, LLaMA, Mistral, Phi, Claude, Gemini, or similar models.
7
Mandatory (Experience 6) - Must have hands-on experience building or implementing Retrieval Augmented Generation (RAG) solutions, vector search, semantic search, or knowledge-based AI applications.
8
Mandatory (Experience 7) - Must have experience with Prompt Engineering and Generative AI frameworks such as LangChain, LangGraph, AI Agents, Azure OpenAI, or similar technologies.
9
Mandatory (Experience 8) - Must have experience developing, consuming, or integrating APIs using Python frameworks such as FastAPI, Flask, or similar technologies.
10
Mandatory (CTC) - The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
11
Preferred (Experience 1) - Experience with LLMOps/MLOps tools for monitoring, evaluation, experimentation, and versioning of AI models.
12
Preferred (Experience 2) - Exposure to Azure OpenAI, Azure Kubernetes Service (AKS), Kubernetes, cloud-native AI deployments, or distributed systems.
13
Preferred (Experience 3) - Experience working with PostgreSQL, MongoDB, Redis, Kafka, or large-scale data platforms.
14
Preferred (Experience 4) - Familiarity with Docker, Kubernetes, cloud platforms, and scalable deployment architecture.
15
Preferred (Company) - Candidates from AI-first startups, product companies, SaaS organizations, fintech, or data-driven technology companies.
16
Mandatory ( Age ) - Candidate Should be Below 28 Years.
17
Mandatory ( Pedigree) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Strong Data Scientist / AI Engineer / Generative AI Engineer profile.
2
Mandatory (Experience 1) - Must have 3+ years of hands-on experience in Data Science, Artificial Intelligence, Machine Learning, Deep Learning, NLP, or Generative AI application development.
3
Mandatory (Experience 2) - Must have strong hands-on experience in Python programming, backend development, API development, and production-grade application support.
4
Mandatory (Experience 3) - Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, or Scikit-learn.
5
Mandatory (Experience 4) - Must have hands-on experience in NLP use cases such as text classification, sentiment analysis, entity recognition (NER), semantic search, embeddings, or document understanding.
6
Mandatory (Experience 5) - Must have experience working with Large Language Models (LLMs) such as GPT, LLaMA, Mistral, Phi, Claude, Gemini, or similar models.
7
Mandatory (Experience 6) - Must have hands-on experience building or implementing Retrieval Augmented Generation (RAG) solutions, vector search, semantic search, or knowledge-based AI applications.
8
Mandatory (Experience 7) - Must have experience with Prompt Engineering and Generative AI frameworks such as LangChain, LangGraph, AI Agents, Azure OpenAI, or similar technologies.
9
Mandatory (Experience 8) - Must have experience developing, consuming, or integrating APIs using Python frameworks such as FastAPI, Flask, or similar technologies.
10
Mandatory (CTC) - The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
11
Preferred (Experience 1) - Experience with LLMOps/MLOps tools for monitoring, evaluation, experimentation, and versioning of AI models.
12
Preferred (Experience 2) - Exposure to Azure OpenAI, Azure Kubernetes Service (AKS), Kubernetes, cloud-native AI deployments, or distributed systems.
13
Preferred (Experience 3) - Experience working with PostgreSQL, MongoDB, Redis, Kafka, or large-scale data platforms.
14
Preferred (Experience 4) - Familiarity with Docker, Kubernetes, cloud platforms, and scalable deployment architecture.
15
Preferred (Company) - Candidates from AI-first startups, product companies, SaaS organizations, fintech, or data-driven technology companies.
16
Mandatory ( Age ) - Candidate Should be Below 28 Years.
17
Mandatory ( Pedigree) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered.

Are you looking to work in the cutting edge area of applying data-science to help global customers get a better insight into their health? If so, read on and apply.
Role Name: Senior Data Scientist
Science Team | Full-Time | In-Office | Bangalore
The Role
The Ultrahuman Science Team builds the algorithms behind the Ring, M1 CGM, blood and urine biomarkers, and Performance Lab assessments. We are hiring a Senior Data Scientist to own those algorithms end to end: from the raw sensor signal to a model that is shipped, monitored, and trusted in users' hands.
This is a build role with real scope. In a typical month you will improve a production algorithm, root-cause a metric users are complaining about, and stand up the data pipeline the next model needs. The common thread is ownership: you take a vague question and return a working answer, without waiting to be handed scope.
What You'll Do
· Own algorithms end to end: sleep staging, activity detection, sensor-derived metrics, and health scores. You frame the problem, build the features, train and evaluate the model, and see it live
· Ship models, not notebooks: you prove a change on our own cohort before it reaches users, and a model is done only when it runs in production and you can tell how it is behaving
· Validate against reference standards: design evaluations against gold standards, reference devices, and study ground truth, and know when a result is real and when it is an artifact
· Own the data layer: cohort extraction, feature pipelines, study data, and raw sensor data, so the next model starts from clean inputs
What This Looks Like in Practice
1. Improving production algorithms - Take an existing production model like sleep staging, root-cause the failure modes against reference data, and ship a fix you can defend with numbers.
2. Building new models - Train an activity classifier on raw sensor data, design the labeled data collection that expands it, and pick the operating point so false positives never erode trust.
3. Proving it before it ships - Run a new steps algorithm against reference-device cohorts, decide with data when it is ready, and monitor how it behaves after rollout.
Who You Are
The two things we can't coach
· High ownership, end to end: you take a problem from a vague question to a shipped model without waiting to be handed scope, and you can point to something you owned from raw data all the way to production
· Hungry for more scope: you have outgrown your current role and want problems biggerthan your title, with the technical depth to be trusted with them
Also important
· You've worked with human health data: wearables, physiological signals, or clinical data.
If your experience is close but not exact, show us why you will ramp fast
· You've built at a startup: or somewhere small enough that nobody handed you clean data, clear specs, or a mature ML platform
· You work like it's 2026: coding agents and AI tooling are part of how you build every day, and you can tell which new capabilities are worth adopting
· You communicate: you can explain a model and its limits to a product manager, an engineer, or a founder, and hold your own with our scientists Core Technical Skills
· Languages and data: Python and SQL daily, comfortable working in a real codebase
· Machine learning: PyTorch or TensorFlow, scikit-learn, and gradient boosting, with the judgment to know which the problem needs
· Advanced machine learning: time series and sequence models, deep learning on continuous physiological signals, and ensembles
· Statistics and evaluation: hypothesis testing, experiment and A/B design, model evaluation, and error analysis against a reference standard
· Scale and cloud: Spark or equivalent on large datasets, and AWS, GCP, or Azure
· Production ML and MLOps: training pipelines, model versioning, deployment, monitoring, and drift detection
· LLMs and agentic systems: fine-tuning and serving models, building agentic pipelines, and using coding agents to move faster
Experience:
- 4 to 5 years building and shipping machine learning systems. We index on what you have shipped and on trajectory, not the exact number of years; if you are a little earlier but have clearly outgrown your current scope, we want to hear from you.
- Bachelor's or higher in engineering, computer science, statistics, or a related field.
How We Work and Who Thrives Here
- The Science team is small and moves fast, and much of the work has no precedent to copy.
- People do their best work here when they are energized by ambiguity, low on ego, quick to adopt a better idea no matter where it comes from, and comfortable owning something before anyone has told them how. If you need a mature data org, clean labelled datasets, and clear guardrails to thrive, this particular role will not be the right fit, and that is worth knowing up front.
What You'll Gain
· Ownership of algorithms that hundreds of thousands of people see every morning
· A dataset most scientists never get to touch: 100M+ nights of sleep and continuous physiological signals at scale
· Direct collaboration with the engineering, product, and design teams building Ultrahuman
Sr.Data Scientist,Python, AI ML
We are looking for a skilled Data Scientist to analyze complex datasets, develop predictive models, and generate actionable insights that support business decisions. The ideal candidate should have strong statistical, analytical, and programming skills, along with hands-on experience in machine learning.
Job Description – Data Scientist (Machine Learning & Forecasting)
About the Role
We are looking for a highly skilled Data Scientist with strong expertise in Machine Learning, Traditional Statistical Modelling, Forecasting, and Predictive Analytics. The ideal candidate will have hands-on experience building and deploying end-to-end ML solutions, working with large datasets, and translating business problems into scalable data science solutions.
The role requires a strong foundation in statistics, predictive modelling, feature engineering, model evaluation, and time-series forecasting, along with the ability to collaborate with cross-functional teams to deliver business impact.
Key Responsibilities
- Design, develop, and deploy Machine Learning models for business-critical use cases.
- Build and optimize traditional ML models such as:
- Linear Regression
- Logistic Regression
- Decision Trees
- Random Forest
- Gradient Boosting (XGBoost, LightGBM, CatBoost)
- Support Vector Machines
- Clustering Algorithms
- Develop forecasting solutions using:
- ARIMA / SARIMA
- Prophet
- Exponential Smoothing
- Time-Series Regression Models
- Perform exploratory data analysis (EDA), feature engineering, and data validation.
- Evaluate model performance using appropriate statistical and business metrics.
- Work with structured and semi-structured datasets from multiple sources.
- Collaborate with business stakeholders to understand requirements and translate them into analytical solutions.
- Build scalable data pipelines and support model deployment in production environments.
- Monitor model performance, identify data drift, and implement model retraining strategies.
- Present insights and recommendations to technical and non-technical stakeholders.
Required Skills & Qualifications
- Bachelor's or Master's degree in Computer Science, Statistics, Mathematics, Data Science, Engineering, or a related quantitative field.
- 5+ years of hands-on experience in Data Science, Machine Learning, and Forecasting.
Technical Skills
Machine Learning
- Strong understanding of supervised and unsupervised learning algorithms.
- Experience with ensemble methods and advanced ML techniques.
- Expertise in model selection, hyperparameter tuning, and performance optimization.
Forecasting & Statistics
- Strong understanding of:
- Time-Series Analysis
- Forecasting Techniques
- Statistical Inference
- Hypothesis Testing
- Probability Distributions
- A/B Testing
Programming
- Advanced proficiency in Python.
- Experience with:
- Pandas
- NumPy
- Scikit-learn
- Statsmodels
- XGBoost / LightGBM
- Prophet
Data & SQL
- Strong SQL skills with experience in complex queries and performance optimization.
- Experience working with large-scale datasets.
Visualization
- Experience with Power BI, Tableau, Matplotlib, Seaborn, or Plotly.
- Cloud & MLOps (Preferred)
- Exposure to AWS, Azure, or GCP.
- Understanding of Docker, Kubernetes, CI/CD, and ML model deployment practices.
Key Competencies
- Strong analytical and problem-solving skills.
- Excellent communication and stakeholder management abilities.
- Ability to work independently in a fast-paced environment.
- Strong business acumen and data-driven decision-making mindset.
Kody Technolab Limited is seeking an experienced AI/ML Engineer to design, develop, and deploy cutting-edge Artificial Intelligence and Machine Learning solutions. The ideal candidate will have
strong expertise in Machine Learning, Deep Learning, Generative AI, LLMs, MLOps, and cloud-based AI deployments.
Key Responsibilities
• Design, develop, and deploy Machine Learning and Deep Learning models for classification, regression, recommendation systems, NLP, Computer Vision, and Generative AI applications.
• Build and maintain end-to-end ML pipelines including data preprocessing, feature engineering, model training, validation, evaluation, and deployment.
• Develop AI solutions using PyTorch, TensorFlow, Scikit-learn, Hugging Face, and related frameworks.
• Work with Large Language Models (LLMs) and foundation models such as GPT, BERT, Llama, Claude, and Stable Diffusion.
• Collaborate with product, engineering, and business teams to translate requirements into scalable AI solutions.
• Optimize model performance, scalability, and reliability for production environments.
• Implement MLOps best practices using tools such as MLflow, Docker, Kubernetes, and Kubeflow.
• Stay updated with emerging trends and research in AI, ML, Deep Learning, and Generative AI.
Required Qualifications
• Bachelor’s or Master’s degree in Computer Science, Data Science, Artificial Intelligence, Mathematics, or a related field.
• 7+ years of hands-on experience in AI/ML product development.
• Strong proficiency in Python and ML frameworks including Scikit-learn, TensorFlow, PyTorch, and Hugging Face.
• Experience with Generative AI, LLMs, GANs, VAEs, diffusion models, and prompt engineering.
• Strong understanding of the ML lifecycle including model training, tuning, deployment, monitoring, and optimization.
• Experience with MLOps tools such as MLflow, Docker, Kubeflow, and CI/CD pipelines.
• Experience with AWS, Azure, or GCP cloud platforms.
• Strong problem-solving and analytical skills.
Preferred Skills
• Fine-tuning and deployment of Large Language Models.
• Experience with RAG (Retrieval Augmented Generation) architectures.
• Contributions to open-source AI projects or research publications.
• Knowledge of model interpretability, data annotation, and feature engineering.
• C++ experience for high-performance AI applications.
Why Join Kody Technolab Limited?
Opportunity to work on innovative AI products, Generative AI solutions, robotics integrations,
and enterprise-scale applications while collaborating with a highly skilled technology team.
Visit the Website to know more about us.
Company Website - Kody Technolab | Deep Tech Company in Robotics & AI Solution
Kody Robots | Robotics Company in India for Autonomous Robots
We are looking for a talented and driven Data Scientist to join our growing Analytics team in India. In this role, you will work at the intersection of advanced machine learning, scalable MLOps infrastructure, and domain-specific healthcare analytics. You will collaborate closely with cross-functional teams to build, deploy, and maintain production-grade ML models that drive real-world impact in clinical trials and healthcare operations.
KEY RESPONSIBILITIES
End-to-End ML Development
• Design, build, and optimize predictive models across the full ML lifecycle—from data ingestion to model serving.
• Conduct rigorous Exploratory Data Analysis (EDA) to surface insights and drive feature engineering decisions.
• Validate model performance using appropriate statistical techniques and domain knowledge.
MLOps & Production Deployment
• Deploy, monitor, and maintain production-grade ML models using Databricks MLFlow endpoints and Unity Catalog.
• Implement CI/CD pipelines for model versioning, experiment tracking, and automated retraining.
• Ensure model reliability, observability, and performance in live production environments.
Language Models & LLM Applications
• Apply transformer-based models (BERT, ClinicalBERT, Trial2Vec) for NLP tasks including classification, NER, and information extraction.
• Build and maintain vector similarity search pipelines for semantic retrieval and recommendation use cases.
• Fine-tune pre-trained models for domain-specific applications in clinical and healthcare contexts.
• Support exploratory work around LLM integration and prompt engineering for internal tooling.
Domain-Driven Analytics
• Apply advanced analytics within complex healthcare and clinical trial datasets—including patient records, trial protocols, and adverse event data.
• Translate ambiguous business problems into structured analytical frameworks with measurable outcomes.
• Partner with domain experts, product managers, and engineering teams to deliver data-driven solutions.
REQUIRED QUALIFICATIONS
Education
• Bachelor’s or Master’s degree in Computer Science, Statistics, Mathematics, Bioinformatics, or a closely related field.
Experience
• 2–4 years of hands-on experience in a data science or machine learning role.
• Demonstrable experience deploying ML models in production environments (not just prototyping).
Technical Skills
• Strong proficiency in Python (pandas, NumPy, scikit-learn, PyTorch / TensorFlow).
• Experience with Databricks, MLFlow (experiment tracking, model registry, endpoints), and Unity Catalog.
• Hands-on experience with BERT-family models and Hugging Face Transformers library.
• Familiarity with vector databases (e.g., FAISS, Pinecone, Weaviate) and embedding-based retrieval.
• Solid understanding of SQL and working with large structured/unstructured datasets.
• Exposure to cloud platforms (AWS / GCP / Azure) and distributed computing frameworks (Spark).
GOOD TO HAVE
• Prior experience with clinical trial data standards (CDISC, CDASH, SDTM) or healthcare ontologies (SNOMED, ICD-10).
• Familiarity with Trial2Vec or similar trial-to-vector embedding approaches.
• Experience with LLM fine-tuning, RAG pipelines, or prompt engineering in a production setting.
• Knowledge of regulatory and compliance considerations in healthcare AI (e.g., FDA guidelines, HIPAA).
• Contributions to open-source ML projects or published research.
THIS ROLE IS NOT FOR YOU IF…
• You have strong SQL/BI skills but limited hands-on ML modelling experience — or you’ve built models only in notebooks without ever deploying them to production.
• Your LLM exposure is limited to API calls and prompt engineering — with no experience fine-tuning models, working with embeddings, or building vector search pipelines.





