Lead Data Scientist / AI/ML Engineer Tier 1 college only Fintech compa at Staffnixcom · Pune · 10 - 12 years · ₹60L - ₹75L / yr · Bootstrapped · Posted 28 Jul 2026

Lead Data Scientist / AI/ML Engineer Tier 1 college only Fintech compa
at Staffnixcom
Strong Lead Data Science, / AI Engineer / Machine Learning Engineer profiles.
2
Mandatory (Experience 1) - Must have 10+ years of experience in Data Science, AI/ML or AI Engineering with hands-on experience building production-grade ML systems.
3
Mandatory (Experience 2) - Must have hands-on experience building AI/ML solutions for Credit Risk, Fraud Risk Management (FRM), Collections & Recovery, with proven delivery of business-impacting AI/ML solutions.
4
Mandatory (Experience 3) - Candidate's Current designation must be Lead or above.
5
Mandatory (Experience 4) - Must have strong experience designing and deploying large-scale distributed Machine Learning systems, including model training, fine-tuning, inference, scalable serving, and production deployment.
6
Mandatory (Experience 5) - Strong programming experience in Python, along with exposure to Spark, Kafka, Kubernetes, APIs/Microservices, CI/CD, Feature Store, Model Registry, and Distributed Computing.
7
Mandatory (Experience 6) - Experience designing and deploying Credit Risk Models, Fraud Detection Models, Graph ML, Early Warning Systems, Portfolio Monitoring, Collections Optimization, Propensity Models, and Recovery Forecasting.
8
Mandatory (Experience 7) - Proven experience leading AI/ML teams, owning end-to-end delivery, mentoring engineers, driving cross-functional execution, and managing production AI platforms.
9
Mandatory (Experience 8) – Must have experience working under BFSI governance, including PII handling, auditability, model governance, compliance, secure-by-design architecture, approval workflows, and model risk management practices.
10
Mandatory ( Education ) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered
11
Mandatory (Age) - Candidate's Age should be below 37 years.
12
Mandatory (CTC) – The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
13
Preferred (Experience 1) - Candidates currently working as Lead / Principal / Engineering Manager / Associate Director / Director in reputed Product, FinTech, Banking, NBFC, or Global Capability Centers will be preferred.
14
Preferred (Experience 2) - Indian professionals currently working overseas (NRI) who are planning to relocate and permanently settle in India are encouraged to apply.
15
Preferred (Experience 4) - Experience building enterprise AI platforms using Graph ML, Vector Databases, LLM-enabled decisioning, distributed training frameworks, and large-scale AI infrastructure.

Similar jobs (10)
Strong Lead Data Science, / AI Engineer / Machine Learning Engineer profiles.
2
Mandatory (Experience 1) - Must have 10+ years of experience in Data Science, AI/ML or AI Engineering with hands-on experience building production-grade ML systems.
3
Mandatory (Experience 2) - Must have hands-on experience building AI/ML solutions for Credit Risk, Fraud Risk Management (FRM), Collections & Recovery, with proven delivery of business-impacting AI/ML solutions.
4
Mandatory (Experience 3) - Candidate's Current designation must be Lead or above.
5
Mandatory (Experience 4) - Must have strong experience designing and deploying large-scale distributed Machine Learning systems, including model training, fine-tuning, inference, scalable serving, and production deployment.
6
Mandatory (Experience 5) - Strong programming experience in Python, along with exposure to Spark, Kafka, Kubernetes, APIs/Microservices, CI/CD, Feature Store, Model Registry, and Distributed Computing.
7
Mandatory (Experience 6) - Experience designing and deploying Credit Risk Models, Fraud Detection Models, Graph ML, Early Warning Systems, Portfolio Monitoring, Collections Optimization, Propensity Models, and Recovery Forecasting.
8
Mandatory (Experience 7) - Proven experience leading AI/ML teams, owning end-to-end delivery, mentoring engineers, driving cross-functional execution, and managing production AI platforms.
9
Mandatory (Experience 8) – Must have experience working under BFSI governance, including PII handling, auditability, model governance, compliance, secure-by-design architecture, approval workflows, and model risk management practices.
10
Mandatory ( Education ) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered
11
Mandatory (Age) - Candidate's Age should be below 37 years.
12
Mandatory (CTC) – The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
13
Preferred (Experience 1) - Candidates currently working as Lead / Principal / Engineering Manager / Associate Director / Director in reputed Product, FinTech, Banking, NBFC, or Global Capability Centers will be preferred.
14
Preferred (Experience 2) - Indian professionals currently working overseas (NRI) who are planning to relocate and permanently settle in India are encouraged to apply.
15
Preferred (Experience 4) - Experience building enterprise AI platforms using Graph ML, Vector Databases, LLM-enabled decisioning, distributed training frameworks, and large-scale AI infrastructure.
Strong Lead Data Science, / AI Engineer / Machine Learning Engineer profiles.
2
Mandatory (Experience 1) - Must have 10+ years of experience in Data Science, AI/ML or AI Engineering with hands-on experience building production-grade ML systems.
3
Mandatory (Experience 2) - Must have hands-on experience building AI/ML solutions for Credit Risk, Fraud Risk Management (FRM), Collections & Recovery, with proven delivery of business-impacting AI/ML solutions.
4
Mandatory (Experience 3) - Candidate's Current designation must be Lead or above.
5
Mandatory (Experience 4) - Must have strong experience designing and deploying large-scale distributed Machine Learning systems, including model training, fine-tuning, inference, scalable serving, and production deployment.
6
Mandatory (Experience 5) - Strong programming experience in Python, along with exposure to Spark, Kafka, Kubernetes, APIs/Microservices, CI/CD, Feature Store, Model Registry, and Distributed Computing.
7
Mandatory (Experience 6) - Experience designing and deploying Credit Risk Models, Fraud Detection Models, Graph ML, Early Warning Systems, Portfolio Monitoring, Collections Optimization, Propensity Models, and Recovery Forecasting.
8
Mandatory (Experience 7) - Proven experience leading AI/ML teams, owning end-to-end delivery, mentoring engineers, driving cross-functional execution, and managing production AI platforms.
9
Mandatory (Experience 8) – Must have experience working under BFSI governance, including PII handling, auditability, model governance, compliance, secure-by-design architecture, approval workflows, and model risk management practices.
10
Mandatory ( Education ) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered
11
Mandatory (Age) - Candidate's Age should be below 37 years.
12
Mandatory (CTC) – The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
13
Preferred (Experience 1) - Candidates currently working as Lead / Principal / Engineering Manager / Associate Director / Director in reputed Product, FinTech, Banking, NBFC, or Global Capability Centers will be preferred.
14
Preferred (Experience 2) - Indian professionals currently working overseas (NRI) who are planning to relocate and permanently settle in India are encouraged to apply.
15
Preferred (Experience 4) - Experience building enterprise AI platforms using Graph ML, Vector Databases, LLM-enabled decisioning, distributed training frameworks, and large-scale AI infrastructure.
Roles & Responsibilities
- Lead AI Product Pods across Credit Risk, Fraud, and Collections functions.
- Build and deploy production-scale Machine Learning systems for lending lifecycle decisioning.
- Own complete ML lifecycle including feature engineering, model training, evaluation, deployment, monitoring, and continuous improvement.
- Design scalable distributed ML infrastructure, feature stores, model registries, and MLOps pipelines.
- Develop AI solutions for underwriting, portfolio risk monitoring, fraud detection, anomaly detection, and recovery optimization.
- Drive model governance, monitoring, explainability, and compliance within BFSI regulatory standards.
- Collaborate with Product, Risk, Engineering, Data, and Business teams to deliver AI-driven business outcomes.
- Define AI platform architecture, operational excellence, SLAs, and incident management practices.
- Build, mentor, and scale high-performing AI Engineering and Data Science teams.
Ideal Candidate
1.Strong Lead Data Science, / AI Engineer / Machine Learning Engineer profiles.
2.Mandatory (Experience 1) - Must have 10+ years of experience in Data Science, AI/ML or AI Engineering with hands-on experience building production-grade ML systems.
3.Mandatory (Experience 2) - Must have hands-on experience building AI/ML solutions for Credit Risk, Fraud Risk Management (FRM), Collections & Recovery, with proven delivery of business-impacting AI/ML solutions.
4.Mandatory (Experience 3) - Candidate's Current designation must be Lead or above.
5.Mandatory (Experience 4) - Must have strong experience designing and deploying large-scale distributed Machine Learning systems, including model training, fine-tuning, inference, scalable serving, and production deployment.
6.Mandatory (Experience 5) - Strong programming experience in Python, along with exposure to Spark, Kafka, Kubernetes, APIs/Microservices, CI/CD, Feature Store, Model Registry, and Distributed Computing
7.Mandatory (Experience 6) - Experience designing and deploying Credit Risk Models, Fraud Detection Models, Graph ML, Early Warning Systems, Portfolio Monitoring, Collections Optimization, Propensity Models, and Recovery Forecasting.
8.Mandatory (Experience 7) - Proven experience leading AI/ML teams, owning end-to-end delivery, mentoring engineers, driving cross-functional execution, and managing production AI platforms.
9.Mandatory (Experience 8) – Must have experience working under BFSI governance, including PII handling, auditability, model governance, compliance, secure-by-design architecture, approval workflows, and model risk management practices.
10.Mandatory ( Education ) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered
11.Mandatory (Age) - Candidate's Age should be below 37 years.
12.Mandatory (CTC) – The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
13.Preferred (Experience 1) - Candidates currently working as Lead / Principal / Engineering Manager / Associate Director / Director in reputed Product, FinTech, Banking, NBFC, or Global Capability Centers will be preferred.
14.Preferred (Experience 2) - Indian professionals currently working overseas (NRI) who are planning to relocate and permanently settle in India are encouraged to apply.
15.Preferred (Experience 4) - Experience building enterprise AI platforms using Graph ML, Vector Databases, LLM-enabled decisioning, distributed training frameworks, and large-scale AI infrastructure.
AuxoAI is hiring a Senior Data Scientist with strong expertise in AI, machine learning engineering (MLE), and generative AI. You will play a leading role in designing, deploying, and scaling production-grade ML systems — including large language model (LLM)-based pipelines, AI copilots, and agentic workflows. This role is ideal for someone who thrives on balancing cutting-edge research with production rigor and loves mentoring while building impact-first AI applications.
Location - Mumbai/Bangalore/Hyderabad/Gurgaon (Hybrid - 3 Days a week in Office)
Responsibilities:
- Own the full ML lifecycle: model design, training, evaluation, deployment
- Design production-ready ML pipelines with CI/CD, testing, monitoring, and drift detection
- Fine-tune LLMs and implement retrieval-augmented generation (RAG) pipelines
- Build agentic workflows for reasoning, planning, and decision-making
- Develop both real-time and batch inference systems using Docker, Kubernetes, and Spark
- Leverage state-of-the-art architectures: transformers, diffusion models, RLHF, and multimodal pipelines
- Collaborate with product and engineering teams to integrate AI models into business applications
- Mentor junior team members and promote MLOps, scalable architecture, and responsible AI best practices
Requirements
- 5+ years of experience in designing, deploying, and scaling ML/DL systems in production
- Proficient in Python and deep learning frameworks such as PyTorch, TensorFlow, or JAX
- Experience with LLM fine-tuning, LoRA/QLoRA, vector search (Weaviate/PGVector), and RAG pipelines
- Familiarity with agent-based development (e.g., ReAct agents, function-calling, orchestration)
- Solid understanding of MLOps: Docker, Kubernetes, Spark, model registries, and deployment workflows
- Strong software engineering background with experience in testing, version control, and APIs
- Proven ability to balance innovation with scalable deployment
- B.S./M.S./Ph.D. in Computer Science, Data Science, or a related field
- Bonus: Open-source contributions, GenAI research, or applied systems at scale
Description
We’re seeking a highly skilled, execution-focused Senior Data Scientist with a minimum of 5 years of experience. This role demands hands-on expertise in building, deploying, and optimizing machine learning models at scale, while working with big data technologies and modern cloud platforms. You will be responsible for driving data-driven solutions from experimentation to production, leveraging advanced tools and frameworks across Python, SQL, Spark, and AWS. The role requires strong technical depth, problem-solving ability, and ownership in delivering business impact through data science.
Responsibilities
- Design, build, and deploy scalable machine learning models into production systems.
- Develop advanced analytics and predictive models using Python, SQL, and popular ML/DL frameworks (Pandas, Scikit-learn, TensorFlow, PyTorch).
- Leverage Databricks, Apache Spark, and Hadoop for large-scale data processing and model training.
- Implement workflows and pipelines using Airflow and AWS EMR for automation and orchestration.
- Collaborate with engineering teams to integrate models into cloud-based applications on AWS.
- Optimize query performance, storage usage, and data pipelines for efficiency.
- Conduct end-to-end experiments, including data preprocessing, feature engineering, model training, validation, and deployment.
- Drive initiatives independently with high ownership and accountability.
- Stay up to date with industry best practices in machine learning, big data, and cloud-native deployments.
Requirements
- Minimum 5 years of experience in Data Science or Applied Machine Learning.
- Strong proficiency in Python, SQL, and ML libraries (Pandas, Scikit-learn, TensorFlow, PyTorch).
- Proven expertise in deploying ML models into production systems.
- Experience with big data platforms (Hadoop, Spark) and distributed data processing.
- Hands-on experience with Databricks, Airflow, and AWS EMR.
- Strong knowledge of AWS cloud services (S3, Lambda, SageMaker, EC2, etc.).
- Solid understanding of query optimization, storage systems, and data pipelines.
- Excellent problem-solving skills, with the ability to design scalable solutions.
- Strong communication and collaboration skills to work in cross-functional teams.
Benefits
- Best-in-class salary: We hire strong talent and compensate accordingly.
- Proximity Talks: Meet and learn from designers, engineers, product leaders, and AI practitioners.
- Continuous learning: Work with a world-class team and stay close to the latest in AI, engineering, and product development.
- High-impact work: Build AI-first systems and products used at scale by global clients.
About Us
Proximity is the trusted technology, design, and consulting partner for some of the biggest Sports, Media, and Entertainment companies in the world. We’re headquartered in San Francisco and have offices in Palo Alto, Dubai, Mumbai, and Bangalore.
Since 2019, Proximity has built high-impact, scalable products used by millions of users every day. Today, we are a global team of engineers, designers, product managers, and experts solving complex problems and building cutting-edge technology at scale.
The Role
Own end-to-end credit & fraud data science: feature engineering from raw bureau JSON ,SMS,DEVICE, scorecard / model development, Business Rule Engine (BRE) design, monitoring, and partnering with product/engineering to put rules live. You will work directly with the existing DS team,Tech,product and founders — decisions are data-backed and debated.
What you will own
- Build and maintain credit scorecards and models for FTB and Repeat Borrowers (Xgboost, Random forest, Support Vector Machine Models, ensemble models, challenger models).
- Engineer features from raw CRIF (or equivalent) bureau JSON — tradelines, enquiries, DPD histories, identity matches — and from raw SMS / FinBox alt-data (collections, rejections, salary, app footprint).
- Design, validate, and ship Models: hard rejects, soft flags, amount caps — with clear lift/capture
/ approval trade-offs.
- Own portfolio risk analytics: vintage / DPD / non-starter / POS bad-rate monitoring; propose tier pauses, cool-offs, and ladder-up changes.
- Build fraud signals (device, SIM/OTP, mule, ring, post-disbursal disappearance) and help prioritise the fraud PRD backlog into production.
- Partner with engineering to productionise features, rules, and models (Watchtower-style shadow underwriting, policy index, monitoring dashboards).
- Challenge and refine existing tier/ladder policy with evidence; communicate clearly to founders and business.
Required experience
- Tenure: 5+ years overall experience in data science/analytics.
- Digital lending: Minimum 3 years hands-on in digital lending/consumer credit (NBFC, fintech lender, digital/STPL/) who has built models themselves.
- Scorecards/models: Built and deployed at least one credit scorecard (first-time borrower or repeat borrower, or combined model) into a live BRE / LOS. Should improve approval–bad-rate trade-offs from production experience.
- Bureau: Parsed and engineered features from raw bureau files (CRIF / CIBIL / Experian JSON or XML) — not only vendor-precomputed attributes.
- Non-starter models: Fraud/non-starter / First Payment default modelling experience in short-tenure lending.
- Limit Assignment: Experience with repeat-borrower ladder / limit-management policies.
- Monitoring and QC: Shadow underwriting/champion–challenger frameworks.
- Alt-data: Worked with SMS / alt-data / device / AA signals for underwriting or fraud (FinBox, similar vendors, or in-house SMS parsing).
- Stack: Strong SQL + Python (pandas, sklearn/Logistic / lightgbm/Xgboost/randomforest, statsmodels). Able to write production-quality notebooks and scripts, not just slide decks.
- Communication: Comfortable debating policy with founders/credit heads using data; owns the "show me the evidence" conversation.
Nice to have:
- Feature stores, Airflow/cron pipelines, S3 + Postgres + DynamoDB.
- Prior Experience: Prior work at a zero-to-one digital lender or STPL product.
What success looks like in 6 months
- A documented feature dictionary from raw bureau + SMS with IV/KS ranking.
- At least one new scorecard/model live with clear expected vs observed bad-rate impact.
- Non-starter / First Payment Defaults monitoring with actionable rule recommendations and clear demonstrated improvements in defaults
- Credible pushback on weak policy ideas — backed by analysis, not opinion.

🚀 We’re Hiring | Data Scientist 🧠📊
Ready to turn data into real-world intelligence? Join us and work on exciting AI/ML & data-driven solutions!
🔹 Experience: 8+ Years
🔹 Must-Have Skills:
🐍 Python | 🤖 Machine Learning | ☁️ Cloud | 🧠 NLP | 📊 Data Visualization
📍 Location: Pune
💼 Work Mode: Work from Office
If you're passionate about Data Science, AI & solving complex business problems, we’d love to hear from you!
📩 Interested? Kindly text
#Hiring #DataScientist #DataScience #MachineLearning #Python #NLP #AI #Cloud #DataVisualization #TechJobs #HiringNow
We are looking for a talented and driven Data Scientist to join our growing Analytics team in India. In this role, you will work at the intersection of advanced machine learning, scalable MLOps infrastructure, and domain-specific healthcare analytics. You will collaborate closely with cross-functional teams to build, deploy, and maintain production-grade ML models that drive real-world impact in clinical trials and healthcare operations.
KEY RESPONSIBILITIES
End-to-End ML Development
• Design, build, and optimize predictive models across the full ML lifecycle—from data ingestion to model serving.
• Conduct rigorous Exploratory Data Analysis (EDA) to surface insights and drive feature engineering decisions.
• Validate model performance using appropriate statistical techniques and domain knowledge.
MLOps & Production Deployment
• Deploy, monitor, and maintain production-grade ML models using Databricks MLFlow endpoints and Unity Catalog.
• Implement CI/CD pipelines for model versioning, experiment tracking, and automated retraining.
• Ensure model reliability, observability, and performance in live production environments.
Language Models & LLM Applications
• Apply transformer-based models (BERT, ClinicalBERT, Trial2Vec) for NLP tasks including classification, NER, and information extraction.
• Build and maintain vector similarity search pipelines for semantic retrieval and recommendation use cases.
• Fine-tune pre-trained models for domain-specific applications in clinical and healthcare contexts.
• Support exploratory work around LLM integration and prompt engineering for internal tooling.
Domain-Driven Analytics
• Apply advanced analytics within complex healthcare and clinical trial datasets—including patient records, trial protocols, and adverse event data.
• Translate ambiguous business problems into structured analytical frameworks with measurable outcomes.
• Partner with domain experts, product managers, and engineering teams to deliver data-driven solutions.
REQUIRED QUALIFICATIONS
Education
• Bachelor’s or Master’s degree in Computer Science, Statistics, Mathematics, Bioinformatics, or a closely related field.
Experience
• 2–4 years of hands-on experience in a data science or machine learning role.
• Demonstrable experience deploying ML models in production environments (not just prototyping).
Technical Skills
• Strong proficiency in Python (pandas, NumPy, scikit-learn, PyTorch / TensorFlow).
• Experience with Databricks, MLFlow (experiment tracking, model registry, endpoints), and Unity Catalog.
• Hands-on experience with BERT-family models and Hugging Face Transformers library.
• Familiarity with vector databases (e.g., FAISS, Pinecone, Weaviate) and embedding-based retrieval.
• Solid understanding of SQL and working with large structured/unstructured datasets.
• Exposure to cloud platforms (AWS / GCP / Azure) and distributed computing frameworks (Spark).
GOOD TO HAVE
• Prior experience with clinical trial data standards (CDISC, CDASH, SDTM) or healthcare ontologies (SNOMED, ICD-10).
• Familiarity with Trial2Vec or similar trial-to-vector embedding approaches.
• Experience with LLM fine-tuning, RAG pipelines, or prompt engineering in a production setting.
• Knowledge of regulatory and compliance considerations in healthcare AI (e.g., FDA guidelines, HIPAA).
• Contributions to open-source ML projects or published research.
THIS ROLE IS NOT FOR YOU IF…
• You have strong SQL/BI skills but limited hands-on ML modelling experience — or you’ve built models only in notebooks without ever deploying them to production.
• Your LLM exposure is limited to API calls and prompt engineering — with no experience fine-tuning models, working with embeddings, or building vector search pipelines.
About the Role
We are looking for a highly skilled Data Scientist with strong expertise in Machine Learning, MLOps, and Generative AI. The ideal candidate will have hands-on experience in building scalable ML models, deploying them in production, and working with modern AI frameworks, including GenAI technologies.
Key Responsibilities
· Design, develop, and deploy machine learning models for real-world business problems
· Work on end-to-end ML lifecycle: data preprocessing, model building, evaluation, deployment, and monitoring
· Implement and manage MLOps pipelines for scalable and reproducible workflows
· Utilize tools like MLflow for experiment tracking, model versioning, and lifecycle management
· Develop and integrate Generative AI (GenAI) solutions such as LLM-based applications
· Collaborate with cross-functional teams (engineering, product, business) to translate requirements into AI solutions
· Optimize model performance and ensure production stability
· Stay updated with the latest advancements in AI/ML and GenAI ecosystems
Required Skills & Qualifications
· 4+ years of experience in Data Science / Machine Learning
· Strong programming skills in Python
· Hands-on experience with ML modeling techniques (supervised, unsupervised, NLP, etc.)
· Solid understanding of MLOps practices and tools
· Experience with MLflow or similar model lifecycle tools
· Practical experience in Generative AI (GenAI), including working with LLMs
· Experience with libraries/frameworks like Scikit-learn, TensorFlow, PyTorch
· Strong understanding of data structures, algorithms, and statistics
· Experience with cloud platforms (AWS/GCP/Azure) is a plus
Good to Have
· Experience with LLM fine-tuning, prompt engineering, or RAG pipelines
· Exposure to Docker, Kubernetes, and CI/CD pipelines
· Knowledge of data engineering workflows
We are looking for a talented and driven Data Scientist to join our growing Analytics team in India. In this role, you will work at the intersection of advanced machine learning, scalable MLOps infrastructure, and domain-specific healthcare analytics. You will collaborate closely with cross-functional teams to build, deploy, and maintain production-grade ML models that drive real-world impact in clinical trials and healthcare operations.
KEY RESPONSIBILITIES
End-to-End ML Development
• Design, build, and optimize predictive models across the full ML lifecycle—from data ingestion to model serving.
• Conduct rigorous Exploratory Data Analysis (EDA) to surface insights and drive feature engineering decisions.
• Validate model performance using appropriate statistical techniques and domain knowledge.
MLOps & Production Deployment
• Deploy, monitor, and maintain production-grade ML models using Databricks MLFlow endpoints and Unity Catalog.
• Implement CI/CD pipelines for model versioning, experiment tracking, and automated retraining.
• Ensure model reliability, observability, and performance in live production environments.
Language Models & LLM Applications
• Apply transformer-based models (BERT, ClinicalBERT, Trial2Vec) for NLP tasks including classification, NER, and information extraction.
• Build and maintain vector similarity search pipelines for semantic retrieval and recommendation use cases.
• Fine-tune pre-trained models for domain-specific applications in clinical and healthcare contexts.
• Support exploratory work around LLM integration and prompt engineering for internal tooling.
Domain-Driven Analytics
• Apply advanced analytics within complex healthcare and clinical trial datasets—including patient records, trial protocols, and adverse event data.
• Translate ambiguous business problems into structured analytical frameworks with measurable outcomes.
• Partner with domain experts, product managers, and engineering teams to deliver data-driven solutions.
REQUIRED QUALIFICATIONS
Education
• Bachelor’s or Master’s degree in Computer Science, Statistics, Mathematics, Bioinformatics, or a closely related field.
Experience
• 2–4 years of hands-on experience in a data science or machine learning role.
• Demonstrable experience deploying ML models in production environments (not just prototyping).
Technical Skills
• Strong proficiency in Python (pandas, NumPy, scikit-learn, PyTorch / TensorFlow).
• Experience with Databricks, MLFlow (experiment tracking, model registry, endpoints), and Unity Catalog.
• Hands-on experience with BERT-family models and Hugging Face Transformers library.
• Familiarity with vector databases (e.g., FAISS, Pinecone, Weaviate) and embedding-based retrieval.
• Solid understanding of SQL and working with large structured/unstructured datasets.
• Exposure to cloud platforms (AWS / GCP / Azure) and distributed computing frameworks (Spark).
GOOD TO HAVE
• Prior experience with clinical trial data standards (CDISC, CDASH, SDTM) or healthcare ontologies (SNOMED, ICD-10).
• Familiarity with Trial2Vec or similar trial-to-vector embedding approaches.
• Experience with LLM fine-tuning, RAG pipelines, or prompt engineering in a production setting.
• Knowledge of regulatory and compliance considerations in healthcare AI (e.g., FDA guidelines, HIPAA).
• Contributions to open-source ML projects or published research.
THIS ROLE IS NOT FOR YOU IF…
• You have strong SQL/BI skills but limited hands-on ML modelling experience — or you’ve built models only in notebooks without ever deploying them to production.
• Your LLM exposure is limited to API calls and prompt engineering — with no experience fine-tuning models, working with embeddings, or building vector search pipelines.





