Data Scientist role @ Zolvit at Zolvit (formerly Vakilsearch) · Bengaluru (Bangalore) · 1 - 5 years · ₹10L - ₹20L / yr · Profitable · Posted 19 Jun 2025
About the Role
We are seeking an innovative Data Scientist specializing in Natural Language Processing (NLP) to join our technology team in Bangalore. The ideal candidate will harness the power of language models and document extraction techniques to transform legal information into accessible, actionable insights for our clients.
Responsibilities
- Develop and implement NLP solutions to automate legal document analysis and extraction
- Create and optimize prompt engineering strategies for large language models
- Design search functionality leveraging semantic understanding of legal documents
- Build document extraction pipelines to process unstructured legal text data
- Develop data visualizations using PowerBI and Tableau to communicate insights
- Collaborate with product and legal teams to enhance our tech-enabled services
- Continuously improve model performance and user experience.
Requirements
- Bachelor's degree in relevant field
- 1-5 years of professional experience in data science, with focus on NLP applications
- Demonstrated experience working with LLM APIs (e.g., OpenAI, Anthropic, )
- Proficiency in prompt engineering and optimization techniques
- Experience with document extraction and information retrieval systems
- Strong skills in data visualization tools, particularly PowerBI and Tableau
- Excellent programming skills in Python and familiarity with NLP libraries
- Strong understanding of legal terminology and document structures (preferred)
- Excellent communication skills in English
What We Offer
- Competitive salary and benefits package
- Opportunity to work at India's largest legal tech company
- Professional growth in the fast-evolving legal technology sector
- Collaborative work environment with industry experts
- Modern office located in Bangalore
- Flexible work arrangements
Qualified candidates are encouraged to apply with a resume highlighting relevant experience with NLP, prompt engineering, and data visualization tools.
Location: Bangalore, India

About Zolvit (formerly Vakilsearch)
About
India's Largest Online Platform For Legal, Tax and Compliance Services. - https://t.co/f4giVXnWD7
About Vakilsearch
Vakilsearch is India's largest online legal, tax and compliance provider. Vakilsearch, through its products and end-to-end workflow automation journey has revolutionized how Start-ups/ Small & Medium Enterprises register, seamlessly run and comply with Government regulations. On our mission to provide one-click access to individuals and businesses for all their legal and professional needs, we have helped over 4 Lac start-ups/ small and medium enterprises to date.
Visit us on www.vakilsearch.com
Vakilsearch in recent news: https://economictimes.indiatimes.com/tech/funding/incorp-india-invests-10-million-in-vakilsearch/articleshow/87352272.cms
Vakilsearch is a people-first organisation that thrives on the enthusiasm of our team to execute our mission to the satisfaction of our customers. Towards this end, we stress on creating an optimal work-life balance and inculcating a strong sense of team spirit that stems from enthusiasm and good vibes. When you work at Vakilsearch, you don't just become an employee, you become family, and we always coalesce around each other to ensure a strong sense of family.
Company video


Connect with the team
Similar jobs (10)
Roles & Responsibilities
- Design and develop intelligent AI-based applications using advanced NLP and LLM techniques to solve real-world business challenges in financial services.
- Build and optimize Retrieval-Augmented Generation (RAG) pipelines leveraging structured and unstructured financial data.
- Integrate and orchestrate LLMs/SLMs for question-answering, summarization, semantic search, and document understanding.
- Develop and maintain RESTful APIs (sync and async) to serve NLP models and chatbot interfaces using frameworks like FastAPI, Flask, etc.
- Should have knowledge of advanced prompting techniques.
- Implement semantic search, hybrid search, and text retrieval systems using Elasticsearch and vector databases (e.g., FAISS, Pinecone, Weaviate).
- Perform NLP tasks such as entity recognition, text classification, intent detection, embedding generation, and sentiment analysis where required.
- Monitor and fine-tune LLM/SLM performance with real-world user data to improve relevance, latency, and accuracy.
- Exposure to LLMOps tools for monitoring, evaluation, and versioning of AI models in production.
- Build, train, and evaluate deep learning models for NLP tasks including classification, NER, summarization, and embedding generation.
- Develop traditional machine learning models (e.g., regression, decision trees, clustering) for structured data analysis and prediction tasks.
- Interact with cross-functional teams to understand system issues and follow up with respective teams to get them fixed.
- Understand and identify areas of improvement across businesses and participate in solution identification and implementation.
- Should be able to work as an Individual Contributor on new and existing projects.
- Positive and problem-solving attitude, must work as an independent contributor.
Ideal Candidate
1.Strong Data Scientist / AI Engineer / Generative AI Engineer profile.
2.Mandatory (Experience 1) - Must have 3+ years of hands-on experience in Data Science, Artificial Intelligence, Machine Learning, Deep Learning, NLP, or Generative AI application development.
3.Mandatory (Experience 2) - Must have strong hands-on experience in Python programming, backend development, API development, and production-grade application support
4.Mandatory (Experience 3) - Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, or Scikit-learn.
5.Mandatory (Experience 4) - Must have hands-on experience in NLP use cases such as text classification, sentiment analysis, entity recognition (NER), semantic search, embeddings, or document understanding.
6.Mandatory (Experience 5) - Must have experience working with Large Language Models (LLMs) such as GPT, LLaMA, Mistral, Phi, Claude, Gemini, or similar models
.7.Mandatory (Experience 6) - Must have hands-on experience building or implementing Retrieval Augmented Generation (RAG) solutions, vector search, semantic search, or knowledge-based AI applications.
8.Mandatory (Experience 7) - Must have experience with Prompt Engineering and Generative AI frameworks such as LangChain, LangGraph, AI Agents, Azure OpenAI, or similar technologies.
9.Mandatory (Experience 8) - Must have experience developing, consuming, or integrating APIs using Python frameworks such as FastAPI, Flask, or similar technologies.
10.Mandatory (CTC) - The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
11.Preferred (Experience 1) - Experience with LLMOps/MLOps tools for monitoring, evaluation, experimentation, and versioning of AI models.
12.Preferred (Experience 2) - Exposure to Azure OpenAI, Azure Kubernetes Service (AKS), Kubernetes, cloud-native AI deployments, or distributed systems.
13.Preferred (Experience 3) - Experience working with PostgreSQL, MongoDB, Redis, Kafka, or large-scale data platforms.
14.Preferred (Experience 4) - Familiarity with Docker, Kubernetes, cloud platforms, and scalable deployment architecture.
15.Preferred (Company) - Candidates from AI-first startups, product companies, SaaS organizations, fintech, or data-driven technology companies.
16.Mandatory ( Age ) - Candidate Should be Below 28 Years.
17.Mandatory ( Pedigree) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered.
Role Overview
As a Data Scientist, you will work with business stakeholders, AI engineers, and domain experts to transform data into actionable insights and intelligent solutions. You will develop machine learning models, perform statistical analysis, and contribute to AI-driven products that create measurable business impact.
Key Responsibilities
Data Science & Machine Learning
- Analyze structured and unstructured data to identify patterns, trends, and business opportunities.
- Perform exploratory data analysis (EDA), feature engineering, and data preparation.
- Develop, evaluate, and optimize machine learning models for prediction, classification, clustering, and forecasting.
- Apply statistical techniques to solve business problems and validate model performance.
- Design and execute experiments to improve model accuracy and business outcomes.
AI Solution Development
- Collaborate with AI Engineers, Data Engineers, and domain experts to build AI-powered solutions.
- Translate business requirements into scalable data science approaches.
- Contribute to Generative AI and advanced analytics initiatives where applicable.
- Document methodologies, model performance, and key findings.
Required Technical Skills
- Strong programming skills in Python and SQL for data analysis, feature engineering, and machine learning.
- Strong understanding of Statistics, Probability, Linear Algebra, and Calculus as applied to machine learning and data science.
- Experience with Exploratory Data Analysis (EDA), data preprocessing, feature engineering, feature selection, and handling missing or imbalanced data.
- Good understanding of Supervised, Unsupervised, and Ensemble Machine Learning algorithms, including their assumptions, strengths, limitations, and appropriate use cases.
- Strong knowledge of Regression, Classification, Clustering, Time Series Forecasting, Dimensionality Reduction, Recommendation Systems, and Anomaly Detection techniques.
- Experience with Model Evaluation, Cross-Validation, Hyperparameter Optimization, Bias-Variance Trade-off, Feature Importance, Explainable AI (XAI), and Performance Metrics.
- Understanding of Statistical Inference, Hypothesis Testing, Probability Distributions, Sampling Techniques, Confidence Intervals, and A/B Testing.
- Experience translating business problems into analytical approaches and developing scalable, data-driven solutions.
- Working knowledge of Generative AI, Large Language Models (LLMs), Prompt Engineering, and Retrieval-Augmented Generation (RAG) is preferred.
Preferred Qualifications
- Bachelor's or master's degree in computer science, Artificial Intelligence, Data Science, Statistics, Mathematics, Engineering, or a related field.
- 2–4 years of experience developing machine learning or data science solutions.
- Experience working on end-to-end data science projects in a business environment.
Nice to Have
- Exposure to Generative AI, LLMs, RAG, or Agentic AI.
- Experience with Computer Vision or Natural Language Processing (NLP).
- Familiarity with cloud-based AI platforms.
- Knowledge of construction, engineering, manufacturing, or industrial domains.
- Participation in hackathons, research, Kaggle competitions, or open-source projects.
Soft Skills
Strong analytical and problem-solving skills, effective communication and collaboration, ownership mindset, adaptability, continuous learning, and a passion for innovation.
Role & Responsibilities
Responsibilities
• Contribute to the development and optimization of enterprise-wide search systems and models.
• Design and implement algorithms to improve indexing, query relevance, and search accuracy.
• Support taxonomy, ontology, and metadata model creation for better search outcomes.
• Collaborate with business units (Loans, Insurance, Investments) to build AI-enabled search features.
• Conduct analysis of user behavior and system metrics to refine search performance.
• Work with engineers, product managers, and designers to deliver integrated search solutions.
• Develop production-grade ML systems for ranking, personalization, and recommendations.
• Participate in proof-of-concept initiatives with internal and external partners.
• Follow best practices in software engineering including CI/CD, testing, and monitoring.
• Keep abreast of emerging developments in AI/ML to apply them in practical solutions.
Ideal Candidate
Strong Data Scientist / AI Engineer / Machine Learning Engineer profiles.
Mandatory (Experience 1) – Must have minimum 5+ years of hands-on experience in Data Science, Machine Learning, Applied AI, NLP, Deep Learning, or Generative AI solutions.
Mandatory (Experience 2) – Must have strong hands-on experience in Python programming, SQL, data analysis, feature engineering, model development, and production-grade ML applications.
Mandatory (Experience 3) – Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, Scikit-learn, or equivalent.
Mandatory (Experience 4) – Must have hands-on experience working on NLP, embeddings, semantic search, text classification, document understanding, recommendation systems, or similar AI/ML use cases.
Mandatory (Experience 5) – Must have experience working with Large Language Models (LLMs) such as GPT, Llama, Mistral, Claude, Gemini, Phi, or similar foundation models.
Mandatory (Experience 6) – Must have hands-on experience building or implementing RAG (Retrieval Augmented Generation) systems, vector search, knowledge retrieval, embeddings, chunking, indexing, or semantic retrieval solutions.
Mandatory (Experience 7) – Must have experience working with Git, CI/CD practices, production environments, and scalable AI/ML systems.
Mandatory (CTC) – The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
Mandatory (Age) - Candidate's Age should be below 30 Years
Preferred (Experience 1) – Experience with MLFlow, Kubeflow, Airflow, Prefect, Feature Stores, Model Registry, or MLOps/LLMOps frameworks.
Preferred (Experience 2) – Experience working with Vector Databases, Spark, PySpark, distributed ML pipelines, large-scale data processing, or real-time ML systems..
Preferred (Experience 3) – Familiarity with Docker, Kubernetes, Azure, AWS, GCP, cloud-native AI deployments, and scalable ML architecture.
Preferred (Company) – Candidates from AI-first startups, Fintech, Banking, Lending, Fraud Analytics, Risk Analytics, Product Companies, SaaS organizations, or data-driven technology companies.
Kindly provide the following details while sending your CV: (Mandatory details)
1) Date of Birth
2) Current Location-
3) Current CTC-
4) Expected CTC-
5) Notice Period-
6) Ready to relocate to Pune?
Regards,
The Supreme Consultancy
Website- https://lnkd.in/eawfxfxU
Strong Data Scientist / AI Engineer / Generative AI Engineer profile.
2
Mandatory (Experience 1) - Must have 3+ years of hands-on experience in Data Science, Artificial Intelligence, Machine Learning, Deep Learning, NLP, or Generative AI application development.
3
Mandatory (Experience 2) - Must have strong hands-on experience in Python programming, backend development, API development, and production-grade application support.
4
Mandatory (Experience 3) - Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, or Scikit-learn.
5
Mandatory (Experience 4) - Must have hands-on experience in NLP use cases such as text classification, sentiment analysis, entity recognition (NER), semantic search, embeddings, or document understanding.
6
Mandatory (Experience 5) - Must have experience working with Large Language Models (LLMs) such as GPT, LLaMA, Mistral, Phi, Claude, Gemini, or similar models.
7
Mandatory (Experience 6) - Must have hands-on experience building or implementing Retrieval Augmented Generation (RAG) solutions, vector search, semantic search, or knowledge-based AI applications.
8
Mandatory (Experience 7) - Must have experience with Prompt Engineering and Generative AI frameworks such as LangChain, LangGraph, AI Agents, Azure OpenAI, or similar technologies.
9
Mandatory (Experience 8) - Must have experience developing, consuming, or integrating APIs using Python frameworks such as FastAPI, Flask, or similar technologies.
10
Mandatory (CTC) - The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
11
Preferred (Experience 1) - Experience with LLMOps/MLOps tools for monitoring, evaluation, experimentation, and versioning of AI models.
12
Preferred (Experience 2) - Exposure to Azure OpenAI, Azure Kubernetes Service (AKS), Kubernetes, cloud-native AI deployments, or distributed systems.
13
Preferred (Experience 3) - Experience working with PostgreSQL, MongoDB, Redis, Kafka, or large-scale data platforms.
14
Preferred (Experience 4) - Familiarity with Docker, Kubernetes, cloud platforms, and scalable deployment architecture.
15
Preferred (Company) - Candidates from AI-first startups, product companies, SaaS organizations, fintech, or data-driven technology companies.
16
Mandatory ( Age ) - Candidate Should be Below 28 Years.
17
Mandatory ( Pedigree) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered.
Strong Data Scientist / AI Engineer / Generative AI Engineer profile.
2
Mandatory (Experience 1) - Must have 3+ years of hands-on experience in Data Science, Artificial Intelligence, Machine Learning, Deep Learning, NLP, or Generative AI application development.
3
Mandatory (Experience 2) - Must have strong hands-on experience in Python programming, backend development, API development, and production-grade application support.
4
Mandatory (Experience 3) - Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, or Scikit-learn.
5
Mandatory (Experience 4) - Must have hands-on experience in NLP use cases such as text classification, sentiment analysis, entity recognition (NER), semantic search, embeddings, or document understanding.
6
Mandatory (Experience 5) - Must have experience working with Large Language Models (LLMs) such as GPT, LLaMA, Mistral, Phi, Claude, Gemini, or similar models.
7
Mandatory (Experience 6) - Must have hands-on experience building or implementing Retrieval Augmented Generation (RAG) solutions, vector search, semantic search, or knowledge-based AI applications.
8
Mandatory (Experience 7) - Must have experience with Prompt Engineering and Generative AI frameworks such as LangChain, LangGraph, AI Agents, Azure OpenAI, or similar technologies.
9
Mandatory (Experience 8) - Must have experience developing, consuming, or integrating APIs using Python frameworks such as FastAPI, Flask, or similar technologies.
10
Mandatory (CTC) - The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
11
Preferred (Experience 1) - Experience with LLMOps/MLOps tools for monitoring, evaluation, experimentation, and versioning of AI models.
12
Preferred (Experience 2) - Exposure to Azure OpenAI, Azure Kubernetes Service (AKS), Kubernetes, cloud-native AI deployments, or distributed systems.
13
Preferred (Experience 3) - Experience working with PostgreSQL, MongoDB, Redis, Kafka, or large-scale data platforms.
14
Preferred (Experience 4) - Familiarity with Docker, Kubernetes, cloud platforms, and scalable deployment architecture.
15
Preferred (Company) - Candidates from AI-first startups, product companies, SaaS organizations, fintech, or data-driven technology companies.
16
Mandatory ( Age ) - Candidate Should be Below 28 Years.
17
Mandatory ( Pedigree) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered.
Company Name: WINIT
Location: Hyderabad
Experience: 0–2 Years
Role Overview
We are looking for a talented AI/ML Trainee to build intelligent automation solutions for Legal Process Outsourcing (LPO). The role involves leveraging Python, SQL, AWS, and Generative AI (LangChain, LLMs, RAG) to automate document-heavy workflows such as contract analysis, data extraction, and legal document processing.
Key Responsibilities
- Develop and deploy AI/ML models for legal document processing and automation.
- Build Python-based pipelines for data processing, automation, and model integration.
- Design and optimize SQL queries for efficient data management and reporting.
- Work with AWS services (S3, Lambda, EC2, Bedrock, etc.) for scalable deployments.
- Develop Generative AI solutions using:
- LangChain
- Large Language Models (LLMs)
- Retrieval-Augmented Generation (RAG)
- Implement document intelligence solutions (classification, summarization, entity extraction).
- Integrate AI models into business applications via APIs.
- Collaborate with cross-functional teams to understand automation needs.
- Ensure model performance, scalability, and data security.
Required Skills
- Strong proficiency in Python
- Good knowledge of SQL.
- Understanding of machine learning fundamentals.
- Hands-on or conceptual knowledge of:
- LangChain
- LLMs
- RAG architectures
- Familiarity with AWS cloud platform.
- Basic understanding of NLP / text processing.
About WINIT
For more information, please visit: www.winitsoftware.com
About the Role:
We are looking for an ideal candidate with 5+ years of experience in Data Science / Machine Learning, with strong hands-on experience in Generative AI, Large Language Models (LLMs), NLP, and AI-powered applications. The candidate should be comfortable working across the complete AI lifecycle—from understanding business requirements and experimenting with models to building, evaluating, deploying, and monitoring production-grade GenAI solutions.
The role requires a combination of strong technical expertise, business understanding, problem-solving ability, and stakeholder management skills.
Key Responsibilities:
Generative AI & LLM
· Design, develop, and deploy Generative AI and LLM-based solutions for enterprise use cases.
· Work with models such as OpenAI, Azure OpenAI, Llama, Mistral, Gemini, or equivalent LLM platforms.
· Develop applications using prompt engineering, structured outputs, function/tool calling, and LLM orchestration.
· Design and implement Retrieval-Augmented Generation (RAG) solutions.
· Work with vector databases and semantic search for enterprise knowledge retrieval.
· Develop and evaluate AI agents and multi-step AI workflows.
· Implement techniques such as prompt optimization, context management, grounding, and hallucination reduction.
· Develop AI solutions for text classification, summarization, information extraction, question answering, document intelligence, and other enterprise use cases.
Machine Learning & Data Science
· Develop and optimize traditional Machine Learning and statistical models where appropriate.
· Perform data exploration, feature engineering, model selection, training, validation, and evaluation.
· Apply appropriate ML and statistical techniques to solve business problems.
· Work with structured, unstructured, and semi-structured data.
· Develop scalable data pipelines to support AI/ML solutions.
· Collaborate with Data Engineers to prepare and manage data for AI applications.
AI Evaluation & Productionization
· Design evaluation frameworks to measure LLM accuracy, relevance, groundedness, toxicity, latency, and cost.
· Implement guardrails and responsible AI practices.
· Monitor model and application performance in production.
· Identify model/data drift and implement appropriate improvement strategies.
· Optimize AI solutions for performance, scalability, reliability, and cost.
· Support deployment and productionization of AI/ML solutions.
· Client & Delivery Responsibilities
· Work closely with the CEO, Delivery team, Solution Architects, Engineering teams, and clients to understand business problems and identify AI opportunities.
· Translate business requirements into practical AI/ML solutions.
· Participate in client discussions, solution presentations, technical workshops, and POCs.
· Develop rapid prototypes and demonstrate the feasibility of GenAI solutions.
· Convert successful POCs into scalable, production-ready applications.
· Provide technical guidance and contribute to AI solution architecture.
· Prepare technical documentation, solution approaches, and project estimates where required.
· Stay current with developments in Generative AI, LLMs, Agentic AI, and AI engineering.
Required Skills:
· 5+ years of hands-on experience in Data Science, Machine Learning, AI, or a related field.
· Strong practical experience in Generative AI and LLM-based applications.
· Strong proficiency in Python.
· Strong understanding of Machine Learning and statistical concepts.
· Hands-on experience with:
o LLMs
o Prompt Engineering
o RAG
o Vector Databases
o Embeddings
o Semantic Search
o LLM Evaluation
o AI Guardrails
· Experience with frameworks/tools such as LangChain, LangGraph, LlamaIndex, or equivalent.
· Experience with APIs and integrating LLMs into enterprise applications.
· Strong SQL and data handling skills.
· Experience working with large and complex datasets.
· Strong understanding of NLP concepts.XX
Technical Skills:
· Experience with OpenAI / Azure OpenAI / AWS Bedrock / Google Vertex AI.
· Experience with vector databases such as Pinecone, Weaviate, Milvus, FAISS, or equivalent.
· Experience with Databricks, Snowflake, or cloud data platforms.
· Experience with Docker and CI/CD.
· Exposure to AWS, Azure, or GCP.
· Experience with ML/AI deployment and MLOps.
· Knowledge of AI security, data privacy, governance, and responsible AI.
· Experience building AI Agents / Agentic AI workflows.
· Experience with multimodal AI is an added advantage
Key Competencies
· Strong analytical and problem-solving ability.
· Ability to translate business problems into practical AI solutions.
· Strong communication and presentation skills.
· Ability to interact confidently with senior stakeholders and clients.
· Strong ownership and delivery mindset.
· Ability to work independently in a fast-paced environment.
- Strong experimentation and innovation mindset.
- Ability to balance technical feasibility, business value, scalability, and cost.
Required Education & Experience:
· Bachelor's or Master's degree in Computer Science, Data Science, Artificial Intelligence, Statistics, Mathematics, Engineering, or a related discipline
Strong Data Scientist / AI Engineer / Generative AI Engineer profile.
2
Mandatory (Experience 1) - Must have 3+ years of hands-on experience in Data Science, Artificial Intelligence, Machine Learning, Deep Learning, NLP, or Generative AI application development.
3
Mandatory (Experience 2) - Must have strong hands-on experience in Python programming, backend development, API development, and production-grade application support.
4
Mandatory (Experience 3) - Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, or Scikit-learn.
5
Mandatory (Experience 4) - Must have hands-on experience in NLP use cases such as text classification, sentiment analysis, entity recognition (NER), semantic search, embeddings, or document understanding.
6
Mandatory (Experience 5) - Must have experience working with Large Language Models (LLMs) such as GPT, LLaMA, Mistral, Phi, Claude, Gemini, or similar models.
7
Mandatory (Experience 6) - Must have hands-on experience building or implementing Retrieval Augmented Generation (RAG) solutions, vector search, semantic search, or knowledge-based AI applications.
8
Mandatory (Experience 7) - Must have experience with Prompt Engineering and Generative AI frameworks such as LangChain, LangGraph, AI Agents, Azure OpenAI, or similar technologies.
9
Mandatory (Experience 8) - Must have experience developing, consuming, or integrating APIs using Python frameworks such as FastAPI, Flask, or similar technologies.
10
Mandatory (CTC) - The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
11
Preferred (Experience 1) - Experience with LLMOps/MLOps tools for monitoring, evaluation, experimentation, and versioning of AI models.
12
Preferred (Experience 2) - Exposure to Azure OpenAI, Azure Kubernetes Service (AKS), Kubernetes, cloud-native AI deployments, or distributed systems.
We are looking for a talented and driven Data Scientist to join our growing Analytics team in India. In this role, you will work at the intersection of advanced machine learning, scalable MLOps infrastructure, and domain-specific healthcare analytics. You will collaborate closely with cross-functional teams to build, deploy, and maintain production-grade ML models that drive real-world impact in clinical trials and healthcare operations.
KEY RESPONSIBILITIES
End-to-End ML Development
• Design, build, and optimize predictive models across the full ML lifecycle—from data ingestion to model serving.
• Conduct rigorous Exploratory Data Analysis (EDA) to surface insights and drive feature engineering decisions.
• Validate model performance using appropriate statistical techniques and domain knowledge.
MLOps & Production Deployment
• Deploy, monitor, and maintain production-grade ML models using Databricks MLFlow endpoints and Unity Catalog.
• Implement CI/CD pipelines for model versioning, experiment tracking, and automated retraining.
• Ensure model reliability, observability, and performance in live production environments.
Language Models & LLM Applications
• Apply transformer-based models (BERT, ClinicalBERT, Trial2Vec) for NLP tasks including classification, NER, and information extraction.
• Build and maintain vector similarity search pipelines for semantic retrieval and recommendation use cases.
• Fine-tune pre-trained models for domain-specific applications in clinical and healthcare contexts.
• Support exploratory work around LLM integration and prompt engineering for internal tooling.
Domain-Driven Analytics
• Apply advanced analytics within complex healthcare and clinical trial datasets—including patient records, trial protocols, and adverse event data.
• Translate ambiguous business problems into structured analytical frameworks with measurable outcomes.
• Partner with domain experts, product managers, and engineering teams to deliver data-driven solutions.
REQUIRED QUALIFICATIONS
Education
• Bachelor’s or Master’s degree in Computer Science, Statistics, Mathematics, Bioinformatics, or a closely related field.
Experience
• 2–4 years of hands-on experience in a data science or machine learning role.
• Demonstrable experience deploying ML models in production environments (not just prototyping).
Technical Skills
• Strong proficiency in Python (pandas, NumPy, scikit-learn, PyTorch / TensorFlow).
• Experience with Databricks, MLFlow (experiment tracking, model registry, endpoints), and Unity Catalog.
• Hands-on experience with BERT-family models and Hugging Face Transformers library.
• Familiarity with vector databases (e.g., FAISS, Pinecone, Weaviate) and embedding-based retrieval.
• Solid understanding of SQL and working with large structured/unstructured datasets.
• Exposure to cloud platforms (AWS / GCP / Azure) and distributed computing frameworks (Spark).
GOOD TO HAVE
• Prior experience with clinical trial data standards (CDISC, CDASH, SDTM) or healthcare ontologies (SNOMED, ICD-10).
• Familiarity with Trial2Vec or similar trial-to-vector embedding approaches.
• Experience with LLM fine-tuning, RAG pipelines, or prompt engineering in a production setting.
• Knowledge of regulatory and compliance considerations in healthcare AI (e.g., FDA guidelines, HIPAA).
• Contributions to open-source ML projects or published research.
THIS ROLE IS NOT FOR YOU IF…
• You have strong SQL/BI skills but limited hands-on ML modelling experience — or you’ve built models only in notebooks without ever deploying them to production.
• Your LLM exposure is limited to API calls and prompt engineering — with no experience fine-tuning models, working with embeddings, or building vector search pipelines.
About the Role
We are looking for enthusiastic LLM Interns to join our team remotely for a 3-month internship. This role is ideal for students or graduates interested in AI, Natural Language Processing (NLP), and Large Language Models (LLMs). You will gain hands-on experience working with cutting-edge AI tools, prompt engineering, and model fine-tuning. While this is an unpaid internship, interns who successfully complete the program will receive a Completion Certificate and a Letter of Recommendation.
Responsibilities
- Research and experiment with LLMs, NLP techniques, and AI frameworks.
- Design, test, and optimize prompts and workflows for different use cases.
- Assist in fine-tuning or integrating LLMs for internal projects.
- Evaluate model outputs and improve accuracy, efficiency, and reliability.
- Collaborate with developers, data scientists, and product managers to implement AI-driven features.
- Document experiments, results, and best practices.
Requirements
- Strong interest in Artificial Intelligence, NLP, and Machine Learning.
- Familiarity with Python and ML libraries (e.g., TensorFlow, PyTorch, Hugging Face Transformers).
- Basic understanding of LLM concepts such as embeddings, fine-tuning, and inference.
- Knowledge of APIs (OpenAI, Anthropic, Hugging Face, etc.) is a plus.
- Good analytical and problem-solving skills.
- Ability to work independently in a remote environment.
What You’ll Gain
- Practical exposure to state-of-the-art AI tools and LLMs.
- Mentorship from AI and software professionals.
- Completion Certificate upon successful completion.
- Letter of Recommendation based on performance.
- Experience to showcase in research projects, academic work, or future AI roles.
Internship Details
- Duration: 3 months
- Location: Remote (Work from Home)
- Stipend: Unpaid
- Perks: Completion Certificate + Letter of Recommendation






