Data Scientist or Senior Machine Learning Engineer at Generative AI Persona platform · Pune · 6 - 7 years · ₹15L - ₹20L / yr · Posted 6 Mar 2026

Data Scientist or Senior Machine Learning Engineer
at Generative AI Persona platform
Description
We are currently hiring for the position of Data Scientist/ Senior Machine Learning Engineer (6–7 years’ experience).
Please find the detailed Job Description attached for your reference. We are looking for candidates with strong experience in:
- Machine Learning model development
- Scalable data pipeline development (ETL/ELT)
- Python and SQL
- Cloud platforms such as Azure/AWS/Databricks
- ML deployment environments (SageMaker, Azure ML, etc.)
Kindly note:
- Location: Pune (Work From Office)
- Immediate joiners preferred
While sharing profiles, please ensure the following details are included:
- Current CTC
- Expected CTC
- Notice Period
- Current Location
- Confirmation on Pune WFO comfort
Must have skills
Machine Learning - 6 years
Python - 6 years
ETL(Extract, Transform, Load) - 6 years
SQL - 6 years
Azure - 6 years

Similar jobs (10)
Description
We’re seeking a highly skilled, execution-focused Senior Data Scientist with a minimum of 5 years of experience. This role demands hands-on expertise in building, deploying, and optimizing machine learning models at scale, while working with big data technologies and modern cloud platforms. You will be responsible for driving data-driven solutions from experimentation to production, leveraging advanced tools and frameworks across Python, SQL, Spark, and AWS. The role requires strong technical depth, problem-solving ability, and ownership in delivering business impact through data science.
Responsibilities
- Design, build, and deploy scalable machine learning models into production systems.
- Develop advanced analytics and predictive models using Python, SQL, and popular ML/DL frameworks (Pandas, Scikit-learn, TensorFlow, PyTorch).
- Leverage Databricks, Apache Spark, and Hadoop for large-scale data processing and model training.
- Implement workflows and pipelines using Airflow and AWS EMR for automation and orchestration.
- Collaborate with engineering teams to integrate models into cloud-based applications on AWS.
- Optimize query performance, storage usage, and data pipelines for efficiency.
- Conduct end-to-end experiments, including data preprocessing, feature engineering, model training, validation, and deployment.
- Drive initiatives independently with high ownership and accountability.
- Stay up to date with industry best practices in machine learning, big data, and cloud-native deployments.
Requirements
- Minimum 5 years of experience in Data Science or Applied Machine Learning.
- Strong proficiency in Python, SQL, and ML libraries (Pandas, Scikit-learn, TensorFlow, PyTorch).
- Proven expertise in deploying ML models into production systems.
- Experience with big data platforms (Hadoop, Spark) and distributed data processing.
- Hands-on experience with Databricks, Airflow, and AWS EMR.
- Strong knowledge of AWS cloud services (S3, Lambda, SageMaker, EC2, etc.).
- Solid understanding of query optimization, storage systems, and data pipelines.
- Excellent problem-solving skills, with the ability to design scalable solutions.
- Strong communication and collaboration skills to work in cross-functional teams.
Benefits
- Best-in-class salary: We hire strong talent and compensate accordingly.
- Proximity Talks: Meet and learn from designers, engineers, product leaders, and AI practitioners.
- Continuous learning: Work with a world-class team and stay close to the latest in AI, engineering, and product development.
- High-impact work: Build AI-first systems and products used at scale by global clients.
About Us
Proximity is the trusted technology, design, and consulting partner for some of the biggest Sports, Media, and Entertainment companies in the world. We’re headquartered in San Francisco and have offices in Palo Alto, Dubai, Mumbai, and Bangalore.
Since 2019, Proximity has built high-impact, scalable products used by millions of users every day. Today, we are a global team of engineers, designers, product managers, and experts solving complex problems and building cutting-edge technology at scale.
Hiring for Data Scientist / Senior Data Scientist
Exp : 4 - 12 yrs
Edu : BE/B.tech/MCA
Work Location : Pune
Notice Period : Immediate - 15 days
Skills :
4+ years of experience in data engineering, data science, or related domains.
Hands-on experience with SQL, Python, and distributed data systems.
Knowledge of machine learning techniques and statistical analysis.
Experience with cloud data platforms (Azure Data Factory, AWS Glue, GCP BigQuery).
Familiarity with DevOps practices and CI/CD for data pipelines.
Platforms & Operations Experience (Preferred)
- Experience working with Azure, AWS, or Google Cloud data tools.
Operational experience with data orchestration tools (Airflow, ADF, Glue).
Understanding of Kubernetes, Docker, or containerized environments.
Hands-on experience with data warehousing platforms (Snowflake, Redshift, BigQuery).
Experience in monitoring, logging, and alerting operations for data workflows.
About the Role
We are looking for a highly skilled Data Scientist with strong expertise in Machine Learning, MLOps, and Generative AI. The ideal candidate will have hands-on experience in building scalable ML models, deploying them in production, and working with modern AI frameworks, including GenAI technologies.
Key Responsibilities
· Design, develop, and deploy machine learning models for real-world business problems
· Work on end-to-end ML lifecycle: data preprocessing, model building, evaluation, deployment, and monitoring
· Implement and manage MLOps pipelines for scalable and reproducible workflows
· Utilize tools like MLflow for experiment tracking, model versioning, and lifecycle management
· Develop and integrate Generative AI (GenAI) solutions such as LLM-based applications
· Collaborate with cross-functional teams (engineering, product, business) to translate requirements into AI solutions
· Optimize model performance and ensure production stability
· Stay updated with the latest advancements in AI/ML and GenAI ecosystems
Required Skills & Qualifications
· 4+ years of experience in Data Science / Machine Learning
· Strong programming skills in Python
· Hands-on experience with ML modeling techniques (supervised, unsupervised, NLP, etc.)
· Solid understanding of MLOps practices and tools
· Experience with MLflow or similar model lifecycle tools
· Practical experience in Generative AI (GenAI), including working with LLMs
· Experience with libraries/frameworks like Scikit-learn, TensorFlow, PyTorch
· Strong understanding of data structures, algorithms, and statistics
· Experience with cloud platforms (AWS/GCP/Azure) is a plus
Good to Have
· Experience with LLM fine-tuning, prompt engineering, or RAG pipelines
· Exposure to Docker, Kubernetes, and CI/CD pipelines
· Knowledge of data engineering workflows
We are looking for a talented and driven Data Scientist to join our growing Analytics team in India. In this role, you will work at the intersection of advanced machine learning, scalable MLOps infrastructure, and domain-specific healthcare analytics. You will collaborate closely with cross-functional teams to build, deploy, and maintain production-grade ML models that drive real-world impact in clinical trials and healthcare operations.
KEY RESPONSIBILITIES
End-to-End ML Development
• Design, build, and optimize predictive models across the full ML lifecycle—from data ingestion to model serving.
• Conduct rigorous Exploratory Data Analysis (EDA) to surface insights and drive feature engineering decisions.
• Validate model performance using appropriate statistical techniques and domain knowledge.
MLOps & Production Deployment
• Deploy, monitor, and maintain production-grade ML models using Databricks MLFlow endpoints and Unity Catalog.
• Implement CI/CD pipelines for model versioning, experiment tracking, and automated retraining.
• Ensure model reliability, observability, and performance in live production environments.
Language Models & LLM Applications
• Apply transformer-based models (BERT, ClinicalBERT, Trial2Vec) for NLP tasks including classification, NER, and information extraction.
• Build and maintain vector similarity search pipelines for semantic retrieval and recommendation use cases.
• Fine-tune pre-trained models for domain-specific applications in clinical and healthcare contexts.
• Support exploratory work around LLM integration and prompt engineering for internal tooling.
Domain-Driven Analytics
• Apply advanced analytics within complex healthcare and clinical trial datasets—including patient records, trial protocols, and adverse event data.
• Translate ambiguous business problems into structured analytical frameworks with measurable outcomes.
• Partner with domain experts, product managers, and engineering teams to deliver data-driven solutions.
REQUIRED QUALIFICATIONS
Education
• Bachelor’s or Master’s degree in Computer Science, Statistics, Mathematics, Bioinformatics, or a closely related field.
Experience
• 2–4 years of hands-on experience in a data science or machine learning role.
• Demonstrable experience deploying ML models in production environments (not just prototyping).
Technical Skills
• Strong proficiency in Python (pandas, NumPy, scikit-learn, PyTorch / TensorFlow).
• Experience with Databricks, MLFlow (experiment tracking, model registry, endpoints), and Unity Catalog.
• Hands-on experience with BERT-family models and Hugging Face Transformers library.
• Familiarity with vector databases (e.g., FAISS, Pinecone, Weaviate) and embedding-based retrieval.
• Solid understanding of SQL and working with large structured/unstructured datasets.
• Exposure to cloud platforms (AWS / GCP / Azure) and distributed computing frameworks (Spark).
GOOD TO HAVE
• Prior experience with clinical trial data standards (CDISC, CDASH, SDTM) or healthcare ontologies (SNOMED, ICD-10).
• Familiarity with Trial2Vec or similar trial-to-vector embedding approaches.
• Experience with LLM fine-tuning, RAG pipelines, or prompt engineering in a production setting.
• Knowledge of regulatory and compliance considerations in healthcare AI (e.g., FDA guidelines, HIPAA).
• Contributions to open-source ML projects or published research.
THIS ROLE IS NOT FOR YOU IF…
• You have strong SQL/BI skills but limited hands-on ML modelling experience — or you’ve built models only in notebooks without ever deploying them to production.
• Your LLM exposure is limited to API calls and prompt engineering — with no experience fine-tuning models, working with embeddings, or building vector search pipelines.
Job Title : Data Engineer – Databricks
Experience : 6+ Years
Location : Noida / Hyderabad / Chennai / Pune / Bengaluru (Hybrid)
Shift : IST (Normal Shift)
Job Summary :
We are seeking an experienced Data Engineer with strong expertise in Databricks, Snowflake, Python, and Spark to build and optimize scalable data pipelines and support AI/ML model deployments. The ideal candidate should have experience working with cloud-based data platforms and preferably possess exposure to the Healthcare domain.
Required Skills :
- Databricks (Preferred)
- Snowflake
- Python
- Apache Spark
- SQL
- Azure Cloud
- Kubernetes
- Apache Airflow
- GitHub & CI/CD Pipelines
- AI/ML Model Deployment
- Data Analytics
Preferred :
- Experience in the Healthcare domain.
- Strong understanding of scalable data engineering architectures and best practices.
We are looking for a talented and driven Data Scientist to join our growing Analytics team in India. In this role, you will work at the intersection of advanced machine learning, scalable MLOps infrastructure, and domain-specific healthcare analytics. You will collaborate closely with cross-functional teams to build, deploy, and maintain production-grade ML models that drive real-world impact in clinical trials and healthcare operations.
KEY RESPONSIBILITIES
End-to-End ML Development
• Design, build, and optimize predictive models across the full ML lifecycle—from data ingestion to model serving.
• Conduct rigorous Exploratory Data Analysis (EDA) to surface insights and drive feature engineering decisions.
• Validate model performance using appropriate statistical techniques and domain knowledge.
MLOps & Production Deployment
• Deploy, monitor, and maintain production-grade ML models using Databricks MLFlow endpoints and Unity Catalog.
• Implement CI/CD pipelines for model versioning, experiment tracking, and automated retraining.
• Ensure model reliability, observability, and performance in live production environments.
Language Models & LLM Applications
• Apply transformer-based models (BERT, ClinicalBERT, Trial2Vec) for NLP tasks including classification, NER, and information extraction.
• Build and maintain vector similarity search pipelines for semantic retrieval and recommendation use cases.
• Fine-tune pre-trained models for domain-specific applications in clinical and healthcare contexts.
• Support exploratory work around LLM integration and prompt engineering for internal tooling.
Domain-Driven Analytics
• Apply advanced analytics within complex healthcare and clinical trial datasets—including patient records, trial protocols, and adverse event data.
• Translate ambiguous business problems into structured analytical frameworks with measurable outcomes.
• Partner with domain experts, product managers, and engineering teams to deliver data-driven solutions.
REQUIRED QUALIFICATIONS
Education
• Bachelor’s or Master’s degree in Computer Science, Statistics, Mathematics, Bioinformatics, or a closely related field.
Experience
• 2–4 years of hands-on experience in a data science or machine learning role.
• Demonstrable experience deploying ML models in production environments (not just prototyping).
Technical Skills
• Strong proficiency in Python (pandas, NumPy, scikit-learn, PyTorch / TensorFlow).
• Experience with Databricks, MLFlow (experiment tracking, model registry, endpoints), and Unity Catalog.
• Hands-on experience with BERT-family models and Hugging Face Transformers library.
• Familiarity with vector databases (e.g., FAISS, Pinecone, Weaviate) and embedding-based retrieval.
• Solid understanding of SQL and working with large structured/unstructured datasets.
• Exposure to cloud platforms (AWS / GCP / Azure) and distributed computing frameworks (Spark).
GOOD TO HAVE
• Prior experience with clinical trial data standards (CDISC, CDASH, SDTM) or healthcare ontologies (SNOMED, ICD-10).
• Familiarity with Trial2Vec or similar trial-to-vector embedding approaches.
• Experience with LLM fine-tuning, RAG pipelines, or prompt engineering in a production setting.
• Knowledge of regulatory and compliance considerations in healthcare AI (e.g., FDA guidelines, HIPAA).
• Contributions to open-source ML projects or published research.
THIS ROLE IS NOT FOR YOU IF…
• You have strong SQL/BI skills but limited hands-on ML modelling experience — or you’ve built models only in notebooks without ever deploying them to production.
• Your LLM exposure is limited to API calls and prompt engineering — with no experience fine-tuning models, working with embeddings, or building vector search pipelines.
Job Description – Data Scientist (Machine Learning & Forecasting)
About the Role
We are looking for a highly skilled Data Scientist with strong expertise in Machine Learning, Traditional Statistical Modelling, Forecasting, and Predictive Analytics. The ideal candidate will have hands-on experience building and deploying end-to-end ML solutions, working with large datasets, and translating business problems into scalable data science solutions.
The role requires a strong foundation in statistics, predictive modelling, feature engineering, model evaluation, and time-series forecasting, along with the ability to collaborate with cross-functional teams to deliver business impact.
Key Responsibilities
- Design, develop, and deploy Machine Learning models for business-critical use cases.
- Build and optimize traditional ML models such as:
- Linear Regression
- Logistic Regression
- Decision Trees
- Random Forest
- Gradient Boosting (XGBoost, LightGBM, CatBoost)
- Support Vector Machines
- Clustering Algorithms
- Develop forecasting solutions using:
- ARIMA / SARIMA
- Prophet
- Exponential Smoothing
- Time-Series Regression Models
- Perform exploratory data analysis (EDA), feature engineering, and data validation.
- Evaluate model performance using appropriate statistical and business metrics.
- Work with structured and semi-structured datasets from multiple sources.
- Collaborate with business stakeholders to understand requirements and translate them into analytical solutions.
- Build scalable data pipelines and support model deployment in production environments.
- Monitor model performance, identify data drift, and implement model retraining strategies.
- Present insights and recommendations to technical and non-technical stakeholders.
Required Skills & Qualifications
- Bachelor's or Master's degree in Computer Science, Statistics, Mathematics, Data Science, Engineering, or a related quantitative field.
- 5+ years of hands-on experience in Data Science, Machine Learning, and Forecasting.
Technical Skills
Machine Learning
- Strong understanding of supervised and unsupervised learning algorithms.
- Experience with ensemble methods and advanced ML techniques.
- Expertise in model selection, hyperparameter tuning, and performance optimization.
Forecasting & Statistics
- Strong understanding of:
- Time-Series Analysis
- Forecasting Techniques
- Statistical Inference
- Hypothesis Testing
- Probability Distributions
- A/B Testing
Programming
- Advanced proficiency in Python.
- Experience with:
- Pandas
- NumPy
- Scikit-learn
- Statsmodels
- XGBoost / LightGBM
- Prophet
Data & SQL
- Strong SQL skills with experience in complex queries and performance optimization.
- Experience working with large-scale datasets.
Visualization
- Experience with Power BI, Tableau, Matplotlib, Seaborn, or Plotly.
- Cloud & MLOps (Preferred)
- Exposure to AWS, Azure, or GCP.
- Understanding of Docker, Kubernetes, CI/CD, and ML model deployment practices.
Key Competencies
- Strong analytical and problem-solving skills.
- Excellent communication and stakeholder management abilities.
- Ability to work independently in a fast-paced environment.
- Strong business acumen and data-driven decision-making mindset.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Title : MLOps Engineer
Mode: Hybrid
Experience : 4 to 7 Years
Location : Hyderabad (Priority)/Bengaluru locations only
Notice Period : Immediate Joiner
Job Summary:
We are looking for a skilled and proactive ML Engineer with strong expertise in Python, Databricks, and Machine Learning model development. The ideal candidate should be proficient in building scalable data pipelines and deploying ML models, with a working knowledge of MLOps principles and tooling. This role offers an opportunity to work on impactful AI/ML initiatives in a collaborative environment.
Key Responsibilities:
• Develop and maintain machine learning pipelines for training, testing, and deploying models
• Design and implement infrastructure for managing and monitoring machine learning models
• Work with data scientists to build scalable, efficient, and automated model training and testing processes
• Collaborate with software engineers to integrate machine learning models into production systems
• Automate and optimize the deployment and scaling of machine learning models in a distributed computing environment
• Monitor and troubleshoot machine learning systems and infrastructure to ensure high availability and performance
• Develop and maintain documentation and best practices for MLOps processes and procedures.
Experience:
Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field
• 3+ years of experience in MLOps or related field, including building and deploying machine learning models at scale
•Proficiency in programming languages such as Python, Java, and C++
•Experience with machine learning frameworks such as TensorFlow, PyTorch, and Keras
• Experience with containerization technologies such as Docker and Kubernetes
• Strong understanding of DevOps principles and practices
• Experience with cloud computing platforms such as AWS, Azure, or Google Cloud
Sr.Data Scientist,Python, AI ML
We are looking for a skilled Data Scientist to analyze complex datasets, develop predictive models, and generate actionable insights that support business decisions. The ideal candidate should have strong statistical, analytical, and programming skills, along with hands-on experience in machine learning.
AuxoAI is hiring a Senior Data Scientist with strong expertise in AI, machine learning engineering (MLE), and generative AI. You will play a leading role in designing, deploying, and scaling production-grade ML systems — including large language model (LLM)-based pipelines, AI copilots, and agentic workflows. This role is ideal for someone who thrives on balancing cutting-edge research with production rigor and loves mentoring while building impact-first AI applications.
Location - Mumbai/Bangalore/Hyderabad/Gurgaon (Hybrid - 3 Days a week in Office)
Responsibilities:
- Own the full ML lifecycle: model design, training, evaluation, deployment
- Design production-ready ML pipelines with CI/CD, testing, monitoring, and drift detection
- Fine-tune LLMs and implement retrieval-augmented generation (RAG) pipelines
- Build agentic workflows for reasoning, planning, and decision-making
- Develop both real-time and batch inference systems using Docker, Kubernetes, and Spark
- Leverage state-of-the-art architectures: transformers, diffusion models, RLHF, and multimodal pipelines
- Collaborate with product and engineering teams to integrate AI models into business applications
- Mentor junior team members and promote MLOps, scalable architecture, and responsible AI best practices
Requirements
- 5+ years of experience in designing, deploying, and scaling ML/DL systems in production
- Proficient in Python and deep learning frameworks such as PyTorch, TensorFlow, or JAX
- Experience with LLM fine-tuning, LoRA/QLoRA, vector search (Weaviate/PGVector), and RAG pipelines
- Familiarity with agent-based development (e.g., ReAct agents, function-calling, orchestration)
- Solid understanding of MLOps: Docker, Kubernetes, Spark, model registries, and deployment workflows
- Strong software engineering background with experience in testing, version control, and APIs
- Proven ability to balance innovation with scalable deployment
- B.S./M.S./Ph.D. in Computer Science, Data Science, or a related field
- Bonus: Open-source contributions, GenAI research, or applied systems at scale






