ML Ops Engineer at Cloudkeeper · Noida · 7 - 12 years · ₹30L - ₹50L / yr · Bootstrapped · Posted 5 Oct 2026

Designation: Lead MLOps Engineer (GPU Optimization)
About CloudKeeper:
CloudKeeper is a cloud cost optimization partner that combines the power of group buying & commitments management, expert cloud consulting & support, and an enhanced visibility & analytics platform to reduce cloud cost & help businesses maximize the value from AWS, Microsoft Azure, & Google Cloud.
A certified AWS Premier Partner, Azure Technology Consulting Partner, Google Cloud Partner, and FinOps Foundation Premier Member, CloudKeeper has helped 350+ global companies save an average of 20% on their cloud bills, modernize their cloud set-up and maximize value — all while maintaining flexibility and avoiding any long-term commitments or cost.
CloudKeeper hived off from TO THE NEW, a digital technology service company with 2500+ employees and an 8-time GPTW winner.
To know more, please visit - https://www.cloudkeeper.com/
Responsibilities:
- Drive R&D and engineering for AI Infrastructure optimization within CloudKeeper's FinOps for AI platform — building the Tuner AI / Commit AI capability on GPU and ML workloads
- Design and build optimization engines for GPU right-sizing, idle shutdown, spot migration with checkpoint/resume automation, inference batching, quantization, and model placement
- Extend the optimization stack to LLM-era workloads — caching, model routing, dynamic batching, prompt optimization, RAG-aware architectures
- Partner with the Lens AI team to translate GPU and ML workload signals into actionable, dollar-quantified optimization recommendations for customers
- Work cross-functionally with product, platform, and customer success teams to ship optimization features end-to-end (data ingestion → optimization engine → customer-facing recommendation)
- Lead technical direction for AI workload optimization, set engineering standards, and mentor the ML / MLOps engineering bench as the AI Infrastructure pillar scales
- (Lead level) Hire, ramp, and grow a team of ML infrastructure engineers as headcount expands
Must Have:
- B.E / B.Tech / M.Tech / MCA with 7+ years of hands-on engineering experience
- Production experience with GPU workloads — has measurably optimized GPU utilization, throughput, or cost in a real production environment, not just academic / lab work
- Strong performance engineering background — must come ready with a concrete optimization story including before/after metrics (latency, throughput, or cost reduction)
- Strong Python + Linux + systems fundamentals
- Solid understanding of the ML model lifecycle — training, serving, inference — able to reason about what is running on the GPU and why
- MLOps fluency — model deployment, monitoring, observability, GPU cluster operations
- Hands-on with cloud GPU instances (AWS P5 / G6, Azure ND series, GCP A3, or equivalent) and Kubernetes-based GPU orchestration (EKS / AKS / GKE GPU node pools, Karpenter, Run:ai, NVIDIA GPU Operator, or similar)
- Familiarity with at least one modern LLM inference framework — vLLM, TGI, Triton, SGLang, Ray Serve, or BentoML
- Strong communication skills — able to translate deep technical optimization into customer / business outcomes
- (Lead level) Experience managing or technically leading a team of 3+ engineers
Good to Have:
- Deep LLM-era optimization expertise — KV caching, semantic caching, model routing, dynamic batching, quantization (FP16 → INT8 → INT4), model distillation, structured outputs
- Familiarity with LLM workload patterns — RAG, agents, embeddings, vector databases (Pinecone, Weaviate, Qdrant)
- CUDA, NCCL, mixed-precision training and inference
- Experience with managed ML training platforms — SageMaker, Azure ML, Vertex AI, Databricks Mosaic
- Exposure to GPU-native clouds — CoreWeave, Lambda Labs, RunPod, Crusoe
- Open source contributions to ML infrastructure projects — vLLM, llama.cpp, TGI, Ray, Triton, KubeRay
- Adjacent experience in cloud cost optimization / FinOps — Spot.io, ScaleOps, Granulate, CAST AI
- Comfort with Agile methodology and modern engineering practices (CI/CD, code review, observability)

Similar jobs (10)
Title : Senior AI Platform / MLOps Engineer
Experience : 6+ years
Work type : Chennai - Work from Office/other locations - Remote
Employment Type : Full Time
Notice Period : Immediate
Work Day :Mon to Fri
Key Responsibilities:
- Install, configure and operate OpenShift, NVIDIA GPU operator, OpenShift AI, and NIM microservices on 12× RTX PRO 6000 across two servers; single-node and HA control-plane topologies
- Serving configuration and tuning: quantized model deployment (FP8/FP4), replica balancing, batching, KV-cache and context management
- Azure GPU build environments: provisioning, cost control, parity with the on-prem stack via pinned container/model versions; cloud-to-factory migration with parity regression
- GitOps CI/CD, container registry, artifact/model versioning, environment promotion; observability and audit wiring (Splunk, Prometheus/Grafana)
- Benchmark automation: load harness, p50/p95/p99 latency, tokens/sec, GPU utilization; the capacity report data pipeline
- Platform upgrade procedure with evaluation-regression gates; deployment runbook as a first-class deliverable
Technical Skills:
- 6+ years infrastructure/platform engineering with 3+ years production Kubernetes; OpenShift experience strongly preferred
- Hands-on GPU inference serving in production: NIM, Triton, vLLM, or TensorRT-LLM — you have sized, deployed, and tuned LLM serving on real GPUs and can talk memory-bandwidth trade-offs
- GitOps fluency (ArgoCD/Flux), infrastructure-as-code, container internals; comfortable in air-gapped/proxy-restricted enterprise networks
- Observability depth: metrics, traces, log pipelines; has built performance test harnesses, not just run them
- Azure or AWS GPU compute operations experience
Strongly preferred
- NVIDIA GPU operator and AI Enterprise stack specifics; KServe; Milvus or pgvector operations; VAST/NFS/S3 storage integration; banking or other regulated-environment delivery
About Ampera:
Ampera Technologies, a purpose driven Digital IT Services with primary focus on supporting our client with their Data, AI / ML, Accessibility and other Digital IT needs. We also ensure that equal opportunities are provided to Persons with Disabilities Talent. Ampera Technologies has its Global Headquarters in Chicago, USA and its Global Delivery Center is based out of Chennai, India. We are actively expanding our Tech Delivery team in Chennai and across India. We offer exciting benefits for our teams, such as 1) Hybrid and Remote work options available, 2) Opportunity to work directly with our Global Enterprise Clients, 3) Opportunity to learn and implement evolving Technologies, 4) Comprehensive healthcare, and 5) Conducive environment for Persons with Disability Talent meeting Physical and Digital Accessibility standards
Greetings!
Hiring For Large Product Based Company!
Role- Mlops Engineer
Experience- 8-12 years
Location- Pune, Nagpur
JD-
- 8-10 years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering
- Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure
Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn

- Bachelor’s degree in computer science, Data Science, Information Systems, or a related field
- 8-10 years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering
- Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure
- Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn
Key Responsibilities:
- Develop and deploy machine learning, deep learning, and NLP models for various business use cases.
- Build end-to-end ML pipelines including data preprocessing, feature engineering, training, evaluation, and production deployment.
- Optimize model performance and ensure scalability in production environments.
- Work closely with data scientists, product teams, and engineers to translate business requirements into AI solutions.
- Conduct data analysis to identify trends and insights.
- Implement MLOps practices for versioning, monitoring, and automating ML workflows.
- Research and evaluate new AI/ML techniques, tools, and frameworks.
- Document system architecture, model design, and development processes.
Required Skills:
- Strong programming skills in Python (NumPy, Pandas, Scikit-learn, TensorFlow, PyTorch, Keras).
- Hands-on experience in building and deploying, finetuning ML/DL models in production.
- Good understanding of machine learning algorithms, neural networks, NLP, and computer vision.
- Experience with REST APIs, Docker, Kubernetes, and cloud platforms (AWS/GCP/Azure).
- Working knowledge of MLOps tools such as MLflow, Airflow, DVC, or Kubeflow.
- Familiarity with data pipelines and big data technologies (Spark, Hadoop) is a plus.
- Strong analytical skills and ability to work with large datasets.
- Excellent communication and problem-solving abilities.
- Experience in deploying models using cloud services (AWS Sagemaker, GCP Vertex AI, etc.).
- Experience in LLM fine-tuning or Generative AI, Voice AI, is an added advantage.
Educational Qualification:
- Bachelor’s or Master’s degree in Computer Science, Data Science, AI, Machine Learning, IT, from IIT/NIT colleges strongly preferred
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.
Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform.
You Will:
- Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines
- Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable
- CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools
- Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms
- Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable
- Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time.
- Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable
- Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users
- Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production.
- Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow.
- Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data
- Technology Evaluation and Innovation: Staying abreast of emerging data technologies and exploring opportunities for innovation to improve the organisation’s data infrastructure
- Troubleshooting and Problem Solving: Diagnosing and resolving complex data-related issues, ensuring the stability and reliability of the data platform
- Perform other duties as assigned
You Have:
- Enterprise SaaS software solutions with high availability and scalability
- Solution handling large scale structured and unstructured data from varied data sources
- Experience in building and maintaining AI/ML Ops platform systems ensuring scalability, reliability, efficiency and security
- Working with Product engineering team to influence designs with data, AI and analytics use cases in mind
- In depth experience in System design, AI/ML Frameworks and tools involving large Petabytes of data with Databricks Lakehouse ecosystem
- AI/MLOps workflows on Databricks , MLFlow, Mosaic AI Agent Framework, Unity Catalog, Vector Search, Knowledge Graph
- Knowledge of AI/ML frameworks like LangChain, LangGraph for AI/ML Ops pipeline integration
- Cloud Platforms: Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP). Experience in AWS hosted data platform is preferable
- Programming languages like Python and SQL
- Modern software engineering practices like Kubernetes, CI/CD, IAC tools (Preferably Terraform), Observability, monitoring and alerting
- Solution Cost Optimisations and design to cost
- Legally eligible to work in India on an ongoing basis
Get to Know Us:
At Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths—because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.
Equal Opportunity Employer:
Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information.
If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.
Job application link : https://grnh.se/z7qx2ehx1us
Job Title: Senior AI/ML Engineer
Company: Timble Technologies Pvt. Ltd
Location: Gurugram (Hybrid)
Experience: 2 TO 5 Years
About Us
Timble Glance is a high-growth AI RegTech and B2B SaaS company catering to top-tier BFSI and enterprise clients. We build cutting-edge systems powering 30+ high-scale APIs for digital identity verification, fraud detection, document intelligence, and compliance automation.
Role Overview
We are looking for a hands-on Senior AI/ML Engineer to design, develop, and productionize high-throughput AI/ML and Generative AI systems. You will own the full lifecycle—from problem formulation and data pipelines to deep learning architectures, RAG systems, LLMOps, and model governance—delivering sub-second latency and high reliability across our enterprise products.
Key Responsibilities
· Model Architecture & Deployment: Design, train, and deploy production-scale ML/Deep Learning and GenAI systems (computer vision, document intelligence, OCR, NLP, fraud risk classification, and LLM applications).
· GenAI & LLM Solutions: Develop robust LLM workflows including prompt engineering, fine-tuning, RAG pipelines, semantic search, vector indexing (Pinecone/Milvus/Chroma), and safety guardrails.
· Pipelines & Engineering: Build performant feature extraction and data pipelines; write modular, vectorized, production-grade Python (NumPy, Pandas) and advanced SQL.
· MLOps & Monitoring: Establish end-to-end MLOps/LLMOps standards—model registries, CI/CD, experiment tracking, drift detection, A/B testing, latency optimization, and cost governance.
· Responsible AI & Security: Ensure model decisions comply with enterprise data security, privacy standards, and auditability required by the BFSI sector.
· Collaboration & Ownership: Translate complex business requirements into technical roadmaps, conduct rigorous code reviews, and mentor junior engineers.
Required Qualifications & Skills
· Education: B.Tech / M.Tech in Computer Science, AI/ML, Mathematics, or a related field—Tier-1 institutes (IIT, IIIT, NIT) strongly preferred.
· Experience: 2+ years of hands-on experience developing, deploying, and maintaining ML/Deep Learning or GenAI models in production environments.
· GenAI & NLP Stack: Hands-on experience with LLMs, embeddings, RAG architectures, and frameworks such as LangChain, LlamaIndex, or Hugging Face.
· Deep Learning Frameworks: Strong proficiency in PyTorch or TensorFlow, with deep knowledge of transformer architectures and modern NLP/CV models.
· Software & Data Engineering: Expert-level Python skills (pytest, Git, OOP, asynchronous programming), solid SQL proficiency, and familiarity with data workflows.
· Deployment & Cloud: Practical exposure to cloud platforms (AWS/GCP), containerization (Docker), API frameworks (FastAPI/Flask), and basic orchestration (Kubernetes).
Preferred Qualifications
· Prior domain experience in Fintech, RegTech, Identity Verification (KYC/AML), Fraud Intelligence, or B2B SaaS.
· Experience optimizing models for low latency and inference cost (e.g., ONNX, TensorRT, model quantization).
· Familiarity with workflow orchestrators such as Airflow, Prefect, or Kubeflow.
Job Description:
We are seeking a versatile and highly skilled Lead AI/ML Engineer with deep expertise in Generative AI (GenAI) and Large Language Models (LLMs). This role requires a leader who can take full ownership of the
AI lifecycle—from initial architectural design to final production execution. You will lead the development of scalable AI-powered applications, demonstrating exceptional execution skills and the ability to deliver high-performance results under pressure in demanding production environments.
Machine Learning & LLM Capability:
End-to-End ML Engineering: Build and manage comprehensive ML pipelines, including data ingestion, preprocessing, training, and evaluation using frameworks like PyTorch, TensorFlow, and Scikit-learn. Advanced LLM Systems: Design and implement sophisticated LLM-based applications such as autonomous agents, chatbots, and complex automation tools.
Generative AI Specialization: Architect and optimize Retrieval-Augmented Generation (RAG) pipelines using vector databases like FAISS, Pinecone, or Weaviate.
Model Optimization: Fine-tune open-source and proprietary models (e.g., LLaMA, GPT) using advanced techniques like LoRA, QLoRA, or instruction tuning.
Agentic Frameworks: Develop complex agentic workflows utilizing frameworks such as LangChain or LlamaIndex.
Prompt Engineering: Implement expert-level prompt engineering, tool/function calling, and structured output generation.
Project Ownership & Execution
Full Lifecycle Ownership: Take complete accountability for the full ML and GenAI lifecycle, spanning data processing, model development, monitoring, and optimization.
Architectural Leadership: Drive strategic architectural decisions for AI platforms, ensuring they are modular, scalable, and maintainable.
Execution Excellence: Write clean, high-performance Python code following strict OOP principles and manage CI/CD pipelines for seamless project execution.
Leadership & Mentoring: Act as a key technical leader, managing stakeholders and mentoring team members to ensure all project milestones are met with quality.
System Integrity: Manage model and prompt versioning, experiment tracking, and comprehensive documentation for all pipelines and workflows.
Performance Under Pressure
Production Reliability: Ensure all AI systems maintain extreme scalability and performance under heavy production workloads, including both batch and real-time processing.
High-Pressure Optimization: Rapidly optimize inference latency and system costs for ML and LLM systems to meet urgent business and technical requirements.
Proactive Problem Solving: Apply strong analytical thinking to address complex challenges such as system drift, hallucinations, and latency in fast-paced environments.
Robust Guardrails: Implement and manage strict evaluation frameworks and feedback loops to maintain system quality under stress.
Qualifications:
Bachelor’s or Master’s degree in Computer Science, AI, ML, or a related field.
Proven expertise in Python, system design, and scalable AI/ML architecture.
Deep knowledge of NLP, Computer Vision, and Deep Learning models.
Hands-on experience with Docker, Kubernetes, MLOps, and major cloud platforms (AWS, GCP, or Azure).
We are looking for an MLOps Engineer to take ML models from notebook to production reliably.
Responsibilities
- Build ML training and deployment pipelines
- Track experiments and models with MLflow
- Run pipelines on Kubeflow or SageMaker
- Monitor model drift and performance
Requirements
- 2+ years in MLOps or ML engineering
- Hands-on with MLflow and Kubeflow or SageMaker
- Experience serving models at scale

Position Overview
The AI Observability Engineer will be instrumental in implementation of scalable, cloud-native solutions to meet the growing needs of our Data & Development team. The successful candidate will demonstrate the ability to abstract complexity and create reusable, scalable patterns that accelerate development. The AI Observability Engineer will build and maintain a robust framework to ensure the reliability and maintainability of DPR Construction's complex AI systems.
Responsibilities
- Standardize observability practices across AI/ML and other development teams including logging, metrics, tracing, and model performance monitoring, ingesting data from multiple platforms
- Lead hands-on implementation of automation-first DevOps and MLOps practices, enabling infrastructure-as-code and consistent, repeatable environment provisioning
- Design and manage intelligent DataOps pipelines with automated data quality monitoring and anomaly detection
- Deploy, maintain and monitor containerized ML workloads
- Extend existing CI/CD pipelines to support automated infrastructure changes and ML workflows
- Implement AI-driven data validation, schema and concept drift detection and metadata management.
- Establish governance frameworks for AI systems, including bias detection, explainability, and auditability
- Extend existing Azure RBAC strategy by automating role and permission management to reduce manual intervention
- Develop automated test suites for model performance, regression, edge cases and bias validation
- Monitor model KPIs (accuracy, precision, recall, latency, calibration)
- Ensure reproducability of experiments and production models
- Act as a technical point of contact for DevOps and MLOps practices, developing reusable patterns, documentation, and proof-of-concepts to drive adoption
Qualifications
- Bachelor’s degree in computer science, Data Science, Information Systems, or a related field
- 5+ years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering
- Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure
- Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn
Job Summary
We are looking for an experienced AI/ML Engineer to design, develop, deploy, and maintain machine learning and AI solutions that address complex business problems. The ideal candidate should have strong hands-on experience in Python, Machine Learning, Generative AI, LLMs, and AI/ML deployment, with the ability to work across the complete AI/ML lifecycle.
Key Responsibilities
- Design, develop, and deploy scalable Machine Learning and AI models for real-world business use cases.
- Build and optimize ML pipelines covering data preparation, feature engineering, model development, evaluation, and deployment.
- Develop solutions using Generative AI, Large Language Models (LLMs), NLP, and deep learning.
- Work with LLMs, prompt engineering, embeddings, vector databases, and Retrieval-Augmented Generation (RAG) architectures.
- Integrate AI/ML models with enterprise applications and APIs.
- Fine-tune and evaluate ML/LLM models based on business requirements.
- Implement MLOps practices for model versioning, deployment, monitoring, and continuous improvement.
- Collaborate with Data Scientists, Software Engineers, Architects, Product Managers, and business stakeholders.
- Conduct model performance evaluation, optimization, and troubleshooting.
- Ensure AI solutions meet requirements around security, scalability, reliability, responsible AI, and data privacy.
- Stay current with emerging AI/ML technologies, frameworks, and industry best practices.
Required Skills
- Strong programming experience in Python.
- Strong understanding of Machine Learning algorithms, statistics, and data structures.
- Hands-on experience with ML frameworks such as TensorFlow, PyTorch, Scikit-learn, or equivalent.
- Experience with Generative AI and LLMs such as OpenAI, Azure OpenAI, Claude, Gemini, or open-source models.
- Strong knowledge of Prompt Engineering, RAG, embeddings, vector databases, and AI agents.
- Experience with NLP, deep learning, or computer vision is an advantage.
- Experience developing and consuming REST APIs and microservices.
- Working knowledge of SQL and NoSQL databases.
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Understanding of Docker, Kubernetes, CI/CD, and MLOps.
- Familiarity with Git and modern software development practices.
Preferred Qualifications
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Data Science, Engineering, or a related field.
- 4 years of relevant experience in AI/ML engineering or a related field.
- Experience building and deploying production-grade AI/ML solutions.
- Enterprise application development experience.
- Experience with Azure OpenAI, AWS Bedrock, Vertex AI, or similar managed AI platforms.
- Experience with LangChain, LlamaIndex, Semantic Kernel, or comparable AI frameworks is a plus.
- Experience with AI/ML model monitoring, evaluation, and optimization.
What You Bring
- Strong problem-solving and analytical skills.
- Ability to translate business requirements into practical AI/ML solutions.
- Strong software engineering and debugging capabilities.
- Ability to work independently as well as collaboratively in a cross-functional environment.
- Good communication skills with the ability to explain complex AI concepts to technical and non-technical stakeholders.
Keywords
AI Engineer | ML Engineer | Machine Learning | Generative AI | LLM | Python | NLP | Deep Learning | RAG | Prompt Engineering | AI Agents | Azure OpenAI | AWS Bedrock | MLOps | TensorFlow | PyTorch | Scikit-learn | Vector Database | Cloud AI
Location: Gurugram
Work mode: Hybrid





