ML Ops Engineer at JK Technosoft Ltd · Bengaluru (Bangalore) · 3 - 5 years · ₹5L - ₹15L / yr · Profitable · Posted 23 Jun 2023

Roles and Responsibilities:
- Design, develop, and maintain the end-to-end MLOps infrastructure from the ground up, leveraging open-source systems across the entire MLOps landscape.
- Creating pipelines for data ingestion, data transformation, building, testing, and deploying machine learning models, as well as monitoring and maintaining the performance of these models in production.
- Managing the MLOps stack, including version control systems, continuous integration and deployment tools, containerization, orchestration, and monitoring systems.
- Ensure that the MLOps stack is scalable, reliable, and secure.
Skills Required:
- 3-6 years of MLOps experience
- Preferably worked in the startup ecosystem
Primary Skills:
- Experience with E2E MLOps systems like ClearML, Kubeflow, MLFlow etc.
- Technical expertise in MLOps: Should have a deep understanding of the MLOps landscape and be able to leverage open-source systems to build scalable, reliable, and secure MLOps infrastructure.
- Programming skills: Proficient in at least one programming language, such as Python, and have experience with data science libraries, such as TensorFlow, PyTorch, or Scikit-learn.
- DevOps experience: Should have experience with DevOps tools and practices, such as Git, Docker, Kubernetes, and Jenkins.
Secondary Skills:
- Version Control Systems (VCS) tools like Git and Subversion
- Containerization technologies like Docker and Kubernetes
- Cloud Platforms like AWS, Azure, and Google Cloud Platform
- Data Preparation and Management tools like Apache Spark, Apache Hadoop, and SQL databases like PostgreSQL and MySQL
- Machine Learning Frameworks like TensorFlow, PyTorch, and Scikit-learn
- Monitoring and Logging tools like Prometheus, Grafana, and Elasticsearch
- Continuous Integration and Continuous Deployment (CI/CD) tools like Jenkins, GitLab CI, and CircleCI
- Explain ability and Interpretability tools like LIME and SHAP

About JK Technosoft Ltd
About
Similar jobs (5)
- Bachelor’s degree in computer science, Data Science, Information Systems, or a related field
- 8-10 years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering
- Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure
- Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn
Greetings!
Hiring For Large Product Based Company!
Role- Mlops Engineer
Experience- 8-12 years
Location- Pune, Nagpur
JD-
- 8-10 years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering
- Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure
Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn
For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.
Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform.
You Will:
- Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines
- Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable
- CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools
- Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms
- Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable
- Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time.
- Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable
- Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users
- Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production.
- Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow.
- Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data
- Technology Evaluation and Innovation: Staying abreast of emerging data technologies and exploring opportunities for innovation to improve the organisation’s data infrastructure
- Troubleshooting and Problem Solving: Diagnosing and resolving complex data-related issues, ensuring the stability and reliability of the data platform
- Perform other duties as assigned
You Have:
- Enterprise SaaS software solutions with high availability and scalability
- Solution handling large scale structured and unstructured data from varied data sources
- Experience in building and maintaining AI/ML Ops platform systems ensuring scalability, reliability, efficiency and security
- Working with Product engineering team to influence designs with data, AI and analytics use cases in mind
- In depth experience in System design, AI/ML Frameworks and tools involving large Petabytes of data with Databricks Lakehouse ecosystem
- AI/MLOps workflows on Databricks , MLFlow, Mosaic AI Agent Framework, Unity Catalog, Vector Search, Knowledge Graph
- Knowledge of AI/ML frameworks like LangChain, LangGraph for AI/ML Ops pipeline integration
- Cloud Platforms: Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP). Experience in AWS hosted data platform is preferable
- Programming languages like Python and SQL
- Modern software engineering practices like Kubernetes, CI/CD, IAC tools (Preferably Terraform), Observability, monitoring and alerting
- Solution Cost Optimisations and design to cost
- Legally eligible to work in India on an ongoing basis
Get to Know Us:
At Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths—because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.
Equal Opportunity Employer:
Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information.
If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.
Job application link : https://grnh.se/z7qx2ehx1us
We are looking for an MLOps Engineer to take ML models from notebook to production reliably.
Responsibilities
- Build ML training and deployment pipelines
- Track experiments and models with MLflow
- Run pipelines on Kubeflow or SageMaker
- Monitor model drift and performance
Requirements
- 2+ years in MLOps or ML engineering
- Hands-on with MLflow and Kubeflow or SageMaker
- Experience serving models at scale
Designation: Lead MLOps Engineer (GPU Optimization)
About CloudKeeper:
CloudKeeper is a cloud cost optimization partner that combines the power of group buying & commitments management, expert cloud consulting & support, and an enhanced visibility & analytics platform to reduce cloud cost & help businesses maximize the value from AWS, Microsoft Azure, & Google Cloud.
A certified AWS Premier Partner, Azure Technology Consulting Partner, Google Cloud Partner, and FinOps Foundation Premier Member, CloudKeeper has helped 350+ global companies save an average of 20% on their cloud bills, modernize their cloud set-up and maximize value — all while maintaining flexibility and avoiding any long-term commitments or cost.
CloudKeeper hived off from TO THE NEW, a digital technology service company with 2500+ employees and an 8-time GPTW winner.
To know more, please visit - https://www.cloudkeeper.com/
Responsibilities:
- Drive R&D and engineering for AI Infrastructure optimization within CloudKeeper's FinOps for AI platform — building the Tuner AI / Commit AI capability on GPU and ML workloads
- Design and build optimization engines for GPU right-sizing, idle shutdown, spot migration with checkpoint/resume automation, inference batching, quantization, and model placement
- Extend the optimization stack to LLM-era workloads — caching, model routing, dynamic batching, prompt optimization, RAG-aware architectures
- Partner with the Lens AI team to translate GPU and ML workload signals into actionable, dollar-quantified optimization recommendations for customers
- Work cross-functionally with product, platform, and customer success teams to ship optimization features end-to-end (data ingestion → optimization engine → customer-facing recommendation)
- Lead technical direction for AI workload optimization, set engineering standards, and mentor the ML / MLOps engineering bench as the AI Infrastructure pillar scales
- (Lead level) Hire, ramp, and grow a team of ML infrastructure engineers as headcount expands
Must Have:
- B.E / B.Tech / M.Tech / MCA with 7+ years of hands-on engineering experience
- Production experience with GPU workloads — has measurably optimized GPU utilization, throughput, or cost in a real production environment, not just academic / lab work
- Strong performance engineering background — must come ready with a concrete optimization story including before/after metrics (latency, throughput, or cost reduction)
- Strong Python + Linux + systems fundamentals
- Solid understanding of the ML model lifecycle — training, serving, inference — able to reason about what is running on the GPU and why
- MLOps fluency — model deployment, monitoring, observability, GPU cluster operations
- Hands-on with cloud GPU instances (AWS P5 / G6, Azure ND series, GCP A3, or equivalent) and Kubernetes-based GPU orchestration (EKS / AKS / GKE GPU node pools, Karpenter, Run:ai, NVIDIA GPU Operator, or similar)
- Familiarity with at least one modern LLM inference framework — vLLM, TGI, Triton, SGLang, Ray Serve, or BentoML
- Strong communication skills — able to translate deep technical optimization into customer / business outcomes
- (Lead level) Experience managing or technically leading a team of 3+ engineers
Good to Have:
- Deep LLM-era optimization expertise — KV caching, semantic caching, model routing, dynamic batching, quantization (FP16 → INT8 → INT4), model distillation, structured outputs
- Familiarity with LLM workload patterns — RAG, agents, embeddings, vector databases (Pinecone, Weaviate, Qdrant)
- CUDA, NCCL, mixed-precision training and inference
- Experience with managed ML training platforms — SageMaker, Azure ML, Vertex AI, Databricks Mosaic
- Exposure to GPU-native clouds — CoreWeave, Lambda Labs, RunPod, Crusoe
- Open source contributions to ML infrastructure projects — vLLM, llama.cpp, TGI, Ray, Triton, KubeRay
- Adjacent experience in cloud cost optimization / FinOps — Spot.io, ScaleOps, Granulate, CAST AI
- Comfort with Agile methodology and modern engineering practices (CI/CD, code review, observability)






