Cutshort logo
For Employers
CLOUDSUFI logo
AI Platform Engineer – DevOps & MLOps/AgentOps (GCP) | Noida Hybrid
AI Platform Engineer – DevOps & MLOps/AgentOps (GCP) | Noida Hybrid

AI Platform Engineer – DevOps & MLOps/AgentOps (GCP) | Noida Hybrid at CLOUDSUFI · Noida · 5 - 9 years · ₹20L - ₹40L / yr · Profitable · Posted 9 Sep 2026

CLOUDSUFI's logo

AI Platform Engineer – DevOps & MLOps/AgentOps (GCP) | Noida Hybrid

Lishta Jain's profile picture
Posted by Lishta Jain
5 - 9 yrs
₹20L - ₹40L / yr
Noida
Skills
Google Cloud Platform (GCP)
CI/CD
MLOps
Agentic AI
skill iconKubernetes
AI Agents
agent deployment



About Us:


CLOUDSUFI, a Google Cloud Premier Partner, is a global leading provider of data-driven digital transformation across cloud-based enterprises. With a global presence and focus on Software & Platforms, Life sciences and Healthcare, Retail, CPG, financial services and supply chain, CLOUDSUFI is positioned to meet customers where they are in their data monetization journey.


Our Values

We are a passionate and empathetic team that prioritizes human values. Our purpose is to elevate the quality of lives for our family, customers, partners and the community.


Equal Opportunity Statement


CLOUDSUFI is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified candidates receive consideration for employment without regard to race, colour, religion, gender, gender identity or expression, sexual orientation and national origin status. We provide equal opportunities in employment, advancement, and all other areas of our workplace. Please explore more at https://www.cloudsufi.com/


About the role


CloudSufi is hiring an AgentOps Engineer to own the runtime, infrastructure, and operational monitoring for an AI agent testing platform deployed for an enterprise client. You will deploy and instrument agents for testing, set up observability and continuous evaluation on live traffic, and manage the cloud infrastructure the platform runs on.


Responsibilities:

- Deploy and instrument AI agents for testing and monitoring.

- Set up observability — traces, logs, and metrics — and continuous evaluation on production traffic.

- Build drift detection and alerting so quality regressions are caught early.

- Own the cloud infrastructure and CI/CD automation supporting the platform.

- Manage environments, access, and operational reliability.


Required qualifications:

- 5+ years in DevOps, MLOps, or platform engineering.

- Strong Google Cloud experience (Cloud Run, Cloud Build, Pub/Sub, Cloud Logging and Monitoring, IAM) or equivalent cloud platform.

- Experience with CI/CD automation and infrastructure as code.

- Familiarity with deploying LLM or AI agent workloads.


Preferred qualifications:

- Experience with OpenTelemetry and AI/agent observability.

- Exposure to Vertex AI, Gemini Enterprise Agent Platform, or ADK.

Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About CLOUDSUFI

Founded :
2019
Type :
Product
Size :
100-500
Stage :
Profitable

About

We exist to eliminate the gap between “Human Intuition” and “Data-Backed Decisions”


Data is the new oxygen, and we believe no organization can live without it. We partner with our customers to get to the core of their problems, enable the data supply chain and help them monetize their data. We make enterprise data dance!


Our work elevates the quality of lives for our family, customers, partners and the community.


The human values that we display in all our interactions are of:


Passion – we are committed in heart and head

Integrity – we are real, honest and, fair

Empathy – we understand business isn’t just B2B, or B2C, it is H2H i.e. Human to Human

Boldness – we have the courage to think and do differently


The CLOUDSUFI Foundation embraces the power of legacy and wisdom of those who have helped laid the foundation for all of us, our seniors. We believe in their abilities and we pledge to equip them, to provide them jobs, and to bring them sufi joy.

Read more

Tech stack

skill iconGo Programming (Golang)
skill iconJava
skill iconMachine Learning (ML)
skill iconData Analytics
skill iconReact.js
skill iconReact Native
Information security
Cyber Security
Data engineering
SAP

Connect with the team

Profile picture
Harmeet Singh
Profile picture
Ayushi Dwivedi
Profile picture
Lishta Jain

Company social profiles

instagramlinkedintwitterfacebook

Similar jobs (10)

Hiring for IT Consulting Firm (MNC)
Hiring for IT Consulting Firm (MNC)
Agency job
via by Sneha k
Pune, Nagpur
5 - 10 yrs
₹20L - ₹30L / yr
MLOps
DevOps
Artificial Intelligence (AI)
skill iconData Science
Data engineering
+5 more

Position Overview 

The AI Observability Engineer will be instrumental in implementation of scalable, cloud-native solutions to meet the growing needs of our Data & Development team. The successful candidate will demonstrate the ability to abstract complexity and create reusable, scalable patterns that accelerate development. The AI Observability Engineer will build and maintain a robust framework to ensure the reliability and maintainability of DPR Construction's complex AI systems. 

 

Responsibilities 

  • Standardize observability practices across AI/ML and other development teams including logging, metrics, tracing, and model performance monitoring, ingesting data from multiple platforms 
  • Lead hands-on implementation of automation-first DevOps and MLOps practices, enabling infrastructure-as-code and consistent, repeatable environment provisioning 
  • Design and manage intelligent DataOps pipelines with automated data quality monitoring and anomaly detection 
  • Deploy, maintain and monitor containerized ML workloads 
  • Extend existing CI/CD pipelines to support automated infrastructure changes and ML workflows 
  • Implement AI-driven data validation, schema and concept drift detection and metadata management. 
  • Establish governance frameworks for AI systems, including bias detection, explainability, and auditability 
  • Extend existing Azure RBAC strategy by automating role and permission management to reduce manual intervention 
  • Develop automated test suites for model performance, regression, edge cases and bias validation 
  • Monitor model KPIs (accuracy, precision, recall, latency, calibration) 
  • Ensure reproducability of experiments and production models 
  • Act as a technical point of contact for DevOps and MLOps practices, developing reusable patterns, documentation, and proof-of-concepts to drive adoption 

Qualifications 

  • Bachelor’s degree in computer science, Data Science, Information Systems, or a related field 
  • 5+ years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering 
  • Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure 
  • Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn 
Read more
company logo
Taher Ujjainwala
Posted by Taher Ujjainwala
Pune
12 - 25 yrs
₹70L - ₹120L / yr (ESOP available)
skill iconPython
skill iconKubernetes
Google Cloud Platform (GCP)
skill iconAmazon Web Services (AWS)
Windows Azure
+15 more

About the Role

We are hiring Staff / Principal Engineers to take full, hands-on ownership of Blitzy's most critical production-grade systems and to deliver high-leverage features that materially improve customer outcomes and engineering velocity. This is the most senior individual contributor role at the company today.

This is not a Senior-plus role, an architecture-only role, or a promotion-track role. We are looking for someone who has already operated at Principal / Staff+ scope in a highly technical environment and expects to spend their time writing, reviewing, and shipping production code.

This role is 100% hands-on. Leverage comes from system ownership, execution quality, and durable technical decisions — not people management or process.


Responsibilities

  • Own mission-critical production systems end-to-end, ensuring correctness, scalability, performance, reliability, and operational excellence.
  • Design, build, and ship high-impact backend systems and features that improve product reliability, performance, and customer value.
  • Architect scalable services and cloud infrastructure using technologies such as Python, REST, gRPC, Kubernetes, and Terraform.
  • Identify and resolve complex technical bottlenecks that limit engineering quality, system performance, or organizational velocity.
  • Build and operate LLM-powered systems and validation loops that evaluate correctness, consistency, durability, and production performance.
  • Design and evolve data architectures incorporating relational, NoSQL, graph, and vector databases to support complex enterprise applications and semantic retrieval.
  • Modernize and improve complex enterprise systems while balancing reliability, maintainability, scalability, and delivery speed.
  • Set and uphold engineering quality standards through hands-on technical leadership, sound technical judgment, and ownership of long-term technical decisions.


Qualifications

  • Direct experience with Python as a primary programming language, backend frameworks, and microservices architectures.
  • Expertise in REST and gRPC, with proficiency in Node.js and JavaScript.
  • Proficiency in GCP, along with experience using at least one additional cloud platform such as AWS or Azure.
  • Advanced knowledge of Kubernetes and Terraform in production environments.
  • Experience operating highly available production systems, including monitoring, scalability, reliability, performance optimization, and operational tooling.
  • Strong knowledge of SQL and NoSQL databases, including PostgreSQL, MySQL, MongoDB, Cassandra, or DynamoDB.
  • Familiarity with graph databases such as Neo4j and vector databases or embedding infrastructure for semantic search and retrieval.
  • Hands-on experience building and operating LLM-powered systems in production, including evaluation, validation, regression testing, tracing, and failure analysis.
  • Working knowledge of LangSmith or comparable LLM observability and evaluation tools; familiarity with OpenAI, Anthropic, or similar model providers is a plus.
  • Ability to contribute across the full stack, with a strong understanding of frontend architecture and the ability to debug, design, and ship across frontend, backend, infrastructure, and AI systems.
  • Understanding of large-scale enterprise software systems, including architecture, integration, deployment, modernization, and long-term maintainability.
  • Proven track record of operating at Staff+, Principal Engineer, or equivalent level, independently driving complex technical initiatives and delivering high-impact outcomes with minimal supervision.


Blitzy is a Cambridge, MA based AI software development platform on a mission to revolutionize the software development life cycle by autonomously building custom software to unlock the next industrial revolution. We're transforming how enterprises build software, turning enterprise requirements into enterprise grade code with an agentic software development platform that can autonomously execute 80% of the quantum of software development work. We're backed by multiple tier 1 investors, and have proven success as founders of previous start-ups.


Our Culture

Who we are:

Led by two pioneering co-founders we are one of the fastest growing companies in the U.S., creating our own category of enterprise autonomous software development. We automate thousands of hours of software development for our customers, which includes strong representation within the Fortune 500.


How we work:

  • We move Blitzy Fast: Time is both our company’s and our clients’ most precious asset. We move quickly and decisively to innovate internally and deliver exceptional software externally.
  • Championship Mindset: We operate like a professional sports team. We win as a team by holding ourselves and each other to high standards, collaborating in-person, and remaining focused on the mission.
  • Passion for Invention: We’re pushing the frontier of what’s possible, requiring constant innovation and iteration.
  • We Work for the Customer: We focus on delivering outsized value to the customers we work with and expanding those relationships into deep, meaningful partnerships.
  • We believe in being ‘everyday athletes’: taking care of ourselves so we can bring our best minds to work. We promote great sleep, movement, and restorative activities for 


Blitzy is an equal opportunity employer committed to building a diverse and inclusive team. We believe different perspectives make us stronger.

Read more
company logo
Swathi S
Posted by Swathi S
Chennai
7 - 12 yrs
₹30L - ₹55L / yr
skill iconAmazon Web Services (AWS)
skill iconPython
CI/CD
DevOps
Platform as a Service (PaaS)
+7 more

Amura’s Vision 


We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.


Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.


Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.


These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.

We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence. 


Role Overview 


We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.


This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.


You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability. 


Key Responsibilities 


Cloud Infrastructure & Platform Engineering (AWS) 

  • Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
  • Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
  • Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
  • Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
  • Build reusable platform templates and shared infrastructure modules. 


AI/ML Infrastructure & MLOps 

  • Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
  • Support GPU-based workloads and optimize compute/storage usage.
  • Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
  • Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines. 


CI/CD, Automation & Developer Productivity 

  • Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
  • Automate deployments, environment provisioning, and release workflows.
  • Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
  • Implement automated patching, scaling, backups, cleanup workflows, and drift detection. 


Containers, Kubernetes & Platform Reliability

  • Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
  • Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
  • Optimize infrastructure for performance, resilience, and cost-efficiency.
  • Implement progressive deployment strategies including blue/green, canary, and rolling deployments. 


Observability, Incident Response & SRE Practices

  • Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
  • Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
  • Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.

FinOps, Cost Governance & Security

  • Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
  • Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
  • Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.

Collaboration, Leadership & Platform Culture

  • Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
  • Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.

Skills & Qualifications


Must-Have:

  • 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
  • Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
  • Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations. 


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations. 


Here are answers to some questions you may have

Where is your office?

Chennai (Velachery)

Work Model

Work from Office – because great stories are built in person!

Do you have an online presence?

https://amura.ai (we are @AmuraHealth on all social media)


Read more
company logo
Sandeep Selvan
Posted by Sandeep Selvan
Bengaluru (Bangalore)
4 - 12 yrs
Best in industry
MLOps
databricks
skill iconMachine Learning (ML)
MLFlow
LangGraph
+4 more

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.


Our India Global Capability Center isn't just supporting global operations—we’re leading global innovation. After scaling rapidly into a best-in-class hub, we deliver the product innovation and enterprise capabilities that accelerate our global growth, profitability, and scale. As we expand Smartsheet India, we’re searching for Senior AI/ML Ops Engineers who crave variety and ownership. You’ll have the opportunity to work across multiple teams and disciplines, building a versatile skillset while solving the complex challenges of a global platform.


You Will:

  • Designing, Developing and overseeing the strategy and architecture of scalable and reliable AI/ML Ops platforms / pipelines
  • Model Deployment: Package and deploy AI/ML services to production, ensuring they are reproducible and interpretable
  • CI/CD Pipeline Development: Design and implement automated CI/CD (Continuous Integration/Continuous Deployment) pipelines to accelerate model deployment using tools
  • Infrastructure Management: Provision and optimize infrastructure for training and serving, utilizing Docker, Kubernetes, or serverless platforms
  • Monitoring & Observability : Implement post-deployment monitoring for model performance, data drift, and latency using tools. Experience in Monte Carlo is preferable
  • Automation: Automate retraining and data pipeline workflows to ensure models stay accurate over time.
  • Manage the deployment of foundation models, fine-tuning workflows, and Retrieval-Augmented Generation (RAG) stacks (Vector DBs, Knowledge Graph. Experience with AWS Bedrock is preferable
  • Resource Optimization: Manage GPU/CPU utilization to minimize cloud costs while maintaining low-latency inference for users
  • Collaboration: Work closely with data scientists, data engineers, and software engineers to bridge the gap between model development and production.
  • Version Control & Governance: Manage versioning for data, code, and models using tools like MLflow.
  • Security & Compliance: Implementing data security measures, ensuring compliance with data governance policies, and protecting sensitive data
  • Technology Evaluation and Innovation: Staying abreast of emerging data technologies and exploring opportunities for innovation to improve the organisation’s data infrastructure
  • Troubleshooting and Problem Solving: Diagnosing and resolving complex data-related issues, ensuring the stability and reliability of the data platform
  • Perform other duties as assigned


You Have:

  • Enterprise SaaS software solutions with high availability and scalability
  • Solution handling large scale structured and unstructured data from varied data sources
  • Experience in building and maintaining AI/ML Ops platform systems ensuring scalability, reliability, efficiency and security
  • Working with Product engineering team to influence designs with data, AI and analytics use cases in mind
  • In depth experience in System design, AI/ML Frameworks and tools involving large Petabytes of data with Databricks Lakehouse ecosystem
  • AI/MLOps workflows on Databricks , MLFlow, Mosaic AI Agent Framework, Unity Catalog, Vector Search, Knowledge Graph
  • Knowledge of AI/ML frameworks like LangChain, LangGraph for AI/ML Ops pipeline integration
  • Cloud Platforms: Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP). Experience in AWS hosted data platform is preferable
  • Programming languages like Python and SQL
  • Modern software engineering practices like Kubernetes, CI/CD, IAC tools (Preferably Terraform), Observability, monitoring and alerting
  • Solution Cost Optimisations and design to cost
  • Legally eligible to work in India on an ongoing basis

 

Get to Know Us:

At Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths—because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.


Equal Opportunity Employer:

Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information. 

If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.



Job application link : https://grnh.se/z7qx2ehx1us

Read more
company logo
Daniel Castellanos
Posted by Daniel Castellanos
Remote only
1 - 10 yrs
₹10L - ₹50L / yr
skill iconKubernetes
ArgoCD
Exim
Postfix

We are looking for an experienced DevOps Engineer to take ownership of production infrastructure, cloud environments, Kubernetes platforms, and infrastructure automation. This is a hands-on role for someone who enjoys solving complex infrastructure challenges and is comfortable being responsible for systems in production.


Key Responsibilities

  • Own and operate production infrastructure, including participating in an on-call rotation and responding to production incidents.
  • Design, operate, and continuously improve Kubernetes clusters in production.
  • Manage and automate infrastructure using Infrastructure as Code, primarily with Terraform.
  • Build, maintain, and optimise cloud infrastructure across AWS, GCP, or Azure.
  • Work extensively with Linux, including system administration, networking, troubleshooting, and system-level configuration.
  • Manage production deployment and GitOps workflows using ArgoCD.
  • Improve infrastructure reliability, scalability, security, monitoring, and operational efficiency.
  • Troubleshoot complex production issues and drive problems through to resolution.
  • Develop automation and processes that reduce manual operational work.

Essential Requirements

  • 4+ years of hands-on experience operating production infrastructure, with personal ownership and responsibility for live systems, including on-call experience.
  • Deep, hands-on Kubernetes experience — you must have operated and managed Kubernetes clusters, rather than simply deploying applications onto clusters managed by another team.
  • Strong experience with Infrastructure as Code, with Terraform strongly preferred. Experience with Pulumi or CloudFormation is also considered.
  • Strong experience with at least one major cloud platform, ideally AWS. Strong GCP or Azure experience is also welcome, provided you are willing to work with AWS.
  • Strong Linux skills and confidence working from the command line, including networking, troubleshooting, system configuration, and performance issues.
  • Production experience with ArgoCD and GitOps-based deployment workflows.
  • Strong troubleshooting and problem-solving skills, with the ability to take ownership of production incidents and infrastructure issues.


Nice to Have

Experience with email infrastructure would be a strong advantage, particularly:

  • Exim
  • IMAP / SMTP
  • Postfix
  • Dovecot
  • General mail server administration and maintenance


Read more
company logo
Lakshit Bagga
Posted by Lakshit Bagga
Remote only
10 - 25 yrs
₹1L - ₹50L / yr
Odoo (OpenERP)


We're Hiring: Agentic Tools Architect

Contract | Remote-first | Global


We're currently recruiting for one of our clients, A company that specializes in acquiring enterprise software businesses and transforming them into AI-native, profitable, scalable operations. They run remote-first, globally distributed teams, and their internal operations platform (Odoo, Redmine, GitHub) currently runs on human-shaped workflows. They're looking for someone to rebuild it for agents.


The Role

Redesign the client's operational tooling so AI agents can create, route, and resolve work alongside people — with minimal oversight. You'll own the roadmap, build the plugins (Rails-first), and set the guardrails.


You'll:

  • Own feature roadmap across Odoo, Redmine, GitHub & adjacent systems
  • Design agent-operable APIs, MCP servers, and automation
  • Build Rails-based Redmine plugins, Odoo modules, GitHub Apps that survive version upgrades
  • Review AI-generated PRs and agent-authored config; set standards for access, integrity, observability

You bring:

  • 10+ years in software/platform engineering with deep SME knowledge of Odoo, Redmine, GitHub, or similar
  • Expert Ruby on Rails, production Redmine plugin experience
  • Real-world agentic dev experience (MCP, LLM tooling, agent workflows)
  • Strong grip on delivery/support/ops business processes


Nice to have: Python + Odoo modules, PostgreSQL, legacy modernization experience, PE-backed/SaaS scale-up background

Read more
company logo
Umama Sayed
Posted by Umama Sayed
Mumbai
5 - 8 yrs
Best in industry
skill iconPython
Large Language Models (LLM)
Artificial Intelligence (AI)
Prompt engineering
LangGraph
+6 more

Senior AI Engineer

Code Generation, Agent Architecture & LLM Systems

📍 Mumbai (On-site) | Full-time | 5+ years


About the Role:

Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.

We are hiring a Senior AI Engineer for a dedicated client engagement focused on building an AI-powered application builder platform - a product where users describe software in plain English and the system generates, previews, and iteratively refines working code.

The mandatory requirement for this role is hands-on production experience shipping LLM-powered systems with agent architectures, with experience in code generation or developer tooling contexts a strong advantage.


The role is product-focused and deeply hands-on. You will own everything between the user's prompt and correct code landing in the project: the agentic loop, code generation pipeline, context management, evaluation suite, and model cost strategy.

You will work alongside the Senior MLOps Engineer who operationalises the infrastructure around your system, and collaborate closely with backend, frontend, and DevOps engineers.


Responsibilities:


Agent Architecture

Design and own the agentic loop for the platform - request interpretation, planning, tool-calling sequence (read file, edit file, run build, search code, install package), and stop conditions.

Make and revisit architectural decisions on single-agent vs. multi-agent designs, including planner/executor splits and dedicated build-repair sub-agents.


Code Generation Pipeline

Own the end-to-end generation flow: task classification, context gathering, planning, targeted edits, verification, and commit.

Implement diff/search-replace-based file editing with fuzzy matching and fallback strategies.

Enforce scope discipline so the agent makes minimal diffs and does not modify code it was not asked to touch.


Self-Repair Loop

Build and tune the automated repair loop that pipes compiler, lint, build, and runtime errors back to the model with retry budgets and model escalation.

This loop is the primary quality lever - the difference between 60-70% and 90%+ build success rates.


Context Management

Build file-relevance retrieval so the agent sees the right files, not the whole codebase: dependency graphs, AST/tree-sitter-based chunking, embeddings, recency signals, and hybrid retrieval.

Implement conversation summarisation and memory for long sessions, and address long-project degradation through codebase summaries and periodic consistency passes.

Own token budgeting and prompt caching strategy.


Prompt Engineering as a Discipline

Own the system prompt and per-task prompt variants (new feature, bug fix, styling change).

Maintain few-shot examples and enforce coding conventions, stack rules, and prohibited behaviours such as no hardcoded secrets and no whole-file rewrites.

Version prompts like code with changelogs and rollback capability.


Evaluation and Quality Measurement

Design and own the evaluation suite: representative test prompts run on every prompt and model change, scored on build success rate, instruction adherence, and output quality including LLM-as-judge and visual/screenshot checks where relevant.

Define regression gates that block quality-degrading changes from shipping.

Treat evals the way engineers treat automated testing: versioned, automated, and tracked over time.

This responsibility is non-negotiable at this level.


Model Strategy and Cost

Design model routing - cheap and fast models for classification and small edits, frontier models for complex generation.

Drive cost optimisation through prompt caching, diff-based edits over full-file rewrites, and tighter context selection.

Track cost per agent run and tokens per task; evaluate new model releases against the eval suite and lead migrations when results justify it.


Safety and Reliability of Agent Behaviour

Defend against prompt injection from user content and fetched web content.

Ensure secrets never appear in generated client code.

Define what the agent's tools may and may not do in collaboration with the platform team.

Contribute to output moderation and abuse-pattern awareness.


Mentorship and Engineering Standards

Run code reviews, define engineering conventions for AI work, and raise the engineering bar across the AI team.

Work closely with the Senior MLOps Engineer on handoff of eval design, prompt configurations, and model routing logic.


Requirements:


Hands-on Production Ownership of LLM-Powered Systems with Agent Architectures (Mandatory)

Must have personally shipped and operated at least one complex production AI system - agentic, multi-step, or code generation - with end-to-end ownership of architecture, evaluation, and cost.

POCs, internal demos, and tutorial-grade work do not qualify.


5+ Years of Professional Software or AI Engineering Experience

With at least 3 years focused on LLM applications, AI engineering, or production AI systems.

Candidates with strong backend backgrounds and a clear, substantive pivot into LLM systems qualify.


Strong Python Proficiency and Service Development

Production-grade Python with FastAPI or equivalent: type hints, async patterns, streaming responses, testing, and packaging.

Not notebook-only.


Depth Across LLM APIs and Agent Systems

Production experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or open-weight models (vLLM, Ollama, Together).

Production experience with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.

Hands-on with tool calling, structured outputs, and multi-step reasoning.


Demonstrated, Systematic Evaluation Practice - Non-Negotiable

Must have built evaluation harnesses that gate production releases, not ad-hoc testing.

Hands-on with at least one of LangSmith, Langfuse, Promptfoo, Ragas, or DeepEval.

Candidates with no systematic answer to evaluation should not be considered at senior level regardless of other strengths.


Cost Discipline for Production AI

Track record of measurable cost optimisation on production AI features.

Able to speak in specifics: cost per request, savings achieved through caching or model routing, context reduction decisions.


AWS Working Knowledge

Hands-on with EC2, S3, IAM, and Docker.

Comfort with CI/CD workflows and deploying AI services.


Awareness of LLM Security Failure Modes

Familiar with prompt injection patterns, understands that system prompt rules alone are insufficient, and has experience with output validation and content safety in production.


Nice to Have

  • Experience with AST/tree-sitter tooling, diff-based editing systems, or compiler-adjacent work
  • MCP server authoring
  • Open-source AI contributions
  • Published technical writing on LLM systems
  • Multi-modal model experience
  • Fine-tuning exposure (LoRA, QLoRA, PEFT)
Read more
company logo
Ritesh Kalvellu
Posted by Ritesh Kalvellu
Bengaluru (Bangalore)
7 - 15 yrs
₹50L - ₹50L / yr (ESOP available)
skill iconPython
Large Language Models (LLM)

About us

MyRico builds personal AI agents for enterprise, the copilots and digital employees that make humans more productive. The MyRico agents sit at the intersection of enterprise memory, high-end security, and an ever-expanding set of capabilities. We're a small team shipping fast, and the product is live with real customers today.


The role

Full-time · Bangalore

You'll own the systems that make an autonomous agent trustworthy in production. This is not just prompt engineering, and it's not model training - it's that and all the engineering layers in between: model steering, memory architecture, orchestration design, latency optimization, deterministic vs non-deterministic systems tradeoff, deployment, and operator tooling.

Concretely, the kind of work you'd have done here last week would include enhancing agent memory systems, runtime and scheduling reliability and predictability, adding new capability and tools to deployed agents, operator tools, live system debugging and benchmarking various models for price, latency and quality. And that is just last week. We are a small nimble startup rapidly working to address customer needs in this growing space, so things change rapidly.


What we're looking for

  • More than 7 years of software engineering, with real production ownership of distributed or stateful systems - you've been paged for something you built and made it not happen again.
  • Strong understanding of LLM based native app building, combining classic and model driven applications to get the best of both. You've built on LLMs beyond demos: agent frameworks, tool use, context management, eval fixtures, and you know why "it worked in the transcript" isn't evidence.
  • Python and shell in production settings; comfortable in TypeScript/Node. You write boring, testable code and prefer the standard library to a new dependency.
  • Systems taste: append-only logs, idempotent reconciliation, fold-the-events state machines, and read-only debugging surfaces feel like home.
  • Evidence discipline: tests before features, claims backed by quoted observations, decisions written down.


Nice to have

  • Experience running the combination of multi-tenant and single-tenant / on-prem-style fleets with ability to handle per-customer isolation, upgrade paths, migration compatibility in both setups.
  • Security instincts for products that touch highly sensitive data and systems, including things like executives' email, calendars, and messages: least privilege, loopback-only services, secrets that never hit a log.
  • You already orchestrate AI coding agents in your own workflow and have opinions about where they break.


How we work

Small team, high trust, written decisions. Designs get adversarial review before code; PRs get automated review driven to zero open findings; features aren't done until verified on a live system. AI agents do a large share of the implementation, your leverage is judgment: framing the problem, freezing the right design, and knowing when the machine is wrong.

Read more
company logo
icon

The recruiter has not been active on this job recently. You may apply but please expect a delayed response.

Remote only
3 - 20 yrs
$20K - $35K / yr
Agentic AI
AI Agents
Multi-agent Systems
Generative AI
skill iconReact.js
+7 more

Senior Agentic Engineer — Ai1 Platform (Operations)

Team: Operations · Type: Full-time · Location: Fully remote (global, async-first)


Who we are

MyZone AI builds and runs Ai1 — a managed, multi-agent AI platform that gives each client their own isolated, always-on team of AI agents working across the channels and tools they already use. We're a small, AI-first, high-trust team that ships like a much bigger one. We work fully remote and async — what you ship and how clearly you write matter more than your hours.

The role

You'll join the Operations team


  • Agentic engineering — complex agent skills and workflows, multi-agent orchestration, custom integrations, and the prompting and guardrails that make autonomous agents reliable.


What you'll do

Build agents that do real work


  • Design complex, reusable skills and multi-step recipes that automate entire business processes — production automations, not demos.
  • Architect multi-agent workflows: delegation, parallel execution, and agents that hand off and reassemble work.
  • Build custom MCP connectors and API integrations for any tool that doesn't have one yet.
  • Engineer robust prompts, human-in-the-loop controls, and built-in QA so automations are safe to run autonomously.
  • Ship automations as self-contained installers that deploy across many client instances with zero rework.


Own application code end-to-end


  • Own one or more services in our modular platform (Node/Express backends, React frontends) — schema, API, and UI.
  • Ship features across CRM, operations, pipeline, reporting, onboarding, assessments, and chatbots.
  • Contribute to our Next.js portal (Next.js 15, TypeScript, Prisma, PostgreSQL).
  • Keep your components healthy: clean migrations, green CI, solid tests.


Set the bar


  • Work clean: feature branches, real PRs, green CI, small changes, semantic versioning.
  • Be security-conscious by default — credentials are treated like credit-card numbers; high-risk actions route through human review.
  • Mentor, set patterns, and raise the team's technical level.

Who we're looking for

Must-haves


  • 5+ years of professional software engineering with strong full-stack depth: Node.js / TypeScript and modern React.
  • Hands-on experience building AI agents in production — tool/function calling, orchestration, memory, and evaluation. You've shipped something that uses an LLM to reliably do real work.
  • Fluency with the Anthropic / Claude API (or equivalent) and real depth in prompt/context engineering.
  • Strong with PostgreSQL and pragmatic API design.
  • Able to own a service end-to-end and ship independently in a low-oversight, async environment.
  • Excellent written communication and sound judgment around security and autonomous actions.


Bonus


  • Built MCP connectors, agent frameworks, or similar orchestration layers.
  • Next.js 15, Prisma, Vercel, monorepos (pnpm / Turborepo).
  • Light DevOps: Linux VPS, PM2, nginx, CI.
  • Used coding agents (Claude Code–style tools) to build real things.
  • Integrations: Slack/Teams apps, Twilio, Google Workspace, GitHub.

How we work

Fully remote, global, async-first — no required location. Small, high-trust, and low-bureaucracy, so you'll own real work and thrive without hand-holding. Everyone here builds automations, versions their work, and maintains what they ship.





MyZone AI is an equal-opportunity employer. We hire the best person for the work, wherever they are.


Read more
company logo
Robin Silverster
Posted by Robin Silverster
icon

The recruiter has not been active on this job recently. You may apply but please expect a delayed response.

Bengaluru (Bangalore)
7 - 11 yrs
₹10L - ₹36L / yr
SRE
Reliability engineering
Google Cloud Platform (GCP)

Location - Bangalore Skill/Experience Expectations: 1. Total Experience 7-11 yrs 2. 3-4 years in managing scalable production environment 3. 2-4 yr experience in managing Google cloud infrastructure 4. proficient in terraform and any programming language 5. Expert in designing and managing observability solutions 6. 5 yr experience in DevOps and SRE practices and troubleshooting critical incidents.

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos