Cutshort logo
For Employers
Its for a IT Service MNC logo
DevOps AI Engineer
Its for a IT Service MNC
DevOps AI Engineer

DevOps AI Engineer at Its for a IT Service MNC · Bengaluru (Bangalore), Mumbai, Chennai · 6 - 10 years · ₹10L - ₹30L / yr · Posted 20 Jul 2026

Freelancer's logo

DevOps AI Engineer

at Its for a IT Service MNC

Agency job
6 - 10 yrs
₹10L - ₹30L / yr
Bengaluru (Bangalore), Mumbai, Chennai
Skills
DevOps
skill icongrafana
Generative AI

Primary Skills: 

Observability - ELK (Elactic/Kibana), Prometheus, Grafana, PromQL

Software and automation - Java, Python/Shell/Bash, Rest-SOAP API, docker containerization, Kubernetes, Kafka

Reliability and DR engineering - Distributed architecture and distributed system fundamentals, micro services, and event-driven architecture.

Cross-team coordination, incident triage and resolution, leadership and stakeholder management.

 

Secondary Skills:

Lang-chain, Langraph, RAG, MCP

Experience with working on LLM's and integrating with the existing applications

Python - FastAPI

Cache - Redis

 

Program Details:

Design, build, and ship LLM-powered and agentic product features that enhance the team efforts and outcomes.

Build agentic AI systems that reason over context, invoke tools, take real actions, and recover gracefully from failure.

Work on integrating the existing AI tools and should know major AI frameworks and libraries. 

Own service reliability and operational governance by defining SLA's, managing error budgets, and reporting reliability (MTTD, MTTR) to leadership for prioritization, risk decisions and planning.

Architect and continuously optimize the observability of platform using Kibana/Elastic (ELF) along with other observability tools like Prometheus, Grafana (dashboards, metrics, alert lifecycle), improving detection quality, reducing noise/toil, and enabling faster triage and measurable uptime improvements.

Engineer advance alerting and automation capabilities with Kibana alerting and anomaly detections and integrating response workflows (routing, runbooks, remediation scripts) to standardize on-call execution and accelerate restoration of services.

Lead incident response for customer-impacting issues across teams-coordination, communications, service restoration, and blameless RCA-then corrective actions that prevent recurrence and reduce operational risk.

Design, automate and validate Disaster Recovery and failover for critical services/journeys (RTO/RPO alignment, DR Drills), ensuring resiliency under failure scenarios and improving recovery.

Consult and partner with application teams by providing production readiness inputs (Resiliency patterns, availability, performance/capacity considerations) and driving platform enhancements that improve stability while optimizing infrastructure and observability spend.


If you are interested for this role, kindly acknowledge this email with your interest.

Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

Similar jobs (10)

company logo
Priya Rawat
Posted by Priya Rawat
icon

The recruiter has not been active on this job recently. You may apply but please expect a delayed response.

Gurugram
4 - 5 yrs
₹8L - ₹10L / yr
RCA
SLA
skill icongrafana
ELKI
SOP

About the Role


We are looking for a proactive and detail-oriented Senior Site Reliability Engineer (SRE) to ensure the reliability, performance, and availability of our applications. The role involves monitoring production systems, troubleshooting issues, and collaborating with cross-functional teams to drive faster resolution and continuous improvement. You will play a key role in maintaining system stability and enhancing observability across our microservices-based platform.


Key Responsibilities


  • Handle MFS application issues by investigating, troubleshooting, and escalating to engineering teams when needed
  • Perform initial root cause analysis (RCA) and support resolution of recurring or moderately complex issues
  • Ensure timely incident resolution in line with SLAs, including proper documentation of fixes and workarounds
  • Identify and analyze system bottlenecks, and assist in deploying fixes via change management processes
  • Collaborate with cross-functional teams (Development, SRE/DevOps, QA, Business) to resolve incidents and improve systems
  • Use observability tools (Grafana, Loki, ELK) to monitor system health, availability, performance, and resiliency
  • Participate in incident/severity calls, ensuring clear communication and coordination
  • Develop and maintain knowledge bases, SOPs, and runbooks for standardized operations and troubleshooting


Required Skills & Experience


  • Strong understanding of Linux/Unix systems for application support
  • Hands-on experience troubleshooting applications in staging and production environments
  • Ability to monitor system performance and identify root causes using logs and metrics
  • Experience working with Kubernetes and microservices-based architectures
  • Proficiency in observability and monitoring tools such as Grafana, Loki, and ELK (Elasticsearch, Logstash, Kibana)
  • Familiarity with CI/CD practices and tools (e.g., Jenkins, GitOps)
  • Experience in API testing and validation using tools like Postman and Swagger/OpenAPI
  • Hands-on experience with PostgreSQL and MongoDB for troubleshooting and ad-hoc reporting
  • Experience with ticketing and documentation tools such as Jira and Confluence
  • Minimum 4+ years of experience in application support or reliability engineering


Education & Certifications


  • Bachelor's degree in Computer Science, Information Technology, or a related field
  • Relevant certifications (Cloud, Kubernetes, Microservices) are a plus


Work Schedule


  • Willingness to work in a 24x7 environment, including weekends and on-call rotations
Read more
Hiring for IT Consulting Firm (MNC)
Hiring for IT Consulting Firm (MNC)
Agency job
via by Sneha k
Pune, Nagpur
5 - 10 yrs
₹20L - ₹30L / yr
MLOps
DevOps
Artificial Intelligence (AI)
skill iconData Science
Data engineering
+5 more

Position Overview 

The AI Observability Engineer will be instrumental in implementation of scalable, cloud-native solutions to meet the growing needs of our Data & Development team. The successful candidate will demonstrate the ability to abstract complexity and create reusable, scalable patterns that accelerate development. The AI Observability Engineer will build and maintain a robust framework to ensure the reliability and maintainability of DPR Construction's complex AI systems. 

 

Responsibilities 

  • Standardize observability practices across AI/ML and other development teams including logging, metrics, tracing, and model performance monitoring, ingesting data from multiple platforms 
  • Lead hands-on implementation of automation-first DevOps and MLOps practices, enabling infrastructure-as-code and consistent, repeatable environment provisioning 
  • Design and manage intelligent DataOps pipelines with automated data quality monitoring and anomaly detection 
  • Deploy, maintain and monitor containerized ML workloads 
  • Extend existing CI/CD pipelines to support automated infrastructure changes and ML workflows 
  • Implement AI-driven data validation, schema and concept drift detection and metadata management. 
  • Establish governance frameworks for AI systems, including bias detection, explainability, and auditability 
  • Extend existing Azure RBAC strategy by automating role and permission management to reduce manual intervention 
  • Develop automated test suites for model performance, regression, edge cases and bias validation 
  • Monitor model KPIs (accuracy, precision, recall, latency, calibration) 
  • Ensure reproducability of experiments and production models 
  • Act as a technical point of contact for DevOps and MLOps practices, developing reusable patterns, documentation, and proof-of-concepts to drive adoption 

Qualifications 

  • Bachelor’s degree in computer science, Data Science, Information Systems, or a related field 
  • 5+ years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering 
  • Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure 
  • Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn 
Read more
company logo
Taher Ujjainwala
Posted by Taher Ujjainwala
Pune
12 - 25 yrs
₹70L - ₹120L / yr (ESOP available)
skill iconPython
skill iconKubernetes
Google Cloud Platform (GCP)
skill iconAmazon Web Services (AWS)
Windows Azure
+15 more

About the Role

We are hiring Staff / Principal Engineers to take full, hands-on ownership of Blitzy's most critical production-grade systems and to deliver high-leverage features that materially improve customer outcomes and engineering velocity. This is the most senior individual contributor role at the company today.

This is not a Senior-plus role, an architecture-only role, or a promotion-track role. We are looking for someone who has already operated at Principal / Staff+ scope in a highly technical environment and expects to spend their time writing, reviewing, and shipping production code.

This role is 100% hands-on. Leverage comes from system ownership, execution quality, and durable technical decisions — not people management or process.


Responsibilities

  • Own mission-critical production systems end-to-end, ensuring correctness, scalability, performance, reliability, and operational excellence.
  • Design, build, and ship high-impact backend systems and features that improve product reliability, performance, and customer value.
  • Architect scalable services and cloud infrastructure using technologies such as Python, REST, gRPC, Kubernetes, and Terraform.
  • Identify and resolve complex technical bottlenecks that limit engineering quality, system performance, or organizational velocity.
  • Build and operate LLM-powered systems and validation loops that evaluate correctness, consistency, durability, and production performance.
  • Design and evolve data architectures incorporating relational, NoSQL, graph, and vector databases to support complex enterprise applications and semantic retrieval.
  • Modernize and improve complex enterprise systems while balancing reliability, maintainability, scalability, and delivery speed.
  • Set and uphold engineering quality standards through hands-on technical leadership, sound technical judgment, and ownership of long-term technical decisions.


Qualifications

  • Direct experience with Python as a primary programming language, backend frameworks, and microservices architectures.
  • Expertise in REST and gRPC, with proficiency in Node.js and JavaScript.
  • Proficiency in GCP, along with experience using at least one additional cloud platform such as AWS or Azure.
  • Advanced knowledge of Kubernetes and Terraform in production environments.
  • Experience operating highly available production systems, including monitoring, scalability, reliability, performance optimization, and operational tooling.
  • Strong knowledge of SQL and NoSQL databases, including PostgreSQL, MySQL, MongoDB, Cassandra, or DynamoDB.
  • Familiarity with graph databases such as Neo4j and vector databases or embedding infrastructure for semantic search and retrieval.
  • Hands-on experience building and operating LLM-powered systems in production, including evaluation, validation, regression testing, tracing, and failure analysis.
  • Working knowledge of LangSmith or comparable LLM observability and evaluation tools; familiarity with OpenAI, Anthropic, or similar model providers is a plus.
  • Ability to contribute across the full stack, with a strong understanding of frontend architecture and the ability to debug, design, and ship across frontend, backend, infrastructure, and AI systems.
  • Understanding of large-scale enterprise software systems, including architecture, integration, deployment, modernization, and long-term maintainability.
  • Proven track record of operating at Staff+, Principal Engineer, or equivalent level, independently driving complex technical initiatives and delivering high-impact outcomes with minimal supervision.


Blitzy is a Cambridge, MA based AI software development platform on a mission to revolutionize the software development life cycle by autonomously building custom software to unlock the next industrial revolution. We're transforming how enterprises build software, turning enterprise requirements into enterprise grade code with an agentic software development platform that can autonomously execute 80% of the quantum of software development work. We're backed by multiple tier 1 investors, and have proven success as founders of previous start-ups.


Our Culture

Who we are:

Led by two pioneering co-founders we are one of the fastest growing companies in the U.S., creating our own category of enterprise autonomous software development. We automate thousands of hours of software development for our customers, which includes strong representation within the Fortune 500.


How we work:

  • We move Blitzy Fast: Time is both our company’s and our clients’ most precious asset. We move quickly and decisively to innovate internally and deliver exceptional software externally.
  • Championship Mindset: We operate like a professional sports team. We win as a team by holding ourselves and each other to high standards, collaborating in-person, and remaining focused on the mission.
  • Passion for Invention: We’re pushing the frontier of what’s possible, requiring constant innovation and iteration.
  • We Work for the Customer: We focus on delivering outsized value to the customers we work with and expanding those relationships into deep, meaningful partnerships.
  • We believe in being ‘everyday athletes’: taking care of ourselves so we can bring our best minds to work. We promote great sleep, movement, and restorative activities for 


Blitzy is an equal opportunity employer committed to building a diverse and inclusive team. We believe different perspectives make us stronger.

Read more
company logo
John Vivek
Posted by John Vivek
icon

The recruiter has not been active on this job recently. You may apply but please expect a delayed response.

Remote only
8 - 25 yrs
₹10L - ₹70L / yr
skill iconJava
Artificial Intelligence (AI)
Generative AI
skill iconSpring Boot
Hibernate (Java)
+1 more

Staff Engineer - AI:

Location : India, Remote

 

Job Description

Egnyte is seeking an experienced Staff Software Engineer to join our Engineering department. The Engineering department builds large distributed components and services that run Egnyte's Cloud Platform. Our code serves billions of requests per day with sub-second latency in a fault-tolerant environment. We process and analyze millions of files and events daily. Some of the responsibilities for this department include Egnyte's Cloud File System, Content Classification, Content Lifecycle Management, User Behavior Analysis, Object Store, Metadata Stores, Search Systems, Recommendations Systems, Synchronization, and intelligent caching of multi-petabyte datasets. We are looking for candidates with a shared passion for building large-scale distributed systems and a keen sense for tackling complexities that come with scaling through multiple orders of magnitude.

In this role, you will (But are not limited to):

  • Design and develop highly scalable and resilient cloud architecture that seamlessly integrates with on-premises systems
  • Drive the team’s goals and technical direction to find and pursue technical opportunities that make Egnyte’s cloud platform more efficient
  • Effectively communicate complex design and architecture details
  • Understand company and industry-wide trends to help develop new technologies
  • Conceptualize, develop, and implement changes that prevent key systems from becoming unreliable, under-utilized, or unsupported
  • Own all aspects of critical software projects from design to implementation, QA, deployment, and monitoring

Qualifications

  • BS, MS, or PhD. in Computer Science or related technical field, or equivalent practical experience
  • 8-15 years of professional experience in engineering with a history of technical innovation
  • Experience providing technical leadership to engineers

Bonus Qualifications (Good to Have)

  • The breadth of knowledge across infrastructure domains, with the ability to reason about everything from data center machine software to database solutions to machine learning infrastructure to front-end web or mobile applications
  • Demonstrated success in designing and developing large-scale, complex systems
  • Expertise with multi-tenant, highly complex, cloud solutions; experience with Hybrid and/or on-premises solutions desired

 

About Egnyte

In a content critical age, Egnyte fuels business growth by enabling content-rich business processes, while also providing organizations with visibility and control over their content assets. Egnyte’s cloud-native content services platform leverages the industry’s leading content intelligence engine to deliver a simple, secure, and vendor-neutral foundation for managing enterprise content across business applications and storage repositories. More than 16,000 customers trust Egnyte to enhance employee productivity, automate data management, and reduce file-sharing cost and complexity. Investors include Google Ventures, Kleiner Perkins, Caufield & Byers, and Goldman Sachs. For more information, visit www.egnyte.com

Read more
MNC
MNC
Agency job
via by Sandhiya b
icon

The recruiter has not been active on this job recently. You may apply but please expect a delayed response.

Bengaluru (Bangalore), Hyderabad
8 - 15 yrs
₹2L - ₹24L / yr
Observability
AppDynamics
Splunk
skill iconPython
Scripting

Observability Engineer (AppDynamics)

Hyderabad

Exp: 8+years of exp

Mandate Skills: AppDynamics, Splunk, Python (Scripting knowledge)

Primary Skill Set:

  • AppDynamics administration
  • SPLOC administration
  • Enterprise monitoring and observability
  • Application performance monitoring
  • Alerting and event management
  • Monitoring strategy and design
  • Platform configuration and governance
  • Python scripting
  • Monitoring automation

 

Secondary Skills:

  • Glassbox monitoring
  • Customer journey observability
  • GenAI concepts for operations       
  • Log, metric, and trace telemetry
  • Dashboarding and visualization


Read more
company logo
Swathi S
Posted by Swathi S
Chennai
7 - 12 yrs
₹30L - ₹55L / yr
skill iconAmazon Web Services (AWS)
skill iconPython
CI/CD
DevOps
Platform as a Service (PaaS)
+7 more

Amura’s Vision 


We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.


Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.


Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.


These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.

We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence. 


Role Overview 


We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.


This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.


You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability. 


Key Responsibilities 


Cloud Infrastructure & Platform Engineering (AWS) 

  • Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
  • Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
  • Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
  • Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
  • Build reusable platform templates and shared infrastructure modules. 


AI/ML Infrastructure & MLOps 

  • Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
  • Support GPU-based workloads and optimize compute/storage usage.
  • Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
  • Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines. 


CI/CD, Automation & Developer Productivity 

  • Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
  • Automate deployments, environment provisioning, and release workflows.
  • Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
  • Implement automated patching, scaling, backups, cleanup workflows, and drift detection. 


Containers, Kubernetes & Platform Reliability

  • Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
  • Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
  • Optimize infrastructure for performance, resilience, and cost-efficiency.
  • Implement progressive deployment strategies including blue/green, canary, and rolling deployments. 


Observability, Incident Response & SRE Practices

  • Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
  • Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
  • Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.

FinOps, Cost Governance & Security

  • Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
  • Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
  • Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.

Collaboration, Leadership & Platform Culture

  • Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
  • Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.

Skills & Qualifications


Must-Have:

  • 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
  • Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
  • Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations. 


Preferred / Nice-to-Have:

  • Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
  • Familiarity with Kafka, Redis, SQS, and event-driven systems.
  • Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
  • AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations. 


Here are answers to some questions you may have

Where is your office?

Chennai (Velachery)

Work Model

Work from Office – because great stories are built in person!

Do you have an online presence?

https://amura.ai (we are @AmuraHealth on all social media)


Read more
company logo
Agency job
via by aarushi Mahajan
Bengaluru (Bangalore)
5 - 8 yrs
₹18L - ₹20L / yr
skill iconJava
skill iconSpring Boot
Spring Cloud
Microservices
Kafka
+6 more

We are looking for a Senior Software Engineer who will be driving architectural direction, multi-team initiatives, and the development of highly scalable distributed systems. You will partner with Product, Architecture, DevOps and cross-functional engineering teams to deliver impactful technology solutions across a complex digital ecosystem. You will independently define and deliver large-scale, multi-phase projects, mentor engineers, and shape technical strategy. The role also includes leveraging LLMs and GenAI-driven tools to enhance engineering productivity, reduce cycle time, and build the next generation of internal developer experience tools.


Technical Leadership • Drive technical strategy and architecture for complex distributed systems and core platform components. • Own end-to-end delivery of high-impact, multi-team engineering initiatives. • Translate ambiguous business requirements into clear, pragmatic technical solutions. Public • Define technical roadmaps, sequence projects into phases, and ensure timely execution across teams.


Software Engineering & System Design • Architect, develop, and optimize microservices using Java, Spring Boot, Spring Cloud. • Design scalable, resilient, secure cloud-native applications using Azure/AWS. • Lead system design reviews, architecture discussions, and technical deep dives. • Solve complex cross-system issues through deep debugging, performance tuning, and root cause analysis. • Build frameworks and shared libraries to accelerate development across squads.

AI & Engineering Productivity • Champion adoption of AI-assisted development practices and LLM/GenAI tools across engineering teams to enhance workflows, improve cycle time, and automate routine tasks. • Define clear guidelines, training, and standards for effective and secure use of AI tools, including prompt engineering, code generation, and review expectations. •


Integrate AI-driven capabilities such as code generation, unit test creation, log analysis, documentation summarization, and code review assistance into the development workflow. • Establish and enforce engineering best practices for AI-generated code, including human-in-the-loop reviews, coding standards, test automation, secure coding, and design-first development. • Measure and optimize engineering productivity using AI-related metrics (cycle time, PR throughput, AI-code adoption, defect escape rate, developer satisfaction) and continuously identify opportunities to reduce bottlenecks with AI.

Collaboration & Cross-Functional Impact • Work closely with Product Owners, Architects, QE, DevOps, and peer engineering teams.

• Communicate technical trade-offs, risks, and decisions clearly to non-technical stakeholders. • Coordinate across teams to resolve dependencies and align on architectural standards.

• Influence technical direction beyond your immediate team. Team Leadership & Mentoring

• Mentor engineers across multiple teams, guiding them in design, coding, debugging, and best practices.

• Participate actively in hiring and talent development.

• Raise the engineering bar through code reviews, design critiques, and technical leadership.

• Help define and uphold engineering best practices, coding standards, and architectural principles. Ownership & Operational Excellence

• Take full ownership of systems from development to deployment, monitoring, and maintenance.

• Ensure engineering deliverables meet performance, reliability, and security criteria.

• Proactively identify areas for improvement and propose scalable, maintainable solutions.

• Lead incident response, post-mortems, and continuous improvement initiatives.  


Main Skills

Experience applying LLMs/GenAI to improve engineering productivity and developer workflows.

• Prompt engineering experience and familiarity with integrating models into applications.

• Experience with Domain-Driven Design, event-driven systems, CQRS, or enterprise integration patterns.

• Familiarity with Linux systems and container-based deployment architectures.

• Experience with application servers such as Tomcat.

• Contributions to internal platforms or open-source projects.

• Strong leadership qualities with the ability to influence without authority.

• High levels of self-motivation, ownership, and drive.

• Critical thinking abilities and a strong problem-solving mindset.

• Excellent team collaboration skills and a positive, proactive attitude.

Deep expertise in Java, J2EE, Spring, Spring Boot, Spring Cloud, and building microservices-based architectures. 

Hands-on experience with Docker, Kubernetes / OpenShift, and CI/CD pipelines.

• Strong knowledge of working with databases: Oracle, MySQL, MongoDB, Cassandra, etc.

• Experience designing and scaling systems on Azure, AWS, or Google Cloud.

• Excellent problem-solving skills, debugging capabilities, and system-level thinking. 

Read more
Remote only
8 - 12 yrs
Best in industry
Terraform
Artificial Intelligence (AI)
IAC
skill iconAmazon Web Services (AWS)
ECS
+6 more


Senior Platform & Site Reliability Engineer

Location: Remote Employment Type: Contract

The Role

This role carries full architectural and operational ownership of the platform layer across a growing SaaS portfolio. The Cloud Architect owns AWS infrastructure standards — VPCs, account structures, networking, and compute design. Everything outside that lane is yours: the CI/CD platform, the observability and reliability stack, the event streaming infrastructure, the deployment pipelines, and the incident engineering model.

Architectural decisions are yours to make and defend, standards are yours to define and enforce, and the reliability of 20+ enterprise SaaS products depends on what you and your team build.

This is an AI-native engineering organisation. Where it is practical and safe to do so, you are expected to use automation and AI-assisted tooling to reduce toil — in CI/CD triage, infrastructure provisioning, observability workflows, and acquisition onboarding. The expectation is not to replace engineering judgement with automation, but to free it up for the problems that genuinely require it.

The Scale You Will Operate At

The portfolio consists of 20+ live, enterprise-grade SaaS solutions running concurrently. Each product serves enterprise customers and processes millions to billions of real-time requests. The architecture is serious: event streaming for real-time data pipelines, batch processing workloads running alongside live transaction flows, and multi-tenant enterprise-grade reliability expectations across every product.

You will design and operate the platform infrastructure that underpins all of it — scaling horizontally as each new acquisition joins the portfolio, without proportionally scaling cost, complexity, or headcount.

What You Will Own

Platform Architecture

  • Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
  • Define, build, and enforce platform standards across portfolio products
  • Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
  • Self-service developer platform so product teams ship without waiting on platform

Event Streaming & Pipeline Infrastructure

  • Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
  • Design and maintain batch processing infrastructure alongside live event flows
  • Ensure pipeline reliability, throughput, and cost are actively managed at scale

CI/CD & Deployment

  • Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
  • Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
  • Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified

Observability & Reliability

  • Own the full observability stack: Grafana, Prometheus, and Loki across all products
  • SLOs and error budgets defined per product; reliability tracked consistently
  • Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen
  • Incident response: on-call design, escalation playbooks, post-mortem facilitation
  • Automated remediation scoped to safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns; novel or ambiguous failures escalate to a human with full context attached

Acquisition Onboarding

  • Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
  • Migration plan and execution for each portfolio company joining the platform — sequenced to avoid disrupting live operations
  • Target: full platform integration within a defined window per acquisition

A Note on Automation

Where automation is safe and failure modes are well understood — routine provisioning, known CI/CD failure classes, secrets rotation, cost anomaly flagging — aggressive automation is expected. Where automation would act on ambiguous signals or carry significant blast radius, human judgement stays in the loop. The goal is to reduce toil on solved problems, not to automate decisions that require engineering expertise.

Platform Stack

Area Stack / Standard IaC Terraform OSS / OpenTofu CI/CD GitHub Actions Event Streaming Architecture and tooling chosen for the workload Observability Grafana, Prometheus, Loki Log Management AWS CloudWatch, Grafana Loki Incident Management OpsGenie (startup tier) or Better Uptime Secrets AWS Secrets Manager / HashiCorp Vault OSS Containers ECS (default), EKS only where justified Cost Monitoring AWS Cost Explorer with custom dashboards What We’re Looking For

  • 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
  • Strong Terraform depth across multi-environment, multi-account setups
  • CI/CD ownership across a multi-product environment with GitHub Actions
  • Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
  • Hands-on Grafana, Prometheus, and Loki in production
  • AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
  • SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture
  • Acquisition or greenfield platform integration experience strongly preferred

How You Work

  • Comfortable operating across multiple products simultaneously — context-switching without dropping standards
  • Cost-efficiency instinct — you optimise spend as a habit, not as a project
  • You treat automation as a tool for eliminating toil, not a substitute for engineering judgement
  • You document decisions, enforce standards through code, and build platforms that other engineers find intuitive to use

Why This Role

The platform function is being built from the ground up. You will have architectural ownership of the entire non-AWS platform layer across a growing portfolio of enterprise SaaS products, with the freedom — and responsibility — to build the reliability and delivery culture of the organisation.

This is not a role that inherits someone else’s decisions and maintains them. Every major architectural choice is still to be made. If you want to build something that lasts and that other engineers depend on, this is the role.

Read more
company logo
Srikanth Bajgur
Posted by Srikanth Bajgur
Bengaluru (Bangalore)
6 - 12 yrs
Best in industry
skill iconJava
Spark
skill iconSpring Boot
RESTful APIs
skill iconReact.js

Join our product development team at Planview as a Senior Software Engineer I and become a pivotal force on the Viz Core Team. This role offers the unique opportunity to shape and lead the development of data-processing pipelines and APIs that are at the heart of our software solutions. These solutions are designed to streamline and enhance the efficiency of software delivery, resonating deeply with software engineers who strive to build better and faster.



At Planview Viz, which is powered by the innovative “Flow Framework,” you will tackle complex data challenges and develop scalable solutions within an AWS cloud environment. Your efforts will be crucial in revolutionizing how businesses harness and interpret vast amounts of workflow data, transforming it into actionable insights that propel organizational efficiency and effectiveness.



Responsibilities (What you'll do)


  • Create and refine powerful data-processing architectures that integrate seamlessly with a diverse array of external tools, enhancing the way software is delivered across industries.
  • Drive operational excellence by critically analyzing problems, defining requirements, and devising robust solutions that push the boundaries of technology.
  • Lead and inspire a team of talented engineers, promoting a culture of ownership, meticulous attention to quality, and proactive problem-solving.
  • Stay at the cutting edge of technology by updating and expanding your team’s knowledge of cloud architectures, data processing, and advanced analytics, ensuring that you and your team remain leaders in technological innovation.
  • Make a direct impact on the efficiency and effectiveness of software delivery worldwide through innovation and leadership.


Qualifications (What you'll bring)


Who We’re Looking For

The ideal candidate is a seasoned professional in cloud software development with a strong foundation in data processing and scalable software systems. You thrive in collaborative environments and are passionate about advancing cloud technology and data architecture to new heights. You are a leader who enjoys mentoring, guiding, and inspiring others.


Preferred Qualifications


Skills, Knowledge, and Expertise


  • A degree in Computer Science, Engineering, or a related field.
  • 6+ years of experience with a modern programming language, with a focus on back-end systems and cloud-based technologies.
  • Strong capability in architecting and designing robust, scalable software systems.
  • Proficiency in AWS or another popular cloud platform.


Additional Qualifications


  • Experience with big-data technologies such as Spark and Apache Kafka.
  • Experience with Java (version 17), Scala, or other JVM programming languages.
  • Experience with tools such as Amazon Redshift and MongoDB.
  • Experience with continuous integration and continuous deployment tools such as GitHub, Jenkins, Travis CI, or similar platforms.
  • Proficiency in test-driven development.
  • Experience with containers and orchestration tools such as Kubernetes.
  • Experience working on remote or distributed teams and projects.

 

Read more
company logo
Agency job
via by Akash Bhatt
Bengaluru (Bangalore), Chennai, Mumbai, Hyderabad, Pune, Gurugram
3 - 10 yrs
₹12L - ₹35L / yr
Linux/Unix
skill iconKubernetes
Monitoring
skill iconDocker
skill iconAmazon Web Services (AWS)
+4 more



We're looking for a Site Reliability Engineer to keep our production systems fast, reliable, and scalable. Sitting at the intersection of software engineering and operations, you'll treat infrastructure as code, automate away toil, and build the observability that lets us catch problems before customers do. You'll own uptime and on-call for critical services, lead incident response and blameless postmortems, and continuously harden the platform against failure. This role suits an engineer who is as comfortable debugging a production incident at 2 a.m. as they are writing the automation that prevents the next one.



Key Responsibilities

  • Own reliability, availability, and performance of production services, including on-call rotation
  • Build and maintain monitoring, alerting, and observability (metrics, logs, traces)
  • Automate deployments, scaling, and operational tasks to reduce manual toil
  • Manage containerized workloads on Kubernetes and cloud infrastructure
  • Design and maintain CI/CD pipelines for safe, frequent releases
  • Lead incident response and drive blameless postmortems with clear follow-ups
  • Perform capacity planning, performance tuning, and cost optimization
  • Define and track SLIs/SLOs and error budgets with product teams


Requirements

  • 3+ years in SRE, DevOps, or production-focused engineering
  • Strong Linux administration and hands-on Kubernetes experience
  • Solid experience with monitoring/observability tools (Prometheus, Grafana, ELK, or similar)
  • Cloud experience with AWS, GCP, or Azure
  • CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation)
  • Proficient scripting in Python and/or Bash


Nice to have

  • Experience with service meshes, Helm, or GitOps (ArgoCD/Flux)
  • Background in high-traffic or distributed systems
Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos