Cutshort logo
For Employers
It is an Product Based Company(Domain- EV Charging) logo
TechOps Engineer- SRE
It is an Product Based Company(Domain- EV Charging)
TechOps Engineer- SRE

TechOps Engineer- SRE at It is an Product Based Company(Domain- EV Charging) · Bengaluru (Bangalore) · 1 - 2 years · ₹8L - ₹10L / yr · Posted 17 Sep 2026

Unique Occupational's logo

TechOps Engineer- SRE

at It is an Product Based Company(Domain- EV Charging)

Agency job
1 - 2 yrs
₹8L - ₹10L / yr
Bengaluru (Bangalore)
Skills
SRE
AWS
Monitorng Tools
Incident Response
Root cause analysis

Job Title: TechOps Engineer

Location: Bengaluru, India (Hybrid)

Employment Type: Full-time

Experience: 1-2 years

Shift Timing: 2 PM to 11 PM IST


Role Overview

We are excited to find a highly engaged engineer who is obsessed with technology that wants to be a part of a “world class” platform SRE team. Engineers must possess an "automation first" mindset, with a relentless focus on documentation, quality, scalability, and reliability using Infrastructure as Code tools. This position will be part of a platform team that is developing exciting products and solutions and playing a key part in driving forward the electrification of transportation.

What you’ll do:  

  • Ensure system reliability, uptime, and performance of global platform.
  • Conduct real-time surveillance of our EV charging systems to proactively identify and mitigate performance issues and anomalies near 24/7 basis. As such, you collaborate with IDT and FMC players to ensure incident detection also happens outside office hours (monitoring shifts among team members subject to duty schedule) 
  • Deliver on change & releases like firmware changes and drive insights & intelligence back into testing processes and tech discussions with the wider organization. 
  •  Successfully deliver and project manage first time right commissioning activities alongside our Engineering Procurement Contract Management (EPCM) partners to successfully bring charge points onto our Charge Point Management System (CPMS).
  • End-to-end EV charger lifecycle management, including deployment, commissioning, monitoring, maintenance, and decommissioning activities.
  • Provide technical guidance and support to DC specialists during the commissioning of EV charging solutions.
  •  Work closely with Shell, Engineering, and IT colleagues to ensure projects are completed on time and to specification.
  • Act as a liaison with the Engineering Procurement Contract Management (EPCM) partner to manage projects from start to finish, ensuring charge points are successfully onboarded on the Charge Point Management System (CPMS).
  • Collaborate with development, operations and support  teams to build scalable and resilient systems.
  • Contribute to incident response, root-cause analysis, and post-mortem reviews, driving continuous improvement.
  • Participate in capacity planning, performance tuning, and resource optimization.
  • Integrate security and compliance best practices into all infrastructure operations.
  • Stay current with emerging SRE tools, frameworks, and cloud technologies to continuously improve reliability practices.
  • Participate in and lead on-call rotations and incident response, conducting detailed postmortems and RCA reports.
  • Flexible to resolve blocking issues during off hours or weekends if required.  

 

What We’re Looking For: 

Basic Qualifications and skills

  • Bachelor’s degree in Engineering , Electrical, ECE, Computer Science, Information Technology, or related field.
  • Overall 1+ years of experience as a Site Reliability Engineer, DevOps/ Technical project coordinator role.
  • Proven experience of DevOps, SRE or Technical Project Coordination with IoT or connected devices based platforms.
  • Experience with incident management and on-call best practices. Provide support to on call engineers.
  • Excellent analytical and problem-solving skills with a proactive mindset. 
  • Expertise with monitoring and observability tools (Dynatrace,Prometheus, Grafana, Zabbix, etc.).
  • Proactively monitor the network, triage performance outliers, and coordinate correction actions to ensure optimal system functionality.
  • Fluency in English (spoken and written). 
  • Successfully recommission or decommission chargers following changes in our network.
  • Responsible for the go-live of the chargers on Shell’s public network following commissioning attempts.

 Note: This role involves managing infrastructure for a global platform operating in over ten countries, requiring effective communication and collaboration across regions. Strong verbal and written communication skills, along with availability and flexibility to resolve blocking issues, are essential to support On-call Engineers. This role may involve EU or US time‑zone shifts based on business requirements. The shift timing will be 2 PM IST to 11 PM IST.  

What We Offer

  • Work with some of the brightest minds in the emerging EV industry.
  • Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
  • Freedom to suggest, implement, and innovate on systems, processes, and technologies.
  • Daily ownership in a high-growth, challenging environment.
  • Flexible work environment with hybrid schedules and virtualization options.
  • Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.



Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

Similar jobs (10)

It is an Product Based Company(Domain- EV Charging)
It is an Product Based Company(Domain- EV Charging)
Agency job
via by Mantasha Naaz
Bengaluru (Bangalore)
0.6 - 2 yrs
₹8L - ₹10L / yr
Site Reliability Engineer
Reliability engineering
skill iconAmazon Web Services (AWS)
Terraform
Ansible
+3 more

Job Title: TechOps Engineer

Location: Bengaluru, India (Hybrid)

Employment Type: Full-time

Experience: 6 Month-2 years (Excluding Internship)

Shift Timing: 2 PM to 11 PM IST


Role Overview

We are excited to find a highly engaged engineer who is obsessed with technology that wants to be a part of a “world class” platform SRE team. Engineers must possess an "automation first" mindset, with a relentless focus on documentation, quality, scalability, and reliability using Infrastructure as Code tools. This position will be part of a platform team that is developing exciting products and solutions and playing a key part in driving forward the electrification of transportation.

What you’ll do:  

  • Ensure system reliability, uptime, and performance of global platform.
  • Conduct real-time surveillance of our EV charging systems to proactively identify and mitigate performance issues and anomalies near 24/7 basis. As such, you collaborate with IDT and FMC players to ensure incident detection also happens outside office hours (monitoring shifts among team members subject to duty schedule) 
  • Deliver on change & releases like firmware changes and drive insights & intelligence back into testing processes and tech discussions with the wider organization. 
  •  Successfully deliver and project manage first time right commissioning activities alongside our Engineering Procurement Contract Management (EPCM) partners to successfully bring charge points onto our Charge Point Management System (CPMS).
  • End-to-end EV charger lifecycle management, including deployment, commissioning, monitoring, maintenance, and decommissioning activities.
  • Provide technical guidance and support to DC specialists during the commissioning of EV charging solutions.
  •  Work closely with Shell, Engineering, and IT colleagues to ensure projects are completed on time and to specification.
  • Act as a liaison with the Engineering Procurement Contract Management (EPCM) partner to manage projects from start to finish, ensuring charge points are successfully onboarded on the Charge Point Management System (CPMS).
  • Collaborate with development, operations and support  teams to build scalable and resilient systems.
  • Contribute to incident response, root-cause analysis, and post-mortem reviews, driving continuous improvement.
  • Participate in capacity planning, performance tuning, and resource optimization.
  • Integrate security and compliance best practices into all infrastructure operations.
  • Stay current with emerging SRE tools, frameworks, and cloud technologies to continuously improve reliability practices.
  • Participate in and lead on-call rotations and incident response, conducting detailed postmortems and RCA reports.
  • Flexible to resolve blocking issues during off hours or weekends if required.  

 

What We’re Looking For: 

Basic Qualifications and skills

  • Bachelor’s degree in Engineering , Electrical, ECE, Computer Science, Information Technology, or related field.
  • Overall 1 years of experience as a Site Reliability Engineer, Technical project coordinator role.
  • Proven experience of SRE or Technical Project Coordination with IoT or connected devices based platforms.
  • Experience with incident management and on-call best practices. Provide support to on call engineers.
  • Excellent analytical and problem-solving skills with a proactive mindset. 
  • Hands-on experience with AWS Cloud and IaC tools such as Terraform or Ansible.
  • Expertise with monitoring and observability tools (Dynatrace,Prometheus, Grafana, Zabbix, etc.).
  • Proactively monitor the network, triage performance outliers, and coordinate correction actions to ensure optimal system functionality.
  • Fluency in English (spoken and written). 
  • Successfully recommission or decommission chargers following changes in our network.
  • Responsible for the go-live of the chargers on Shell’s public network following commissioning attempts.

 Note: This role involves managing infrastructure for a global platform operating in over ten countries, requiring effective communication and collaboration across regions. Strong verbal and written communication skills, along with availability and flexibility to resolve blocking issues, are essential to support On-call Engineers. This role may involve EU or US time‑zone shifts based on business requirements. The shift timing will be 2 PM IST to 11 PM IST.  


What We Offer

  • Work with some of the brightest minds in the emerging EV industry.
  • Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
  • Freedom to suggest, implement, and innovate on systems, processes, and technologies.
  • Daily ownership in a high-growth, challenging environment.
  • Flexible work environment with hybrid schedules and virtualization options.
  • Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.


Read more
It is an Product Based Company(Domain- EV Charging)
It is an Product Based Company(Domain- EV Charging)
Agency job
via by Mantasha Naaz
Bengaluru (Bangalore)
2 - 4 yrs
₹10L - ₹14L / yr
SRE
Reliability engineering
Incident management
24/7

Platform Engineer

Location: Bengaluru, India (Hybrid)

Employment Type: Full-time

Experience: 2-4 years

About Compnay

This is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.

Operating over 70,000 charge points globally, this is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.


Role Overview

What you’ll do:

  • Ensure system reliability, uptime, and performance of global platform.
  • Conduct real-time surveillance of our EV charging systems to proactively identify and mitigate performance issues and anomalies near 24/7 basis. As such, you collaborate with IDT and FMC players to ensure incident detection also happens outside office hours (monitoring shifts among team members subject to duty schedule).
  • Deliver on change & releases like firmware changes and drive insights & intelligence back into testing processes and tech discussions with the wider organization.
  • Successfully deliver and project manage first time right commissioning activities alongside our Engineering Procurement Contract Management (EPCM) partners to successfully bring charge points onto our Charge Point Management System (CPMS).
  • Provide technical guidance and support to DC specialists during the commissioning of EV charging solutions.
  • Work closely with Shell, Engineering, and IT colleagues to ensure projects are completed on time and to specification.
  • Act as a liaison with the Engineering Procurement Contract Management (EPCM) partner to manage projects from start to finish, ensuring charge points are successfully onboarded on the Charge Point Management System (CPMS).
  • Collaborate with development, operations and support teams to build scalable and resilient systems.
  • Contribute to incident response, root-cause analysis, and post-mortem reviews, driving continuous improvement.
  • Participate in capacity planning, performance tuning, and resource optimization.
  • Integrate security and compliance best practices into all infrastructure operations.
  • Stay current with emerging SRE tools, frameworks, and cloud technologies to continuously improve reliability practices.
  • Participate in and lead on-call rotations and incident response, conducting detailed postmortems and RCA reports.
  • Flexible to resolve blocking issues during off hours or weekends if required.

What We’re Looking For:

Basic Qualifications and Skills

  • Bachelor’s degree in Engineering, Electrical, ECE, Computer Science, Information Technology, or related field.
  • 2–4 years of overall experience with at least 1+ years of experience as a Site Reliability Engineer, DevOps Engineer, or Technical Project Coordinator.
  • Proven experience of DevOps, SRE or Technical Project Coordination with IoT or connected devices-based platforms.
  • Experience with incident management and on-call best practices. Provide support to on-call engineers.
  • Excellent analytical and problem-solving skills with a proactive mindset.
  • Expertise with monitoring and observability tools (Dynatrace, Prometheus, Grafana, Zabbix, etc.).
  • Solid understanding of cloud platforms (AWS) and AWS native services (EKS, EC2, S3, RDS, Lambda).
  • Proactively monitor the network, triage performance outliers, and coordinate correction actions to ensure optimal system functionality.
  • Fluency in English (spoken and written).
  • Successfully recommission or decommission chargers following changes in our network.
  • Responsible for the go-live of the chargers on Shell’s public network following commissioning attempts.

Additional Information

  • This role involves managing infrastructure for a global platform operating in over ten countries, requiring effective communication and collaboration across regions.
  • Strong verbal and written communication skills, along with availability and flexibility to resolve blocking issues, are essential to support on-call engineers.
  • This role may involve EU or US time-zone shifts based on business requirements.
  • Shift timing: 2 PM IST to 11 PM IST.

What is required to be successful in this role:

  • Global platform experience (B2C or B2B).
  • AWS native service experience.
  • Firmware deployment and cloud cost optimization experience.
  • Strong exposure to monitoring and alerts.
  • Experience with firmware rollout, IoT devices onboarding and offboarding will be an added advantage.
  • Experience as an SRE or DevOps Engineer with some exposure to Project Management or Technical Project Management in IoT-based projects will be helpful.  

What We Offer

  • Work with some of the brightest minds in the emerging EV industry.
  • Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
  • Freedom to suggest, implement, and innovate on systems, processes, and technologies.
  • Daily ownership in a high-growth, challenging environment.
  • Flexible work environment with hybrid schedules and virtualization options.
  • Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.


Read more
company logo
Harsha Mehrotra
Posted by Harsha Mehrotra
Remote only
3 - 5 yrs
Best in industry
CI/CD
skill iconPostgreSQL
skill iconPython
skill iconAmazon Web Services (AWS)
Artificial Intelligence (AI)
+2 more

SRE / Success Engineering role focused on production operations, reliability, AWS infrastructure, monitoring, incident management, and platform support for the ZT platform.


Core responsibilities include:

  • Production monitoring and debugging of live systems.
  • Incident investigation, troubleshooting, and problem resolution.
  • AWS cloud infrastructure support and maintenance.
  • Deployment and operational support activities.
  • Supporting a 24x7 production environment.
  • Working with GitHub-based development workflows.
  • Technical debt remediation and platform improvements.
  • Customer issue investigation and support.
  • Security and compliance-related work, including FedRAMP initiatives.


Preferred Skills:

AWS (especially S3 and EC2)

Strong debugging and troubleshooting skills

Site Reliability Engineering (SRE) experience

GitHub experience

Basic software development skills

TypeScript/JavaScript knowledge

C# preferred

AI experience is a plus.


Candidate should be a hands-on engineer with strong AWS, SRE, operational ownership, production support, and debugging capabilities, rather than a pure application or full-stack developer.

Read more
company logo
Pradeep Chandkiran
Posted by Pradeep Chandkiran
Bengaluru (Bangalore)
3 - 5 yrs
₹8L - ₹14L / yr
skill iconKubernetes
Terraform
skill iconAmazon Web Services (AWS)

Location: Bangalore preferred / Hybrid as applicable

Experience: 3+ years

Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline

Salary: Above market standards, flexible for the right candidate

Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations


About FrontM

FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.

The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.


Role Summary

As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.

This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.


Key Responsibilities

Cloud Infrastructure & DevOps Architecture (≈45%)

· Own, maintain and improve AWS cloud infrastructure for FrontM platforms

· Create and maintain Terraform scripts for infrastructure deployment and management

· Manage Kubernetes workloads deployed within AWS EKS

· Support multi-zone AWS infrastructure design for availability, resilience and scale

· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda

· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap

CI/CD, Operations & Platform Reliability (≈35%)

· Build, maintain and improve CI/CD pipelines for backend and platform services

· Oversee technical operations with hands-on administration, monitoring and release support

· Ensure continuous server uptime, stability, performance and maintainability

· Debug, respond to and restore system outages in production and staging environments

· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io

· Support backend stability, scale and performance across Node.js, Java and related services

Security, Networking & Production Support (≈20%)

· Maintain AWS security configurations, access controls and monitoring practices

· Support complex networking requirements across multi-domain SaaS implementations

· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users

· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements

· Document operational procedures, incident findings and technical support steps clearly


Required Technical Skills

Cloud Infrastructure & AWS

· Strong hands-on experience with AWS infrastructure and cloud operations

· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda

· Experience with AWS security setup, monitoring and multi-zone infrastructure

· Ability to manage infrastructure using Terraform

Kubernetes, CI/CD & Observability

· Strong experience with Kubernetes, preferably AWS EKS

· Extensive CI/CD and DevOps experience

· Experience with infrastructure observability and application monitoring tools

· Ability to diagnose production bottlenecks, server failures and performance issues

Backend, Networking & SaaS Operations

· Experience supporting Node.js, Java and backend system procedures for stability and scale

· Good understanding of APIs, integrations and backend service dependencies

· Experience with complex networking and multi-domain SaaS implementations

· Ability to troubleshoot technical issues with non-technical end users

Nice to Have

· Experience with MongoDB clusters in MongoDB Atlas

Personal Attributes

· Strong ownership mindset for uptime, reliability and production stability

· Practical problem-solving approach with the ability to act quickly during incidents

· Clear written and spoken communication in English

· Ability to work independently and coordinate with senior management when required

· Comfortable working in fast-moving engineering teams

· Attention to detail in security, monitoring, documentation and operational processes


Why join FrontM?

Long-Term Career Growth

Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.

Engineering Challenges That Matter

Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.

Broad Technical Ownership

Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.


Apply now

Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.

Read more
company logo
Sakshi Kochar
Posted by Sakshi Kochar
icon

The recruiter has not been active on this job recently. You may apply but please expect a delayed response.

Remote only
3.5 - 6 yrs
₹6L - ₹18L / yr
Terraform
Serverless
skill iconAmazon Web Services (AWS)
DevSecOps

Mactores is a trusted leader among businesses in providing modern data platform solutions. Since 2008, Mactores have been enabling businesses to accelerate their value through automation by providing End-to-End Data Solutions that are automated, agile, and secure. We collaborate with customers to strategize, navigate, and accelerate an ideal path forward with a digital transformation via assessments, migration, or modernization.


You will be part of the DevOps engineers' team, managing large customer deployments including Linux and Windows Administration, Large Enterprise Application, and Big Data Workloads. You will have broad business and technology expertise coupled with a background in professional services and client-facing skills. You are passionate about the best practices of cloud deployment and ensuring the customer expectation is set and met appropriately. You will help us build scalable, efficient cloud infrastructure.

 

You’ll implement monitoring for automated system health checks. Lastly, you’ll build our CI pipeline, and train and guide the team in DevOps practices.  If you love to solve problems using your skills, then come join the Team Mactores. We have a casual and fun office environment that actively steers clear of rigid "corporate" culture, focuses on productivity and creativity, and allows you to be part of a world-class team while still being yourself.


What you will do?

  • Application migration projects from on-premises to AWS.
  • Database (RDBMS, NoSQL, DW, Hadoop) migration projects from on-premises to AWS. 
  • Automate operational and server provisioning workflows using AWS CFT on AWS.
  • Share the responsibility for deploying releases and conducting other operations maintenance.
  • Enhance operations infrastructures such as Jenkins clusters, Bitbucket, monitoring tools (Consul), and metrics tools such as Graphite and Grafana.
  • Provide operational support for the rest of the Engineering team.
  • Help migrate our remaining dedicated hardware infrastructure to the cloud.
  • Establish and maintain operational best practices.


What do you have?

  • 2+ years of experience in using Terraform for IaaC.
  • 2+ years of configuration management and engineering for large scale customers, ideally supporting an Agile development process.
  • 2+ years of Linux Administration experience.
  • Deep understanding of version control systems (git), including branching and merging strategies.
  • Experience working with cloud platforms (AWS/EC2/ ECS/ RDS/ CloudFormation, Cloudwatch, etc.) and cloud automation tools (Ansible, Chef).
  • Experience with software build tools (Maven, Gradle) and continuous integration tools (Jenkins).
  • Must have supported Java-based applications in a production environment.
  • Experience with Linux environments and scripting languages - bash, python, Groovy.
  • Experience in supporting Node.js in production is a plus.
  • Knowledge of service discovery tools such as Consul is a plus.
  • Comfortable working late evening hours, which is when most patching occurs.
  • You are extremely proactive at identifying ways to improve things and make them more reliable.


You will be preferred if

  • You are AWS DevOps Pro or AWS SA Pro Certified


Read more
Hiring for DevOps lead
Hiring for DevOps lead
Agency job
via by rincy v
Remote only
7 - 20 yrs
₹12L - ₹18L / yr
DevOps
DataDog
RUM

Hiring: DevOps Lead

📍 Kochi / Trivandrum / Remote

💼 Full-time

🕐 General Shift | Australian Overlap

We are looking for an experienced DevOps Lead to join our team.


🔹 Key Responsibilities

Implement and continually improve the observability platform using Datadog, particularly from a user experience perspective.

Work across engineering squads as a virtual team member, supporting their DevOps, infrastructure, and observability requirements.

Configure and manage Datadog RUM, Session Replay, Distributed Tracing, and APM.

Support campaign readiness activities, including load and performance testing.

Participate in gamedays and incident response activities for production systems.

Liaise closely with the managed infrastructure provider on infrastructure requirements and activities.


🔹 Essential Skills & Requirements

✅ 7+ years of relevant DevOps / Cloud / Observability experience

✅ Strong hands-on experience with Datadog

✅ Strong experience with RUM, Session Replay, Distributed Tracing, and APM

✅ Solid experience with AWS

✅ Strong Infrastructure-as-Code experience using CloudFormation and AWS CDK

✅ Good working knowledge of GitHub Actions and AWS CodePipeline

✅ Real incident response experience on high-traffic systems

✅ Experience with Load & Performance Testing

✅ Strong problem-solving and communication skills

✅ Ability to work closely with infrastructure partners and internal platform teams


🔹 Skills - Good to Have

⭐ Experience with high-traffic platforms

⭐ Experience with campaign readiness and gamedays

⭐ Infrastructure partner management

⭐ Advanced AWS observability

⭐ Performance engineering


📌 Experience: 7+ Years

📌 Work Location: Kochi / Trivandrum / Remote

📌 Shift: General Shift with Australian Overlap

📌 Remote: Mandatory 1 week at office


📩 Interested candidates can share their resume

Read more
Remote only
8 - 12 yrs
Best in industry
Terraform
Artificial Intelligence (AI)
IAC
skill iconAmazon Web Services (AWS)
ECS
+6 more


Senior Platform & Site Reliability Engineer

Location: Remote Employment Type: Contract

The Role

This role carries full architectural and operational ownership of the platform layer across a growing SaaS portfolio. The Cloud Architect owns AWS infrastructure standards — VPCs, account structures, networking, and compute design. Everything outside that lane is yours: the CI/CD platform, the observability and reliability stack, the event streaming infrastructure, the deployment pipelines, and the incident engineering model.

Architectural decisions are yours to make and defend, standards are yours to define and enforce, and the reliability of 20+ enterprise SaaS products depends on what you and your team build.

This is an AI-native engineering organisation. Where it is practical and safe to do so, you are expected to use automation and AI-assisted tooling to reduce toil — in CI/CD triage, infrastructure provisioning, observability workflows, and acquisition onboarding. The expectation is not to replace engineering judgement with automation, but to free it up for the problems that genuinely require it.

The Scale You Will Operate At

The portfolio consists of 20+ live, enterprise-grade SaaS solutions running concurrently. Each product serves enterprise customers and processes millions to billions of real-time requests. The architecture is serious: event streaming for real-time data pipelines, batch processing workloads running alongside live transaction flows, and multi-tenant enterprise-grade reliability expectations across every product.

You will design and operate the platform infrastructure that underpins all of it — scaling horizontally as each new acquisition joins the portfolio, without proportionally scaling cost, complexity, or headcount.

What You Will Own

Platform Architecture

  • Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
  • Define, build, and enforce platform standards across portfolio products
  • Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
  • Self-service developer platform so product teams ship without waiting on platform

Event Streaming & Pipeline Infrastructure

  • Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
  • Design and maintain batch processing infrastructure alongside live event flows
  • Ensure pipeline reliability, throughput, and cost are actively managed at scale

CI/CD & Deployment

  • Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
  • Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
  • Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified

Observability & Reliability

  • Own the full observability stack: Grafana, Prometheus, and Loki across all products
  • SLOs and error budgets defined per product; reliability tracked consistently
  • Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen
  • Incident response: on-call design, escalation playbooks, post-mortem facilitation
  • Automated remediation scoped to safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns; novel or ambiguous failures escalate to a human with full context attached

Acquisition Onboarding

  • Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
  • Migration plan and execution for each portfolio company joining the platform — sequenced to avoid disrupting live operations
  • Target: full platform integration within a defined window per acquisition

A Note on Automation

Where automation is safe and failure modes are well understood — routine provisioning, known CI/CD failure classes, secrets rotation, cost anomaly flagging — aggressive automation is expected. Where automation would act on ambiguous signals or carry significant blast radius, human judgement stays in the loop. The goal is to reduce toil on solved problems, not to automate decisions that require engineering expertise.

Platform Stack

Area Stack / Standard IaC Terraform OSS / OpenTofu CI/CD GitHub Actions Event Streaming Architecture and tooling chosen for the workload Observability Grafana, Prometheus, Loki Log Management AWS CloudWatch, Grafana Loki Incident Management OpsGenie (startup tier) or Better Uptime Secrets AWS Secrets Manager / HashiCorp Vault OSS Containers ECS (default), EKS only where justified Cost Monitoring AWS Cost Explorer with custom dashboards What We’re Looking For

  • 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
  • Strong Terraform depth across multi-environment, multi-account setups
  • CI/CD ownership across a multi-product environment with GitHub Actions
  • Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
  • Hands-on Grafana, Prometheus, and Loki in production
  • AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
  • SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture
  • Acquisition or greenfield platform integration experience strongly preferred

How You Work

  • Comfortable operating across multiple products simultaneously — context-switching without dropping standards
  • Cost-efficiency instinct — you optimise spend as a habit, not as a project
  • You treat automation as a tool for eliminating toil, not a substitute for engineering judgement
  • You document decisions, enforce standards through code, and build platforms that other engineers find intuitive to use

Why This Role

The platform function is being built from the ground up. You will have architectural ownership of the entire non-AWS platform layer across a growing portfolio of enterprise SaaS products, with the freedom — and responsibility — to build the reliability and delivery culture of the organisation.

This is not a role that inherits someone else’s decisions and maintains them. Every major architectural choice is still to be made. If you want to build something that lasts and that other engineers depend on, this is the role.

Read more
Service Co
Service Co
Agency job
via by Rishika Teja
Pune
6 - 8 yrs
₹15L - ₹24L / yr
skill iconAmazon Web Services (AWS)
Amazon EC2
Amazon S3
Amazon VPC
Amazon EKS
+8 more

Hiring for Senior Devops Engineer - AWS


Exp : 6 - 8 yrs

Edu : BE/B.Tech

Work Location : Pune WFO

Notice Period : Immediate - 15 days


Skills :

5+ years of hands-on experience in DevOps Engineering and Infrastructure Management.


Exp with AWS Cloud, including EC2, S3, VPC, EKS, Lambda, IAM, CloudWatch, and other core AWS services.


Exp in Terraform for Infrastructure as Code (IaC).


Exp with Kubernetes (EKS) and Helm for container orchestration and deployment.


Good knowledge of Monitoring & Observability tools such as Grafana and Prometheus.

Read more
A US-based cybersecurity startup
A US-based cybersecurity startup
Agency job
via by Ariba Khan
Remote
2 - 5 yrs
Best in industry
skill iconAmazon Web Services (AWS)
CI/CD
IaC
skill iconGitHub

About the company

The client is a US-based cybersecurity startup prevents, detects, and responds to software supply chain attacks by analyzing behavior across the full software development lifecycle for both developers and AI coding agents. They are building a vertical AI agent for supply chain security across three pillars: securing AI agents on developer machines, OSS package security, and CI/CD security, covering the entire agentic pipeline from dev environment to cloud.


Founded by ex-Microsoft, 21 years & ex-Uber, Microsoft, Plaid founders, they are a 16-person team working on hard problems at the intersection of security, AI, and open source. 


Why this role is exciting 

They are at the forefront of supply chain security research and product development. They were the first to detect several major supply chain attacks in 2025 and 2026, including the axios npm compromise and tj-actions. Their research is regularly cited by Bloomberg, TechCrunch, Hacker News, and Dark Reading. The US Cybersecurity and Infrastructure Security Agency (CISA) has published advisories citing the company. 


Beyond their enterprise customers, the company has been adopted by more than 15,000 open-source projects, including projects from Microsoft, Google, Amazon, and Datadog. 

This is a zero-to-one platform role. Today the AWS cost, service health visibility, and their release process are owned in pieces by engineers who are also shipping product. You will own all three end to end, as the first dedicated platform hire, with meaningful early-stage equity upside. 


Their stack 

GitHub.com for code repositories. GitHub Actions for CI/CD pipelines. AWS serverless for hosting our services: Lambda, API Gateway, DynamoDB, and S3. Backend services are written in Golang. 


What you'll do 

  • Own AWS cost. Build real visibility into where our cloud spend goes, then drive it down. Put guardrails and anomaly alerting in place so cost regressions get caught in days, not at the end of the month. 
  • Build service health visibility. Define and instrument the metrics that tell us whether the platform is healthy. Build the dashboards, set the SLOs, and wire up alerting that pages on real problems and stays quiet otherwise. 
  • Own the release process. Design and implement how code gets from a pull request to production, including staged rollouts, rollback, release approvals, and change management that satisfies our enterprise customers and our ISO 27001 obligations without slowing engineers down. 
  • Improve stability. Run infrastructure as code, tighten our incident response and on-call practices, and drive postmortems to actual fixes. 
  • Secure our own pipelines. We sell CI/CD and cloud security, so our own build and deploy path should be the best reference implementation we have. You will run it that way. 


What we're looking for 

  • 2 to 5 years of experience with strong engineering fundamentals. 
  • Demonstrated AWS cost reduction. We want to hear what the spend was, what you did, and what it became. 
  • Hands-on depth with AWS serverless in production, not just in theory. 
  • Experience implementing metrics, dashboards, and alerting for service health, and evidence that it changed how the team operated. 
  • Experience building or significantly improving CI/CD and release processes, ideally with GitHub Actions. 
  • Comfortable writing code. Golang or Python, plus infrastructure as code. 
  • An AI-native mindset, with prior hands-on experience building with or operating AI agents. 
  • Early-stage startup experience. You are extremely hands-on and can drive a problem end to end without a team behind you. 
  • A security background is a plus but not required. 
Read more
company logo
Agency job
via by Soundarya Valli Chintapalli
Hyderabad
3 - 8 yrs
₹8L - ₹18L / yr
Linux/Unix
  • Linux troubleshooting
  • Hands-on AWS
  • Production/Application Support
  • Bash/Shell/Python
  • Monitoring/log analysis
  • Incident resolution
  • Application deployment/support
  • Basic networking and database knowledge
  • Production/batch support exposure
  • Willingness for rotational weekend/critical production support


Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos