Cutshort logo
For Employers
Client acquires and transforms into AI native scalable businesses. logo
Senior Platform & Site Reliability Engineer
Client acquires and transforms into AI native scalable businesses.
Senior Platform & Site Reliability Engineer

Senior Platform & Site Reliability Engineer at Client acquires and transforms into AI native scalable businesses. · Remote only · 8 - 12 years · Remote only · Posted 21 Sep 2026

HyrHub's logo

Senior Platform & Site Reliability Engineer

at Client acquires and transforms into AI native scalable businesses.

Agency job
via HyrHub
8 - 12 yrs
Best in industry
Remote only
Skills
Terraform
skill icongrafana
prometheus
skill iconAmazon Web Services (AWS)

What You Will Own 

Platform Architecture 

  • Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure 
  • Define, build, and enforce platform standards across portfolio products 
  • Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed 
  • Self-service developer platform so product teams ship without waiting on platform 

 

Event Streaming & Pipeline Infrastructure 

  • Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines 
  • Design and maintain batch processing infrastructure alongside live event flows 
  • Ensure pipeline reliability, throughput, and cost are actively managed at scale 

 

CI/CD & Deployment 

  • Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products 
  • Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures 
  • Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified 

 

Observability & Reliability 

  • Own the full observability stack: Grafana, Prometheus, and Loki across all products 
  • SLOs and error budgets defined per product; reliability tracked consistently 
  • Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen 
  • Incident response: on-call design, escalation playbooks, post-mortem facilitation 
  • Automated remediation scoped to a defined set of safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns. Novel or ambiguous failures escalate to a human with full context attached 

 

Acquisition Onboarding 

  • Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture 
  • Migration plan and execution for each portfolio company joining the platform.
  • Target: full platform integration within a defined window per acquisition 

What We're Looking For 


Experience & Background 

  • 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time 
  • Strong Terraform depth across multi-environment, multi-account setups 
  • CI/CD ownership across a multi-product environment with GitHub Actions 
  • Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management 
  • Hands-on Grafana, Prometheus, and Loki in production 
  • AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer 
  • SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture 
  • Acquisition or greenfield platform integration experience strongly preferred 


 

Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

Similar jobs (10)

Remote only
8 - 12 yrs
Best in industry
Terraform
Artificial Intelligence (AI)
IAC
skill iconAmazon Web Services (AWS)
ECS
+6 more


Senior Platform & Site Reliability Engineer

Location: Remote Employment Type: Contract

The Role

This role carries full architectural and operational ownership of the platform layer across a growing SaaS portfolio. The Cloud Architect owns AWS infrastructure standards — VPCs, account structures, networking, and compute design. Everything outside that lane is yours: the CI/CD platform, the observability and reliability stack, the event streaming infrastructure, the deployment pipelines, and the incident engineering model.

Architectural decisions are yours to make and defend, standards are yours to define and enforce, and the reliability of 20+ enterprise SaaS products depends on what you and your team build.

This is an AI-native engineering organisation. Where it is practical and safe to do so, you are expected to use automation and AI-assisted tooling to reduce toil — in CI/CD triage, infrastructure provisioning, observability workflows, and acquisition onboarding. The expectation is not to replace engineering judgement with automation, but to free it up for the problems that genuinely require it.

The Scale You Will Operate At

The portfolio consists of 20+ live, enterprise-grade SaaS solutions running concurrently. Each product serves enterprise customers and processes millions to billions of real-time requests. The architecture is serious: event streaming for real-time data pipelines, batch processing workloads running alongside live transaction flows, and multi-tenant enterprise-grade reliability expectations across every product.

You will design and operate the platform infrastructure that underpins all of it — scaling horizontally as each new acquisition joins the portfolio, without proportionally scaling cost, complexity, or headcount.

What You Will Own

Platform Architecture

  • Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
  • Define, build, and enforce platform standards across portfolio products
  • Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
  • Self-service developer platform so product teams ship without waiting on platform

Event Streaming & Pipeline Infrastructure

  • Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
  • Design and maintain batch processing infrastructure alongside live event flows
  • Ensure pipeline reliability, throughput, and cost are actively managed at scale

CI/CD & Deployment

  • Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
  • Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
  • Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified

Observability & Reliability

  • Own the full observability stack: Grafana, Prometheus, and Loki across all products
  • SLOs and error budgets defined per product; reliability tracked consistently
  • Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen
  • Incident response: on-call design, escalation playbooks, post-mortem facilitation
  • Automated remediation scoped to safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns; novel or ambiguous failures escalate to a human with full context attached

Acquisition Onboarding

  • Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
  • Migration plan and execution for each portfolio company joining the platform — sequenced to avoid disrupting live operations
  • Target: full platform integration within a defined window per acquisition

A Note on Automation

Where automation is safe and failure modes are well understood — routine provisioning, known CI/CD failure classes, secrets rotation, cost anomaly flagging — aggressive automation is expected. Where automation would act on ambiguous signals or carry significant blast radius, human judgement stays in the loop. The goal is to reduce toil on solved problems, not to automate decisions that require engineering expertise.

Platform Stack

Area Stack / Standard IaC Terraform OSS / OpenTofu CI/CD GitHub Actions Event Streaming Architecture and tooling chosen for the workload Observability Grafana, Prometheus, Loki Log Management AWS CloudWatch, Grafana Loki Incident Management OpsGenie (startup tier) or Better Uptime Secrets AWS Secrets Manager / HashiCorp Vault OSS Containers ECS (default), EKS only where justified Cost Monitoring AWS Cost Explorer with custom dashboards What We’re Looking For

  • 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
  • Strong Terraform depth across multi-environment, multi-account setups
  • CI/CD ownership across a multi-product environment with GitHub Actions
  • Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
  • Hands-on Grafana, Prometheus, and Loki in production
  • AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
  • SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture
  • Acquisition or greenfield platform integration experience strongly preferred

How You Work

  • Comfortable operating across multiple products simultaneously — context-switching without dropping standards
  • Cost-efficiency instinct — you optimise spend as a habit, not as a project
  • You treat automation as a tool for eliminating toil, not a substitute for engineering judgement
  • You document decisions, enforce standards through code, and build platforms that other engineers find intuitive to use

Why This Role

The platform function is being built from the ground up. You will have architectural ownership of the entire non-AWS platform layer across a growing portfolio of enterprise SaaS products, with the freedom — and responsibility — to build the reliability and delivery culture of the organisation.

This is not a role that inherits someone else’s decisions and maintains them. Every major architectural choice is still to be made. If you want to build something that lasts and that other engineers depend on, this is the role.

Read more
Service Co
Pune
6 - 11 yrs
₹15L - ₹26L / yr
AWS Iaas
Platform as a Service (PaaS)
AWS VPS
Amazon EKS

Key Skills:


• Bachelor's or Master's degree in Computer Science or related field.


• Minimum 5 years of experience in Platform Engineering, DevOps, or Cloud Infrastructure Engineering.


• Experience migrating data and systems between AWS IaaS and PaaS.


• Experience operating and supporting applications using AWS VPC, EKS, and related services for multi-account operations.


• Experience developing fast and reliable Continuous Integration/Continuous Deployment (CI/CD) workflows used by hundreds of application teams.


• Experience administering and troubleshooting Operating Systems such as Linux, Windows, and MacOS.


• Professional Certifications in AWS Networks, CNCF Technologies, or Kubernetes.


• Experience using and configuring observability tools such as ELK, Prometheus/Grafana, AWS CloudWatch, and Jaeger.


• Experience of applied GitOps principles using ArgoCD or Flux.


• Public examples of code you've worked on with other people using any of these technologies:


o Configuration management/Infrastructure as Code (IAC) tools, such as AWS CDK, AWS CloudFormation, Terraform, Ansible, or Puppet.


o Systems solutions in one or more programming languages, such as Golang, Python, Java.


o Build, Release, Deploy or Ops Workflows using Bamboo, Argo Project, or GitHub Actions.

Read more
company logo
Lakshit Bagga
Posted by Lakshit Bagga
Remote only
8 - 30 yrs
₹1L - ₹60L / yr
DevOps
skill iconAmazon Web Services (AWS)
prometheus
skill icongrafana
Terraform

We're hiring a Cloud Architect (Contract) to work with our Equity Partners who builds profitable growth by acquiring and operating enterprise software companies. Refining a proprietary operating model across 40+ acquisitions and two decades of hands-on experience, now supercharged by our patented agentic AI platform . In this role, you'll take full architectural control of our CI/CD, observability, and event streaming infrastructure, build the standards every new acquisition plugs into, and use AI-assisted automation to keep 20+ products reliable without proportionally scaling headcount.


Job title: Cloud/Platform Architect (SRE)

Type: Global Remote | Contract


What You Bring

  • 8–12 years in platform engineering, DevOps, or SRE, with growing ownership over time
  • Deep Terraform experience across multi-account, multi-env setups
  • Real production experience with event streaming at scale
  • Hands-on Grafana, Prometheus, Loki, and strong AWS depth (ECS, EKS, IAM, VPC, RDS)
  • SRE fundamentals: SLOs, error budgets, on-call design, post-mortems
  • Bonus: acquisition or greenfield platform-building experience


Roles and Responsibilities

  • Own everything outside core AWS infra: CI/CD, observability, event streaming, deployment, incidents
  • Define the standards every future acquisition will plug into
  • Keep 20+ enterprise products running at serious scale (millions–billions of requests)
  • Build self-service tooling so product teams never wait on you
  • Use AI/automation to kill toil — not to replace engineering judgement


Ready to build the platform that scales an entire portfolio? — let's connect.

Read more
company logo
Atharva K
Posted by Atharva K
Pune
7 - 10 yrs
₹30L - ₹45L / yr
skill iconAmazon Web Services (AWS)
IDP
Terraform

Platform Engineering Lead (For client company)

Location: Pune, India

Experience: 7+ years 


What Success Looks Like

  • Engineering teams ship faster with confidence and built-in guardrails.
  • Cloud cost, security, and reliability are predictable, measurable, and well-managed.
  • CI/CD pipelines are trusted, standardized, and production-ready.
  • Platform decisions reduce cognitive load instead of introducing unnecessary process. 

Scope & Expectations

This is a hands-on leadership role combining architecture and implementation.

You will:

  • Build, not just review.
  • Own the platform roadmap—not just infrastructure tickets.
  • Act as a force multiplier for product engineering teams rather than becoming a bottleneck.
  • Drive platform strategy while remaining deeply involved in execution.

Key Responsibilities

Platform & Cloud Architecture

  • Own Zoop's platform and cloud architecture across GCP and AWS.
  • Design reusable, opinionated platform patterns instead of one-off infrastructure.
  • Build and evolve Zoop's Internal Developer Platform (IDP), including:
  • Self-service environments
  • Golden paths (paved roads)
  • Standardized templates
  • Built-in engineering guardrails
  • Lead Kubernetes and cloud-native adoption at scale.
  • Drive infrastructure automation using Terraform, Pulumi, or similar Infrastructure-as-Code (IaC) tools.

CI/CD, Reliability & Developer Experience

  • Establish robust CI/CD practices with quality gates and production readiness.
  • Improve deployment safety through automation and testing.
  • Define and monitor:
  • Golden Signals
  • SLIs
  • SLOs
  • Incident response processes
  • Reduce operational toil and improve developer productivity.
  • Make observability a first-class capability using cost-efficient monitoring systems.
  • Build an observability platform that multiple engineering teams can easily integrate into their applications.

Security, Privacy & Compliance

  • Build security-by-default into infrastructure and deployment pipelines.
  • Lead implementation and continuous compliance for:
  • DPDP Act (India)
  • ISO 27001:2022
  • SOC 2 Type II
  • Implement:
  • Zero Trust architecture
  • Least-privilege access
  • Secure data isolation

FinOps & Cloud Optimization

  • Make cloud costs transparent and accountable across engineering teams.
  • Establish FinOps practices including:
  • Budgets
  • Cost alerts
  • Optimization routines
  • Drive build-vs-buy decisions using clear ROI analysis.

AI, Data & MLOps Foundations

  • Build secure and scalable foundations for AI and MLOps workloads.
  • Define guardrails for AI systems and sensitive data handling.

Leadership & Collaboration

  • Partner closely with engineering teams to align infrastructure strategy with product goals.
  • Mentor engineers and guide teams through technical change.
  • Balance long-term platform initiatives with practical execution.

What We're Looking For

Experience

  • 7+ years of experience building and operating production infrastructure.
  • Experience scaling engineering platforms in high-growth or regulated companies.
  • Strong hands-on expertise in:
  • Kubernetes and the cloud-native ecosystem
  • Service Mesh technologies
  • Policy Engines
  • GCP, AWS (Azure exposure is a plus)
  • Terraform and Infrastructure as Code

Engineering & Operations

  • Strong understanding of SDLC and modern CI/CD systems (Jenkins, GitOps, etc.).
  • Experience with observability tools such as:
  • Grafana
  • Prometheus
  • New Relic
  • Comfortable reading and contributing to production systems written in:
  • Go
  • Python
  • Node.js

Security & Compliance

  • Practical experience implementing ISO 27001 and SOC 2 controls.
  • Strong understanding of:
  • Data protection
  • Privacy
  • Identity and access management
  • Security best practices

Mindset

We're looking for someone who is:

  • Action-oriented with sound engineering judgment.
  • Analytical, cost-conscious, and reliability-focused.
  • Collaborative, calm under pressure, and open to feedback.
  • Comfortable challenging decisions and explaining trade-offs when necessary.

Nice to Have

  • Experience in fintech, identity, or other regulated industries.
  • Built Internal Developer Platforms (IDPs) or shared infrastructure tooling.
  • Contributions to open-source projects.







Read more
company logo
Bengaluru (Bangalore)
4 - 10 yrs
₹20L - ₹40L / yr
Linux/Unix
skill iconAmazon Web Services (AWS)
Terraform
Ansible
Observability

About the Team

SecurITe’s mission is to build an Agentic-AI driven security platform that protects critical infrastructure from modern cyber threats. Our focus is on delivering highly performant, resilient, and intelligent network security systems that help defenders stay ahead of adversaries.

About the Role

We’re looking for an exceptional Platform Engineer to help lead the development of our next-generation cybersecurity AI platform. This is a rare opportunity to shape how agentic AI transforms the future of cyber defense.

As a Platform Engineer, you will design, build, and operate the foundational infrastructure, deployment systems, and developer platforms that power our cybersecurity products across cloud and on-premise environments. You will work at the intersection of infrastructure engineering, cloud-native technologies, automation, reliability, and security to enable scalable and resilient product delivery.

You’ll collaborate closely with AI/ML, backend, security, QA, and product engineering teams to create self-service platforms, deployment pipelines, observability systems, and operational tooling that accelerate innovation while maintaining enterprise-grade reliability and security.

This role is ideal for Linux platform engineers and system specialists who excel at solving complex system challenges, automating wherever possible, and building resilient platforms that enable teams to move faster with confidence.

Why This Role Matters

Cybersecurity is undergoing a fundamental shift. AI is no longer an enhancement—it’s becoming the core engine of how detection, investigation, and response are executed. As our Platform Engineer, you will architect and build the infrastructure, automation, deployment, and operational systems that make this transformation real.

Your work will directly influence the scalability, reliability, and security of our AI-driven cybersecurity platform across both cloud and enterprise on-premise deployments. You’ll help establish the operational backbone that enables rapid innovation, secure product delivery, and resilient large-scale deployments in mission-critical environments.

This is a chance to solve novel technical challenges involving distributed systems, hybrid infrastructure, observability, automation, and secure software delivery while shaping how defenders outpace modern attackers.

What You’ll Do

●    Design, build, and own AWS infrastructure from the ground up (VPC architecture, EC2 fleet management, IAM, networking, security groups)

●    Administer and harden AlmaLinux VMs across production, staging, and dev environments

●    Build automation for provisioning, patching, and configuration management (infrastructure-as-code, config management tooling)

●    Design and implement observability: monitoring, logging, alerting, and on-call-worthy SLAs from scratch

●    Lead incident response — diagnosis, RCA, and post-incident documentation — with no dedicated ops team to escalate to

●    Make and document build-vs-buy and architecture decisions as the product and team scale

●    Work directly with founders/engineering to translate ambiguous asks into scoped technical plans

Impact You’ll Have

●    Accelerate engineering velocity through scalable developer platforms and automation

●    Improve deployment reliability, platform uptime, and operational efficiency

●    Enable secure and scalable AI-driven cybersecurity workloads

●    Reduce operational overhead through infrastructure automation and self-service systems

●    Help establish enterprise-grade cloud and on-premise deployment capabilities

●    Enhance product resiliency, observability, and operational excellence

●    Shape the long-term platform architecture powering next-generation cybersecurity products

●    Enable rapid and secure delivery of critical security innovations to customers

Required Experience

●    4+ years hands-on Linux administration (RHEL-family strongly preferred — AlmaLinux, CentOS, RHEL)

●    Deep Linux internals: systemd, networking, storage/LVM, process/resource management, kernel-level troubleshooting

●    Real AWS architecture experience — not just operating existing infra, but designing it (VPC, EC2, IAM, security groups, networking)

●    Demonstrated ability to scope and solve ambiguous problems independently, without a runbook or senior engineer to defer to

●    Scripting/automation proficiency (Python and/or Bash) beyond one-off scripts — built tooling that runs unattended

●    Track record of end-to-end ownership: has designed, built, and operated a system (not just contributed to one)

●    Clear, proactive communicator — documents decisions and explains reasoning without being asked

Required Skills & Qualifications

●    Strong Linux system administration and troubleshooting skills

●    Redhat certifications

●    Strong understanding of networking fundamentals, security, and distributed systems

●    Proficiency with Docker, and container orchestration

●    Experience with Terraform, Ansible, or similar infrastructure automation tools

●    Strong scripting or programming skills in Python, Bash, or Go

●    Knowledge of observability tools such as Prometheus, Grafana, ELK, or OpenTelemetry

●    Understanding of platform security best practices and secure infrastructure design

●    Familiarity with virtualization technologies and hybrid infrastructure environments

●    Strong problem-solving and debugging abilities

●    Excellent communication and collaboration skills

●    Ability to thrive in fast-paced startup environments

Nice to Have

●    Configuration management/automation at scale (Ansible, AWX, Terraform, or similar)

●    Monitoring/observability stack experience (Prometheus, Grafana, Zabbix, ELK, Datadog)

●    Container experience (Docker; Kubernetes a plus but not core to this VM-based stack)

●    Experience in a startup or small-team environment where infra was built from scratch

●    Security/compliance exposure (vulnerability remediation, hardening, SSO/access control)

The Mindset

Problem Solver

You thrive on complex, ambiguous challenges and engineer elegant solutions.

Ownership-Driven

You take initiative, move fast, and deliver outcomes without hand-holding.

Continuous Learner

You stay ahead of the curve in AI, ML, cloud-native technologies, and emerging infrastructure trends.

Startup DNA

You excel in fast-moving environments where priorities evolve and impact is immediate.

 

Read more
A US-based cybersecurity startup
A US-based cybersecurity startup
Agency job
via by Ariba Khan
Remote
2 - 5 yrs
Best in industry
skill iconAmazon Web Services (AWS)
CI/CD
IaC
skill iconGitHub

About the company

The client is a US-based cybersecurity startup prevents, detects, and responds to software supply chain attacks by analyzing behavior across the full software development lifecycle for both developers and AI coding agents. They are building a vertical AI agent for supply chain security across three pillars: securing AI agents on developer machines, OSS package security, and CI/CD security, covering the entire agentic pipeline from dev environment to cloud.


Founded by ex-Microsoft, 21 years & ex-Uber, Microsoft, Plaid founders, they are a 16-person team working on hard problems at the intersection of security, AI, and open source. 


Why this role is exciting 

They are at the forefront of supply chain security research and product development. They were the first to detect several major supply chain attacks in 2025 and 2026, including the axios npm compromise and tj-actions. Their research is regularly cited by Bloomberg, TechCrunch, Hacker News, and Dark Reading. The US Cybersecurity and Infrastructure Security Agency (CISA) has published advisories citing the company. 


Beyond their enterprise customers, the company has been adopted by more than 15,000 open-source projects, including projects from Microsoft, Google, Amazon, and Datadog. 

This is a zero-to-one platform role. Today the AWS cost, service health visibility, and their release process are owned in pieces by engineers who are also shipping product. You will own all three end to end, as the first dedicated platform hire, with meaningful early-stage equity upside. 


Their stack 

GitHub.com for code repositories. GitHub Actions for CI/CD pipelines. AWS serverless for hosting our services: Lambda, API Gateway, DynamoDB, and S3. Backend services are written in Golang. 


What you'll do 

  • Own AWS cost. Build real visibility into where our cloud spend goes, then drive it down. Put guardrails and anomaly alerting in place so cost regressions get caught in days, not at the end of the month. 
  • Build service health visibility. Define and instrument the metrics that tell us whether the platform is healthy. Build the dashboards, set the SLOs, and wire up alerting that pages on real problems and stays quiet otherwise. 
  • Own the release process. Design and implement how code gets from a pull request to production, including staged rollouts, rollback, release approvals, and change management that satisfies our enterprise customers and our ISO 27001 obligations without slowing engineers down. 
  • Improve stability. Run infrastructure as code, tighten our incident response and on-call practices, and drive postmortems to actual fixes. 
  • Secure our own pipelines. We sell CI/CD and cloud security, so our own build and deploy path should be the best reference implementation we have. You will run it that way. 


What we're looking for 

  • 2 to 5 years of experience with strong engineering fundamentals. 
  • Demonstrated AWS cost reduction. We want to hear what the spend was, what you did, and what it became. 
  • Hands-on depth with AWS serverless in production, not just in theory. 
  • Experience implementing metrics, dashboards, and alerting for service health, and evidence that it changed how the team operated. 
  • Experience building or significantly improving CI/CD and release processes, ideally with GitHub Actions. 
  • Comfortable writing code. Golang or Python, plus infrastructure as code. 
  • An AI-native mindset, with prior hands-on experience building with or operating AI agents. 
  • Early-stage startup experience. You are extremely hands-on and can drive a problem end to end without a team behind you. 
  • A security background is a plus but not required. 
Read more
company logo
Pradeep Chandkiran
Posted by Pradeep Chandkiran
Bengaluru (Bangalore)
3 - 5 yrs
₹8L - ₹14L / yr
skill iconKubernetes
Terraform
skill iconAmazon Web Services (AWS)

Location: Bangalore preferred / Hybrid as applicable

Experience: 3+ years

Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline

Salary: Above market standards, flexible for the right candidate

Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations


About FrontM

FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.

The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.


Role Summary

As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.

This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.


Key Responsibilities

Cloud Infrastructure & DevOps Architecture (≈45%)

· Own, maintain and improve AWS cloud infrastructure for FrontM platforms

· Create and maintain Terraform scripts for infrastructure deployment and management

· Manage Kubernetes workloads deployed within AWS EKS

· Support multi-zone AWS infrastructure design for availability, resilience and scale

· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda

· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap

CI/CD, Operations & Platform Reliability (≈35%)

· Build, maintain and improve CI/CD pipelines for backend and platform services

· Oversee technical operations with hands-on administration, monitoring and release support

· Ensure continuous server uptime, stability, performance and maintainability

· Debug, respond to and restore system outages in production and staging environments

· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io

· Support backend stability, scale and performance across Node.js, Java and related services

Security, Networking & Production Support (≈20%)

· Maintain AWS security configurations, access controls and monitoring practices

· Support complex networking requirements across multi-domain SaaS implementations

· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users

· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements

· Document operational procedures, incident findings and technical support steps clearly


Required Technical Skills

Cloud Infrastructure & AWS

· Strong hands-on experience with AWS infrastructure and cloud operations

· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda

· Experience with AWS security setup, monitoring and multi-zone infrastructure

· Ability to manage infrastructure using Terraform

Kubernetes, CI/CD & Observability

· Strong experience with Kubernetes, preferably AWS EKS

· Extensive CI/CD and DevOps experience

· Experience with infrastructure observability and application monitoring tools

· Ability to diagnose production bottlenecks, server failures and performance issues

Backend, Networking & SaaS Operations

· Experience supporting Node.js, Java and backend system procedures for stability and scale

· Good understanding of APIs, integrations and backend service dependencies

· Experience with complex networking and multi-domain SaaS implementations

· Ability to troubleshoot technical issues with non-technical end users

Nice to Have

· Experience with MongoDB clusters in MongoDB Atlas

Personal Attributes

· Strong ownership mindset for uptime, reliability and production stability

· Practical problem-solving approach with the ability to act quickly during incidents

· Clear written and spoken communication in English

· Ability to work independently and coordinate with senior management when required

· Comfortable working in fast-moving engineering teams

· Attention to detail in security, monitoring, documentation and operational processes


Why join FrontM?

Long-Term Career Growth

Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.

Engineering Challenges That Matter

Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.

Broad Technical Ownership

Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.


Apply now

Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.

Read more
company logo
Remote, Pune
5 - 10 yrs
₹20L - ₹32L / yr
Infrastructure Platform Engineer
skill iconAmazon Web Services (AWS)
Terraform
AWS CloudFormation
skill iconKubernetes
+13 more

Job Title : SDE 3 – Infrastructure Platform Engineer

Experience : 5.5 to 8.5 Years

Number of Positions : 2

Employment Type : C2H (Contract to Hire)

Work Mode : Remote during contractual period → 5 Days WFO after conversion

Contract Duration : 3 Months

Post-Conversion Location : Pune

Notice Period : Immediate Joiners / Serving Notice Period / Up to 15 Days preferred

(Candidates officially serving a 30-day notice period may also be considered if they are on the bench and have a negotiable joining date)


Role Overview :

We are looking for an experienced SDE 3 – Infrastructure Platform Engineer to design, build, and operate scalable, secure, and highly reliable cloud infrastructure and internal platform capabilities.


The ideal candidate will have strong hands-on experience in Cloud Infrastructure, Infrastructure as Code (IaC), CI/CD, Docker, Kubernetes, automation, observability, networking, and distributed systems.


Mandatory Skills : AWS / Azure / GCP, Terraform / CloudFormation, Kubernetes, Docker, CI/CD, Platform / Infrastructure Engineering, Python / Go / Java / Ruby, Networking, Cloud Security, Distributed Systems, Scalability & Reliability, Strong Coding & Automation.


Key Responsibilities :

  • Design and maintain scalable, highly available infrastructure on AWS / GCP / Azure.
  • Build and manage Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools.
  • Develop automation for infrastructure provisioning, deployments, monitoring, and operations.
  • Manage and optimize Docker and Kubernetes workloads.
  • Build internal platform tools to improve developer productivity and engineering efficiency.
  • Implement monitoring, logging, alerting, and observability solutions.
  • Participate in incident response, RCA, postmortems, and reliability improvements.
  • Design and improve CI/CD pipelines and deployment automation.
  • Contribute to system design, architecture discussions, scalability, security, and cost optimization.
  • Collaborate with application, data, and product engineering teams.


Required Skills :

  • 5.5 to 8.5 years of experience in Infrastructure / Platform Engineering or similar roles.
  • Strong hands-on experience with AWS, GCP, or Azure.
  • Strong expertise in Terraform / CloudFormation.
  • Experience with CI/CD, Docker, and Kubernetes.
  • Strong programming skills in at least one of:
  • Python, Go, Java, or Ruby.
  • Good understanding of networking, cloud security, distributed systems, scalability, and reliability.
  • Experience working with production infrastructure and highly available systems.
  • Strong troubleshooting and problem-solving skills.


Nice to Have :

  • Experience with SRE practices and production on-call ownership.
  • Experience in fintech, payments, banking, or transaction-heavy systems.
  • Knowledge of cloud security, compliance, or FinOps/cost optimization.
  • Experience building internal developer platforms or productivity tools.
  • Previous product company experience.


Interview Process :

Round 1 : Take-Home Coding Assignment – Submit within 48 hours

Round 2 : Coding Assignment Discussion – 1 Hour

Round 3 : Technical Managerial Round – 30 Minutes


Note : The take-home coding assignment is mandatory. Candidates should be comfortable completing and submitting the assignment within 48 hours before proceeding.


Ideal Candidate :

Strong Platform / Infrastructure Engineer with hands-on experience in :

Cloud + Terraform / CloudFormation + Kubernetes + CI/CD + Programming + SRE / Production Operations


Pure DevOps profiles without strong coding and platform engineering experience are not preferred.

Read more
company logo
Bhawna Khemani
Posted by Bhawna Khemani
Bengaluru (Bangalore)
5 - 9 yrs
₹19L - ₹30L / yr
skill iconNodeJS (Node.js)
skill iconReact.js
CI/CD
skill iconGitHub
Azure Data Factory
+2 more

What you'll bring


Design, build, and maintain developer platform capabilities that improve day-to-day engineering workflows across teams.

  • Own and evolve CI/CD pipelines and deployment workflows, improving speed, reliability, and feedback loops.
  • Build internal tools and services (CLIs, automations, dashboards) that support developer productivity and operational excellence.
  • Improve local development environments, including dependency management and consistent developer setup.
  • Drive adoption of engineering standards (quality, security, performance, observability) through templates, guardrails, and documentation.
  • Partner with teams to identify friction points and implement scalable improvements (including self-service tooling).
  • Continuously improve documentation, onboarding material, and internal enablement content for developers

What you'll bring

Do you fit the profile? 

 

  Technical Proficiency:

  • Strong hands-on experience with JavaScript/TypeScript and building modern web applications and tooling.
  • Hands-on experience with Generative AI and Agentic AI technologies, including designing, building, integrating, and leveraging AI-powered solutions to enhance products, platforms, or engineering productivity.
  • Solid backend experience (e.g., Node.js, or equivalent platform services experience) and comfort building APIs and integrations.
  • Strong experience with CI/CD systems (e.g., GitHub Actions, Azure DevOps) and modern delivery practices.
  • Familiarity with cloud platforms (Azure preferred) and containerized development (Docker/Kubernetes).
  • Experience designing scalable developer tooling and platform patterns (templates, shared libraries).
  • Strong understanding of software craftsmanship: Clean Code, SOLID principles, maintainability, and documentation.
  • Experience with testing strategies (unit, integration, end-to-end) and quality automation.
  • Exposure to observability concepts and tooling (metrics, logs, traces; e.g., Grafana/Kibana/Application Insights).

    Nice to Have:

  • Experience with monorepos, build systems, and performance optimization for large front-end codebases.
  • Experience with platform security practices (secrets management , dependency scanning, security).
  • Experience with developer portals, internal documentation systems, and knowledge management practices.
  • Familiarity with event-driven architecture and integration patterns.

Essential Soft Skills:

  • Strong analytical and problem-solving skills; able to identify root causes and propose pragmatic solutions.
  • Excellent communication skills; can explain technical topics clearly to diverse stakeholders.
  • Comfortable collaborating in multicultural, globally distributed teams.
  • Proactive mindset and strong ownership - you see platform gaps and take responsibility for improving them.
  • Ability to balance long-term platform strategy with short-term delivery needs and developer pain points.
  • Critical thinking and a proactive mindset when supporting internal stakeholders.
Read more
Service Co
Service Co
Agency job
via by Rishika Teja
Pune
6 - 8 yrs
₹15L - ₹24L / yr
skill iconAmazon Web Services (AWS)
Amazon EC2
Amazon S3
Amazon VPC
Amazon EKS
+8 more

Hiring for Senior Devops Engineer - AWS


Exp : 6 - 8 yrs

Edu : BE/B.Tech

Work Location : Pune WFO

Notice Period : Immediate - 15 days


Skills :

5+ years of hands-on experience in DevOps Engineering and Infrastructure Management.


Exp with AWS Cloud, including EC2, S3, VPC, EKS, Lambda, IAM, CloudWatch, and other core AWS services.


Exp in Terraform for Infrastructure as Code (IaC).


Exp with Kubernetes (EKS) and Helm for container orchestration and deployment.


Good knowledge of Monitoring & Observability tools such as Grafana and Prometheus.

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos