Cutshort logo
For Employers
bangalore based startup logo
Site Reliability Engineer
bangalore based startup
Site Reliability Engineer

Site Reliability Engineer at bangalore based startup · Remote only · 3 - 5 years · ₹5L - ₹15L / yr · Remote only · Posted 24 Nov 2022

Cornertree Consulting's logo

Site Reliability Engineer

at bangalore based startup

Agency job
3 - 5 yrs
₹5L - ₹15L / yr
Remote only
Skills
AWS CloudFormation
skill iconAmazon Web Services (AWS)
Reliability engineering
Reliability analysis
Production support
DevOps
Solid experience in Internet Infrastructure, including at least 4 - 5 years in cloud based rapid delivery environment.
Experience automating systems engineering tasks.
Experience in fast-paced and dynamic SRE or Production Support engineering teams
A proven track record of managing successful complex internet-based product platforms/architectures.
Experience building metrics and monitoring platforms and defining alerting strategies.
Strong analytical ability with a focus on making data driven decisions.
Capable of technical deep-dives, yet verbally and cognitively agile enough to hold their own in a strategy discussion with senior technical or executive leadership
Experience working in a managed services environment.
Good communication skills, both written and oral.
Solid understanding of Engineering, DevOps and cloud computing fundamentals.
Good understanding of cloud services including AWS.
Strong automation and CI / CD experience.
Solid experience with containerized applications/orchestration and serverless functions.
GitHub, CD/CI tools experience.
If I asked your previous team members about you, they would say you were a great leader and they would very much welcome an opportunity to work for you once again.
Experience in high SLA environments.
Computer Science, Engineering or Sciences degree required or equivalent work experience.
Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

Similar jobs (10)

EDM NETWORK
AHLOUCHE AHLOUCHE
Posted by AHLOUCHE AHLOUCHE
Remote only
1 - 10 yrs
$10K - $20K / yr (ESOP available)
User Experience (UX) Design

Key Responsibilities

Platform Monitoring and Reliability

  • Monitor the health and performance of EDM’s advertiser, publisher, broker, and internal platforms.
  • Maintain monitoring, alerting, logging, and system-health dashboards.
  • Investigate platform outages, degraded performance, failed transactions, delayed data, and integration errors.
  • Respond to production incidents and coordinate resolutions with the appropriate engineers and vendors.
  • Perform root-cause analysis and document corrective and preventive actions.
  • Help maintain defined uptime, response-time, recovery-time, and system-reliability targets.
  • Identify recurring problems and recommend permanent solutions.

Cloud Infrastructure and Systems Operations

  • Maintain and support cloud infrastructure, servers, databases, networks, storage, and production environments.
  • Support development, staging, and production environments.
  • Assist with infrastructure scaling, system upgrades, patching, backups, and disaster recovery.
  • Monitor cloud usage and help control infrastructure and technology costs.
  • Maintain access controls, service accounts, certificates, domain configurations, and environment variables.
  • Ensure production systems are properly documented and recoverable.

Deployment and Release Support

  • Support safe and consistent application deployments.
  • Maintain or improve continuous integration and deployment workflows.
  • Coordinate release schedules, deployment validation, rollback procedures, and post-release monitoring.
  • Help engineering teams identify configuration or infrastructure problems before releases reach production.
  • Maintain deployment documentation, technical checklists, and change logs.
  • Reduce manual deployment work through automation.

API and Integration Support

  • Monitor and troubleshoot third-party APIs, webhooks, postbacks, dialer connections, CRM integrations, payment systems, tracking platforms, and compliance services.
  • Investigate failed lead deliveries, missing postbacks, duplicate records, delayed reporting, and authentication problems.
  • Support ping-post, real-time bidding, call-routing, SIP, and data-transfer workflows.
  • Work with advertisers, publishers, and vendors to diagnose technical integration problems.
  • Create clear documentation for common integration methods and troubleshooting procedures.
  • Develop alerts that identify integration failures before clients report them.

Call and Lead Operations

  • Monitor call-routing, tracking, recording, attribution, and disposition systems.
  • Investigate calls that fail to route, connect, record, track, or report correctly.
  • Support number provisioning, routing rules, caps, schedules, geographic restrictions, buyer availability, and failover logic.
  • Validate that leads, calls, and transactions are properly attributed to the correct advertiser, publisher, campaign, and payout.
  • Assist with discrepancies involving call duration, billable events, conversions, payouts, and reporting.
  • Help protect revenue by identifying technical leakage and delivery failures.

Data and Reporting Support

  • Monitor data pipelines, scheduled jobs, reporting processes, and database performance.
  • Investigate discrepancies between platform reporting, billing records, payment records, and third-party systems.
  • Write and maintain database queries for troubleshooting, validation, and operational reporting.
  • Assist with data corrections using controlled and documented procedures.
  • Support dashboards and operational alerts for revenue, margin, consumption, conversion, and platform activity.
  • Maintain appropriate controls around production data access and modification.

Security and Access Management

  • Support role-based access controls, multifactor authentication, audit logging, encryption, and secure system configuration.
  • Provision and remove employee, contractor, client, and vendor access.
  • Monitor suspicious activity and report potential security incidents.
  • Assist with vulnerability remediation, security reviews, access audits, and incident-response procedures.
  • Protect consumer, advertiser, publisher, employee, and company information.
  • Follow company policies for credentials, production access, sensitive data, and change management.

Automation and Process Improvement

  • Automate repetitive operational tasks using scripts, workflows, APIs, and infrastructure tools.
  • Reduce manual work associated with monitoring, deployments, reporting, reconciliation, onboarding, and support.
  • Build internal tools that improve visibility and response times.
  • Identify operational bottlenecks that affect revenue, margin, client satisfaction, or employee productivity.
  • Maintain clear runbooks and standard operating procedures for recurring technical tasks.

Technical Support and Documentation

  • Serve as an escalation point for complex platform and integration issues.
  • Translate technical problems into clear explanations for nontechnical teams.
  • Create and maintain architecture diagrams, system inventories, runbooks, troubleshooting guides, and incident reports.
  • Track incidents and technical requests through completion.
  • Document known issues, temporary workarounds, permanent resolutions, and system dependencies.
  • Participate in an on-call rotation for urgent production incidents.

First 90-Day Priorities

The successful candidate will be expected to:

  • Learn EDM’s platforms, infrastructure, integrations, call-routing systems, reporting processes, and revenue workflows.
  • Document critical systems, dependencies, credentials ownership, vendor contacts, and escalation procedures.
  • Review existing monitoring, alerting, backups, access controls, and deployment procedures.
  • Establish baseline metrics for uptime, incident volume, response time, recovery time, deployment success, and integration failures.
  • Resolve high-priority recurring production and integration issues.
  • Improve alerting for call-routing failures, API errors, delayed data, failed jobs, and reporting discrepancies.
  • Create runbooks for the company’s most common and highest-risk technical incidents.
  • Identify at least three meaningful automation or cost-saving opportunities.
  • Participate in production support and demonstrate ownership of incidents through resolution.

Performance Expectations

Success will be measured by:

  • Platform uptime and reliability
  • Mean time to acknowledge and resolve incidents
  • Reduction in recurring production problems
  • Deployment success and rollback rates
  • API, postback, webhook, and call-routing reliability
  • Reporting and data accuracy
  • Backup and recovery readiness
  • Quality of technical documentation
  • Security and access-control compliance
  • Reduction in manual operational work
  • Infrastructure costs relative to platform volume
  • Responsiveness to internal teams, clients, and technical partners

Required Qualifications

  • Three or more years of experience in operations engineering, DevOps, site reliability engineering, cloud infrastructure, systems administration, or production support.
  • Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.
  • Experience supporting Linux-based production environments.
  • Working knowledge of networking, DNS, SSL certificates, firewalls, load balancing, and application security.
  • Experience with relational databases and SQL.
  • Experience troubleshooting REST APIs, webhooks, authentication, and third-party integrations.
  • Familiarity with monitoring, logging, alerting, and incident-management tools.
  • Experience with scripting languages such as Python, Bash, JavaScript, or PowerShell.
  • Understanding of source control, deployment pipelines, and release management.
  • Strong troubleshooting, documentation, and communication skills.
  • Ability to prioritize incidents based on business and revenue impact.
  • Availability to participate in an on-call rotation.

Preferred Qualifications

  • Experience in ad-tech, mar-tech, affiliate marketing, lead generation, pay-per-call, telecommunications, or SaaS.
  • Familiarity with SIP, VoIP, dialers, call-tracking platforms, routing systems, and phone-number provisioning.
  • Experience with containers, infrastructure as code, and automated deployment tools.
  • Experience with Docker, Kubernetes, Terraform, GitHub Actions, or similar technologies.
  • Familiarity with payment processing, usage-based billing, reconciliation, and commission systems.
  • Experience working with real-time bidding, ping-post, lead distribution, or high-volume transactional systems.
  • Understanding of TCPA-related controls, consent records, DNC suppression, data privacy, or regulated marketing environments.
  • Experience with security audits, disaster-recovery testing, and compliance documentation.

Ideal Candidate

The ideal candidate:

  • Takes ownership instead of waiting for someone else to fix the problem.
  • Remains calm and methodical during high-impact incidents.
  • Understands the difference between applying a temporary fix and eliminating a root cause.
  • Communicates technical problems clearly and directly.
  • Recognizes that production reliability is a business and revenue responsibility.
  • Automates repetitive work whenever practical.
  • Documents systems so the company is not dependent on one person’s memory.
  • Protects security and stability without creating unnecessary bureaucracy.
  • Is comfortable working in a fast-moving entrepreneurial environment.
  • Can manage competing priorities while maintaining attention to detail.


Read more
Deltek
Remote only
3 - 5 yrs
Best in industry
CI/CD
skill iconPostgreSQL
skill iconPython
skill iconAmazon Web Services (AWS)
Artificial Intelligence (AI)
+2 more

SRE / Success Engineering role focused on production operations, reliability, AWS infrastructure, monitoring, incident management, and platform support for the ZT platform.


Core responsibilities include:

  • Production monitoring and debugging of live systems.
  • Incident investigation, troubleshooting, and problem resolution.
  • AWS cloud infrastructure support and maintenance.
  • Deployment and operational support activities.
  • Supporting a 24x7 production environment.
  • Working with GitHub-based development workflows.
  • Technical debt remediation and platform improvements.
  • Customer issue investigation and support.
  • Security and compliance-related work, including FedRAMP initiatives.


Preferred Skills:

AWS (especially S3 and EC2)

Strong debugging and troubleshooting skills

Site Reliability Engineering (SRE) experience

GitHub experience

Basic software development skills

TypeScript/JavaScript knowledge

C# preferred

AI experience is a plus.


Candidate should be a hands-on engineer with strong AWS, SRE, operational ownership, production support, and debugging capabilities, rather than a pure application or full-stack developer.

Read more
Agami Tech
at Agami Tech
3 recruiters
Digish Shah
Posted by Digish Shah
Mumbai
1 - 3 yrs
Best in industry
RHCSA
Linux/Unix
Cloud Computing
VMWare
Firewall
+5 more

Location : Mumbai


Big Picture (The Opportunity) :

Are you looking for an opportunity to advance your Career? & If you are able to maintain a positive attitude even when everything goes wrong, if you are detail oriented and self motivated with a passion to learn and improve your skills and knowledge, we have a perfect job for you !

What do we want from you ? (Our Expectations) :

  • Zero to 2 years experience in Linux Operating System.
  • Flexible working hours - able to support occasional nights, weekends, and call-ins, able to quickly adapt to a constantly changing faced paced environment. Open to travel to sites.
  • An ideal person who is excited and motivated about running and supporting a production - grade critical infrastructure and looks for opportunities to improve processes with automation.


What are you required to do ? (Your Responsibilities) :

  • Proactively maintain and develop all linux infrastructure technology to maintain a 24*7*365 uptime service.
  • Engineering of systems administration-related various solutions for our various SAAS products and projects as well as operational needs.
  • Proactively monitoring system performance and capacity planning.
  • Providing technical support to customers for Applications, Operating systems, and networking.
  • Will also be the first point of contact for our clients, where installations are placed on a permanent basis, for basic troubleshooting and problem solving.
  • Fault finding, analysis and logging information for reporting of performance exceptions.
  • Maintain best practices on managing systems and services across all environments.

 

 

Skills & Qualification Required (Add Value) :

  • Graduate - Preferred to have Bachelor‘s Degree in Engineering, Computer Science or related field.
  • He / She should be familiar with the installation and configuration of Linux operating systems and setup and operation of TCP/IP networking on Linux systems also familiar with Internet concepts including SMTP, IMAP, POP, HTTP, DNS, LDAP and related protocols.
  • You should possess excellent communication skills .
  • Knowledge of Email concepts, Helpdesk Concept , VoIP Concept , Cloud computing will be an added advantage.
Read more
It is an Product Based Company(Domain- EV Charging)
It is an Product Based Company(Domain- EV Charging)
Agency job
via Unique Occupational by Mantasha Naaz
Bengaluru (Bangalore)
6 - 8 yrs
₹18L - ₹20L / yr
SRE
Reliability engineering
on call Support
Incident management
skill iconAmazon Web Services (AWS)

Job Title: Senior Site Reliability Engineer 

Location: Bengaluru, India (Hybrid)

Employment Type: Full-time

Experience: 6+ years

About Compnay

It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.

Operating over 70,000 charge points globally, It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.

We value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.

Role Overview

We are seeking a skilled and proactive Site Reliability Engineer (SRE) to join our growing team. In this role, you will be responsible for maintaining system reliability, scalability, and performance across our EV charging platforms. You will collaborate closely with development and operations teams to build resilient, automated, and observable systems.

Key Responsibilities

  • Ensure high availability, performance, and reliability of production systems
  • Design, implement, and manage scalable infrastructure solutions
  • Build and maintain CI/CD pipelines for efficient software delivery
  • Monitor system health using observability tools and respond to incidents proactively
  • Automate operational processes using scripting and Infrastructure as Code (IaC)
  • Manage containerized environments using Docker and Kubernetes
  • Collaborate with cross-functional teams to improve system architecture and resilience
  • Participate in on-call rotations and incident management processes
  • Continuously optimize cloud infrastructure for cost, performance, and scalability

Required Qualifications & Skills

  • Bachelor’s degree in Computer Science, IT, or related field
  • 4+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure roles
  • Strong experience with containerization (Docker) and orchestration (Kubernetes)
  • Proficiency in Linux administration, networking, and system security
  • Hands-on experience with cloud platforms, especially AWS (EKS, EC2, S3, RDS, Lambda)
  • Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or similar
  • Knowledge of Infrastructure as Code tools (Terraform, AWS CloudFormation, Ansible)
  • Proficiency in scripting languages (Python, Bash, or PowerShell)
  • Experience with monitoring tools like Dynatrace, Prometheus, Grafana, or Zabbix
  • Solid understanding of system architecture, microservices, and SaaS/PaaS models
  • Strong analytical and problem-solving skills   

What We Offer

  • Work with some of the brightest minds in the emerging EV industry.
  • Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
  • Freedom to suggest, implement, and innovate on systems, processes, and technologies.
  • Daily ownership in a high-growth, challenging environment.
  • Flexible work environment with hybrid schedules and virtualization options.
  • Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.


Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by Farook sharief
Hyderabad
7 - 11 yrs
₹2L - ₹15L / yr
Network
Reliability engineering
skill iconAmazon Web Services (AWS)
Windows Azure

Job Summary

We are looking for an experienced Site Reliability Engineer (SRE) with strong expertise in networking, DNS, load balancing, and hybrid cloud environments. The candidate will be responsible for maintaining the reliability, availability, performance, and scalability of production infrastructure and services.

The ideal candidate should have strong troubleshooting skills and experience working across network, cloud, infrastructure, and application environments.

Key Responsibilities

  • Monitor and maintain the availability and reliability of production systems and services.
  • Troubleshoot complex network, infrastructure, and application connectivity issues.
  • Manage and troubleshoot DNS services, DNS resolution, records, and configuration issues.
  • Configure, manage, and troubleshoot Load Balancers and traffic routing.
  • Work with Layer 4 and Layer 7 networking and understand TCP/IP, HTTP/HTTPS, routing, and network connectivity.
  • Support hybrid cloud environments involving on-premises infrastructure and public cloud platforms.
  • Troubleshoot connectivity between on-premises data centers and cloud environments.
  • Participate in production incidents, troubleshooting, root cause analysis (RCA), and problem management.
  • Develop automation scripts and tools to reduce manual operational activities.
  • Configure and maintain monitoring, alerting, and observability solutions.
  • Work closely with Network, Cloud, DevOps, Security, and Application teams.
  • Participate in on-call support and resolve production issues within defined SLAs.
  • Document infrastructure, troubleshooting procedures, incident reports, and operational processes.
  • Identify opportunities to improve system reliability, performance, and scalability.
Read more
It is an Product Based Company(Domain- EV Charging)
It is an Product Based Company(Domain- EV Charging)
Agency job
via Unique Occupational by Mantasha Naaz
Bengaluru (Bangalore)
0.6 - 2 yrs
₹8L - ₹10L / yr
skill iconAmazon Web Services (AWS)
Terraform
On call Support
Incident management
Reliability engineering

Jr Platform Engineer

Location: Bengaluru, India (Hybrid)

Employment Type: Full-time

Experience: 0.6-2 years

Shift Timing: 2 PM to 11 PM IST


About Company

It is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.

Operating over 70,000 charge points globally,It is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.

At this company, we value purpose-driven individuals who want to make a meaningful impact and help create a cleaner, smarter, and more connected world.


Role Overview

It is seeking a  TechOps Engineer! We are excited to find a highly engaged engineer who is obsessed with technology that wants to be a part of a “world class” platform SRE team. It engineers must possess an "automation first" mindset, with a relentless focus on documentation, quality, scalability, and reliability using Infrastructure as Code tools. This position will be part of a platform team that is developing exciting products and solutions and playing a key part in driving forward the electrification of transportation.


What you’ll do:  

  • Ensure system reliability, uptime, and performance of global platform.
  • Conduct real-time surveillance of our EV charging systems to proactively identify and mitigate performance issues and anomalies near 24/7 basis. As such, you collaborate with IDT and FMC players to ensure incident detection also happens outside office hours (monitoring shifts among team members subject to duty schedule) 
  • Deliver on change & releases like firmware changes and drive insights & intelligence back into testing processes and tech discussions with the wider organization. 
  •  Successfully deliver and project manage first time right commissioning activities alongside our Engineering Procurement Contract Management (EPCM) partners to successfully bring charge points onto our Charge Point Management System (CPMS).
  • End-to-end EV charger lifecycle management, including deployment, commissioning, monitoring, maintenance, and decommissioning activities.
  • Provide technical guidance and support to DC specialists during the commissioning of EV charging solutions.
  •  Work closely with Shell, Engineering, and IT colleagues to ensure projects are completed on time and to specification.
  • Act as a liaison with the Engineering Procurement Contract Management (EPCM) partner to manage projects from start to finish, ensuring charge points are successfully onboarded on the Charge Point Management System (CPMS).
  • Collaborate with development, operations and support  teams to build scalable and resilient systems.
  • Contribute to incident response, root-cause analysis, and post-mortem reviews, driving continuous improvement.
  • Participate in capacity planning, performance tuning, and resource optimization.
  • Integrate security and compliance best practices into all infrastructure operations.
  • Stay current with emerging SRE tools, frameworks, and cloud technologies to continuously improve reliability practices.
  • Participate in and lead on-call rotations and incident response, conducting detailed postmortems and RCA reports.
  • Flexible to resolve blocking issues during off hours or weekends if required.  

 

What We’re Looking For: 

Basic Qualifications and skills

  • Bachelor’s degree in Engineering , Electrical, ECE, Computer Science, Information Technology, or related field.
  • Overall 1+ years of experience as a Site Reliability Engineer, DevOps/ Technical project coordinator role.
  • Proven experience of DevOps, SRE, or Technical Project Coordination with IoT or connected devices based platforms.
  • Hands-on experience with cloud platforms such as AWS and Infrastructure as Code (IaC) tools such as Terraform.  
  • Experience with incident management and on-call best practices. Provide support to on call engineers.
  • Excellent analytical and problem-solving skills with a proactive mindset. 
  • Expertise with monitoring and observability tools (Dynatrace,Prometheus, Grafana, Zabbix, etc.).
  • Proactively monitor the network, triage performance outliers, and coordinate correction actions to ensure optimal system functionality.
  • Fluency in English (spoken and written). 
  • Successfully recommission or decommission chargers following changes in our network.
  • Responsible for the go-live of the chargers on Shell’s public network following commissioning attempts.

 Note: This role involves managing infrastructure for a global platform operating in over ten countries, requiring effective communication and collaboration across regions. Strong verbal and written communication skills, along with availability and flexibility to resolve blocking issues, are essential to support On-call Engineers. This role may involve EU or US time‑zone shifts based on business requirements. The shift timing will be 2 PM IST to 11 PM IST.  


What We Offer

  • Work with some of the brightest minds in the emerging EV industry.
  • Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
  • Freedom to suggest, implement, and innovate on systems, processes, and technologies.
  • Daily ownership in a high-growth, challenging environment.
  • Flexible work environment with hybrid schedules and virtualization options.
  • Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.


Read more
icon

The recruiter has not been active on this job recently. You may apply but please expect a delayed response.

Remote only
8 - 12 yrs
Best in industry
Terraform
Artificial Intelligence (AI)
IAC
skill iconAmazon Web Services (AWS)
ECS
+6 more


Senior Platform & Site Reliability Engineer

Location: Remote Employment Type: Contract

The Role

This role carries full architectural and operational ownership of the platform layer across a growing SaaS portfolio. The Cloud Architect owns AWS infrastructure standards — VPCs, account structures, networking, and compute design. Everything outside that lane is yours: the CI/CD platform, the observability and reliability stack, the event streaming infrastructure, the deployment pipelines, and the incident engineering model.

Architectural decisions are yours to make and defend, standards are yours to define and enforce, and the reliability of 20+ enterprise SaaS products depends on what you and your team build.

This is an AI-native engineering organisation. Where it is practical and safe to do so, you are expected to use automation and AI-assisted tooling to reduce toil — in CI/CD triage, infrastructure provisioning, observability workflows, and acquisition onboarding. The expectation is not to replace engineering judgement with automation, but to free it up for the problems that genuinely require it.

The Scale You Will Operate At

The portfolio consists of 20+ live, enterprise-grade SaaS solutions running concurrently. Each product serves enterprise customers and processes millions to billions of real-time requests. The architecture is serious: event streaming for real-time data pipelines, batch processing workloads running alongside live transaction flows, and multi-tenant enterprise-grade reliability expectations across every product.

You will design and operate the platform infrastructure that underpins all of it — scaling horizontally as each new acquisition joins the portfolio, without proportionally scaling cost, complexity, or headcount.

What You Will Own

Platform Architecture

  • Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
  • Define, build, and enforce platform standards across portfolio products
  • Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
  • Self-service developer platform so product teams ship without waiting on platform

Event Streaming & Pipeline Infrastructure

  • Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
  • Design and maintain batch processing infrastructure alongside live event flows
  • Ensure pipeline reliability, throughput, and cost are actively managed at scale

CI/CD & Deployment

  • Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
  • Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
  • Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified

Observability & Reliability

  • Own the full observability stack: Grafana, Prometheus, and Loki across all products
  • SLOs and error budgets defined per product; reliability tracked consistently
  • Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen
  • Incident response: on-call design, escalation playbooks, post-mortem facilitation
  • Automated remediation scoped to safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns; novel or ambiguous failures escalate to a human with full context attached

Acquisition Onboarding

  • Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
  • Migration plan and execution for each portfolio company joining the platform — sequenced to avoid disrupting live operations
  • Target: full platform integration within a defined window per acquisition

A Note on Automation

Where automation is safe and failure modes are well understood — routine provisioning, known CI/CD failure classes, secrets rotation, cost anomaly flagging — aggressive automation is expected. Where automation would act on ambiguous signals or carry significant blast radius, human judgement stays in the loop. The goal is to reduce toil on solved problems, not to automate decisions that require engineering expertise.

Platform Stack

Area Stack / Standard IaC Terraform OSS / OpenTofu CI/CD GitHub Actions Event Streaming Architecture and tooling chosen for the workload Observability Grafana, Prometheus, Loki Log Management AWS CloudWatch, Grafana Loki Incident Management OpsGenie (startup tier) or Better Uptime Secrets AWS Secrets Manager / HashiCorp Vault OSS Containers ECS (default), EKS only where justified Cost Monitoring AWS Cost Explorer with custom dashboards What We’re Looking For

  • 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
  • Strong Terraform depth across multi-environment, multi-account setups
  • CI/CD ownership across a multi-product environment with GitHub Actions
  • Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
  • Hands-on Grafana, Prometheus, and Loki in production
  • AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
  • SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture
  • Acquisition or greenfield platform integration experience strongly preferred

How You Work

  • Comfortable operating across multiple products simultaneously — context-switching without dropping standards
  • Cost-efficiency instinct — you optimise spend as a habit, not as a project
  • You treat automation as a tool for eliminating toil, not a substitute for engineering judgement
  • You document decisions, enforce standards through code, and build platforms that other engineers find intuitive to use

Why This Role

The platform function is being built from the ground up. You will have architectural ownership of the entire non-AWS platform layer across a growing portfolio of enterprise SaaS products, with the freedom — and responsibility — to build the reliability and delivery culture of the organisation.

This is not a role that inherits someone else’s decisions and maintains them. Every major architectural choice is still to be made. If you want to build something that lasts and that other engineers depend on, this is the role.

Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by Akilandeswari Panneerselvam
Hyderabad
5 - 10 yrs
₹6L - ₹10L / yr
SRE
Production support
Cloud Computing
skill iconKubernetes
Linux/Unix
+1 more

Site Reliability Engineer (SRE) / Production Support Engineer

Experience: 5–10 Years

Location: Hyderabad

Work Mode: Face-to-Face Drive

Shift: Rotational Shifts

Job Description

Looking for an experienced SRE / Production Support Engineer with strong experience in application and production support, incident management, monitoring, troubleshooting, and cloud operations.

Key Skills

Production Support, Incident Management, Splunk, APM, SLI/SLO, Cloud, Kubernetes, Docker, Terraform, Linux/Windows Administration, Shell Scripting and Python.

Good understanding of production deployments, batch monitoring, network/load balancing, and troubleshooting is required.

Candidates from SRE, Production Support, Application Support, Cloud Operations, or DevOps backgrounds are preferred.

Read more
FrontM Limited
Pradeep Chandkiran
Posted by Pradeep Chandkiran
Bengaluru (Bangalore)
3 - 5 yrs
₹8L - ₹14L / yr
skill iconKubernetes
Terraform
skill iconAmazon Web Services (AWS)

Location: Bangalore preferred / Hybrid as applicable

Experience: 3+ years

Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline

Salary: Above market standards, flexible for the right candidate

Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations


About FrontM

FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.

The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.


Role Summary

As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.

This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.


Key Responsibilities

Cloud Infrastructure & DevOps Architecture (≈45%)

· Own, maintain and improve AWS cloud infrastructure for FrontM platforms

· Create and maintain Terraform scripts for infrastructure deployment and management

· Manage Kubernetes workloads deployed within AWS EKS

· Support multi-zone AWS infrastructure design for availability, resilience and scale

· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda

· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap

CI/CD, Operations & Platform Reliability (≈35%)

· Build, maintain and improve CI/CD pipelines for backend and platform services

· Oversee technical operations with hands-on administration, monitoring and release support

· Ensure continuous server uptime, stability, performance and maintainability

· Debug, respond to and restore system outages in production and staging environments

· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io

· Support backend stability, scale and performance across Node.js, Java and related services

Security, Networking & Production Support (≈20%)

· Maintain AWS security configurations, access controls and monitoring practices

· Support complex networking requirements across multi-domain SaaS implementations

· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users

· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements

· Document operational procedures, incident findings and technical support steps clearly


Required Technical Skills

Cloud Infrastructure & AWS

· Strong hands-on experience with AWS infrastructure and cloud operations

· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda

· Experience with AWS security setup, monitoring and multi-zone infrastructure

· Ability to manage infrastructure using Terraform

Kubernetes, CI/CD & Observability

· Strong experience with Kubernetes, preferably AWS EKS

· Extensive CI/CD and DevOps experience

· Experience with infrastructure observability and application monitoring tools

· Ability to diagnose production bottlenecks, server failures and performance issues

Backend, Networking & SaaS Operations

· Experience supporting Node.js, Java and backend system procedures for stability and scale

· Good understanding of APIs, integrations and backend service dependencies

· Experience with complex networking and multi-domain SaaS implementations

· Ability to troubleshoot technical issues with non-technical end users

Nice to Have

· Experience with MongoDB clusters in MongoDB Atlas

Personal Attributes

· Strong ownership mindset for uptime, reliability and production stability

· Practical problem-solving approach with the ability to act quickly during incidents

· Clear written and spoken communication in English

· Ability to work independently and coordinate with senior management when required

· Comfortable working in fast-moving engineering teams

· Attention to detail in security, monitoring, documentation and operational processes


Why join FrontM?

Long-Term Career Growth

Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.

Engineering Challenges That Matter

Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.

Broad Technical Ownership

Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.


Apply now

Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.

Read more
MNC
MNC
Agency job
via VY SYSTEMS PRIVATE LIMITED by aafia parveen
Hyderabad
7 - 11 yrs
₹2L - ₹15L / yr
Reliability engineering
skill iconAmazon Web Services (AWS)
Windows Azure
Network
DNS

SRE – Network / DNS / Load Balancer

Experience: 7–11 Years

Location: Hyderabad

Work Mode: WFO

Availability: Immediate Joiner

Job Description:

  • Strong experience in Site Reliability Engineering (SRE) with focus on infrastructure and application reliability.
  • Hands-on experience with Network, DNS and Load Balancer troubleshooting and administration.
  • Monitor system performance, availability, latency and infrastructure health.
  • Troubleshoot network connectivity, DNS resolution, routing and load-balancing issues.
  • Experience with load balancers such as F5, BIG-IP, HAProxy or similar technologies.
  • Good understanding of TCP/IP, HTTP/HTTPS, LAN/WAN, SSL/TLS and networking concepts.
  • Experience with DNS technologies such as BIND, Infoblox or equivalent.
  • Work on incident management, root cause analysis and problem resolution.
  • Collaborate with application, network, cloud and infrastructure teams to resolve production issues.
  • Experience with monitoring and alerting tools such as Splunk, Grafana, Prometheus, AppDynamics or similar tools.
  • Strong troubleshooting, production support and communication skills.
  • Willingness to work from Hyderabad office (WFO) and join immediately. 


Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos