Senior Platform engineer at The World’s Largest Sports Community to Book Venues, Find Tr · Bengaluru (Bangalore) · 4 - 8 years · ₹15L - ₹35L / yr · Posted 3 Jul 2026

Senior Platform engineer
at The World’s Largest Sports Community to Book Venues, Find Tr
SENIOR PLATFORM ENGINEER
LOCATION: BENGALURU
It is a leading global sports activity platform with over 5 million
users across 5 countries, that connects players, venues, trainers, and
brands in one full stack ecosystem across 50+ sports. We operate at
the intersection of consumer technology, social networking,
marketplaces & real-world logistics and are building a scalable, multi-
vertical sports-tech venture with a focus on long-term sustainability
and global dominance.
Senior Platform Engineer
The Opportunity: This is a high-ownership, foundational role. You will set the direction for how we
run our systems — incident response, reliability, security, quality and cost — and build the systems
and practices.
Roles and Responsibilities:
● Infrastructure: Be the accountable owner of our Cloud Infra across AWS and GCP —
including ECS, RDS, ElastiCache, networking (VPC), and IAM services
● Observability & SLOs: Define SLIs/SLOs and error budgets and build the metrics, logs, tracing,
and alerting that catch problems before customers do.
● Reliability engineering: Incident management & on-call. Own end to end incident process —
severity levels, paging, an incident bridge, and act as incident commander.
● Cost/FinOps: Own cloud cost — right-sizing, commitments (Savings Plans / RIs), cost-
allocation tagging, and resource lifecycle — balancing efficiency against performance.
● Security: Drive security best practice — least-privilege IAM, MFA/SSO, secrets management
www.playo.co
and rotation, network segmentation, detective controls, and encryption.
Requirements:
● 4 - 6 years experience with end-to-end ownership of production including platform,
infrastructure and SRE responsibilities
● Hands-on depth across AWS - compute, data stores, networking, and IAM, and fluency with
the AWS Well-Architected Framework and its trade-offs.
● Deep observability understanding — SLIs/SLOs/error budgets, metrics/logs/traces, and
alerting that pages on real user-facing symptoms
● Passion for sports and fitness
● Excitement about working in a fast-paced startup environment

Similar jobs (10)
Platform Engineering Lead (For client company)
Location: Pune, India
Experience: 7+ years
What Success Looks Like
- Engineering teams ship faster with confidence and built-in guardrails.
- Cloud cost, security, and reliability are predictable, measurable, and well-managed.
- CI/CD pipelines are trusted, standardized, and production-ready.
- Platform decisions reduce cognitive load instead of introducing unnecessary process.
Scope & Expectations
This is a hands-on leadership role combining architecture and implementation.
You will:
- Build, not just review.
- Own the platform roadmap—not just infrastructure tickets.
- Act as a force multiplier for product engineering teams rather than becoming a bottleneck.
- Drive platform strategy while remaining deeply involved in execution.
Key Responsibilities
Platform & Cloud Architecture
- Own Zoop's platform and cloud architecture across GCP and AWS.
- Design reusable, opinionated platform patterns instead of one-off infrastructure.
- Build and evolve Zoop's Internal Developer Platform (IDP), including:
- Self-service environments
- Golden paths (paved roads)
- Standardized templates
- Built-in engineering guardrails
- Lead Kubernetes and cloud-native adoption at scale.
- Drive infrastructure automation using Terraform, Pulumi, or similar Infrastructure-as-Code (IaC) tools.
CI/CD, Reliability & Developer Experience
- Establish robust CI/CD practices with quality gates and production readiness.
- Improve deployment safety through automation and testing.
- Define and monitor:
- Golden Signals
- SLIs
- SLOs
- Incident response processes
- Reduce operational toil and improve developer productivity.
- Make observability a first-class capability using cost-efficient monitoring systems.
- Build an observability platform that multiple engineering teams can easily integrate into their applications.
Security, Privacy & Compliance
- Build security-by-default into infrastructure and deployment pipelines.
- Lead implementation and continuous compliance for:
- DPDP Act (India)
- ISO 27001:2022
- SOC 2 Type II
- Implement:
- Zero Trust architecture
- Least-privilege access
- Secure data isolation
FinOps & Cloud Optimization
- Make cloud costs transparent and accountable across engineering teams.
- Establish FinOps practices including:
- Budgets
- Cost alerts
- Optimization routines
- Drive build-vs-buy decisions using clear ROI analysis.
AI, Data & MLOps Foundations
- Build secure and scalable foundations for AI and MLOps workloads.
- Define guardrails for AI systems and sensitive data handling.
Leadership & Collaboration
- Partner closely with engineering teams to align infrastructure strategy with product goals.
- Mentor engineers and guide teams through technical change.
- Balance long-term platform initiatives with practical execution.
What We're Looking For
Experience
- 7+ years of experience building and operating production infrastructure.
- Experience scaling engineering platforms in high-growth or regulated companies.
- Strong hands-on expertise in:
- Kubernetes and the cloud-native ecosystem
- Service Mesh technologies
- Policy Engines
- GCP, AWS (Azure exposure is a plus)
- Terraform and Infrastructure as Code
Engineering & Operations
- Strong understanding of SDLC and modern CI/CD systems (Jenkins, GitOps, etc.).
- Experience with observability tools such as:
- Grafana
- Prometheus
- New Relic
- Comfortable reading and contributing to production systems written in:
- Go
- Python
- Node.js
Security & Compliance
- Practical experience implementing ISO 27001 and SOC 2 controls.
- Strong understanding of:
- Data protection
- Privacy
- Identity and access management
- Security best practices
Mindset
We're looking for someone who is:
- Action-oriented with sound engineering judgment.
- Analytical, cost-conscious, and reliability-focused.
- Collaborative, calm under pressure, and open to feedback.
- Comfortable challenging decisions and explaining trade-offs when necessary.
Nice to Have
- Experience in fintech, identity, or other regulated industries.
- Built Internal Developer Platforms (IDPs) or shared infrastructure tooling.
- Contributions to open-source projects.
Platform Engineer
Location: Hyderabad, Telangana — On-site
Experience: 3–6 Years
Employment Type: Full-time
Hyderabad-based mobility-tech startup building the technology infrastructure behind student transportation.
We operate a real-time transportation platform that brings together student tracking, routing, parent notifications, driver applications, operations dashboards, and cloud infrastructure to make student mobility safer, more reliable, and easier to manage.
As we scale, we’re looking for a Platform Engineer who can take ownership of the infrastructure and platform layer that powers these systems.
The Role
As a Platform Engineer, you will own the systems that enable our engineering teams to build, deploy, scale, monitor, and operate reliable production services.
This is an early-stage startup role with significant ownership. You will work closely with engineering and product teams to build infrastructure from the ground up, improve deployment velocity, strengthen reliability, and ensure our platform can scale with the business.
You should be comfortable moving between cloud infrastructure, Kubernetes, CI/CD, observability, security, and distributed systems.
What You'll Do
- Design, build, and maintain scalable cloud infrastructure on AWS
- Own production infrastructure across EC2, IAM, RDS, networking, monitoring, and deployments
- Build and maintain Docker and Kubernetes environments for production workloads
- Develop and improve CI/CD pipelines for reliable and rapid deployments
- Manage infrastructure as code using Terraform
- Establish infrastructure standards for scalability, security, reliability, and cost efficiency
- Monitor production systems and proactively identify performance and reliability issues
- Build observability around applications and infrastructure, including metrics, logs, alerts, and incident monitoring
- Troubleshoot production issues and conduct root-cause analysis (RCA)
- Work with backend engineers to design infrastructure for Java/Kotlin-based distributed systems
- Support highly available services involving real-time tracking, routing, notifications, and operational workflows
- Improve deployment processes, release reliability, rollback strategies, and disaster recovery
- Identify infrastructure bottlenecks and continuously improve platform performance
- Help establish engineering practices around reliability, security, and operational excellence
- Work across the stack when required and take end-to-end ownership of infrastructure problems
What We're Looking For
Must Have
- 3–6 years of experience in Platform Engineering, DevOps, SRE, Infrastructure Engineering, or a closely related role
- Strong hands-on experience with AWS
- Strong understanding of EC2, IAM, RDS, networking, monitoring, and production deployments
- Experience with Docker and Kubernetes
- Strong experience building and managing CI/CD pipelines
- Hands-on experience with Terraform / Infrastructure as Code
- Strong Linux and networking fundamentals
- Experience troubleshooting production systems and performing RCA
- Understanding of distributed systems, scalability, availability, and system design
- Experience working with backend services built using Java/Kotlin or similar technologies
- Ability to independently own infrastructure problems from design → implementation → deployment → monitoring
Good to Have
- Experience with Redis, PostgreSQL, or MongoDB
- Experience with microservices architecture
- Experience with AWS security and IAM best practices
- Experience building observability and alerting systems
- Experience with multi-cloud environments such as AWS, Azure, or GCP
- Experience working in an early-stage startup
- Experience with real-time systems, IoT, location services, or mobility platforms
- Open-source contributions or meaningful personal engineering projects
What Makes This Role Different
At ZeroMoblt, you won't be working within a large infrastructure team where responsibilities are narrowly defined.
You'll have the opportunity to:
- Own critical infrastructure decisions
- Build platform capabilities from 0 → 1
- Work directly with engineering and product teams
- Solve real-world scalability and reliability problems
- Influence architecture and engineering practices
- See your work directly impact a platform serving 10,000+ students
- Work in a fast-moving environment with minimal bureaucracy
We're looking for someone who enjoys ownership, ambiguity, and solving problems independently.
This role may not be the right fit if you prefer highly structured processes, narrowly defined responsibilities, or large-company environments with multiple layers of ownership.
Why ZeroMoblt?
- High ownership and autonomy
- Direct exposure to product and engineering decisions
- Opportunity to build infrastructure at an early-stage mobility startup
- Work on real-time transportation and location-based systems
- Hyderabad-based, on-site team
Job Title : SDE 3 – Infrastructure Platform Engineer
Experience : 5.5 to 8.5 Years
Number of Positions : 2
Employment Type : C2H (Contract to Hire)
Work Mode : Remote during contractual period → 5 Days WFO after conversion
Contract Duration : 3 Months
Post-Conversion Location : Pune
Notice Period : Immediate Joiners / Serving Notice Period / Up to 15 Days preferred
(Candidates officially serving a 30-day notice period may also be considered if they are on the bench and have a negotiable joining date)
Role Overview :
We are looking for an experienced SDE 3 – Infrastructure Platform Engineer to design, build, and operate scalable, secure, and highly reliable cloud infrastructure and internal platform capabilities.
The ideal candidate will have strong hands-on experience in Cloud Infrastructure, Infrastructure as Code (IaC), CI/CD, Docker, Kubernetes, automation, observability, networking, and distributed systems.
Mandatory Skills : AWS / Azure / GCP, Terraform / CloudFormation, Kubernetes, Docker, CI/CD, Platform / Infrastructure Engineering, Python / Go / Java / Ruby, Networking, Cloud Security, Distributed Systems, Scalability & Reliability, Strong Coding & Automation.
Key Responsibilities :
- Design and maintain scalable, highly available infrastructure on AWS / GCP / Azure.
- Build and manage Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools.
- Develop automation for infrastructure provisioning, deployments, monitoring, and operations.
- Manage and optimize Docker and Kubernetes workloads.
- Build internal platform tools to improve developer productivity and engineering efficiency.
- Implement monitoring, logging, alerting, and observability solutions.
- Participate in incident response, RCA, postmortems, and reliability improvements.
- Design and improve CI/CD pipelines and deployment automation.
- Contribute to system design, architecture discussions, scalability, security, and cost optimization.
- Collaborate with application, data, and product engineering teams.
Required Skills :
- 5.5 to 8.5 years of experience in Infrastructure / Platform Engineering or similar roles.
- Strong hands-on experience with AWS, GCP, or Azure.
- Strong expertise in Terraform / CloudFormation.
- Experience with CI/CD, Docker, and Kubernetes.
- Strong programming skills in at least one of:
- Python, Go, Java, or Ruby.
- Good understanding of networking, cloud security, distributed systems, scalability, and reliability.
- Experience working with production infrastructure and highly available systems.
- Strong troubleshooting and problem-solving skills.
Nice to Have :
- Experience with SRE practices and production on-call ownership.
- Experience in fintech, payments, banking, or transaction-heavy systems.
- Knowledge of cloud security, compliance, or FinOps/cost optimization.
- Experience building internal developer platforms or productivity tools.
- Previous product company experience.
Interview Process :
Round 1 : Take-Home Coding Assignment – Submit within 48 hours
Round 2 : Coding Assignment Discussion – 1 Hour
Round 3 : Technical Managerial Round – 30 Minutes
Note : The take-home coding assignment is mandatory. Candidates should be comfortable completing and submitting the assignment within 48 hours before proceeding.
Ideal Candidate :
Strong Platform / Infrastructure Engineer with hands-on experience in :
Cloud + Terraform / CloudFormation + Kubernetes + CI/CD + Programming + SRE / Production Operations
Pure DevOps profiles without strong coding and platform engineering experience are not preferred.

Key Skills:
• Bachelor's or Master's degree in Computer Science or related field.
• Minimum 5 years of experience in Platform Engineering, DevOps, or Cloud Infrastructure Engineering.
• Experience migrating data and systems between AWS IaaS and PaaS.
• Experience operating and supporting applications using AWS VPC, EKS, and related services for multi-account operations.
• Experience developing fast and reliable Continuous Integration/Continuous Deployment (CI/CD) workflows used by hundreds of application teams.
• Experience administering and troubleshooting Operating Systems such as Linux, Windows, and MacOS.
• Professional Certifications in AWS Networks, CNCF Technologies, or Kubernetes.
• Experience using and configuring observability tools such as ELK, Prometheus/Grafana, AWS CloudWatch, and Jaeger.
• Experience of applied GitOps principles using ArgoCD or Flux.
• Public examples of code you've worked on with other people using any of these technologies:
o Configuration management/Infrastructure as Code (IAC) tools, such as AWS CDK, AWS CloudFormation, Terraform, Ansible, or Puppet.
o Systems solutions in one or more programming languages, such as Golang, Python, Java.
o Build, Release, Deploy or Ops Workflows using Bamboo, Argo Project, or GitHub Actions.
About the Team
SecurITe’s mission is to build an Agentic-AI driven security platform that protects critical infrastructure from modern cyber threats. Our focus is on delivering highly performant, resilient, and intelligent network security systems that help defenders stay ahead of adversaries.
About the Role
We’re looking for an exceptional Platform Engineer to help lead the development of our next-generation cybersecurity AI platform. This is a rare opportunity to shape how agentic AI transforms the future of cyber defense.
As a Platform Engineer, you will design, build, and operate the foundational infrastructure, deployment systems, and developer platforms that power our cybersecurity products across cloud and on-premise environments. You will work at the intersection of infrastructure engineering, cloud-native technologies, automation, reliability, and security to enable scalable and resilient product delivery.
You’ll collaborate closely with AI/ML, backend, security, QA, and product engineering teams to create self-service platforms, deployment pipelines, observability systems, and operational tooling that accelerate innovation while maintaining enterprise-grade reliability and security.
This role is ideal for Linux platform engineers and system specialists who excel at solving complex system challenges, automating wherever possible, and building resilient platforms that enable teams to move faster with confidence.
Why This Role Matters
Cybersecurity is undergoing a fundamental shift. AI is no longer an enhancement—it’s becoming the core engine of how detection, investigation, and response are executed. As our Platform Engineer, you will architect and build the infrastructure, automation, deployment, and operational systems that make this transformation real.
Your work will directly influence the scalability, reliability, and security of our AI-driven cybersecurity platform across both cloud and enterprise on-premise deployments. You’ll help establish the operational backbone that enables rapid innovation, secure product delivery, and resilient large-scale deployments in mission-critical environments.
This is a chance to solve novel technical challenges involving distributed systems, hybrid infrastructure, observability, automation, and secure software delivery while shaping how defenders outpace modern attackers.
What You’ll Do
● Design, build, and own AWS infrastructure from the ground up (VPC architecture, EC2 fleet management, IAM, networking, security groups)
● Administer and harden AlmaLinux VMs across production, staging, and dev environments
● Build automation for provisioning, patching, and configuration management (infrastructure-as-code, config management tooling)
● Design and implement observability: monitoring, logging, alerting, and on-call-worthy SLAs from scratch
● Lead incident response — diagnosis, RCA, and post-incident documentation — with no dedicated ops team to escalate to
● Make and document build-vs-buy and architecture decisions as the product and team scale
● Work directly with founders/engineering to translate ambiguous asks into scoped technical plans
Impact You’ll Have
● Accelerate engineering velocity through scalable developer platforms and automation
● Improve deployment reliability, platform uptime, and operational efficiency
● Enable secure and scalable AI-driven cybersecurity workloads
● Reduce operational overhead through infrastructure automation and self-service systems
● Help establish enterprise-grade cloud and on-premise deployment capabilities
● Enhance product resiliency, observability, and operational excellence
● Shape the long-term platform architecture powering next-generation cybersecurity products
● Enable rapid and secure delivery of critical security innovations to customers
Required Experience
● 4+ years hands-on Linux administration (RHEL-family strongly preferred — AlmaLinux, CentOS, RHEL)
● Deep Linux internals: systemd, networking, storage/LVM, process/resource management, kernel-level troubleshooting
● Real AWS architecture experience — not just operating existing infra, but designing it (VPC, EC2, IAM, security groups, networking)
● Demonstrated ability to scope and solve ambiguous problems independently, without a runbook or senior engineer to defer to
● Scripting/automation proficiency (Python and/or Bash) beyond one-off scripts — built tooling that runs unattended
● Track record of end-to-end ownership: has designed, built, and operated a system (not just contributed to one)
● Clear, proactive communicator — documents decisions and explains reasoning without being asked
Required Skills & Qualifications
● Strong Linux system administration and troubleshooting skills
● Redhat certifications
● Strong understanding of networking fundamentals, security, and distributed systems
● Proficiency with Docker, and container orchestration
● Experience with Terraform, Ansible, or similar infrastructure automation tools
● Strong scripting or programming skills in Python, Bash, or Go
● Knowledge of observability tools such as Prometheus, Grafana, ELK, or OpenTelemetry
● Understanding of platform security best practices and secure infrastructure design
● Familiarity with virtualization technologies and hybrid infrastructure environments
● Strong problem-solving and debugging abilities
● Excellent communication and collaboration skills
● Ability to thrive in fast-paced startup environments
Nice to Have
● Configuration management/automation at scale (Ansible, AWX, Terraform, or similar)
● Monitoring/observability stack experience (Prometheus, Grafana, Zabbix, ELK, Datadog)
● Container experience (Docker; Kubernetes a plus but not core to this VM-based stack)
● Experience in a startup or small-team environment where infra was built from scratch
● Security/compliance exposure (vulnerability remediation, hardening, SSO/access control)
The Mindset
Problem Solver
You thrive on complex, ambiguous challenges and engineer elegant solutions.
Ownership-Driven
You take initiative, move fast, and deliver outcomes without hand-holding.
Continuous Learner
You stay ahead of the curve in AI, ML, cloud-native technologies, and emerging infrastructure trends.
Startup DNA
You excel in fast-moving environments where priorities evolve and impact is immediate.
Senior Platform & Site Reliability Engineer
Location: Remote Employment Type: Contract
The Role
This role carries full architectural and operational ownership of the platform layer across a growing SaaS portfolio. The Cloud Architect owns AWS infrastructure standards — VPCs, account structures, networking, and compute design. Everything outside that lane is yours: the CI/CD platform, the observability and reliability stack, the event streaming infrastructure, the deployment pipelines, and the incident engineering model.
Architectural decisions are yours to make and defend, standards are yours to define and enforce, and the reliability of 20+ enterprise SaaS products depends on what you and your team build.
This is an AI-native engineering organisation. Where it is practical and safe to do so, you are expected to use automation and AI-assisted tooling to reduce toil — in CI/CD triage, infrastructure provisioning, observability workflows, and acquisition onboarding. The expectation is not to replace engineering judgement with automation, but to free it up for the problems that genuinely require it.
The Scale You Will Operate At
The portfolio consists of 20+ live, enterprise-grade SaaS solutions running concurrently. Each product serves enterprise customers and processes millions to billions of real-time requests. The architecture is serious: event streaming for real-time data pipelines, batch processing workloads running alongside live transaction flows, and multi-tenant enterprise-grade reliability expectations across every product.
You will design and operate the platform infrastructure that underpins all of it — scaling horizontally as each new acquisition joins the portfolio, without proportionally scaling cost, complexity, or headcount.
What You Will Own
Platform Architecture
- Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
- Define, build, and enforce platform standards across portfolio products
- Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
- Self-service developer platform so product teams ship without waiting on platform
Event Streaming & Pipeline Infrastructure
- Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
- Design and maintain batch processing infrastructure alongside live event flows
- Ensure pipeline reliability, throughput, and cost are actively managed at scale
CI/CD & Deployment
- Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
- Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
- Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified
Observability & Reliability
- Own the full observability stack: Grafana, Prometheus, and Loki across all products
- SLOs and error budgets defined per product; reliability tracked consistently
- Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses, not a blank screen
- Incident response: on-call design, escalation playbooks, post-mortem facilitation
- Automated remediation scoped to safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns; novel or ambiguous failures escalate to a human with full context attached
Acquisition Onboarding
- Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
- Migration plan and execution for each portfolio company joining the platform — sequenced to avoid disrupting live operations
- Target: full platform integration within a defined window per acquisition
A Note on Automation
Where automation is safe and failure modes are well understood — routine provisioning, known CI/CD failure classes, secrets rotation, cost anomaly flagging — aggressive automation is expected. Where automation would act on ambiguous signals or carry significant blast radius, human judgement stays in the loop. The goal is to reduce toil on solved problems, not to automate decisions that require engineering expertise.
Platform Stack
Area Stack / Standard IaC Terraform OSS / OpenTofu CI/CD GitHub Actions Event Streaming Architecture and tooling chosen for the workload Observability Grafana, Prometheus, Loki Log Management AWS CloudWatch, Grafana Loki Incident Management OpsGenie (startup tier) or Better Uptime Secrets AWS Secrets Manager / HashiCorp Vault OSS Containers ECS (default), EKS only where justified Cost Monitoring AWS Cost Explorer with custom dashboards What We’re Looking For
- 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
- Strong Terraform depth across multi-environment, multi-account setups
- CI/CD ownership across a multi-product environment with GitHub Actions
- Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
- Hands-on Grafana, Prometheus, and Loki in production
- AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
- SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture
- Acquisition or greenfield platform integration experience strongly preferred
How You Work
- Comfortable operating across multiple products simultaneously — context-switching without dropping standards
- Cost-efficiency instinct — you optimise spend as a habit, not as a project
- You treat automation as a tool for eliminating toil, not a substitute for engineering judgement
- You document decisions, enforce standards through code, and build platforms that other engineers find intuitive to use
Why This Role
The platform function is being built from the ground up. You will have architectural ownership of the entire non-AWS platform layer across a growing portfolio of enterprise SaaS products, with the freedom — and responsibility — to build the reliability and delivery culture of the organisation.
This is not a role that inherits someone else’s decisions and maintains them. Every major architectural choice is still to be made. If you want to build something that lasts and that other engineers depend on, this is the role.
About Searce
Searce is a global, AI-native, engineering-led technology consultancy and a Premier Google
Cloud Partner — recognized as the Google Cloud Workplace AI Transformation Partner of the
Year, APAC (2026). With 20+ years of experience and 3,000+ clients across 10+ countries, we
help businesses stay ahead of the cloud curve.
The Role
We're looking for a Lead Cloud Security & Reliability Engineer with deep GCP expertise to own
end-to-end cloud reliability and security forAPAC enterprise clients. As Lead, you'll set the architectural direction, mentor your squad, and drive measurable client outcomes across multi-
cloud environments.
What You'll Do
Own Client Delivery — Lead 24x7 GCP cloud operations forAPAC clients. Define SLO frameworks and ensure adherence.
Architect Solutions — Design scalable, secure GCP-primary architectures with multi-cloud awareness.
Drive Reliability — Lead incident response, RCA, and long-term remediation across production systems.
Mentor & Elevate — Coach and grow a squad of Senior CSREs.
Drive FinOps — Own cloud cost governance and optimization with quantified impact.
Be the Expert — Represent Searce's technical depth in global client conversations.
What We're Looking For
Experience
7–12 years total with 5+ years on GCP cloud infrastructure
Strong background in Cloud Managed Services / MSP environments
Proven experience leading a team in client-facing delivery
Multi-cloud exposure (AWS/Azure secondary) preferred
Technical Skills (Must-Have)
- GCP: GKE, IAM, VPC, Cloud Monitoring, Stackdriver, KMS — demonstrated in work
- experience
- Kubernetes: GKE — production cluster management, Helm
- IaC: Terraform — module-level, reusable frameworks
- Observability: Prometheus, Grafana, Thanos or equivalent
- Security: IAM, Zero-trust, DevSecOps, CSPM tools
- Scripting: Python or Go
- FinOps: GCP cost governance demonstrated
Nice to Have
- GCP Professional Cloud Architect / Pro DevOps Engineer certification
- AWS / Azure secondary experience
- CKA (Certified Kubernetes Administrator)
- ITIL / change management awareness
- APAC client delivery experience
Why Searce?
🏆 Google Cloud Partner of the Year — APAC 2026
🌍 Work with APAC enterprise clients across multiple industries
🤖 AI-first, engineering-led culture
📈 Lead-level ownership with real career growth
🤝 HAPPIER values — Humble, Adaptable, Positive, Passionate, Innovative, Excellence,
Responsible
Location: Bangalore preferred / Hybrid as applicable
Experience: 3+ years
Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline
Salary: Above market standards, flexible for the right candidate
Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations
About FrontM
FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.
The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.
Role Summary
As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.
This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.
Key Responsibilities
Cloud Infrastructure & DevOps Architecture (≈45%)
· Own, maintain and improve AWS cloud infrastructure for FrontM platforms
· Create and maintain Terraform scripts for infrastructure deployment and management
· Manage Kubernetes workloads deployed within AWS EKS
· Support multi-zone AWS infrastructure design for availability, resilience and scale
· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap
CI/CD, Operations & Platform Reliability (≈35%)
· Build, maintain and improve CI/CD pipelines for backend and platform services
· Oversee technical operations with hands-on administration, monitoring and release support
· Ensure continuous server uptime, stability, performance and maintainability
· Debug, respond to and restore system outages in production and staging environments
· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io
· Support backend stability, scale and performance across Node.js, Java and related services
Security, Networking & Production Support (≈20%)
· Maintain AWS security configurations, access controls and monitoring practices
· Support complex networking requirements across multi-domain SaaS implementations
· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users
· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements
· Document operational procedures, incident findings and technical support steps clearly
Required Technical Skills
Cloud Infrastructure & AWS
· Strong hands-on experience with AWS infrastructure and cloud operations
· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
· Experience with AWS security setup, monitoring and multi-zone infrastructure
· Ability to manage infrastructure using Terraform
Kubernetes, CI/CD & Observability
· Strong experience with Kubernetes, preferably AWS EKS
· Extensive CI/CD and DevOps experience
· Experience with infrastructure observability and application monitoring tools
· Ability to diagnose production bottlenecks, server failures and performance issues
Backend, Networking & SaaS Operations
· Experience supporting Node.js, Java and backend system procedures for stability and scale
· Good understanding of APIs, integrations and backend service dependencies
· Experience with complex networking and multi-domain SaaS implementations
· Ability to troubleshoot technical issues with non-technical end users
Nice to Have
· Experience with MongoDB clusters in MongoDB Atlas
Personal Attributes
· Strong ownership mindset for uptime, reliability and production stability
· Practical problem-solving approach with the ability to act quickly during incidents
· Clear written and spoken communication in English
· Ability to work independently and coordinate with senior management when required
· Comfortable working in fast-moving engineering teams
· Attention to detail in security, monitoring, documentation and operational processes
Why join FrontM?
Long-Term Career Growth
Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.
Engineering Challenges That Matter
Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.
Broad Technical Ownership
Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.
Apply now
Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.
Senior Cloud Site Reliability Engineer (CSRE) – Azure
About Searce:
Searce is an AI-native, engineering-led modern technology consultancy that empowers
clients to futurify their businesses by delivering real, intelligent business outcomes. As a
trusted partner for over 3,000 clients globally, Searce specializes in cloud modernization,
data engineering, applied AI, and robust cloud platform security. Driven by a "HAPPIER"
cultural mindset and our proprietary evlos problem-solving framework, we eliminate
bureaucratic fluff to build working prototypes fast and scale enterprise production
environments intelligently. We don't just fix systems; we leverage multi-cloud technologies
to transform client operations into distinct competitive advantages.
Position Overview:
We are looking for a high-caliber Senior or Lead Cloud Site Reliability Engineer (CSRE) to
architect, secure, and stabilize next-generation hybrid and multi-cloud environments.
Operating at the intersection of infrastructure design, security compliance, and production
operations, you will serve as the technical Subject Matter Expert (SME) across GCP, Azure,
and AWS.
Whether optimizing a microservice mesh on GKE, tuning autoscaling on AKS, or driving a
massive disaster recovery drill across AWS regions, your focus will be absolute reliability. For
the Lead path, you will couple this deep engineering toolkit with stakeholder management
and mentorship to drive an elite operational culture.
Experience & Level Expectation:
Years of Experience: 3 to 10 years of intensive, hands-on production operations
experience in a dedicated DevOps, Cloud Platform Engineering, or SRE role.
Associate level (3-5 Years): Expected to show flawless execution of IaC, advanced
triaging of infrastructure failures, and ownership of the CI/CD and deployment
lifecycles.
Intermediate level (5-10 Years): Expected to take architectural ownership, serve as
primary Incident Commander for complex outages, design cross-cloud governance
frameworks, and act as a reliable bridge between technical teams and client
leadership.
Key Responsibilities & Role Expectations:
Multi-Cloud Platforms & Orchestration: Design, configure, and maintain
production-grade Kubernetes clusters across major platforms (AKS).
Manage advanced network routing, service meshes (e.g., Istio), and multi-tenant
isolation.
Infrastructure as Code (IaC) & GitOps: Build declarative, enterprise-grade, reusable
infrastructure components using Terraform or Crossplane. Standardize automated
environment provisioning to eliminate configuration drift across multi-branch
environments.
Incident Management & Reliability (SRE): Own and optimize the production on-call
rotation. Lead rapid mitigation strategies for Sev-1/Sev-2 system outages, reducing
Mean Time to Recovery (MTTR) through centralized log and metric correlation.
Root Cause Analysis (RCA): Facilitate rigorous, blameless post-incident reviews to
identify core architectural vulnerabilities and establish long-term fixes preventing
recurrence.
Lifecycle, Patching & Upgrades: Plan and execute zero-downtime cluster upgrades,
operating system patching strategies (Linux/Windows), database lifecycle updates,
and multi-region Disaster Recovery (DR) failover drills.
Core Core Operations & Legacy Integration: Manage enterprise-level hybrid
networking architecture (VPCs, Firewalls, Load Balancers, DNS routing, and DHCP
configurations) while effectively connecting cloud native services to legacy
infrastructures like Active Directory.
Security & Governance: Embed Zero Trust policies, secure secrets management
(Secrets Manager/Key Vault), and continuous vulnerability patching into the
automated SDLC pipeline.
Required Technical Skills:
- Microsoft Azure: Azure Virtual Machines, Virtual Networks, Azure Active Directory, Azure Update Management.
- Containers & Orchestration
- Production-level management of GKE, AKS, and EKS.
- Advanced mastery of Docker, Helm, Kubernetes StatefulSets, Pod Disruption
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Location - Bangalore Skill/Experience Expectations: 1. Total Experience 7-11 yrs 2. 3-4 years in managing scalable production environment 3. 2-4 yr experience in managing Google cloud infrastructure 4. proficient in terraform and any programming language 5. Expert in designing and managing observability solutions 6. 5 yr experience in DevOps and SRE practices and troubleshooting critical incidents.






