Sr.DevOps Engineer at B2B2C tech Web3 startup · Bengaluru (Bangalore) · 4 - 10 years · ₹15L - ₹35L / yr · Posted 23 Nov 2022
Our Client is a B2B2C tech Web3 startup founded by founders - IITB Graduates who are experienced in retail, ecommerce and fintech.
Vision: Client aims to change the way that customers, creators, and retail investors interact and transact at brands of all shapes and sizes. Essentially, becoming the Web3 version of brands driven social ecommerce & investment platform. We have two broader development phases to achieve our mission.
candidate will be responsible for automating the deployment of cloud infrastructure and services to
support application development and hosting (architecting, engineering, deploying, and operationally
managing the underlying logical and physical cloud computing infrastructure).
Location: Bangalore
Reporting Manager: VP, Engineering
Job Description:
● Collaborate with teams to build and deliver solutions implementing serverless,
microservice-based, IaaS, PaaS, and containerized architectures in GCP/AWS environments.
● Responsible for deploying highly complex, distributed transaction processing systems.
● Work on continuous improvement of the products through innovation and learning. Someone with
a knack for benchmarking and optimization
● Hiring, developing, and cultivating a high and reliable cloud support team
● Building and operating complex CI/CD pipelines at scale
● Work with GCP Services, Private Service Connect, Cloud Run, Cloud Functions, Pub/Sub, Cloud
Storage, Networking in general
● Collaborate with Product Management and Product Engineering teams to drive excellence in
Google Cloud products and features.
● Ensures efficient data storage and processing functions in accordance with company security
policies and best practices in cloud security.
● Ensuring scaled database setup/montioring with near zero downtime
Key Skills:
● Hands-on software development experience in Python, NodeJS, or Java
● 5+ years of Linux/Unix Administration monitoring, reliability, and security of Linux-based, online,
high-traffic services and Web/eCommerce properties
● 5+ years of production experience in large-scale cloud-based Infrastructure (GCP preferred)
● Strong experience with Log Analysis and Monitoring tools such as CloudWatch, Splunk,
Dynatrace, Nagios, etc.
● Hands-on experience with AWS Cloud – EC2, S3 Buckets, RDS
● Hands-on experience with Infrastructure as a Code (e.g., cloud formation, ARM, Terraform,
Ansible, Chef, Puppet) and Version control tools
● Hands-on experience with configuration management (Chef/Ansible)
● Experience in designing High Availability infrastructure and planning for Disaster Recovery
solutions
● Knowledgeable in networking and security Familiar with GCP services (in Databases, Containers,
Compute, stores, etc) with comfort in GCP serverless technologies
● Exposure to Linkerd, Helm charts, and Ambassador is mandatory
● Experience with Big Data on GCP BigQuery, Pub/Sub, Dataproc, and Dataflow is plus
● Experience with Groovy and writing Jenkinsfile
● Experience with time-series and document databases (e.g. Elasticsearch, InfluxDB, Prometheus)
● Experience with message buses (e.g. Apache Kafka, NATS)
Regards
Team Merito

Similar jobs (10)
Platform Engineering Lead (For client company)
Location: Pune, India
Experience: 7+ years
What Success Looks Like
- Engineering teams ship faster with confidence and built-in guardrails.
- Cloud cost, security, and reliability are predictable, measurable, and well-managed.
- CI/CD pipelines are trusted, standardized, and production-ready.
- Platform decisions reduce cognitive load instead of introducing unnecessary process.
Scope & Expectations
This is a hands-on leadership role combining architecture and implementation.
You will:
- Build, not just review.
- Own the platform roadmap—not just infrastructure tickets.
- Act as a force multiplier for product engineering teams rather than becoming a bottleneck.
- Drive platform strategy while remaining deeply involved in execution.
Key Responsibilities
Platform & Cloud Architecture
- Own Zoop's platform and cloud architecture across GCP and AWS.
- Design reusable, opinionated platform patterns instead of one-off infrastructure.
- Build and evolve Zoop's Internal Developer Platform (IDP), including:
- Self-service environments
- Golden paths (paved roads)
- Standardized templates
- Built-in engineering guardrails
- Lead Kubernetes and cloud-native adoption at scale.
- Drive infrastructure automation using Terraform, Pulumi, or similar Infrastructure-as-Code (IaC) tools.
CI/CD, Reliability & Developer Experience
- Establish robust CI/CD practices with quality gates and production readiness.
- Improve deployment safety through automation and testing.
- Define and monitor:
- Golden Signals
- SLIs
- SLOs
- Incident response processes
- Reduce operational toil and improve developer productivity.
- Make observability a first-class capability using cost-efficient monitoring systems.
- Build an observability platform that multiple engineering teams can easily integrate into their applications.
Security, Privacy & Compliance
- Build security-by-default into infrastructure and deployment pipelines.
- Lead implementation and continuous compliance for:
- DPDP Act (India)
- ISO 27001:2022
- SOC 2 Type II
- Implement:
- Zero Trust architecture
- Least-privilege access
- Secure data isolation
FinOps & Cloud Optimization
- Make cloud costs transparent and accountable across engineering teams.
- Establish FinOps practices including:
- Budgets
- Cost alerts
- Optimization routines
- Drive build-vs-buy decisions using clear ROI analysis.
AI, Data & MLOps Foundations
- Build secure and scalable foundations for AI and MLOps workloads.
- Define guardrails for AI systems and sensitive data handling.
Leadership & Collaboration
- Partner closely with engineering teams to align infrastructure strategy with product goals.
- Mentor engineers and guide teams through technical change.
- Balance long-term platform initiatives with practical execution.
What We're Looking For
Experience
- 7+ years of experience building and operating production infrastructure.
- Experience scaling engineering platforms in high-growth or regulated companies.
- Strong hands-on expertise in:
- Kubernetes and the cloud-native ecosystem
- Service Mesh technologies
- Policy Engines
- GCP, AWS (Azure exposure is a plus)
- Terraform and Infrastructure as Code
Engineering & Operations
- Strong understanding of SDLC and modern CI/CD systems (Jenkins, GitOps, etc.).
- Experience with observability tools such as:
- Grafana
- Prometheus
- New Relic
- Comfortable reading and contributing to production systems written in:
- Go
- Python
- Node.js
Security & Compliance
- Practical experience implementing ISO 27001 and SOC 2 controls.
- Strong understanding of:
- Data protection
- Privacy
- Identity and access management
- Security best practices
Mindset
We're looking for someone who is:
- Action-oriented with sound engineering judgment.
- Analytical, cost-conscious, and reliability-focused.
- Collaborative, calm under pressure, and open to feedback.
- Comfortable challenging decisions and explaining trade-offs when necessary.
Nice to Have
- Experience in fintech, identity, or other regulated industries.
- Built Internal Developer Platforms (IDPs) or shared infrastructure tooling.
- Contributions to open-source projects.
About the Team
SecurITe’s mission is to build an Agentic-AI driven security platform that protects critical infrastructure from modern cyber threats. Our focus is on delivering highly performant, resilient, and intelligent network security systems that help defenders stay ahead of adversaries.
About the Role
We’re looking for an exceptional Platform Engineer to help lead the development of our next-generation cybersecurity AI platform. This is a rare opportunity to shape how agentic AI transforms the future of cyber defense.
As a Platform Engineer, you will design, build, and operate the foundational infrastructure, deployment systems, and developer platforms that power our cybersecurity products across cloud and on-premise environments. You will work at the intersection of infrastructure engineering, cloud-native technologies, automation, reliability, and security to enable scalable and resilient product delivery.
You’ll collaborate closely with AI/ML, backend, security, QA, and product engineering teams to create self-service platforms, deployment pipelines, observability systems, and operational tooling that accelerate innovation while maintaining enterprise-grade reliability and security.
This role is ideal for Linux platform engineers and system specialists who excel at solving complex system challenges, automating wherever possible, and building resilient platforms that enable teams to move faster with confidence.
Why This Role Matters
Cybersecurity is undergoing a fundamental shift. AI is no longer an enhancement—it’s becoming the core engine of how detection, investigation, and response are executed. As our Platform Engineer, you will architect and build the infrastructure, automation, deployment, and operational systems that make this transformation real.
Your work will directly influence the scalability, reliability, and security of our AI-driven cybersecurity platform across both cloud and enterprise on-premise deployments. You’ll help establish the operational backbone that enables rapid innovation, secure product delivery, and resilient large-scale deployments in mission-critical environments.
This is a chance to solve novel technical challenges involving distributed systems, hybrid infrastructure, observability, automation, and secure software delivery while shaping how defenders outpace modern attackers.
What You’ll Do
● Design, build, and own AWS infrastructure from the ground up (VPC architecture, EC2 fleet management, IAM, networking, security groups)
● Administer and harden AlmaLinux VMs across production, staging, and dev environments
● Build automation for provisioning, patching, and configuration management (infrastructure-as-code, config management tooling)
● Design and implement observability: monitoring, logging, alerting, and on-call-worthy SLAs from scratch
● Lead incident response — diagnosis, RCA, and post-incident documentation — with no dedicated ops team to escalate to
● Make and document build-vs-buy and architecture decisions as the product and team scale
● Work directly with founders/engineering to translate ambiguous asks into scoped technical plans
Impact You’ll Have
● Accelerate engineering velocity through scalable developer platforms and automation
● Improve deployment reliability, platform uptime, and operational efficiency
● Enable secure and scalable AI-driven cybersecurity workloads
● Reduce operational overhead through infrastructure automation and self-service systems
● Help establish enterprise-grade cloud and on-premise deployment capabilities
● Enhance product resiliency, observability, and operational excellence
● Shape the long-term platform architecture powering next-generation cybersecurity products
● Enable rapid and secure delivery of critical security innovations to customers
Required Experience
● 4+ years hands-on Linux administration (RHEL-family strongly preferred — AlmaLinux, CentOS, RHEL)
● Deep Linux internals: systemd, networking, storage/LVM, process/resource management, kernel-level troubleshooting
● Real AWS architecture experience — not just operating existing infra, but designing it (VPC, EC2, IAM, security groups, networking)
● Demonstrated ability to scope and solve ambiguous problems independently, without a runbook or senior engineer to defer to
● Scripting/automation proficiency (Python and/or Bash) beyond one-off scripts — built tooling that runs unattended
● Track record of end-to-end ownership: has designed, built, and operated a system (not just contributed to one)
● Clear, proactive communicator — documents decisions and explains reasoning without being asked
Required Skills & Qualifications
● Strong Linux system administration and troubleshooting skills
● Redhat certifications
● Strong understanding of networking fundamentals, security, and distributed systems
● Proficiency with Docker, and container orchestration
● Experience with Terraform, Ansible, or similar infrastructure automation tools
● Strong scripting or programming skills in Python, Bash, or Go
● Knowledge of observability tools such as Prometheus, Grafana, ELK, or OpenTelemetry
● Understanding of platform security best practices and secure infrastructure design
● Familiarity with virtualization technologies and hybrid infrastructure environments
● Strong problem-solving and debugging abilities
● Excellent communication and collaboration skills
● Ability to thrive in fast-paced startup environments
Nice to Have
● Configuration management/automation at scale (Ansible, AWX, Terraform, or similar)
● Monitoring/observability stack experience (Prometheus, Grafana, Zabbix, ELK, Datadog)
● Container experience (Docker; Kubernetes a plus but not core to this VM-based stack)
● Experience in a startup or small-team environment where infra was built from scratch
● Security/compliance exposure (vulnerability remediation, hardening, SSO/access control)
The Mindset
Problem Solver
You thrive on complex, ambiguous challenges and engineer elegant solutions.
Ownership-Driven
You take initiative, move fast, and deliver outcomes without hand-holding.
Continuous Learner
You stay ahead of the curve in AI, ML, cloud-native technologies, and emerging infrastructure trends.
Startup DNA
You excel in fast-moving environments where priorities evolve and impact is immediate.
Hiring Platform Engineer
Exp: 6 -- 10 yrs
Edu : BE/B.tech/MCA
Work Location : Pune
Skills :
Platform monitoring ,Incident trouble shooting, Incident recovery, openshift ,kubernetes.
2 years of IT operations, infrastructure, cloud or application support experience.
Exp in Linux command-line knowledge.
Exp in networking knowledge including IP addressing, DNS, ports and connectivity troubleshooting.

Platform Engineer
Location: Bengaluru, India (Hybrid)
Employment Type: Full-time
Experience: 2-4 years
About Compnay
This is driving the electric mobility revolution through cutting-edge software, infrastructure, and professional services. Our technology empowers utilities, cities, fleets, transit agencies, and automakers to deploy EV charging infrastructure at scale safely, efficiently, and sustainably. With a global footprint spanning three continents and operations in 13 countries, we are passionate about shaping the future of sustainable transport.
Operating over 70,000 charge points globally, this is driving the transition toward cleaner, smarter, and more efficient mobility. The India team serves as a critical operational hub, supporting global platforms focused on decarbonization, digitalization, and scalable infrastructure growth.
Role Overview
What you’ll do:
- Ensure system reliability, uptime, and performance of global platform.
- Conduct real-time surveillance of our EV charging systems to proactively identify and mitigate performance issues and anomalies near 24/7 basis. As such, you collaborate with IDT and FMC players to ensure incident detection also happens outside office hours (monitoring shifts among team members subject to duty schedule).
- Deliver on change & releases like firmware changes and drive insights & intelligence back into testing processes and tech discussions with the wider organization.
- Successfully deliver and project manage first time right commissioning activities alongside our Engineering Procurement Contract Management (EPCM) partners to successfully bring charge points onto our Charge Point Management System (CPMS).
- Provide technical guidance and support to DC specialists during the commissioning of EV charging solutions.
- Work closely with Shell, Engineering, and IT colleagues to ensure projects are completed on time and to specification.
- Act as a liaison with the Engineering Procurement Contract Management (EPCM) partner to manage projects from start to finish, ensuring charge points are successfully onboarded on the Charge Point Management System (CPMS).
- Collaborate with development, operations and support teams to build scalable and resilient systems.
- Contribute to incident response, root-cause analysis, and post-mortem reviews, driving continuous improvement.
- Participate in capacity planning, performance tuning, and resource optimization.
- Integrate security and compliance best practices into all infrastructure operations.
- Stay current with emerging SRE tools, frameworks, and cloud technologies to continuously improve reliability practices.
- Participate in and lead on-call rotations and incident response, conducting detailed postmortems and RCA reports.
- Flexible to resolve blocking issues during off hours or weekends if required.
What We’re Looking For:
Basic Qualifications and Skills
- Bachelor’s degree in Engineering, Electrical, ECE, Computer Science, Information Technology, or related field.
- 2–4 years of overall experience with at least 1+ years of experience as a Site Reliability Engineer, DevOps Engineer, or Technical Project Coordinator.
- Proven experience of DevOps, SRE or Technical Project Coordination with IoT or connected devices-based platforms.
- Experience with incident management and on-call best practices. Provide support to on-call engineers.
- Excellent analytical and problem-solving skills with a proactive mindset.
- Expertise with monitoring and observability tools (Dynatrace, Prometheus, Grafana, Zabbix, etc.).
- Solid understanding of cloud platforms (AWS) and AWS native services (EKS, EC2, S3, RDS, Lambda).
- Proactively monitor the network, triage performance outliers, and coordinate correction actions to ensure optimal system functionality.
- Fluency in English (spoken and written).
- Successfully recommission or decommission chargers following changes in our network.
- Responsible for the go-live of the chargers on Shell’s public network following commissioning attempts.
Additional Information
- This role involves managing infrastructure for a global platform operating in over ten countries, requiring effective communication and collaboration across regions.
- Strong verbal and written communication skills, along with availability and flexibility to resolve blocking issues, are essential to support on-call engineers.
- This role may involve EU or US time-zone shifts based on business requirements.
- Shift timing: 2 PM IST to 11 PM IST.
What is required to be successful in this role:
- Global platform experience (B2C or B2B).
- AWS native service experience.
- Firmware deployment and cloud cost optimization experience.
- Strong exposure to monitoring and alerts.
- Experience with firmware rollout, IoT devices onboarding and offboarding will be an added advantage.
- Experience as an SRE or DevOps Engineer with some exposure to Project Management or Technical Project Management in IoT-based projects will be helpful.
What We Offer
- Work with some of the brightest minds in the emerging EV industry.
- Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
- Freedom to suggest, implement, and innovate on systems, processes, and technologies.
- Daily ownership in a high-growth, challenging environment.
- Flexible work environment with hybrid schedules and virtualization options.
- Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.
Job Title : SDE 3 – Infrastructure Platform Engineer
Experience : 5.5 to 8.5 Years
Number of Positions : 2
Employment Type : C2H (Contract to Hire)
Work Mode : Remote during contractual period → 5 Days WFO after conversion
Contract Duration : 3 Months
Post-Conversion Location : Pune
Notice Period : Immediate Joiners / Serving Notice Period / Up to 15 Days preferred
(Candidates officially serving a 30-day notice period may also be considered if they are on the bench and have a negotiable joining date)
Role Overview :
We are looking for an experienced SDE 3 – Infrastructure Platform Engineer to design, build, and operate scalable, secure, and highly reliable cloud infrastructure and internal platform capabilities.
The ideal candidate will have strong hands-on experience in Cloud Infrastructure, Infrastructure as Code (IaC), CI/CD, Docker, Kubernetes, automation, observability, networking, and distributed systems.
Mandatory Skills : AWS / Azure / GCP, Terraform / CloudFormation, Kubernetes, Docker, CI/CD, Platform / Infrastructure Engineering, Python / Go / Java / Ruby, Networking, Cloud Security, Distributed Systems, Scalability & Reliability, Strong Coding & Automation.
Key Responsibilities :
- Design and maintain scalable, highly available infrastructure on AWS / GCP / Azure.
- Build and manage Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools.
- Develop automation for infrastructure provisioning, deployments, monitoring, and operations.
- Manage and optimize Docker and Kubernetes workloads.
- Build internal platform tools to improve developer productivity and engineering efficiency.
- Implement monitoring, logging, alerting, and observability solutions.
- Participate in incident response, RCA, postmortems, and reliability improvements.
- Design and improve CI/CD pipelines and deployment automation.
- Contribute to system design, architecture discussions, scalability, security, and cost optimization.
- Collaborate with application, data, and product engineering teams.
Required Skills :
- 5.5 to 8.5 years of experience in Infrastructure / Platform Engineering or similar roles.
- Strong hands-on experience with AWS, GCP, or Azure.
- Strong expertise in Terraform / CloudFormation.
- Experience with CI/CD, Docker, and Kubernetes.
- Strong programming skills in at least one of:
- Python, Go, Java, or Ruby.
- Good understanding of networking, cloud security, distributed systems, scalability, and reliability.
- Experience working with production infrastructure and highly available systems.
- Strong troubleshooting and problem-solving skills.
Nice to Have :
- Experience with SRE practices and production on-call ownership.
- Experience in fintech, payments, banking, or transaction-heavy systems.
- Knowledge of cloud security, compliance, or FinOps/cost optimization.
- Experience building internal developer platforms or productivity tools.
- Previous product company experience.
Interview Process :
Round 1 : Take-Home Coding Assignment – Submit within 48 hours
Round 2 : Coding Assignment Discussion – 1 Hour
Round 3 : Technical Managerial Round – 30 Minutes
Note : The take-home coding assignment is mandatory. Candidates should be comfortable completing and submitting the assignment within 48 hours before proceeding.
Ideal Candidate :
Strong Platform / Infrastructure Engineer with hands-on experience in :
Cloud + Terraform / CloudFormation + Kubernetes + CI/CD + Programming + SRE / Production Operations
Pure DevOps profiles without strong coding and platform engineering experience are not preferred.
About the role
You will be a pivotal member of the leadership team, responsible for driving the company's product and platform strategy while overseeing engineering excellence across AI, data, platform, DevOps and infrastructure. This is a mission-critical role for an engineering leader who is both strategic and hands-on, and who can build scalable technologies for global enterprise clients in a fast-paced innovation ecosystem.
Reports to: CEO · Location: Bengaluru, India — hybrid, 3 days a week in office
What you will do
Technology & product leadership
- Define and drive the technology vision, architecture and long-term platform roadmap.
- Oversee the architecture, design and delivery of highly scalable enterprise systems.
- Ensure engineering excellence, velocity and reliability across the product lifecycle.
Engineering & platform management
- Lead platform engineering, product technology, DevOps, infrastructure and Quality Engineering.
- Build robust cloud-native systems using the Azure, AWS and GCP ecosystems.
- Oversee operational effectiveness, including uptime, production reliability and cost optimisation.
Innovation & AI strategy
- Spearhead Mission AI by developing scalable, production-grade AI/ML and GenAI capabilities.
- Own the GenAI/LLM solutions architecture.
- Core specialisation in AI agents and autonomous workflows, data extraction and intelligent automation.
- Direct hands-on architectural oversight of large language models, applied AI and multi-layered deep learning products.
- Extensive expertise in building intelligent document automation systems — similar in complexity to patent parsing, legal workflow automation and structured decision intelligence tools.
Technical leadership
- A proven track record overseeing product architecture, tech strategy and cross-functional engineering execution across core full-stack and AI platforms.
- Collaborate with executive leadership on business strategy, client requirements and product delivery.
- Build, mentor and scale high-performing engineering teams with a growth mindset.
- Establish a strong technology culture grounded in ownership, innovation and continuous learning.
What success looks like
- Mission AI: build a world-class AI system leveraging GenAI, ML and enterprise-grade data engineering.
- Hyper-scaling: architect and scale the platform to match global industry leaders in the IP space.
- Culture building: develop a strong engineering organisation with high ownership, performance and innovation DNA.
Qualifications & experience
- Preferably under 15 years of enterprise software engineering experience across B2B SaaS or technology-first companies; fewer is fine for an exceptional engineer.
- A proven track record of taking early-stage AI/ML prototypes and scaling them into robust enterprise SaaS platforms featuring multi-agent orchestration, complex workflows and decision intelligence.
- Experienced in building and mentoring agile, lean startup teams of full-stack and machine learning engineers from the ground up.
- Proven leadership in defining and executing technology strategy and platform roadmaps.
- Extensive cloud-native engineering experience with Azure, AWS and GCP.
Technical expertise
- Strong full-stack engineering background (Java, Python, JavaScript frameworks).
- Expertise with JS frameworks such as React, Angular and Node.js.
- Experience building and scaling distributed systems and microservices.
- Strong knowledge of databases (SQL, NoSQL), data modelling and unstructured data management.
- Strong understanding of Agile methodologies and tools (Atlassian, Git, CI/CD pipelines).
Behavioural & leadership competencies
- Product and delivery management expertise, end to end, including delivery and customer support.
- Excellent communication, with the ability to influence executive stakeholders.
- High technical proficiency combined with strong business acumen.
- Strong analytical and decision-making skills.
Job Description:
- Infrastructure Management: Design, implement, and manage scalable, reliable, and secure cloud infrastructure using AWS, GCP, and/or Azure.
- CI/CD Pipelines: Develop and maintain continuous integration and continuous deployment (CI/CD) pipelines to streamline the development lifecycle.
- Automation: Automate infrastructure provisioning, configuration management, and application deployment processes.
- Monitoring and Performance: Implement monitoring, logging, and alerting solutions to ensure system health, performance, and reliability.
- Security: Ensure the security of cloud infrastructure and applications, including identity management and compliance with industry standards.
- Collaboration: Work closely with client and development teams to integrate DevOps practices and deliver high-quality software.
- Documentation: Maintain comprehensive documentation of infrastructure, configurations, and processes.
- Innovation: Stay current with emerging technologies and industry trends, integrating them into the DevOps strategy as appropriate.
Qualifications:
- Education: Bachelor's degree in Computer Science, Information Technology, or a related field.
- Experience: 7 - 10 years of overall experience with relevant experience of at least 7 years in DevOps and served as a lead or senior engineer.

Key Skills:
• Bachelor's or Master's degree in Computer Science or related field.
• Minimum 5 years of experience in Platform Engineering, DevOps, or Cloud Infrastructure Engineering.
• Experience migrating data and systems between AWS IaaS and PaaS.
• Experience operating and supporting applications using AWS VPC, EKS, and related services for multi-account operations.
• Experience developing fast and reliable Continuous Integration/Continuous Deployment (CI/CD) workflows used by hundreds of application teams.
• Experience administering and troubleshooting Operating Systems such as Linux, Windows, and MacOS.
• Professional Certifications in AWS Networks, CNCF Technologies, or Kubernetes.
• Experience using and configuring observability tools such as ELK, Prometheus/Grafana, AWS CloudWatch, and Jaeger.
• Experience of applied GitOps principles using ArgoCD or Flux.
• Public examples of code you've worked on with other people using any of these technologies:
o Configuration management/Infrastructure as Code (IAC) tools, such as AWS CDK, AWS CloudFormation, Terraform, Ansible, or Puppet.
o Systems solutions in one or more programming languages, such as Golang, Python, Java.
o Build, Release, Deploy or Ops Workflows using Bamboo, Argo Project, or GitHub Actions.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Location - Bangalore Skill/Experience Expectations: 1. Total Experience 7-11 yrs 2. 3-4 years in managing scalable production environment 3. 2-4 yr experience in managing Google cloud infrastructure 4. proficient in terraform and any programming language 5. Expert in designing and managing observability solutions 6. 5 yr experience in DevOps and SRE practices and troubleshooting critical incidents.
Lead Cloud Reliability Engineer
Job Responsibilities
● Lead and manage the Cloud Reliability teams to provide strong Managed Services support to end-customers.
● Isolate, troubleshoot and resolve issues reported by CMS clients in their cloud environment
● Drive the communication with the customer providing details about the issue, current steps, next plan of action, ETA
● Gather client's requirements related to use of specic cloud services and provide assistance in seing them up and resolving issues
● Create SOPs and knowledge articles for use by the L1 teams to resolve common issues
● Identify recurring issues, perform root cause analysis and propose/implement preventive actions
● Follow change management procedure to identify, record and implement changes
● Plan and deploy OS, security patches in Windows/Linux environment and upgrade k8s clusters
● Identify the recurring manual activities and contribute to automation
● Provide technical guidance and educate team members on development and operations. Monitor metrics and develop ways to improve.
● System troubleshooting and problem-solving across plaorm and application domains. Ability to use a wide variety of open-source technologies and cloud services.
● Build, maintain, and monitor conguration standards.
● Ensuring critical system security through using best-in-class cloud security solutions.
Qualifications
● 4-7 years experience in Cloud Infrastructure and Operations domains and IT operational experience preferably in a global enterprise environment.
● Specialize in one or two cloud deployment platforms: AWS, GCP
● Hands on experience with AWS/GCP services (EKS, ECS, EC2, VPC, RDS, Lambda, GKE, Compute Engine)
● Understanding of one or more programming languages (Python, JavaScript, Ruby, Java, .Net)
● Logging and Monitoring tools (ELK, Stackdriver, CloudWatch)
● Knowledge on Conguration Management tools such as Ansible, Terraform, Puppet, Chef
● Experience working with deployment and orchestration technologies (such as Docker, Kubernetes, Mesos)
● Good analytical, communication, problem solving, and learning skills.
● Knowledge on programming against cloud plaorms such as Google Cloud Platform and lean development methodologies.
● Strong service aitude and a commitment to quality.
● Willingness to work in shifts.






