Cutshort logo
For Employers
Infra360 Solutions Pvt Ltd logo
Senior DevOps Engineer (SRE)
Senior DevOps Engineer (SRE)

Senior DevOps Engineer (SRE) at Infra360 Solutions Pvt Ltd · Gurugram · 3 - 8 years · ₹10L - ₹15L / yr · Bootstrapped · Posted 13 Dec 2024

Infra360 Solutions Pvt Ltd's logo

Senior DevOps Engineer (SRE)

HR Infra360's profile picture
Posted by HR Infra360
3 - 8 yrs
₹10L - ₹15L / yr
Gurugram
Skills
skill iconDocker
skill iconKubernetes
DevOps
skill iconAmazon Web Services (AWS)
Windows Azure
Google Cloud Platform (GCP)
helm

Please Apply - https://zrec.in/7EYKe?source=CareerSite


About Us

Infra360 Solutions is a services company specializing in Cloud, DevSecOps, Security, and Observability solutions. We help technology companies adapt DevOps culture in their organization by focusing on long-term DevOps roadmap. We focus on identifying technical and cultural issues in the journey of successfully implementing the DevOps practices in the organization and work with respective teams to fix issues to increase overall productivity. We also do training sessions for the developers and make them realize the importance of DevOps. We provide these services - DevOps, DevSecOps, FinOps, Cost Optimizations, CI/CD, Observability, Cloud Security, Containerization, Cloud Migration, Site Reliability, Performance Optimizations, SIEM and SecOps, Serverless automation, Well-Architected Review, MLOps, Governance, Risk & Compliance. We do assessments of technology architecture, security, governance, compliance, and DevOps maturity model for any technology company and help them optimize their cloud cost, streamline their technology architecture, and set up processes to improve the availability and reliability of their website and applications. We set up tools for monitoring, logging, and observability. We focus on bringing the DevOps culture to the organization to improve its efficiency and delivery.


Job Description

Job Title:             Senior DevOps Engineer / SRE

Department:       Technology

Location:             Gurgaon

Work Mode:         On-site

Working Hours:   10 AM - 7 PM 

Terms:                 Permanent

Experience:      4-6 years

Education:           B.Tech/MCA

Notice Period:     Immediately

​

About Us

At Infra360.io, we are a next-generation cloud consulting and services company committed to delivering comprehensive, 360-degree solutions for cloud, infrastructure, DevOps, and security. We partner with clients to transform and optimize their technology landscape, ensuring resilience, scalability, cost efficiency and innovation.

Our core services include Cloud Strategy, Site Reliability Engineering (SRE), DevOps, Cloud Security Posture Management (CSPM), and related Managed Services. We specialize in driving operational excellence across multi-cloud environments, helping businesses achieve their goals with agility and reliability.

We thrive on ownership, collaboration, problem-solving, and excellence, fostering an environment where innovation and continuous learning are at the forefront. Join us as we expand and redefine what’s possible in cloud technology and infrastructure.


Role Summary

We are seeking a Senior DevOps Engineer (SRE) to manage and optimize large-scale, mission-critical production systems. The ideal candidate will have a strong problem-solving mindset, extensive experience in troubleshooting, and expertise in scaling, automating, and enhancing system reliability. This role requires hands-on proficiency in tools like Kubernetes, Terraform, CI/CD, and cloud platforms (AWS, GCP, Azure), along with scripting skills in Python or Go. The candidate will drive observability and monitoring initiatives using tools like Prometheus, Grafana, and APM solutions (Datadog, New Relic, OpenTelemetry).

Strong communication, incident management skills, and a collaborative approach are essential. Experience in team leadership and multi-client engagement is a plus.


Ideal Candidate Profile


  • Solid 4-6 years of experience as an SRE and DevOps with a proven track record of handling large-scale production environments
  • Bachelor's or Master's degree in Computer Science, Engineering, or a related field
  • Strong Hands-on experience with managing Large Scale Production Systems
  • Strong Production Troubleshooting Skills and handling high-pressure situations.
  • Strong Experience with Databases (PostgreSQL, MongoDB, ElasticSearch, Kafka)
  • Worked on making production systems more Scalable, Highly Available and Fault-tolerant
  • Hands-on experience with ELK or other logging and observability tools
  • Hands-on experience with Prometheus, Grafana & Alertmanager and on-call processes like Pagerduty
  • Problem-Solving Mindset
  • Strong with skills - K8s, Terraform, Helm, ArgoCD, AWS/GCP/Azure etc
  • Good with Python/Go Scripting Automation
  • Strong with fundamentals like DNS, Networking, Linux
  • Experience with APM tools like - Newrelic, Datadog, OpenTelemetry
  • Good experience with Incident Response, Incident Management, Writing detailed RCAs
  • Experience with Applications best practices in making apps more reliable and fault-tolerant
  • Strong leadership skills and the ability to mentor team members and provide guidance on best practices.
  • Able to manage multiple clients and take ownership of client issues.
  • Experience with Git and coding best practices


Good to have

  • Team-leading Experience
  • Multiple Client Handling
  • Requirements gathering from clients
  • Good Communication


Key Responsibilities


  1. Design and Development:
  2. Architect, design, and develop high-quality, scalable, and secure cloud-based software solutions.
  3. Collaborate with product and engineering teams to translate business requirements into technical specifications.
  4. Write clean, maintainable, and efficient code, following best practices and coding standards.
  5. Cloud Infrastructure:
  6. Develop and optimise cloud-native applications, leveraging cloud services like AWS, Azure, or Google Cloud Platform (GCP).
  7. Implement and manage CI/CD pipelines for automated deployment and testing.
  8. Ensure the security, reliability, and performance of cloud infrastructure.
  9. Technical Leadership:
  10. Mentor and guide junior engineers, providing technical leadership and fostering a collaborative team environment.
  11. Participate in code reviews, ensuring adherence to best practices and high-quality code delivery.
  12. Lead technical discussions and contribute to architectural decisions.
  13. Problem Solving and Troubleshooting:
  14. Identify, diagnose, and resolve complex software and infrastructure issues.
  15. Perform root cause analysis for production incidents and implement preventative measures.
  16. Continuous Improvement:
  17. Stay up-to-date with the latest industry trends, tools, and technologies in cloud computing and software engineering.
  18. Contribute to the continuous improvement of development processes, tools, and methodologies.
  19. Drive innovation by experimenting with new technologies and solutions to enhance the platform.
  20. Collaboration:
  21. Work closely with DevOps, QA, and other teams to ensure smooth integration and delivery of software releases.
  22. Communicate effectively with stakeholders, including technical and non-technical team members.
  23. Client Interaction & Management: 
  24. Will serve as a direct point of contact for multiple clients.
  25. Able to handle the unique technical needs and challenges of two or more clients concurrently. 
  26. Involve both direct interaction with clients and internal team coordination.
  27. Production Systems Management: 
  28. Must have extensive experience in managing, monitoring, and debugging production environments. 
  29. Will work on troubleshooting complex issues and ensure that production systems are running smoothly with minimal downtime.
Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About Infra360 Solutions Pvt Ltd

Founded :
2022
Type :
Services
Size :
20-100
Stage :
Bootstrapped

About

At Infra360.io, we are a next-generation cloud consulting and services company committed to delivering comprehensive, 360-degree solutions for Cloud, Infrastructure, DevOps, MLOps and Security. We partner with clients to modernize and optimize their cloud, ensuring resilience, scalability, cost efficiency and innovation.


We thrive on ownership, collaboration, problem-solving, and excellence, fostering an environment where innovation and continuous learning are at the forefront. Join us as we expand and redefine what’s possible in cloud technology and infrastructure.


Read more

Candid answers by the company

What does the company do?
What is the location preference of jobs?

Our core services include Cloud Strategy, Site Reliability Engineering (SRE), DevOps, Cloud Security Posture Management (CSPM), and related Managed Services. We specialize in driving operational excellence across multi-cloud environments, helping businesses achieve their goals with agility and reliability.

Company social profiles

bloglinkedin

Similar jobs (10)

Staffnixcom
Mayank Choudhary
Posted by Mayank Choudhary
Gurugram, Hyderabad
6 - 8 yrs
₹40L - ₹50L / yr
DevOps

Strong Hands-On DevOps Engineer Profile with deep Kubernetes and multi-cloud experience

2

Mandatory (Experience 1): Must have 6+ years of hands-on DevOps experience, deploying, operating, troubleshooting, and scaling enterprise SaaS environments.

3

Mandatory (Experience 2): Must be able to own end-to-end customer deployments across cloud and enterprise environments (the role covers cloud, BYOC, air-gapped, and data-center deployments).

4

Mandatory (Tech skill 1): Must have strong hands-on Kubernetes experience — troubleshooting, networking, workloads, storage, RBAC, and security

5

Mandatory (Tech skill 2): Must have hands-on Helm experience — deployment, templating, and troubleshooting.

6

Mandatory (Tech skill 3): Must have hands-on Terraform experience — infrastructure provisioning and automation

7

Mandatory (Tech skill 4): Must have strong Linux and cloud infrastructure troubleshooting skills

8

Mandatory (Tech skill 5): Must have hands-on experience with observability — metrics, logs, traces, and alerting

9

Mandatory (Tech skill 6): Must have hands-on experience with one or more of AWS / GCP / Azure / OCI

10

Mandatory (Tech skill 7): Must have a strong incident-management and production-troubleshooting mindset, with a "fix first, optimize later" approach to customer escalations.

11

Mandatory (Company): B2B SaaS product companies

12

Preferred (Enterprise deployments): Prior experience with BYOC, air-gapped, or restricted-network environments.

13

Preferred (Observability tools): Prometheus / Grafana / OpenTelemetry / ELK / OpenSearch / Datadog

14

Preferred (Other): GitOps / CI-CD; Python or Golang for automation; multi-tenant SaaS infrastructure scaling; exposure to AI/ML pipeline deployments or iPaaS / reverse ETL connectors; SRE concepts (SLIs/SLOs, DR, high availability

Read more
Gurugram
5 - 10 yrs
₹12L - ₹18L / yr
DevOps
Reliability engineering
skill iconAmazon Web Services (AWS)
Terraform
Ansible
+18 more

Job Title : DevOps Engineer / Site Reliability Engineer (SRE)

Experience : 5+ Years

Location : Gurugram, Haryana

Work Mode : On-site (Full-time)


About the Role :

We are looking for a skilled DevOps Engineer with 5+ years of experience in cloud infrastructure, CI/CD, automation, Kubernetes, and Site Reliability Engineering (SRE). The ideal candidate will be responsible for building scalable cloud infrastructure, automating deployments, improving system reliability, and ensuring high availability across production environments.


Mandatory Skills :

AWS, Terraform, Ansible, CloudFormation, Jenkins, GitLab CI, GitHub Actions, Docker, Kubernetes, Helm, Python, Bash, Grafana, Prometheus, ELK Stack, CloudWatch, New Relic, SRE, CI/CD, Infrastructure as Code (IaC), Linux


Key Responsibilities :

  • Design, deploy, and manage cloud infrastructure primarily on AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB, Auto Scaling, Lambda).
  • Build and maintain Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation.
  • Develop and optimize CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
  • Deploy and manage containerized applications using Docker, Kubernetes, and Helm.
  • Implement monitoring and observability using Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
  • Drive SRE practices by defining SLIs, SLOs, SLAs, handling production incidents, conducting RCA, and improving system reliability.
  • Automate operational tasks using Python, Bash, and Groovy scripting.
  • Collaborate with Development, QA, Security, and Operations teams to ensure reliable and secure software delivery.


Required Skills & Qualifications :

  • Bachelor's degree in Computer Science, IT, Electronics, or a related field.
  • 5+ years of experience in DevOps, SRE, or Cloud Infrastructure.
  • Strong expertise in AWS, with exposure to Azure/GCP.
  • Hands-on experience with Terraform, Ansible, CloudFormation, Docker, Kubernetes, Helm, Jenkins, GitLab CI, GitHub Actions, and Git.
  • Strong scripting skills in Python and Bash.
  • Experience with monitoring tools such as Grafana, Prometheus, ELK Stack, CloudWatch, and New Relic.
  • Good understanding of Linux, networking, SQL, and cloud security best practices.


Preferred Skills :

  • Experience with multi-cloud environments and DevSecOps practices.
  • Knowledge of disaster recovery, automation, and microservices architecture.
  • Strong troubleshooting, communication, and problem-solving skills.
Read more
TalentXO
tabbasum shaikh
Posted by tabbasum shaikh
Hyderabad, Gurugram
6 - 8 yrs
₹30L - ₹40L / yr
skill iconKubernetes
helm
Terraform
Linux/Unix
AWS
+2 more

Roles & Responsibilities

  1. Own end-to-end customer deployments across cloud, BYOC, air-gapped, and data-center environments
  2. Deploy and operate production environments using Kubernetes, Helm, and Terraform
  3. Troubleshoot complex Kubernetes, networking, infrastructure, application, and deployment issues
  4. Build and maintain observability, monitoring, alerting, and reliability
  5. Handle customer escalations and drive issues to resolution with a "fix first, optimize later" mindset
  6. Automate deployment and operational workflows
  7. Work closely with Engineering, Product, and Customer teams to resolve production challenges
  8. Participate in on-call and shift rotations, including critical customer escalations outside standard working hours

Ideal Candidate

1.Strong Hands-On DevOps Engineer Profile with deep Kubernetes and multi-cloud experience

2.Mandatory (Experience 1): Must have 6+ years of hands-on DevOps experience, deploying, operating, troubleshooting, and scaling enterprise SaaS environments.

3.Mandatory (Experience 2): Must be able to own end-to-end customer deployments across cloud and enterprise environments (the role covers cloud, BYOC, air-gapped, and data-center deployments).

4.Mandatory (Tech skill 1): Must have strong hands-on Kubernetes experience — troubleshooting, networking, workloads, storage, RBAC, and security

5.Mandatory (Tech skill 2): Must have hands-on Helm experience — deployment, templating, and troubleshooting.

6.Mandatory (Tech skill 3): Must have hands-on Terraform experience — infrastructure provisioning and automation

7.Mandatory (Tech skill 4): Must have strong Linux and cloud infrastructure troubleshooting skills

8.Mandatory (Tech skill 5): Must have hands-on experience with observability — metrics, logs, traces, and alerting

9.Mandatory (Tech skill 6): Must have hands-on experience with one or more of AWS / GCP / Azure / OCI

10.Mandatory (Tech skill 7): Must have a strong incident-management and production-troubleshooting mindset, with a "fix first, optimize later" approach to customer escalations.

11.Mandatory (Company): B2B SaaS product companies

12.Preferred (Enterprise deployments): Prior experience with BYOC, air-gapped, or restricted-network environments.

13.Preferred (Observability tools): Prometheus / Grafana / OpenTelemetry / ELK / OpenSearch / Datadog

14.Preferred (Other): GitOps / CI-CD; Python or Golang for automation; multi-tenant SaaS infrastructure scaling; exposure to AI/ML pipeline deployments or iPaaS / reverse ETL connectors; SRE concepts (SLIs/SLOs, DR, high availability).

Read more
Gurugram
4 - 8 yrs
₹6L - ₹15L / yr
DevOps
Linux administration
skill iconAmazon Web Services (AWS)
skill iconDocker
skill iconKubernetes
+20 more

Job Title : DevOps Engineer / Site Reliability Engineer (SRE)

Experience : 4+ Years

Location : Gurugram, Sector 48, Haryana (On-site)

Employment Type : Full-Time

Working Days : Monday to Saturday (1st & 3rd Saturday Off)


About the Role :

We are looking for a hands-on DevOps Engineer / Site Reliability Engineer (SRE) with strong experience in Linux, AWS, Kubernetes, Docker, CI/CD, Infrastructure as Code, and production application deployments.

The ideal candidate should have real-world production experience, excellent troubleshooting skills, and the ability to manage both infrastructure and application-level issues.


Mandatory Skills :

Linux, AWS, Docker, Kubernetes, Terraform, Ansible, Jenkins, GitHub Actions, GitLab CI/CD, CI/CD, Infrastructure as Code (IaC), Python, Bash, Git, Grafana, Prometheus, ELK, CloudWatch, New Relic, SRE (SLI/SLO/SLA), Networking (DNS, HTTP/HTTPS, TCP/IP, Load Balancer), Production Application Deployment & Troubleshooting


Key Responsibilities :

  • Manage and maintain AWS cloud infrastructure.
  • Build and optimize CI/CD pipelines using Jenkins, GitHub Actions, or GitLab CI.
  • Deploy, monitor, and troubleshoot applications across production environments.
  • Automate infrastructure using Terraform and Ansible.
  • Manage Docker containers and Kubernetes clusters.
  • Monitor systems using Grafana, Prometheus, ELK, CloudWatch, and New Relic.
  • Perform Linux server administration and troubleshooting.
  • Handle production incidents, Root Cause Analysis (RCA), and improve system reliability.
  • Collaborate with development teams to support application releases and automation.


Required Qualifications :

  • Bachelor's degree in Computer Science or related field.
  • 4+ years of hands-on experience in DevOps / SRE.
  • Strong Linux administration and production troubleshooting skills.
  • Experience with AWS and modern DevOps toolchains.
  • Hands-on experience with application deployment and production support.


What We're Looking For :

  • Strong practical Linux and cloud knowledge.
  • Real production experience with application deployments.
  • Ability to troubleshoot both infrastructure and application issues.
  • Experience handling live production incidents.
  • Excellent communication and problem-solving skills.
  • Candidates should be comfortable with scenario-based technical discussions and demonstrate genuine hands-on expertise.


Interview Process :

  1. HR Screening
  2. Technical Round
  3. Client Technical Round
  4. Final Discussion


Note : The interview will focus on practical hands-on experience in Linux, AWS, Kubernetes, Docker, CI/CD, Infrastructure as Code, application deployment, production troubleshooting, and real-world DevOps scenarios.

Read more
Qoves Inc
Agency job
via Qoves Inc by Jessie James Acusar
Remote only
5 - 10 yrs
Best in industry
skill iconKubernetes
CI/CD
Linux administration
Terraform

About the Role


The non-negotiable is deep, hands-on infrastructure expertise spanning hybrid cloud and self-hosted systems. You will architect and maintain our unique infrastructure combining AWS CDN, bare metal servers, and Kubernetes clusters. This is not just maintenance work: you will establish company-wide DevOps policy, build security-hardened environments, and create the documentation and processes that scale with us. You will design our CI/CD pipelines, implement zero-trust networking, and guide the technical team on how to keep our production systems robustly online. This role is hands-on, autonomous, and sets the standard for how we approach infrastructure as the company grows.


What You'll Build

  •  Hybrid Infrastructure Management: Architect and maintain our unique infrastructure spanning AWS CDN, bare metal servers, and Kubernetes clusters for computationally intensive facial analysis workloads.
  • CI/CD Pipeline Architecture: Design and implement CircleCI or Jenkins pipelines with comprehensive build testing, versioning, and change logging.
  • Zero-Trust Networking: Build and maintain mesh topology networks using Tailscale or Wireguard to securely connect our hybrid infrastructure.
  • Security-First Culture: Establish and enforce security policies including key rotation, access controls, compliance frameworks, and employee security management.
  • Infrastructure as Code: Document and codify all infrastructure decisions, creating repeatable, auditable deployments.
  • Containerization Strategy: Implement and optimize Docker/K8s deployments for our AI/ML workloads.
  • Cost Optimization: Continue our approach of strategic compute placement using owned, rented or borrowed infrastructure where it makes financial sense without sacrificing security or reliability.
  • Observability & Monitoring: Implement comprehensive logging, monitoring, and alerting across our distributed systems.


What We're Looking For

  •  5+ years of DevOps or infrastructure engineering experience, with a track record of building from scratch
  • Hybrid infrastructure expertise: experience managing both cloud (AWS) and self-hosted infrastructure, understanding the tradeoffs and security risks of each
  • Kubernetes production experience: deep knowledge of cluster design, operations, and scaling
  • Networking mastery: strong understanding of VPCs, mesh networks, VPNs, and zero-trust architectures
  • Security-first mindset: experience with security compliance, key management, IAM policies, and hardening production systems
  • CI/CD expertise: hands-on experience building robust pipelines for build testing before deployment
  • Infrastructure as Code: proficiency with Terraform, Ansible, or similar tools
  • Scripting and automation: strong Python, Bash, or Go skills for tooling and automation
  • Policy and documentation: ability to establish best practices and document them clearly for team adoption
  • Leadership mentality: comfortable setting standards and directing technical decisions, not just executing them

Nice to Have

  • Experience architecting infrastructure for AI/ML workloads
  • Background in a fast-moving startup or scale-up environment
  • H ands-on experience with cost optimization across cloud and on-premises infrastructure


Why Join

  • Opportunity to solve real healthcare problems with cutting-edge technology
  • Well-funded startup with a strong market presence
  • Work with advanced AI technology in a healthcare context
  • Collaborate with a talented team in a fast-paced environment
  • Competitive salary with equity options
  • Performance and quarterly bonuses
  • Professional development opportunities


Compensation and Logistics

  • Remote, full-time
  • Reports to: Head of Engineering
  • Competitive based on experience
Read more
FrontM Limited
Pradeep Chandkiran
Posted by Pradeep Chandkiran
Bengaluru (Bangalore)
3 - 5 yrs
₹8L - ₹14L / yr
skill iconKubernetes
Terraform
skill iconAmazon Web Services (AWS)

Location: Bangalore preferred / Hybrid as applicable

Experience: 3+ years

Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline

Salary: Above market standards, flexible for the right candidate

Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations


About FrontM

FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.

The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.


Role Summary

As a Senior DevOps Engineer, you will take ownership of FrontM’s AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontM’s production systems.

This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.


Key Responsibilities

Cloud Infrastructure & DevOps Architecture (≈45%)

· Own, maintain and improve AWS cloud infrastructure for FrontM platforms

· Create and maintain Terraform scripts for infrastructure deployment and management

· Manage Kubernetes workloads deployed within AWS EKS

· Support multi-zone AWS infrastructure design for availability, resilience and scale

· Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda

· Contribute to DevOps architecture planning in line with FrontM’s platform roadmap

CI/CD, Operations & Platform Reliability (≈35%)

· Build, maintain and improve CI/CD pipelines for backend and platform services

· Oversee technical operations with hands-on administration, monitoring and release support

· Ensure continuous server uptime, stability, performance and maintainability

· Debug, respond to and restore system outages in production and staging environments

· Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io

· Support backend stability, scale and performance across Node.js, Java and related services

Security, Networking & Production Support (≈20%)

· Maintain AWS security configurations, access controls and monitoring practices

· Support complex networking requirements across multi-domain SaaS implementations

· Troubleshoot network, infrastructure and access issues with internal teams and customer-side users

· Work with backend teams to support API integrations and infrastructure abstractions for complex requirements

· Document operational procedures, incident findings and technical support steps clearly


Required Technical Skills

Cloud Infrastructure & AWS

· Strong hands-on experience with AWS infrastructure and cloud operations

· Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda

· Experience with AWS security setup, monitoring and multi-zone infrastructure

· Ability to manage infrastructure using Terraform

Kubernetes, CI/CD & Observability

· Strong experience with Kubernetes, preferably AWS EKS

· Extensive CI/CD and DevOps experience

· Experience with infrastructure observability and application monitoring tools

· Ability to diagnose production bottlenecks, server failures and performance issues

Backend, Networking & SaaS Operations

· Experience supporting Node.js, Java and backend system procedures for stability and scale

· Good understanding of APIs, integrations and backend service dependencies

· Experience with complex networking and multi-domain SaaS implementations

· Ability to troubleshoot technical issues with non-technical end users

Nice to Have

· Experience with MongoDB clusters in MongoDB Atlas

Personal Attributes

· Strong ownership mindset for uptime, reliability and production stability

· Practical problem-solving approach with the ability to act quickly during incidents

· Clear written and spoken communication in English

· Ability to work independently and coordinate with senior management when required

· Comfortable working in fast-moving engineering teams

· Attention to detail in security, monitoring, documentation and operational processes


Why join FrontM?

Long-Term Career Growth

Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.

Engineering Challenges That Matter

Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.

Broad Technical Ownership

Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.


Apply now

Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.

Read more
YOYO AI
YOYO AI
Agency job
via Hello Edge by Saquib Mundagnur
Bengaluru (Bangalore)
5 - 8 yrs
₹18L - ₹24L / yr
CI/CD
skill iconAmazon Web Services (AWS)
skill iconJenkins
ArgoCD

As a DevOps Engineer at YOYO, you'll own the infrastructure and delivery backbone that keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and observability that let a small, fast-moving team ship confidently and you'll keep our AI and data workloads reliable and affordable at scale. This is a hands-on role with real ownership: you won't be maintaining someone else's setup, you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make deployment boring, incidents rare, and scaling a non-event.


If you are Interested DM me on LinkedIn - Saquib Mundagnur



 What You'll Own


CI/CD & developer experience - Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible automated testing, rollbacks, and release controls.Cloud infrastructure & IaC - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow.


Containers & orchestration - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch processing for the speech pipeline.


Reliability & observability (SRE) - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting, on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients.


Data & pipeline infrastructure - Support the infrastructure behind large-scale, edge-to-cloud data movement and processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware.


Security & compliance - Bake security into the platform: secrets management, IAM/least-privilege, encryption in transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security requirements) for a product that handles sensitive customer conversations.


Cost & scale - Own cloud cost visibility and optimization; make scaling decisions that balance reliability and spend. 



What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production systems at meaningful scale.

- Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and Infrastructure-as-Code (**Terraform** or equivalent).

- Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI, Jenkins, Argo, or similar).

- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong automation-first mindset.

- Real experience with **observability** (Prometheus/Grafana, ELK, Datadog, OpenTelemetry, or similar) and running incident response / on-call.

- A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team

Read more
Euphoric Thought Technologies
Bengaluru (Bangalore)
3 - 5 yrs
₹6L - ₹12L / yr
skill iconAmazon Web Services (AWS)
DevOps
skill iconKubernetes
Terraform
CI/CD
+1 more

Job Summary :

We are looking for a proactive and skilled DevOps Engineer to join our team and play a key role in building, managing, and scaling infrastructure for high-performance systems. The ideal candidate will have hands-on experience with Kubernetes, Docker, Python scripting, cloud platforms, and DevOps practices around CI/CD, monitoring, and incident response.

Key Responsibilities :

- Design, build, and maintain scalable, reliable, and secure infrastructure on cloud platforms such as AWS.

- Implement Infrastructure as Code (IaC) using tools like Terraform, Cloud Formation, or similar.

- Manage Kubernetes clusters, configure namespaces, services, deployments, and auto scaling. CI/CD & Release Management

- Build and optimize CI/CD pipelines for automated testing, building, and deployment of services.

- Collaborate with developers to ensure smooth and frequent deployments to production.

- Manage versioning and rollback strategies for critical deployments.

- Containerization & Orchestration using Kubernetes.

- Containerize applications using Docker, and manage them using Kubernetes.

- Write automation scripts using Python or Shell for infrastructure tasks, monitoring, and deployment flows.

- Develop utilities and tools to enhance operational efficiency and reliability.

- Monitoring & Incident Management

- Analyze system performance and implement infrastructure scaling strategies based on load and usage trends.

- Optimize application and system performance through proactive monitoring and configuration tuning.

Desired Skills and Experience :

- Experience Required - 6+ yrs.

- Hands-on experience on cloud services like AWS, EKS etc.

- Ability to design a good cloud solution.

- Strong Linux troubleshooting, Shell Scripting, Kubernetes, Docker, Ansible, Jenkins Skills.

- Design and implement the CI/CD pipeline following the best industry practices using open-source tools.

- Use knowledge and research to constantly modernize our applications and infrastructure stacks.

- Be a team player and strong problem-solver to work with a diverse team.

- Having good communication skills.

Read more
Gurugram
4 - 10 yrs
₹4L - ₹10L / yr
DevOps
Site Reliability Engineer (SRE)
skill iconAmazon Web Services (AWS)
skill iconDocker
skill iconKubernetes
+14 more

🚀 Job Title : DevOps Engineer / Site Reliability Engineer (SRE)

Experience Level : 4+ Years

Location : Gurugram Sector 48, Haryana (On-site)

Employment Type : Full Time Opportunity


About the Role :

We are looking for a proactive DevOps / Site Reliability Engineer (SRE) with around 4 years of hands-on experience designing, automating, and scaling cloud infrastructure and CI/CD delivery pipelines.

In this role, you will bridge the gap between development and operations. You will be responsible for orchestrating containerized applications, automating infrastructure via Code (IaC), establishing SRE best practices (SLIs, SLOs, SLAs), and ensuring maximum uptime, resiliency, and operational efficiency across multi-cloud environments (AWS/Azure/GCP).


Mandatory Skills :

AWS, Kubernetes, Docker, Terraform, Ansible, Jenkins, GitLab CI/CD, GitHub Actions, Python, Bash, CI/CD, Infrastructure as Code (IaC), Grafana, Prometheus, ELK, New Relic, CloudWatch, SRE, SLI/SLO/SLA, Linux


Key Responsibilities :

1. Cloud Infrastructure & Infrastructure as Code (IaC) :

  • Provision, configure, and maintain scalable, high-availability infrastructure on multi-cloud platforms, primarily AWS (EC2, VPC, IAM, S3, RDS, Route53, ALB/ASG, Lambda, EBS).
  • Build, deploy, and manage Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation to enforce consistency and eliminate configuration drift.
  • Execute disaster recovery (DR) planning, automated failover / failback mechanisms, and chaos engineering exercises to validate system resiliency.

2. CI/CD, Automation & Development :

  • Design, end-to-end maintain, and optimize robust CI/CD pipelines using Jenkins, GitLab CI, and GitHub Actions.
  • Automate release pipelines, versioning, branching strategies, and approval gates using Groovy, Python, and Bash scripting. Integrate automated code quality and security scanning tools (SonarQube, Black Duck, or Fortify) directly into delivery pipelines.
  • Develop custom tools, scripts, or microservices (e.g., Python / Node.js) to automate manual operational tasks and operational toil.

3. Containerization & Orchestration :

  • Onboard and orchestrate containerized microservices utilizing Docker and Kubernetes (including Helm charts).
  • Ensure high availability, auto-scaling, resource management, and fault tolerance for Kubernetes pod deployments.

4. Observability, SRE & Incident Management :

  • Drive Site Reliability Engineering (SRE) maturity by establishing, tracking, and reporting SLIs, SLOs, and SLAs with cross-functional engineering teams.
  • Build, configure, and manage full-stack observability tools : Grafana, Prometheus, New Relic, Elasticsearch / Logstash / Kibana (ELK), Sentry, and AWS CloudWatch.
  • Set up real-time alerting, custom metric dashboards, and automated log rotation / pruning scripts.
  • Handle production incidents, lead Root Cause Analysis (RCA) investigations, and implement preventive measures to reduce Mean Time to Resolution (MTTR).


Required Qualifications & Skills :

  • Education : Bachelor’s Degree in Electronics and Communication Engineering, Computer Science, or a related technical field.
  • Experience : ~4 years of experience in DevOps, SRE, or Cloud System Administration roles.
  • Cloud & Infrastructure : Hands-on experience with AWS (Core services like EC2, S3, VPC, RDS, IAM, Lambda, Auto Scaling) and exposure to Azure / GCP.
  • CI/CD & Version Control : Proficiency with Jenkins, GitLab CI, GitHub Actions, and Git workflows.
  • Containerization : Core proficiency in Docker and Kubernetes cluster management / onboarding.
  • Infrastructure as Code : Expertise in Ansible, Terraform, or AWS CloudFormation.
  • Scripting & Languages : Strong hands-on automation skills with Python, Bash, and foundational knowledge of Node.js, Java or C++.
  • Observability & Logging : Strong experience with Grafana, Prometheus, New Relic, ELK stack, or Splunk.
  • Database & SQL : Familiarity with relational databases (MySQL, RDS) for monitoring setup and operational analytics.
Read more
MindBridge
Silfa Rodrigues
Posted by Silfa Rodrigues
HSR Bangalore
6 - 8 yrs
₹18L - ₹24L / yr
CI/CD
skill iconAmazon Web Services (AWS)
Microsoft Windows Azure
Google Cloud Platform (GCP)
Terraform
+3 more

 The Role 

As a **DevOps Engineer** you'll own the infrastructure and delivery backbone that 

keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and 

observability that let a small, fast-moving team ship confidently — and you'll keep our AI and 

data workloads reliable and affordable at scale. 

This is a hands-on role with real ownership: you won't be maintaining someone else's setup, 

you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make 

deployment boring, incidents rare, and scaling a non-event. --- 


What You'll Own 

**CI/CD & developer experience** 

- Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with 

confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible 

automated testing, rollbacks, and release controls. 

**Cloud infrastructure & IaC** - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or 

similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow. 

**Containers & orchestration** - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and 

resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch 

processing for the speech pipeline. 

**Reliability & observability (SRE)** - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting, 

on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients. 

**Data & pipeline infrastructure** - Support the infrastructure behind large-scale, edge-to-cloud data movement and 

processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware. 

**Security & compliance** - Bake security into the platform: secrets management, IAM/least-privilege, encryption in 

transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security 

requirements) for a product that handles sensitive customer conversations. 

**Cost & scale** - Own cloud cost visibility and optimization; make scaling decisions that balance reliability 

and spend. --- 


What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production 

systems at meaningful scale. - Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and 

Infrastructure-as-Code (**Terraform** or equivalent). - Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI, 

Jenkins, Argo, or similar). 

- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong 

automation-first mindset. - Real experience with **observability** (Prometheus/Grafana, ELK, Datadog, 

OpenTelemetry, or similar) and running incident response / on-call. - A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team. 


Bonus Points - Experience running **ML/AI or GPU workloads** in production (inference serving, batch 

pipelines, model deployment). - Experience with data-intensive infrastructure — streaming/queues (Kafka, SQS), data 

pipelines, or large object/audio storage. - Exposure to **edge devices / IoT fleets**, OTA updates, or high-volume device-to-cloud 

ingestion. - Experience with compliance/security frameworks (SOC 2, ISO 27001, DPDP). - FinOps / cloud cost-optimization experience. - Early-stage startup experience. --- 


Why Join - Own infrastructure that's already live with leading retail brands and growing fast — real 

scale, real impact. - Work across genuinely interesting workloads: speech AI, GPU inference, large-scale data, 

and edge-to-cloud ingestion. - Small team, high ownership, direct line to engineering leadership — your decisions ship. - Build the platform foundation of a category-defining product from an

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos