5 Monitoring Jobs in Chennai | Monitoring Job openings in Chennai
Apply to 5+ Monitoring Jobs in Chennai on CutShort.io. Explore the latest Monitoring Job opportunities across top companies like Google, Amazon & Adobe.
Bengaluru (Bangalore), Chennai, Mumbai, Hyderabad, Pune, Gurugram · 3 - 10 years · ₹12L - ₹35L / yr · Posted 21 Sep 2026
We're looking for a Site Reliability Engineer to keep our production systems fast, reliable, and scalable. Sitting at the intersection of software engineering and operations, you'll treat infrastructure as code, automate away toil, and build the observability that lets us catch problems before customers do. You'll own uptime and on-call for critical services, lead incident response and blameless postmortems, and continuously harden the platform against failure. This role suits an engineer who is as comfortable debugging a production incident at 2 a.m. as they are writing the automation that prevents the next one.
Key Responsibilities
- Own reliability, availability, and performance of production services, including on-call rotation
- Build and maintain monitoring, alerting, and observability (metrics, logs, traces)
- Automate deployments, scaling, and operational tasks to reduce manual toil
- Manage containerized workloads on Kubernetes and cloud infrastructure
- Design and maintain CI/CD pipelines for safe, frequent releases
- Lead incident response and drive blameless postmortems with clear follow-ups
- Perform capacity planning, performance tuning, and cost optimization
- Define and track SLIs/SLOs and error budgets with product teams
Requirements
- 3+ years in SRE, DevOps, or production-focused engineering
- Strong Linux administration and hands-on Kubernetes experience
- Solid experience with monitoring/observability tools (Prometheus, Grafana, ELK, or similar)
- Cloud experience with AWS, GCP, or Azure
- CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation)
- Proficient scripting in Python and/or Bash
Nice to have
- Experience with service meshes, Helm, or GitOps (ArgoCD/Flux)
- Background in high-traffic or distributed systems
Chennai · 8 - 12 years · ₹50L - ₹75L / yr · Posted 14 Aug 2024
Role & Responsibilities:
AWS Cloud Management:
• Lead the design, deployment, and management of AWS cloud infrastructure to ensure
scalability, security, and reliability.
• Oversee the implementation of best practices for cloud resource utilization.
Automated Provisioning:
• Drive the development and maintenance of automated provisioning processes for
infrastructure deployment, leveraging tools such as Terraform and Packer.
• Continuously enhance deployment workflows to optimize efficiency.
Financial Operations (FinOps):
• Implement and champion FinOps practices to optimize cloud costs and resource
utilization.
• Conduct regular cost analysis and identify opportunities for cost savings without
compromising performance.
Infrastructure as Code (IaC):
• Collaborate with teams to implement and maintain IaC scripts for infrastructure
configuration and deployment.
• Ensure version control and consistency in infrastructure code across projects.
Team Leadership:
• Lead and mentor a team of DevOps engineers, providing technical guidance and
support.\
• Foster a collaborative and innovative team culture focused on continuous improvement.
Continuous Integration/Continuous Deployment (CI/CD):
• Drive the implementation and maintenance of CI/CD pipelines to automate software
delivery processes.
• Ensure seamless and reliable application deployments across environments.
Monitoring and Optimization:
• Implement monitoring solutions for cloud resources and applications.
• Proactively identify and address performance bottlenecks, ensuring optimal system
performance.
Chennai · 4 - 5 years · ₹4L - ₹6L / yr · Profitable · Posted 20 Feb 2023
JOB Description
- Monitoring entire infrastructure of Olam using various monitoring tools like
- SCOM, SolarWinds, Telegraph, OEM.
- Monitoring various types of alerts like
- CPU Utilization
- Memory Utilization
- Database related alerts
- DR Replication issues
- Backup Failure Alerts
- Exchange Mail Queue Threshold Alerts
- Service Mailbox quota breach alert
- Adobe Experience Manager / Site 24/7 Alerts
- Application URL Alerting
- Scheduling Maintenance Mode for planned Activity.
- Daily repeat CI analysis of events/alerts/incident and raising proactive problem tickets which helps in reduction of major incident.
- Handling Major Incidents, Driving the major incident bridge, sending communication about major incident to stake holders.
- CMDB Inventory Management – Onboarding and Offboarding of Device's are commissioned/decommissioned.
- Coordinating with Service Provider for MPLS related outage
- Daily follow ups with Regional and internal teams to ensure all the node are up and running fine.
Hyderabad, Mumbai, Chennai, Bengaluru (Bangalore) · 2 - 7 years · ₹5L - ₹15L / yr · Bootstrapped · Posted 6 Feb 2023
This role is for Work from the office.
Job Description
Roles & Responsibilities
- Work across the entire landscape that spans network, compute, storage, databases, applications, and business domain
- Use the Big Data and AI-driven features of vuSmartMaps to provide solutions that will enable customers to improve the end-user experience for their applications
- Create detailed designs, solutions and validate with internal engineering and customer teams, and establish a good network of relationships with customers and experts
- Understand the application architecture and transaction-level workflow to identify touchpoints and metrics to be monitored and analyzed
- Analytics and analysis of data and provide insights and recommendations
- Constantly stay ahead in communicating with customers. Manage planning and execution of platform implementation at customer sites.
- Work with the product team in developing new features, identifying solution gaps, etc.
- Interest and aptitude in learning new technologies - Big Data, no SQL databases, Elastic Search, Mongo DB, DevOps.
Skills & Experience
- At least 2+ years of experience in IT Infrastructure Management
- Experience in working with large-scale IT infra, including applications, databases, and networks.
- Experience in working with monitoring tools, automation tools
- Hands-on experience in Linux and scripting.
- Knowledge/Experience in the following technologies will be an added plus: ElasticSearch, Kafka, Docker Containers, MongoDB, Big Data, SQL databases, ELK stack, REST APIs, web services, and JMX.
Remote, Chennai · 0 - 4 years · ₹1L - ₹3L / yr · Profitable · Remote friendly · Posted 28 Oct 2020


