1 Application server Jobs in Hyderabad | Application server Job openings in Hyderabad
Apply to 1 Application server Jobs in Hyderabad on CutShort.io. Explore the latest Application server Job opportunities across top companies like Google, Amazon & Adobe.
Hyderabad · 5 - 10 years · ₹4L - ₹15L / yr · Profitable · Posted 7 Oct 2026
ob Summary
We are looking for an experienced Site Reliability Engineer (SRE) / Production Support Engineer with strong hands-on experience in application and production support, incident management, monitoring, cloud operations, automation, and infrastructure technologies.
The ideal candidate will be responsible for ensuring the availability, reliability, performance, and stability of production applications and infrastructure. The role involves troubleshooting critical production issues, monitoring applications and infrastructure, supporting deployments, managing incidents, and driving automation and operational improvements.
Key Responsibilities
- Provide L2/L3 Application and Production Support for critical business applications.
- Monitor production applications, infrastructure, batch jobs, and system health.
- Handle and troubleshoot critical production incidents, ensuring timely resolution and minimal business impact.
- Participate in Incident, Problem, and Change Management processes.
- Perform root-cause analysis (RCA) for recurring and major production issues.
- Troubleshoot issues related to applications, networks, load balancers, databases, operating systems, and infrastructure.
- Monitor application and infrastructure performance using Splunk, APM, and other monitoring tools.
- Create and maintain Splunk queries, dashboards, alerts, and operational monitoring.
- Support production deployments, including Blue-Green and Canary deployment strategies.
- Work with cloud infrastructure and perform day-to-day Cloud Operations activities.
- Manage and troubleshoot containerized applications using Docker and Kubernetes.
- Work with Terraform / Infrastructure as Code (IaC) for infrastructure provisioning and automation.
- Support Linux and Windows server administration.
- Develop and maintain Shell scripts and Python automation scripts to reduce manual operational activities.
- Monitor and analyze SLIs, SLOs, Error Budgets, and Burn Rates.
- Identify reliability risks and proactively implement solutions to improve system availability and performance.
- Collaborate with Development, DevOps, Infrastructure, Network, Database, and Cloud teams during production incidents.
- Leverage GenAI tools such as GitHub Copilot, Claude, or similar tools to improve troubleshooting, automation, documentation, and operational efficiency.
- Participate in on-call/shift-based production support as required by business and customer needs.
Mandatory / Key Skills
- Production / Application Support
- Incident Management
- Production Monitoring & Batch Monitoring
- Splunk – Queries, Dashboards & Monitoring
- APM / Application Performance Monitoring
- SLI / SLO / Error Budget / Burn Rate
- Cloud Operations
- Kubernetes
- Docker
- Terraform / Infrastructure as Code
- Linux & Windows Administration
- Shell Scripting
- Python Scripting / Automation
- Production Deployment Support
- Blue-Green & Canary Deployments
- Network, Load Balancing & Database Troubleshooting

