SRE Application Support at MNC · Bengaluru (Bangalore), Hyderabad · 8 - 12 years · ₹2L - ₹22L / yr · Posted 16 Sep 2026

Application Production Support with SRE, Linux/Unix, Splunk/AppD/Grafana, Troubleshooting
WFO-Immediate
8 to 12 Yrs
Bangalore/Hyderabad

Similar jobs (10)
Job Description
We are looking for an experienced Application Support Engineer with strong expertise in production support, networking, DNS, Linux/Unix, and application monitoring.
Key Responsibilities
- Provide application and production support for business-critical applications.
- Monitor application performance, availability, and system health.
- Troubleshoot issues related to Network, DNS, Linux, and Unix.
- Analyze application and system logs to identify and resolve production issues.
- Monitor applications using Splunk, Grafana, and AppDynamics.
- Identify and resolve incidents within defined timelines.
- Perform root cause analysis and support issue resolution.
- Coordinate with technical teams for incident investigation and escalation.
- Conduct application health checks and proactively identify potential issues.
- Maintain proper documentation of incidents, troubleshooting steps, and resolutions.
- Follow standard production support and incident management processes.
Required Skills
- 7–11 years of experience in Application/Production Support.
- Strong knowledge of Networking and DNS.
- Hands-on experience with Linux/Unix.
- Experience with monitoring tools such as Splunk, Grafana, and AppDynamics.
- Strong troubleshooting and analytical skills.
- Good understanding of incident management and production support.
- Excellent communication and coordination skills.
- Ability to work in a fast-paced production support environment.
Role Summary:
We are looking for an experienced Application Production Support Engineer with strong expertise in application support, incident and change management, Linux/Unix, SQL, Oracle, and monitoring tools. The candidate will be responsible for maintaining application availability, troubleshooting production issues, monitoring system performance, and coordinating with technical and business stakeholders.
Key Responsibilities
- Provide L2/L3 production support for business-critical applications.
- Monitor applications and infrastructure using Splunk, Grafana, and AppDynamics.
- Analyze and resolve production incidents within defined SLAs.
- Perform incident, problem, change, and service request management.
- Troubleshoot application issues across Linux/Unix, SQL, and Oracle environments.
- Perform SQL queries and database-level troubleshooting to identify application issues.
- Analyze application logs, alerts, and performance metrics to identify root causes.
- Coordinate with development, database, infrastructure, and other technical teams for issue resolution.
- Participate in Root Cause Analysis (RCA) and implement corrective/preventive actions.
- Support application deployments, releases, and production changes.
- Ensure effective communication with business users and stakeholders during critical incidents.
- Identify recurring issues and drive problem management and service improvement initiatives.
- Maintain support documentation, knowledge articles, and operational procedures.
- Participate in on-call/shift support as required.
Mandatory Skills
- 6+ years of experience in Application Production Support.
- Strong experience in Incident & Change Management.
- Hands-on experience with Linux/Unix.
- Good knowledge of SQL and Oracle database support.
- Experience with monitoring and observability tools:
- Splunk
- Grafana
- AppDynamics
- Strong troubleshooting and problem-solving skills.
- Good understanding of application monitoring, logs, alerts, and performance analysis.
- Strong stakeholder management and communication skills
We are looking for an experienced Application Support Engineer with strong knowledge of Network, DNS, Linux/Unix, and application monitoring tools. The candidate will be responsible for production support, incident troubleshooting, monitoring, and ensuring application availability and stability.
Responsibilities:
- Provide L2/L3 application and production support for critical applications.
- Troubleshoot issues related to Network, DNS, Linux/Unix, and application connectivity.
- Monitor application health, performance, and availability using Splunk, Grafana, AppDynamics, or similar tools.
- Analyze logs, alerts, and performance metrics to identify and resolve incidents.
- Handle incident, problem, and change management activities as per ITIL processes.
- Perform root cause analysis (RCA) and implement preventive measures.
- Coordinate with Network, Infrastructure, Development, and other support teams for issue resolution.
- Participate in production deployments, maintenance activities, and on-call support.
Required Skills:
- Strong experience in Application/Production Support.
- Good knowledge of Linux/Unix administration and troubleshooting.
- Strong understanding of Networking concepts and DNS.
- Hands-on experience with Splunk, Grafana, AppDynamics, or similar monitoring tools.
- Good understanding of TCP/IP, HTTP/HTTPS, load balancing, and connectivity troubleshooting.
- Experience with incident, problem, and change management.
- Strong troubleshooting, communication, and stakeholder management skills.
- Good knowledge of SQL and application log analysis is preferred.
We are looking for an experienced Application Support Engineer with strong expertise in Linux/Unix, Networking, Routing, Load Balancing, and Production Support. The ideal candidate should have hands-on experience troubleshooting application and infrastructure issues using enterprise monitoring and observability tools such as Splunk, Grafana, and AppDynamics.
Key Responsibilities
- Provide L2/L3 production support for enterprise applications and infrastructure.
- Troubleshoot application, Linux/Unix, network, routing, and connectivity-related issues.
- Monitor application and infrastructure health using Splunk, Grafana, and AppDynamics.
- Analyze alerts, logs, performance metrics, and system behavior to identify issues.
- Troubleshoot routing, load balancing, connectivity, and network-related problems.
- Perform incident investigation, troubleshooting, and root cause analysis.
- Coordinate with Network, Infrastructure, Application, and other technical teams for issue resolution.
- Handle production incidents and ensure timely resolution within defined SLAs.
- Participate in problem management and identify recurring issues.
- Perform application health checks and proactively identify potential failures.
- Maintain troubleshooting guides, runbooks, and operational documentation.
- Support planned changes, deployments, and maintenance activities.
Must-Have Skills
Application Support
Linux / Unix
Networking
Routing
Load Balancing
Splunk
Grafana
AppDynamics
Production Support
Incident Management
Troubleshooting
Required Skills & Tools
- 4–6 years of experience in Application / Production Support.
- Strong hands-on experience with Linux/Unix administration and troubleshooting.
- Good understanding of TCP/IP, networking, routing, and connectivity concepts.
- Practical experience troubleshooting load balancing issues.
- Hands-on experience with monitoring and observability tools such as Splunk, Grafana, or AppDynamics.
- Strong log analysis and troubleshooting skills.
- Experience handling production incidents and working within SLA-driven environments.
- Strong communication and coordination skills.
Preferred Experience
- Exposure to ITIL processes including Incident, Problem, and Change Management.
- Experience with scripting/automation using Shell or Python.
- Exposure to cloud environments is an added advantage.
- Experience supporting enterprise-scale applications and infrastructure.
Key Competencies
- Strong analytical and troubleshooting skills.
- Ability to work under pressure during critical production incidents.
- Strong ownership and accountability.
- Excellent communication and stakeholder management.
- Ability to work collaboratively with global technical teams.
Qualifications
Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline.
SRE / Success Engineering role focused on production operations, reliability, AWS infrastructure, monitoring, incident management, and platform support for the ZT platform.
Core responsibilities include:
- Production monitoring and debugging of live systems.
- Incident investigation, troubleshooting, and problem resolution.
- AWS cloud infrastructure support and maintenance.
- Deployment and operational support activities.
- Supporting a 24x7 production environment.
- Working with GitHub-based development workflows.
- Technical debt remediation and platform improvements.
- Customer issue investigation and support.
- Security and compliance-related work, including FedRAMP initiatives.
Preferred Skills:
AWS (especially S3 and EC2)
Strong debugging and troubleshooting skills
Site Reliability Engineering (SRE) experience
GitHub experience
Basic software development skills
TypeScript/JavaScript knowledge
C# preferred
AI experience is a plus.
Candidate should be a hands-on engineer with strong AWS, SRE, operational ownership, production support, and debugging capabilities, rather than a pure application or full-stack developer.
Job Summary:
We are looking for a Senior SRE/DevOps Engineer with strong experience in site reliability, automation, monitoring, observability, and production support. The candidate will be responsible for ensuring the reliability, availability, security, and performance of enterprise platforms.
Key Responsibilities:
- Own reliability, availability, security, and performance of enterprise browser platforms.
- Manage access and identity controls, Group Policy, and SAML/SSO integrations.
- Handle CI/CD deployments, configuration management, monitoring, and health checks.
- Develop automation and operational workflows using Python and Bash.
- Perform performance monitoring and production troubleshooting.
- Implement and maintain observability frameworks.
- Work with monitoring tools such as Datadog, Splunk, Dynatrace, Prometheus, and Grafana.
- Participate in incident management, RCA, and continuous improvement activities.
- Support highly available production environments in a 24/7 shift model.
- Collaborate with application, infrastructure, security, and operations teams.
Mandatory Skills:
- 10+ years of experience in SRE / DevOps.
- Strong hands-on experience in Python and Bash scripting.
- Experience with CI/CD and configuration management.
- Strong knowledge of monitoring and observability.
- Hands-on experience with Datadog, Splunk, Dynatrace, Prometheus, or Grafana.
- Experience with SAML/SSO, Identity & Access Management, and Group Policy.
- Banking domain experience.
- Linux troubleshooting
- Hands-on AWS
- Production/Application Support
- Bash/Shell/Python
- Monitoring/log analysis
- Incident resolution
- Application deployment/support
- Basic networking and database knowledge
- Production/batch support exposure
- Willingness for rotational weekend/critical production support
SRE – Network / DNS / Load Balancer
Experience: 7–11 Years
Location: Hyderabad
Work Mode: WFO
Availability: Immediate Joiner
Job Description:
- Strong experience in Site Reliability Engineering (SRE) with focus on infrastructure and application reliability.
- Hands-on experience with Network, DNS and Load Balancer troubleshooting and administration.
- Monitor system performance, availability, latency and infrastructure health.
- Troubleshoot network connectivity, DNS resolution, routing and load-balancing issues.
- Experience with load balancers such as F5, BIG-IP, HAProxy or similar technologies.
- Good understanding of TCP/IP, HTTP/HTTPS, LAN/WAN, SSL/TLS and networking concepts.
- Experience with DNS technologies such as BIND, Infoblox or equivalent.
- Work on incident management, root cause analysis and problem resolution.
- Collaborate with application, network, cloud and infrastructure teams to resolve production issues.
- Experience with monitoring and alerting tools such as Splunk, Grafana, Prometheus, AppDynamics or similar tools.
- Strong troubleshooting, production support and communication skills.
- Willingness to work from Hyderabad office (WFO) and join immediately.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Role Summary
We are seeking a proactive and technically skilled Python Application Support Engineer to join our Technical Operations team. This role is crucial for ensuring the stability and reliability of our mission-critical, Python-based applications. You will be responsible for timely incident resolution, deep-dive troubleshooting, implementing permanent fixes, and driving operational efficiency through automation.
🔑 Key Responsibilities
Technical Troubleshooting & Incident Management
- Incident Resolution: Serve as the primary point of contact for complex Level 2 and Level 3 production incidents, diagnosing root causes and resolving issues across our Python application stack.
- Deep-Dive Analysis: Utilize log analysis tools (e.g., Splunk, ELK Stack) and monitoring platforms (e.g., Prometheus, Grafana) to quickly identify and address anomalies in application behavior.
- Code Debugging: Analyze, debug, and fix application issues directly within the Python codebase, including Flask/Django services, worker queues, and custom scripts.
- Database Health: Troubleshoot performance issues and conduct basic SQL/NoSQL query tuning and health checks (e.g., for PostgreSQL, MongoDB, or Redis).
Operational Excellence & Automation
- Monitoring & Alerting: Continuously refine and optimize application monitoring, alerting, and logging configurations to improve mean time to detect (MTTD) and mean time to resolve (MTTR).
- Python Automation: Develop, maintain, and enhance automated scripts (primarily in Python) to streamline routine operational tasks, reporting, health checks, and system recovery processes.
- Documentation: Create and maintain comprehensive documentation, runbooks, and knowledge base articles for application support procedures and recurring issues.
Collaboration & Prevention
- Cross-Functional Fixes: Collaborate closely with the Development and DevOps teams to provide clear technical feedback on recurring issues and implement permanent, scalable solutions.
- Proactive Maintenance: Identify potential system bottlenecks, performance degradation points, and areas prone to failure, recommending and implementing preventative measures.
⚙️ Required Qualifications
- Experience: 3 to 5 years of professional experience in Application Support, Production Support, Site Reliability Engineering (SRE), or a similar technical role.
- Python Expertise (Mandatory): Strong hands-on experience with Python scripting and programming, including the ability to read, debug, and modify application code.
- Operating Systems: Proficient working knowledge of Linux/Unix environments and shell scripting.
- Databases: Solid experience with relational (e.g., PostgreSQL, MySQL) and/or NoSQL (e.g., MongoDB, Redis) databases, focusing on query analysis and performance.
We're looking for a Site Reliability Engineer to keep our production systems fast, reliable, and scalable. Sitting at the intersection of software engineering and operations, you'll treat infrastructure as code, automate away toil, and build the observability that lets us catch problems before customers do. You'll own uptime and on-call for critical services, lead incident response and blameless postmortems, and continuously harden the platform against failure. This role suits an engineer who is as comfortable debugging a production incident at 2 a.m. as they are writing the automation that prevents the next one.
Key Responsibilities
- Own reliability, availability, and performance of production services, including on-call rotation
- Build and maintain monitoring, alerting, and observability (metrics, logs, traces)
- Automate deployments, scaling, and operational tasks to reduce manual toil
- Manage containerized workloads on Kubernetes and cloud infrastructure
- Design and maintain CI/CD pipelines for safe, frequent releases
- Lead incident response and drive blameless postmortems with clear follow-ups
- Perform capacity planning, performance tuning, and cost optimization
- Define and track SLIs/SLOs and error budgets with product teams
Requirements
- 3+ years in SRE, DevOps, or production-focused engineering
- Strong Linux administration and hands-on Kubernetes experience
- Solid experience with monitoring/observability tools (Prometheus, Grafana, ELK, or similar)
- Cloud experience with AWS, GCP, or Azure
- CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation)
- Proficient scripting in Python and/or Bash
Nice to have
- Experience with service meshes, Helm, or GitOps (ArgoCD/Flux)
- Background in high-traffic or distributed systems







