Observability Engineer at MNC · Bengaluru (Bangalore), Hyderabad · 7 - 15 years · ₹25L - ₹30L / yr · Posted 5 Oct 2026

Job Summary
We are looking for a Senior Observability Engineer with strong expertise in ThousandEyes, Splunk, and Monitoring & Observability Engineering. The candidate should have experience implementing and enhancing monitoring solutions, along with basic Python automation scripting.
Primary Skills
- ThousandEyes
- Splunk
- Monitoring & Observability Engineering
- Oracle Database
- SQL
- MongoDB
- Python Automation
Key Responsibilities
- Implement, configure, and enhance enterprise monitoring and observability solutions.
- Work extensively with ThousandEyes for network and application performance monitoring.
- Develop and maintain monitoring dashboards, alerts, and reports using Splunk.
- Monitor application, infrastructure, network, and database performance.
- Work with Oracle Database, SQL, and MongoDB for monitoring and troubleshooting.
- Develop basic Python automation scripts to improve monitoring and operational efficiency.
- Analyze performance issues and support troubleshooting and root cause analysis.
- Collaborate with application, infrastructure, network, and database teams to resolve observability-related issues.
- Continuously improve monitoring coverage, alerting, and operational processes.
Secondary Skills
- GenAI Concepts
- SaaS Architecture Concepts
- Cloud Application Architecture
Ideal Candidate Profile
- 7+ years of experience in Observability / Monitoring Engineering.
- Strong hands-on experience with ThousandEyes and Splunk.
- Good understanding of databases including Oracle, SQL, and MongoDB.
- Basic hands-on experience with Python automation.
- Understanding of application, infrastructure, and network monitoring.
- Exposure to GenAI, SaaS, or Cloud Application Architecture is an added advantage.

Similar jobs (10)
Job Summary
We are looking for an experienced Observability Engineer with strong expertise in ThousandEyes, Splunk, database technologies, and Python automation. The candidate will be responsible for monitoring application, network, and infrastructure performance, identifying issues, and developing automation solutions to improve system visibility, reliability, and operational efficiency.
Required Skills
- Observability
- ThousandEyes / Cisco ThousandEyes
- Splunk
- OracleDB
- MongoDB
- SQL
- Python Automation
Roles & Responsibilities
- Implement and support observability and monitoring solutions across applications, networks, and infrastructure.
- Work with ThousandEyes for network and digital experience monitoring.
- Configure and maintain monitoring dashboards, alerts, and performance metrics.
- Use Splunk for log analysis, monitoring, troubleshooting, and reporting.
- Work with OracleDB, MongoDB, and SQL for data analysis and troubleshooting.
- Develop Python scripts for monitoring and operational automation.
- Analyze performance issues and identify root causes across applications and infrastructure.
- Collaborate with application, infrastructure, and network teams to resolve incidents.
- Improve monitoring, alerting, and automation processes.
- Prepare performance and monitoring reports and provide actionable insights.
Mandatory Skills
- 7+ years of relevant experience
- Strong experience in Observability / Monitoring
- Hands-on experience with ThousandEyes
- Splunk
- Python Automation
- SQL
- OracleDB / MongoDB
Job Description
We are looking for an Azure Cloud & Observability Engineer with strong experience in Azure infrastructure and enterprise monitoring tools such as Splunk, Grafana, and AppDynamics.
Responsibilities
- Design, deploy, and manage Azure cloud infrastructure and services.
- Monitor application and infrastructure performance using Splunk, Grafana, and AppDynamics.
- Configure dashboards, alerts, health rules, and monitoring metrics.
- Perform log analysis, troubleshooting, and root-cause analysis for production issues.
- Implement observability solutions for applications, cloud infrastructure, and services.
- Automate monitoring and operational activities using scripting.
- Support incident, problem, and change management processes.
- Collaborate with development, DevOps, and SRE teams to improve system reliability.
- Maintain monitoring standards, documentation, and operational procedures.
Primary Skills
- Microsoft Azure
- Splunk
- Grafana
- AppDynamics
- Cloud Monitoring & Observability
- Application Performance Monitoring (APM)
- Log Analysis & Troubleshooting
Secondary Skills
- Azure Monitor / Log Analytics
- Azure VMs, Storage, Networking
- Linux
- Python / PowerShell / Shell Scripting
- CI/CD
- Git
- ITIL / ServiceNow
Job Description
We are looking for an Azure Cloud & Observability Engineer with strong experience in Azure infrastructure and enterprise monitoring tools such as Splunk, Grafana, and AppDynamics.
Responsibilities
- Design, deploy, and manage Azure cloud infrastructure and services.
- Monitor application and infrastructure performance using Splunk, Grafana, and AppDynamics.
- Configure dashboards, alerts, health rules, and monitoring metrics.
- Perform log analysis, troubleshooting, and root-cause analysis for production issues.
- Implement observability solutions for applications, cloud infrastructure, and services.
- Automate monitoring and operational activities using scripting.
- Support incident, problem, and change management processes.
- Collaborate with development, DevOps, and SRE teams to improve system reliability.
- Maintain monitoring standards, documentation, and operational procedures.
Primary Skills
- Microsoft Azure
- Splunk
- Grafana
- AppDynamics
- Cloud Monitoring & Observability
- Application Performance Monitoring (APM)
- Log Analysis & Troubleshooting
Secondary Skills
- Azure Monitor / Log Analytics
- Azure VMs, Storage, Networking
- Linux
- Python / PowerShell / Shell Scripting
- CI/CD
- Git
- ITIL / ServiceNow
Key Responsibilities
- Design, deploy, configure, upgrade, and administer enterprise-scale Dynatrace environments.
- Deploy and manage OneAgent, ActiveGate, extensions, and monitoring configurations across application and infrastructure environments.
- Implement observability for applications, APIs, microservices, databases, containers, Kubernetes, cloud platforms, and traditional infrastructure.
- Configure service detection, process groups, management zones, tags, naming rules, metrics, logs, traces, and topology.
- Develop operational and executive dashboards, notebooks, reports, SLOs, and alerting strategies.
- Configure and optimize Davis AI problem detection, anomaly detection, baselines, and root-cause analysis.
- Implement Real User Monitoring (RUM), Synthetic Monitoring, Session Replay, distributed tracing, and log monitoring as required.
- Analyze application performance issues, service dependencies, transaction traces, response times, resource utilization, and infrastructure bottlenecks.
- Lead troubleshooting of complex production performance and availability incidents using Dynatrace telemetry.
- Reduce alert noise through effective event correlation, thresholds, anomaly-detection configuration, and monitoring standards.
- Integrate Dynatrace with enterprise platforms such as ServiceNow, Jira, PagerDuty, Splunk, CI/CD pipelines, and collaboration/notification tools.
- Automate Dynatrace configuration and deployment using APIs, configuration-as-code, scripting, and DevOps tooling.
- Work closely with application, infrastructure, cloud, SRE, DevOps, production support, and operations teams to define observability requirements.
- Establish Dynatrace monitoring standards, reusable configurations, governance, and best practices.
- Perform platform health checks, capacity assessments, license/usage optimization, and monitoring coverage reviews.
- Create technical documentation, runbooks, architecture diagrams, troubleshooting guides, and operational procedures.
- Mentor junior engineers and provide technical leadership for observability initiatives.
Required Technical Skills
- 5–8+ years of overall IT experience with significant experience in application/infrastructure monitoring or observability.
- 3–5+ years of hands-on Dynatrace experience in enterprise environments.
- Strong knowledge of:
- Dynatrace OneAgent and ActiveGate
- Application Performance Monitoring (APM)
- Infrastructure Monitoring
- Distributed Tracing
- Real User Monitoring (RUM)
- Synthetic Monitoring
- Log Monitoring and Analytics
- Metrics, traces, logs, and events
- Davis AI and automated root-cause analysis
- Dashboards, SLOs, alerting, and anomaly detection
- Dynatrace APIs and automation
- Experience monitoring Java/JVM, .NET, web applications, APIs, microservices, and databases.
- Experience with Kubernetes, Docker, OpenShift, or other container platforms.
- Working knowledge of at least one major cloud platform: AWS, Microsoft Azure, or Google Cloud Platform.
- Strong understanding of application architecture, HTTP/HTTPS, REST APIs, networking, operating systems, and middleware.
- Experience with Linux and Windows environments.
- Scripting/automation experience using Python, PowerShell, Bash, Ansible, Terraform, or similar technologies.
- Experience integrating monitoring platforms with ITSM, incident management, and DevOps tools.
- Strong analytical and production troubleshooting skills.
Preferred Skills
- Experience with Dynatrace Grail, DQL (Dynatrace Query Language), OpenPipeline, and the latest Dynatrace platform capabilities.
- Knowledge of OpenTelemetry (OTel) and modern telemetry standards.
- Experience implementing observability for large-scale Kubernetes and cloud-native environments.
- Familiarity with SRE practices, including SLIs, SLOs, error budgets, and observability-driven incident management.
- Knowledge of additional monitoring platforms such as Splunk, AppDynamics, Datadog, New Relic, Grafana, Prometheus, or ELK.
- Experience with infrastructure-as-code and configuration-as-code approaches.
- Dynatrace certifications are preferred.
Professional Skills
- Strong problem-solving and root-cause analysis capabilities.
- Ability to troubleshoot complex application and infrastructure performance issues independently.
- Strong written and verbal communication skills.
- Ability to work effectively with application owners, developers, SREs, infrastructure teams, and senior stakeholders.
- Ability to translate business and operational requirements into observability solutions.
- Experience working in enterprise production environments with incident, problem, and change-management processes.
- Ability to mentor engineers and drive technical standards across teams.
Education & Certification
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent professional experience.
- Dynatrace Associate/Professional-level certification is desirable.
- Cloud, Kubernetes, ITIL, SRE, or DevOps certifications are advantageous.
Experience Level
Senior Engineer
- Overall Experience: 5–8+ years
- Dynatrace/Observability Experience: 3–5+ years
- Expected proficiency: Advanced hands-on implementation, administration, troubleshooting, automation, and solution design
Job Summary:
We are looking for a Senior SRE/DevOps Engineer with strong experience in site reliability, automation, monitoring, observability, and production support. The candidate will be responsible for ensuring the reliability, availability, security, and performance of enterprise platforms.
Key Responsibilities:
- Own reliability, availability, security, and performance of enterprise browser platforms.
- Manage access and identity controls, Group Policy, and SAML/SSO integrations.
- Handle CI/CD deployments, configuration management, monitoring, and health checks.
- Develop automation and operational workflows using Python and Bash.
- Perform performance monitoring and production troubleshooting.
- Implement and maintain observability frameworks.
- Work with monitoring tools such as Datadog, Splunk, Dynatrace, Prometheus, and Grafana.
- Participate in incident management, RCA, and continuous improvement activities.
- Support highly available production environments in a 24/7 shift model.
- Collaborate with application, infrastructure, security, and operations teams.
Mandatory Skills:
- 10+ years of experience in SRE / DevOps.
- Strong hands-on experience in Python and Bash scripting.
- Experience with CI/CD and configuration management.
- Strong knowledge of monitoring and observability.
- Hands-on experience with Datadog, Splunk, Dynatrace, Prometheus, or Grafana.
- Experience with SAML/SSO, Identity & Access Management, and Group Policy.
- Banking domain experience.
Hiring: DevOps Lead
📍 Kochi / Trivandrum / Remote
💼 Full-time
🕐 General Shift | Australian Overlap
We are looking for an experienced DevOps Lead to join our team.
🔹 Key Responsibilities
Implement and continually improve the observability platform using Datadog, particularly from a user experience perspective.
Work across engineering squads as a virtual team member, supporting their DevOps, infrastructure, and observability requirements.
Configure and manage Datadog RUM, Session Replay, Distributed Tracing, and APM.
Support campaign readiness activities, including load and performance testing.
Participate in gamedays and incident response activities for production systems.
Liaise closely with the managed infrastructure provider on infrastructure requirements and activities.
🔹 Essential Skills & Requirements
✅ 7+ years of relevant DevOps / Cloud / Observability experience
✅ Strong hands-on experience with Datadog
✅ Strong experience with RUM, Session Replay, Distributed Tracing, and APM
✅ Solid experience with AWS
✅ Strong Infrastructure-as-Code experience using CloudFormation and AWS CDK
✅ Good working knowledge of GitHub Actions and AWS CodePipeline
✅ Real incident response experience on high-traffic systems
✅ Experience with Load & Performance Testing
✅ Strong problem-solving and communication skills
✅ Ability to work closely with infrastructure partners and internal platform teams
🔹 Skills - Good to Have
⭐ Experience with high-traffic platforms
⭐ Experience with campaign readiness and gamedays
⭐ Infrastructure partner management
⭐ Advanced AWS observability
⭐ Performance engineering
📌 Experience: 7+ Years
📌 Work Location: Kochi / Trivandrum / Remote
📌 Shift: General Shift with Australian Overlap
📌 Remote: Mandatory 1 week at office
📩 Interested candidates can share their resume
Job Description
We are looking for an experienced Application Support Engineer with strong expertise in production support, networking, DNS, Linux/Unix, and application monitoring.
Key Responsibilities
- Provide application and production support for business-critical applications.
- Monitor application performance, availability, and system health.
- Troubleshoot issues related to Network, DNS, Linux, and Unix.
- Analyze application and system logs to identify and resolve production issues.
- Monitor applications using Splunk, Grafana, and AppDynamics.
- Identify and resolve incidents within defined timelines.
- Perform root cause analysis and support issue resolution.
- Coordinate with technical teams for incident investigation and escalation.
- Conduct application health checks and proactively identify potential issues.
- Maintain proper documentation of incidents, troubleshooting steps, and resolutions.
- Follow standard production support and incident management processes.
Required Skills
- 7–11 years of experience in Application/Production Support.
- Strong knowledge of Networking and DNS.
- Hands-on experience with Linux/Unix.
- Experience with monitoring tools such as Splunk, Grafana, and AppDynamics.
- Strong troubleshooting and analytical skills.
- Good understanding of incident management and production support.
- Excellent communication and coordination skills.
- Ability to work in a fast-paced production support environment.

Position Overview
The AI Observability Engineer will be instrumental in implementation of scalable, cloud-native solutions to meet the growing needs of our Data & Development team. The successful candidate will demonstrate the ability to abstract complexity and create reusable, scalable patterns that accelerate development. The AI Observability Engineer will build and maintain a robust framework to ensure the reliability and maintainability of DPR Construction's complex AI systems.
Responsibilities
- Standardize observability practices across AI/ML and other development teams including logging, metrics, tracing, and model performance monitoring, ingesting data from multiple platforms
- Lead hands-on implementation of automation-first DevOps and MLOps practices, enabling infrastructure-as-code and consistent, repeatable environment provisioning
- Design and manage intelligent DataOps pipelines with automated data quality monitoring and anomaly detection
- Deploy, maintain and monitor containerized ML workloads
- Extend existing CI/CD pipelines to support automated infrastructure changes and ML workflows
- Implement AI-driven data validation, schema and concept drift detection and metadata management.
- Establish governance frameworks for AI systems, including bias detection, explainability, and auditability
- Extend existing Azure RBAC strategy by automating role and permission management to reduce manual intervention
- Develop automated test suites for model performance, regression, edge cases and bias validation
- Monitor model KPIs (accuracy, precision, recall, latency, calibration)
- Ensure reproducability of experiments and production models
- Act as a technical point of contact for DevOps and MLOps practices, developing reusable patterns, documentation, and proof-of-concepts to drive adoption
Qualifications
- Bachelor’s degree in computer science, Data Science, Information Systems, or a related field
- 5+ years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering
- Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure
- Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn
Application Production Support with SRE, Linux/Unix, Splunk/AppD/Grafana, Troubleshooting
WFO-Immediate
8 to 12 Yrs
Bangalore/Hyderabad
Site Reliability Engineer (SRE) / DevOps Engineer (Walk-In Drive)
Location: Gurgaon Experience: 3–6 Years
About the Role
We are looking for a hands-on Site Reliability Engineer (SRE) / DevOps Engineer with strong programming and automation skills.
The role will initially involve development and automation work, helping the engineer build a strong understanding of the applications and platform. Over time, the role will expand into broader DevOps and SRE responsibilities, including CI/CD, cloud infrastructure, observability, production reliability, incident management, and operational automation.
The ideal candidate should be comfortable working with both application code and production systems and should use engineering and automation to improve reliability and reduce manual effort.
Key Responsibilities
· Develop and enhance internal applications, automation tools, APIs, utilities, and platform capabilities using Python.
· Write clean, maintainable, testable, and production-ready code.
· Participate in code reviews, debugging, testing, and technical discussions.
· Build, maintain, and improve CI/CD pipelines and automated deployment processes.
· Work with Docker and Kubernetes for application deployment and operations.
· Support on prem and cloud-based application and infrastructure deployments.
· Maintain reliable, scalable, secure, and highly available production environments.
· Implement and manage monitoring, logging, alerting, and observability solutions.
· Contribute to defining and tracking SLIs, SLOs, and error budgets.
· Troubleshoot application and production issues and perform Root Cause Analysis (RCA).
· Identify recurring operational problems and address them through automation and engineering improvements.
· Support incident response, change management, deployment governance, and disaster recovery practices.
· Maintain runbooks, SOPs, incident documentation, and technical documentation.
· Collaborate with Engineering, Product, Platform, Security, Operations, and external teams.
Technical Skills
· Strong hands-on experience with Python for development and automation.
· Experience developing scripts, APIs, integrations, utilities, or backend services.
· Good understanding of software engineering principles, debugging, logging, testing, and exception handling.
· Experience with REST APIs, JSON, Git, pull requests, and code reviews.
· Strong knowledge of Linux/Unix environments and basic Windows administration.
· Good understanding of networking concepts including DNS, TCP/IP, HTTP/HTTPS, load balancing, and firewalls.
· Experience with at least one cloud platform: AWS, Azure, or GCP.
· Hands-on experience with Docker and Kubernetes.
· Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or equivalent.
· Familiarity with Infrastructure-as-Code tools such as Terraform is preferred.
· Experience with monitoring and observability tools such as Grafana, Prometheus, Power BI, or equivalent.
· Ability to analyze logs, metrics, alerts, and traces for troubleshooting.
· Understanding of SRE concepts including SLIs, SLOs, availability, reliability, error budgets, and RCA.
· Experience with JIRA, ServiceNow, and Confluence is desirable.
Preferred Experience
· 3–6 years of experience in SRE, DevOps, Platform Engineering, Cloud Engineering, or related roles.
· Strong Python development or automation experience.
· Experience supporting applications across development, deployment, and production environments.
· Exposure to cloud-native, distributed, or production-grade systems.
· Understanding of security and compliance best practices.
· Familiarity with AI-assisted engineering tools such as GitHub Copilot, Claude Code, or similar tools.
Soft Skills
· Strong analytical and troubleshooting skills.
· Engineering and automation mindset.
· Good written and verbal communication skills.
· Effective cross-functional collaboration.
· Ownership-driven approach to problem solving.
· Ability to remain structured during production incidents.






