Senior Data Dog Engineer at VDart Technology · Bengaluru (Bangalore) · 7 - 12 years · ₹20L - ₹30L / yr · Profitable · Posted 21 May 2026

Senior Datadog Platform Engineer / Observability Developer
We are looking for an experienced Datadog Engineer to manage and enhance enterprise observability and monitoring platforms across Azure Cloud and Azure DevOps environments.
Required Experience
- 6–10 years of overall IT experience
- 4–6 years of strong hands-on Datadog experience
- Experience with Azure Cloud, Azure DevOps, CI/CD, and enterprise integrations
- Strong knowledge of automation, observability, and platform engineering
Key Responsibilities
- Configure and manage Datadog platform, dashboards, alerts, logging, APM, and observability solutions
- Develop automation and reusable frameworks using Datadog APIs, scripting, and CI/CD pipelines, Azure Cloud services, Azure DevOps, Jira, Kubernetes, and enterprise applications
- Implement Monitoring-as-Code and Dashboard-as-Code standards
- Support onboarding, telemetry integration, operational workflows, and platform optimization
Preferred Skills
- Datadog platform engineering and integrations
- Azure Cloud and Azure DevOps
- Python, PowerShell, YAML, JSON, APIs
- CI/CD pipelines and automation
- Kubernetes / AKS and cloud-native platforms
Soft Skills
- Strong communication and stakeholder management
- Good troubleshooting and problem-solving skills
- Ability to work independently and drive ownership
- Strong collaboration and documentation skills

About VDart Technology
About
Company social profiles
Similar jobs (10)
Hiring: DevOps Lead
📍 Kochi / Trivandrum / Remote
💼 Full-time
🕐 General Shift | Australian Overlap
We are looking for an experienced DevOps Lead to join our team.
🔹 Key Responsibilities
Implement and continually improve the observability platform using Datadog, particularly from a user experience perspective.
Work across engineering squads as a virtual team member, supporting their DevOps, infrastructure, and observability requirements.
Configure and manage Datadog RUM, Session Replay, Distributed Tracing, and APM.
Support campaign readiness activities, including load and performance testing.
Participate in gamedays and incident response activities for production systems.
Liaise closely with the managed infrastructure provider on infrastructure requirements and activities.
🔹 Essential Skills & Requirements
✅ 7+ years of relevant DevOps / Cloud / Observability experience
✅ Strong hands-on experience with Datadog
✅ Strong experience with RUM, Session Replay, Distributed Tracing, and APM
✅ Solid experience with AWS
✅ Strong Infrastructure-as-Code experience using CloudFormation and AWS CDK
✅ Good working knowledge of GitHub Actions and AWS CodePipeline
✅ Real incident response experience on high-traffic systems
✅ Experience with Load & Performance Testing
✅ Strong problem-solving and communication skills
✅ Ability to work closely with infrastructure partners and internal platform teams
🔹 Skills - Good to Have
⭐ Experience with high-traffic platforms
⭐ Experience with campaign readiness and gamedays
⭐ Infrastructure partner management
⭐ Advanced AWS observability
⭐ Performance engineering
📌 Experience: 7+ Years
📌 Work Location: Kochi / Trivandrum / Remote
📌 Shift: General Shift with Australian Overlap
📌 Remote: Mandatory 1 week at office
📩 Interested candidates can share their resume
Job Description
We are looking for an Azure Cloud & Observability Engineer with strong experience in Azure infrastructure and enterprise monitoring tools such as Splunk, Grafana, and AppDynamics.
Responsibilities
- Design, deploy, and manage Azure cloud infrastructure and services.
- Monitor application and infrastructure performance using Splunk, Grafana, and AppDynamics.
- Configure dashboards, alerts, health rules, and monitoring metrics.
- Perform log analysis, troubleshooting, and root-cause analysis for production issues.
- Implement observability solutions for applications, cloud infrastructure, and services.
- Automate monitoring and operational activities using scripting.
- Support incident, problem, and change management processes.
- Collaborate with development, DevOps, and SRE teams to improve system reliability.
- Maintain monitoring standards, documentation, and operational procedures.
Primary Skills
- Microsoft Azure
- Splunk
- Grafana
- AppDynamics
- Cloud Monitoring & Observability
- Application Performance Monitoring (APM)
- Log Analysis & Troubleshooting
Secondary Skills
- Azure Monitor / Log Analytics
- Azure VMs, Storage, Networking
- Linux
- Python / PowerShell / Shell Scripting
- CI/CD
- Git
- ITIL / ServiceNow
Key Responsibilities
- Design, deploy, configure, upgrade, and administer enterprise-scale Dynatrace environments.
- Deploy and manage OneAgent, ActiveGate, extensions, and monitoring configurations across application and infrastructure environments.
- Implement observability for applications, APIs, microservices, databases, containers, Kubernetes, cloud platforms, and traditional infrastructure.
- Configure service detection, process groups, management zones, tags, naming rules, metrics, logs, traces, and topology.
- Develop operational and executive dashboards, notebooks, reports, SLOs, and alerting strategies.
- Configure and optimize Davis AI problem detection, anomaly detection, baselines, and root-cause analysis.
- Implement Real User Monitoring (RUM), Synthetic Monitoring, Session Replay, distributed tracing, and log monitoring as required.
- Analyze application performance issues, service dependencies, transaction traces, response times, resource utilization, and infrastructure bottlenecks.
- Lead troubleshooting of complex production performance and availability incidents using Dynatrace telemetry.
- Reduce alert noise through effective event correlation, thresholds, anomaly-detection configuration, and monitoring standards.
- Integrate Dynatrace with enterprise platforms such as ServiceNow, Jira, PagerDuty, Splunk, CI/CD pipelines, and collaboration/notification tools.
- Automate Dynatrace configuration and deployment using APIs, configuration-as-code, scripting, and DevOps tooling.
- Work closely with application, infrastructure, cloud, SRE, DevOps, production support, and operations teams to define observability requirements.
- Establish Dynatrace monitoring standards, reusable configurations, governance, and best practices.
- Perform platform health checks, capacity assessments, license/usage optimization, and monitoring coverage reviews.
- Create technical documentation, runbooks, architecture diagrams, troubleshooting guides, and operational procedures.
- Mentor junior engineers and provide technical leadership for observability initiatives.
Required Technical Skills
- 5–8+ years of overall IT experience with significant experience in application/infrastructure monitoring or observability.
- 3–5+ years of hands-on Dynatrace experience in enterprise environments.
- Strong knowledge of:
- Dynatrace OneAgent and ActiveGate
- Application Performance Monitoring (APM)
- Infrastructure Monitoring
- Distributed Tracing
- Real User Monitoring (RUM)
- Synthetic Monitoring
- Log Monitoring and Analytics
- Metrics, traces, logs, and events
- Davis AI and automated root-cause analysis
- Dashboards, SLOs, alerting, and anomaly detection
- Dynatrace APIs and automation
- Experience monitoring Java/JVM, .NET, web applications, APIs, microservices, and databases.
- Experience with Kubernetes, Docker, OpenShift, or other container platforms.
- Working knowledge of at least one major cloud platform: AWS, Microsoft Azure, or Google Cloud Platform.
- Strong understanding of application architecture, HTTP/HTTPS, REST APIs, networking, operating systems, and middleware.
- Experience with Linux and Windows environments.
- Scripting/automation experience using Python, PowerShell, Bash, Ansible, Terraform, or similar technologies.
- Experience integrating monitoring platforms with ITSM, incident management, and DevOps tools.
- Strong analytical and production troubleshooting skills.
Preferred Skills
- Experience with Dynatrace Grail, DQL (Dynatrace Query Language), OpenPipeline, and the latest Dynatrace platform capabilities.
- Knowledge of OpenTelemetry (OTel) and modern telemetry standards.
- Experience implementing observability for large-scale Kubernetes and cloud-native environments.
- Familiarity with SRE practices, including SLIs, SLOs, error budgets, and observability-driven incident management.
- Knowledge of additional monitoring platforms such as Splunk, AppDynamics, Datadog, New Relic, Grafana, Prometheus, or ELK.
- Experience with infrastructure-as-code and configuration-as-code approaches.
- Dynatrace certifications are preferred.
Professional Skills
- Strong problem-solving and root-cause analysis capabilities.
- Ability to troubleshoot complex application and infrastructure performance issues independently.
- Strong written and verbal communication skills.
- Ability to work effectively with application owners, developers, SREs, infrastructure teams, and senior stakeholders.
- Ability to translate business and operational requirements into observability solutions.
- Experience working in enterprise production environments with incident, problem, and change-management processes.
- Ability to mentor engineers and drive technical standards across teams.
Education & Certification
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent professional experience.
- Dynatrace Associate/Professional-level certification is desirable.
- Cloud, Kubernetes, ITIL, SRE, or DevOps certifications are advantageous.
Experience Level
Senior Engineer
- Overall Experience: 5–8+ years
- Dynatrace/Observability Experience: 3–5+ years
- Expected proficiency: Advanced hands-on implementation, administration, troubleshooting, automation, and solution design
🚀 Hiring: Senior Azure Platform Engineer | Azure | Terraform | Kubernetes
📍 Location: India
💼 Employment: Full-Time | Long-Term
🎯 Experience: 10+ Years | 8+ Years Hands-on Azure
🔑 What We’re Looking For
▪️ Strong expertise in Azure Cloud Platform & Landing Zone Architecture
▪️ Advanced hands-on experience with Terraform & reusable IaC modules
▪️ Expertise in Kubernetes, Helm & container platforms
▪️ Strong understanding of Azure Networking, Entra ID, RBAC & Azure Policy
▪️ Experience with GitHub, GitHub Actions / Azure Pipelines & CI/CD
▪️ Hands-on GitOps & ArgoCD experience
▪️ Strong observability skills with Prometheus, Grafana & Azure Monitor
▪️ Experience designing and supporting microservices architectures
▪️ Knowledge of security, policy-as-code and IaC security scanning
▪️ Exposure to Azure Arc / Azure Local / Azure Stack HCI is a plus
Cloud Expertise(Azure):
• Strong understanding of cloud services and resources like AI services, webapp, database, including monitoring tools like Azure Monitor and Log Analytics.
• Experience with Infrastructure as Code (IaC) tools such as Arm template / Bicep/Terraform.
• Deep understanding of Networking concepts(DNS, DHCP , Hub and Spoke).
• Understanding on policies and security aspects of cloud.
Kubernetes & Helm:
• In-depth knowledge of Kubernetes concepts such as pods, services, ingress, config maps, and secrets.
• Understand of Kubernetes templates and its deployment.
• Proficiency with Helm/ Kustomize or equivalent for Kubernetes package management and deployment automation.
• Implement Kubernetes best practices, including security, networking, and scaling.
• Concepts of Docker and Containers
CI/CD & Programming:
• Hands-on experience with YAML-based CI/CD pipelines (e.g., Azure DevOps, GitHub Actions).
• Familiarity with scripting and automation tools such as PowerShell, Azure CLI, or Bash.
• Proven skill in python programming and concepts.
Monitoring and Observability:
Expertise in creating and managing Grafana dashboards for visualizing metrics and logs.
• Knowledge of Log Analytics & Azure Application Insights for performance monitoring and tracing.
Location - Bangalore Skill/Experience Expectations: 1. Total Experience 7-11 yrs 2. 3-4 years in managing scalable production environment 3. 2-4 yr experience in managing Google cloud infrastructure 4. proficient in terraform and any programming language 5. Expert in designing and managing observability solutions 6. 5 yr experience in DevOps and SRE practices and troubleshooting critical incidents.
Observability Engineer (AppDynamics)
Hyderabad
Exp: 8+years of exp
Mandate Skills: AppDynamics, Splunk, Python (Scripting knowledge)
Primary Skill Set:
- AppDynamics administration
- SPLOC administration
- Enterprise monitoring and observability
- Application performance monitoring
- Alerting and event management
- Monitoring strategy and design
- Platform configuration and governance
- Python scripting
- Monitoring automation
Secondary Skills:
- Glassbox monitoring
- Customer journey observability
- GenAI concepts for operations
- Log, metric, and trace telemetry
- Dashboarding and visualization

Position Overview
The AI Observability Engineer will be instrumental in implementation of scalable, cloud-native solutions to meet the growing needs of our Data & Development team. The successful candidate will demonstrate the ability to abstract complexity and create reusable, scalable patterns that accelerate development. The AI Observability Engineer will build and maintain a robust framework to ensure the reliability and maintainability of DPR Construction's complex AI systems.
Responsibilities
- Standardize observability practices across AI/ML and other development teams including logging, metrics, tracing, and model performance monitoring, ingesting data from multiple platforms
- Lead hands-on implementation of automation-first DevOps and MLOps practices, enabling infrastructure-as-code and consistent, repeatable environment provisioning
- Design and manage intelligent DataOps pipelines with automated data quality monitoring and anomaly detection
- Deploy, maintain and monitor containerized ML workloads
- Extend existing CI/CD pipelines to support automated infrastructure changes and ML workflows
- Implement AI-driven data validation, schema and concept drift detection and metadata management.
- Establish governance frameworks for AI systems, including bias detection, explainability, and auditability
- Extend existing Azure RBAC strategy by automating role and permission management to reduce manual intervention
- Develop automated test suites for model performance, regression, edge cases and bias validation
- Monitor model KPIs (accuracy, precision, recall, latency, calibration)
- Ensure reproducability of experiments and production models
- Act as a technical point of contact for DevOps and MLOps practices, developing reusable patterns, documentation, and proof-of-concepts to drive adoption
Qualifications
- Bachelor’s degree in computer science, Data Science, Information Systems, or a related field
- 5+ years of experience in DevOps, MLOps, Data Engineering, Software Engineering or Site Reliability Engineering
- Strong understanding of cloud infrastructure and experience working with at least one major cloud provider, preferably Azure
- Proficiency in at least one objected-oriented programming language, preferably python with hands-on experience in ml frameworks like TensorFlow, PyTorch or Scikit-learn
DevOps, Cloud Engineering, or related roles Strong hands-on experience with Microsoft Azure (VMs, App Services, AKS, Storage, Networking, etc.) Expertise in Azure DevOps (Pipelines, Repos, Boards, Artifacts) Experience with Infrastructure as Code tools such as ARM, Bicep, or Terraform Proficiency in scripting (PowerShell, Bash, or Python) Hands-on experience with containerization and orchestration (Docker, Kubernetes/AKS) Strong experience with monitoring and logging tools within Azure ecosystem Good understanding of cloud security, networking, and governance in Azure Experience in managing production environments and handling escalations

Key Skills:
• Bachelor's or Master's degree in Computer Science or related field.
• Minimum 5 years of experience in Platform Engineering, DevOps, or Cloud Infrastructure Engineering.
• Experience migrating data and systems between AWS IaaS and PaaS.
• Experience operating and supporting applications using AWS VPC, EKS, and related services for multi-account operations.
• Experience developing fast and reliable Continuous Integration/Continuous Deployment (CI/CD) workflows used by hundreds of application teams.
• Experience administering and troubleshooting Operating Systems such as Linux, Windows, and MacOS.
• Professional Certifications in AWS Networks, CNCF Technologies, or Kubernetes.
• Experience using and configuring observability tools such as ELK, Prometheus/Grafana, AWS CloudWatch, and Jaeger.
• Experience of applied GitOps principles using ArgoCD or Flux.
• Public examples of code you've worked on with other people using any of these technologies:
o Configuration management/Infrastructure as Code (IAC) tools, such as AWS CDK, AWS CloudFormation, Terraform, Ansible, or Puppet.
o Systems solutions in one or more programming languages, such as Golang, Python, Java.
o Build, Release, Deploy or Ops Workflows using Bamboo, Argo Project, or GitHub Actions.





