Infrastructure Engineer at Kutumb · Bengaluru (Bangalore) · 2 - 4 years · ₹15L - ₹30L / yr · Raised funding · Posted 20 May 2022

Kutumb is the first and largest communities platform for Bharat. We are growing at an exponential trajectory. More than 1 Crore users use Kutumb to connect with their community. We are backed by world-class VCs and angel investors. We are growing and looking for exceptional Infrastructure Engineers to join our Engineering team.
More on this here - https://kutumbapp.com/why-join-us.html">https://kutumbapp.com/why-join-us.html
We’re excited if you have:
- Recent experience designing and building unified observability platforms that enable companies to use the sometimes-overwhelming amount of available data (metrics, logs, and traces) to determine quickly if their application or service is operating as desired
- Expertise in deploying and using open-source observability tools in large-scale environments, including Prometheus, Grafana, ELK (ElasticSearch + Logstash + Kibana), Jaeger, Kiali, and/or Loki
- Familiarity with open standards like OpenTelemetry, OpenTracing, and OpenMetrics
- Familiarity with Kubernetes and Istio as the architecture on which the observability platform runs, and how they integrate and scale. Additionally, the ability to contribute improvements back to the joint platform for the benefit of all teams
- Demonstrated customer engagement and collaboration skills to curate custom dashboards and views, and identify and deploy new tools, to meet their requirements
- The drive and self-motivation to understand the intricate details of a complex infrastructure environment
- Using CICD tools to automatically perform canary analysis and roll out changes after passing automated gates (think Argo & keptn)
- Hands-on experience working with AWS
- Bonus points for knowledge of ETL pipelines and Big data architecture
- Great problem-solving skills & takes pride in your work
- Enjoys building scalable and resilient systems, with a focus on systems that are robust by design and suitably monitored
- Abstracting all of the above into as simple of an interface as possible (like Knative) so developers don't need to know about it unless they choose to open the escape hatch
What you’ll be doing:
- Design and build automation around the chosen tools to make onboarding new services easy for developers (dashboards, alerts, traces, etc)
- Demonstrate great communication skills in working with technical and non-technical audiences
- Contribute new open-source tools and/or improvements to existing open-source tools back to the CNCF ecosystem
Tools we use:
Kops, Argo, Prometheus/ Loki/ Grafana, Kubernetes, AWS, MySQL/ PostgreSQL, Apache Druid, Cassandra, Fluentd, Redis, OpenVPN, MongoDB, ELK
What we offer:
- High pace of learning
- Opportunity to build the product from scratch
- High autonomy and ownership
- A great and ambitious team to work with
- Opportunity to work on something that really matters
- Top of the class market salary and meaningful ESOP ownership

About Kutumb
About
Kutumb is the first and largest communities platform for Bharat. We are growing at an exponential trajectory. More than 1 Crore users use Kutumb to connect with their community. We are backed by world-class VCs and angel investors.
Connect with the team
Similar jobs (9)
Job Summary
We are looking for a Senior Observability Engineer with strong expertise in ThousandEyes, Splunk, and Monitoring & Observability Engineering. The candidate should have experience implementing and enhancing monitoring solutions, along with basic Python automation scripting.
Primary Skills
- ThousandEyes
- Splunk
- Monitoring & Observability Engineering
- Oracle Database
- SQL
- MongoDB
- Python Automation
Key Responsibilities
- Implement, configure, and enhance enterprise monitoring and observability solutions.
- Work extensively with ThousandEyes for network and application performance monitoring.
- Develop and maintain monitoring dashboards, alerts, and reports using Splunk.
- Monitor application, infrastructure, network, and database performance.
- Work with Oracle Database, SQL, and MongoDB for monitoring and troubleshooting.
- Develop basic Python automation scripts to improve monitoring and operational efficiency.
- Analyze performance issues and support troubleshooting and root cause analysis.
- Collaborate with application, infrastructure, network, and database teams to resolve observability-related issues.
- Continuously improve monitoring coverage, alerting, and operational processes.
Secondary Skills
- GenAI Concepts
- SaaS Architecture Concepts
- Cloud Application Architecture
Ideal Candidate Profile
- 7+ years of experience in Observability / Monitoring Engineering.
- Strong hands-on experience with ThousandEyes and Splunk.
- Good understanding of databases including Oracle, SQL, and MongoDB.
- Basic hands-on experience with Python automation.
- Understanding of application, infrastructure, and network monitoring.
- Exposure to GenAI, SaaS, or Cloud Application Architecture is an added advantage.
Job Description
We are looking for an Azure Cloud & Observability Engineer with strong experience in Azure infrastructure and enterprise monitoring tools such as Splunk, Grafana, and AppDynamics.
Responsibilities
- Design, deploy, and manage Azure cloud infrastructure and services.
- Monitor application and infrastructure performance using Splunk, Grafana, and AppDynamics.
- Configure dashboards, alerts, health rules, and monitoring metrics.
- Perform log analysis, troubleshooting, and root-cause analysis for production issues.
- Implement observability solutions for applications, cloud infrastructure, and services.
- Automate monitoring and operational activities using scripting.
- Support incident, problem, and change management processes.
- Collaborate with development, DevOps, and SRE teams to improve system reliability.
- Maintain monitoring standards, documentation, and operational procedures.
Primary Skills
- Microsoft Azure
- Splunk
- Grafana
- AppDynamics
- Cloud Monitoring & Observability
- Application Performance Monitoring (APM)
- Log Analysis & Troubleshooting
Secondary Skills
- Azure Monitor / Log Analytics
- Azure VMs, Storage, Networking
- Linux
- Python / PowerShell / Shell Scripting
- CI/CD
- Git
- ITIL / ServiceNow
Job Description
We are looking for an Azure Cloud & Observability Engineer with strong experience in Azure infrastructure and enterprise monitoring tools such as Splunk, Grafana, and AppDynamics.
Responsibilities
- Design, deploy, and manage Azure cloud infrastructure and services.
- Monitor application and infrastructure performance using Splunk, Grafana, and AppDynamics.
- Configure dashboards, alerts, health rules, and monitoring metrics.
- Perform log analysis, troubleshooting, and root-cause analysis for production issues.
- Implement observability solutions for applications, cloud infrastructure, and services.
- Automate monitoring and operational activities using scripting.
- Support incident, problem, and change management processes.
- Collaborate with development, DevOps, and SRE teams to improve system reliability.
- Maintain monitoring standards, documentation, and operational procedures.
Primary Skills
- Microsoft Azure
- Splunk
- Grafana
- AppDynamics
- Cloud Monitoring & Observability
- Application Performance Monitoring (APM)
- Log Analysis & Troubleshooting
Secondary Skills
- Azure Monitor / Log Analytics
- Azure VMs, Storage, Networking
- Linux
- Python / PowerShell / Shell Scripting
- CI/CD
- Git
- ITIL / ServiceNow
Key Responsibilities
- Design, deploy, configure, upgrade, and administer enterprise-scale Dynatrace environments.
- Deploy and manage OneAgent, ActiveGate, extensions, and monitoring configurations across application and infrastructure environments.
- Implement observability for applications, APIs, microservices, databases, containers, Kubernetes, cloud platforms, and traditional infrastructure.
- Configure service detection, process groups, management zones, tags, naming rules, metrics, logs, traces, and topology.
- Develop operational and executive dashboards, notebooks, reports, SLOs, and alerting strategies.
- Configure and optimize Davis AI problem detection, anomaly detection, baselines, and root-cause analysis.
- Implement Real User Monitoring (RUM), Synthetic Monitoring, Session Replay, distributed tracing, and log monitoring as required.
- Analyze application performance issues, service dependencies, transaction traces, response times, resource utilization, and infrastructure bottlenecks.
- Lead troubleshooting of complex production performance and availability incidents using Dynatrace telemetry.
- Reduce alert noise through effective event correlation, thresholds, anomaly-detection configuration, and monitoring standards.
- Integrate Dynatrace with enterprise platforms such as ServiceNow, Jira, PagerDuty, Splunk, CI/CD pipelines, and collaboration/notification tools.
- Automate Dynatrace configuration and deployment using APIs, configuration-as-code, scripting, and DevOps tooling.
- Work closely with application, infrastructure, cloud, SRE, DevOps, production support, and operations teams to define observability requirements.
- Establish Dynatrace monitoring standards, reusable configurations, governance, and best practices.
- Perform platform health checks, capacity assessments, license/usage optimization, and monitoring coverage reviews.
- Create technical documentation, runbooks, architecture diagrams, troubleshooting guides, and operational procedures.
- Mentor junior engineers and provide technical leadership for observability initiatives.
Required Technical Skills
- 5–8+ years of overall IT experience with significant experience in application/infrastructure monitoring or observability.
- 3–5+ years of hands-on Dynatrace experience in enterprise environments.
- Strong knowledge of:
- Dynatrace OneAgent and ActiveGate
- Application Performance Monitoring (APM)
- Infrastructure Monitoring
- Distributed Tracing
- Real User Monitoring (RUM)
- Synthetic Monitoring
- Log Monitoring and Analytics
- Metrics, traces, logs, and events
- Davis AI and automated root-cause analysis
- Dashboards, SLOs, alerting, and anomaly detection
- Dynatrace APIs and automation
- Experience monitoring Java/JVM, .NET, web applications, APIs, microservices, and databases.
- Experience with Kubernetes, Docker, OpenShift, or other container platforms.
- Working knowledge of at least one major cloud platform: AWS, Microsoft Azure, or Google Cloud Platform.
- Strong understanding of application architecture, HTTP/HTTPS, REST APIs, networking, operating systems, and middleware.
- Experience with Linux and Windows environments.
- Scripting/automation experience using Python, PowerShell, Bash, Ansible, Terraform, or similar technologies.
- Experience integrating monitoring platforms with ITSM, incident management, and DevOps tools.
- Strong analytical and production troubleshooting skills.
Preferred Skills
- Experience with Dynatrace Grail, DQL (Dynatrace Query Language), OpenPipeline, and the latest Dynatrace platform capabilities.
- Knowledge of OpenTelemetry (OTel) and modern telemetry standards.
- Experience implementing observability for large-scale Kubernetes and cloud-native environments.
- Familiarity with SRE practices, including SLIs, SLOs, error budgets, and observability-driven incident management.
- Knowledge of additional monitoring platforms such as Splunk, AppDynamics, Datadog, New Relic, Grafana, Prometheus, or ELK.
- Experience with infrastructure-as-code and configuration-as-code approaches.
- Dynatrace certifications are preferred.
Professional Skills
- Strong problem-solving and root-cause analysis capabilities.
- Ability to troubleshoot complex application and infrastructure performance issues independently.
- Strong written and verbal communication skills.
- Ability to work effectively with application owners, developers, SREs, infrastructure teams, and senior stakeholders.
- Ability to translate business and operational requirements into observability solutions.
- Experience working in enterprise production environments with incident, problem, and change-management processes.
- Ability to mentor engineers and drive technical standards across teams.
Education & Certification
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent professional experience.
- Dynatrace Associate/Professional-level certification is desirable.
- Cloud, Kubernetes, ITIL, SRE, or DevOps certifications are advantageous.
Experience Level
Senior Engineer
- Overall Experience: 5–8+ years
- Dynatrace/Observability Experience: 3–5+ years
- Expected proficiency: Advanced hands-on implementation, administration, troubleshooting, automation, and solution design
Job Summary
We are looking for an experienced Observability Engineer with strong expertise in ThousandEyes, Splunk, database technologies, and Python automation. The candidate will be responsible for monitoring application, network, and infrastructure performance, identifying issues, and developing automation solutions to improve system visibility, reliability, and operational efficiency.
Required Skills
- Observability
- ThousandEyes / Cisco ThousandEyes
- Splunk
- OracleDB
- MongoDB
- SQL
- Python Automation
Roles & Responsibilities
- Implement and support observability and monitoring solutions across applications, networks, and infrastructure.
- Work with ThousandEyes for network and digital experience monitoring.
- Configure and maintain monitoring dashboards, alerts, and performance metrics.
- Use Splunk for log analysis, monitoring, troubleshooting, and reporting.
- Work with OracleDB, MongoDB, and SQL for data analysis and troubleshooting.
- Develop Python scripts for monitoring and operational automation.
- Analyze performance issues and identify root causes across applications and infrastructure.
- Collaborate with application, infrastructure, and network teams to resolve incidents.
- Improve monitoring, alerting, and automation processes.
- Prepare performance and monitoring reports and provide actionable insights.
Mandatory Skills
- 7+ years of relevant experience
- Strong experience in Observability / Monitoring
- Hands-on experience with ThousandEyes
- Splunk
- Python Automation
- SQL
- OracleDB / MongoDB
Hiring Platform Engineer
Exp: 6 -- 10 yrs
Edu : BE/B.tech/MCA
Work Location : Pune
Skills :
Platform monitoring ,Incident trouble shooting, Incident recovery, openshift ,kubernetes.
2 years of IT operations, infrastructure, cloud or application support experience.
Exp in Linux command-line knowledge.
Exp in networking knowledge including IP addressing, DNS, ports and connectivity troubleshooting.
Hiring: DevOps Lead
📍 Kochi / Trivandrum / Remote
💼 Full-time
🕐 General Shift | Australian Overlap
We are looking for an experienced DevOps Lead to join our team.
🔹 Key Responsibilities
Implement and continually improve the observability platform using Datadog, particularly from a user experience perspective.
Work across engineering squads as a virtual team member, supporting their DevOps, infrastructure, and observability requirements.
Configure and manage Datadog RUM, Session Replay, Distributed Tracing, and APM.
Support campaign readiness activities, including load and performance testing.
Participate in gamedays and incident response activities for production systems.
Liaise closely with the managed infrastructure provider on infrastructure requirements and activities.
🔹 Essential Skills & Requirements
✅ 7+ years of relevant DevOps / Cloud / Observability experience
✅ Strong hands-on experience with Datadog
✅ Strong experience with RUM, Session Replay, Distributed Tracing, and APM
✅ Solid experience with AWS
✅ Strong Infrastructure-as-Code experience using CloudFormation and AWS CDK
✅ Good working knowledge of GitHub Actions and AWS CodePipeline
✅ Real incident response experience on high-traffic systems
✅ Experience with Load & Performance Testing
✅ Strong problem-solving and communication skills
✅ Ability to work closely with infrastructure partners and internal platform teams
🔹 Skills - Good to Have
⭐ Experience with high-traffic platforms
⭐ Experience with campaign readiness and gamedays
⭐ Infrastructure partner management
⭐ Advanced AWS observability
⭐ Performance engineering
📌 Experience: 7+ Years
📌 Work Location: Kochi / Trivandrum / Remote
📌 Shift: General Shift with Australian Overlap
📌 Remote: Mandatory 1 week at office
📩 Interested candidates can share their resume
AWS / Kubernetes / OpenShift / Linux – Job Description
Role: Cloud / DevOps Engineer
Experience: 4–8 Years
Location: Bangalore
Mandatory Skills: AWS, Kubernetes, OpenShift, Linux
Job Summary
Looking for a Cloud/DevOps Engineer with strong hands-on experience in AWS, Kubernetes, OpenShift, and Linux. The candidate will be responsible for managing cloud infrastructure, container platforms, deployments, monitoring, and troubleshooting.
Key Responsibilities
- Manage and support AWS cloud infrastructure and services.
- Deploy and manage applications using Kubernetes and OpenShift.
- Perform Linux administration, troubleshooting, and issue resolution.
- Monitor application and infrastructure health and resolve production issues.
- Work on containerized applications and deployments.
- Support scaling, configuration, upgrades, and maintenance of Kubernetes/OpenShift environments.
- Troubleshoot cloud, container, and Linux-related issues.
- Collaborate with development and operations teams for deployments and releases.
- Follow security, availability, and operational best practices.
Mandatory Skills
- Strong hands-on experience in AWS
- Good experience with Kubernetes
- Hands-on experience in OpenShift
- Strong Linux administration and troubleshooting skills
- Experience in cloud/container environment support
- Good production troubleshooting and incident management skills
Job Description
The engineer will provide direct support to service development teams using the platform and maintain/develop platform components across CI pipelines, tool integrations, deployment architecture, monitoring, documentation, and self-service initiatives.
Responsibilities
- Support onboarding and technical discussions with development teams.
- Provide technical guidance and documentation.
- Support development engineers on CI/CD pipelines, deployments, logging and monitoring.
- Improve platform code, processes and documentation.
- Provide production infrastructure/operations support, including on-call support.
- Research product requirements and new technology rollouts.
Core Mandatory Skills
Kubernetes / Amazon EKS, Terraform, Terragrunt, AWS Cloud, Docker, CI/CD Pipelines, GitOps, Python / Go / Java, Cloud Operations / DevOps, Logging / Monitoring / Tracing, Incident Management







