Primary Skills:
Observability - ELK (Elactic/Kibana), Prometheus, Grafana, PromQL
Software and automation - Java, Python/Shell/Bash, Rest-SOAP API, docker containerization, Kubernetes, Kafka
Reliability and DR engineering - Distributed architecture and distributed system fundamentals, micro services, and event-driven architecture.
Cross-team coordination, incident triage and resolution, leadership and stakeholder management.
Secondary Skills:
Lang-chain, Langraph, RAG, MCP
Experience with working on LLM's and integrating with the existing applications
Python - FastAPI
Cache - Redis
Program Details:
Design, build, and ship LLM-powered and agentic product features that enhance the team efforts and outcomes.
Build agentic AI systems that reason over context, invoke tools, take real actions, and recover gracefully from failure.
Work on integrating the existing AI tools and should know major AI frameworks and libraries.
Own service reliability and operational governance by defining SLA's, managing error budgets, and reporting reliability (MTTD, MTTR) to leadership for prioritization, risk decisions and planning.
Architect and continuously optimize the observability of platform using Kibana/Elastic (ELF) along with other observability tools like Prometheus, Grafana (dashboards, metrics, alert lifecycle), improving detection quality, reducing noise/toil, and enabling faster triage and measurable uptime improvements.
Engineer advance alerting and automation capabilities with Kibana alerting and anomaly detections and integrating response workflows (routing, runbooks, remediation scripts) to standardize on-call execution and accelerate restoration of services.
Lead incident response for customer-impacting issues across teams-coordination, communications, service restoration, and blameless RCA-then corrective actions that prevent recurrence and reduce operational risk.
Design, automate and validate Disaster Recovery and failover for critical services/journeys (RTO/RPO alignment, DR Drills), ensuring resiliency under failure scenarios and improving recovery.
Consult and partner with application teams by providing production readiness inputs (Resiliency patterns, availability, performance/capacity considerations) and driving platform enhancements that improve stability while optimizing infrastructure and observability spend.
If you are interested for this role, kindly acknowledge this email with your interest.