Oracle Cloud Infrastructure (OCI) L3 / Technical Lead at SSA Company · Mumbai · 8 - 15 years · ₹10L - ₹25L / yr · Posted 12 Apr 2026

About the Role
We are seeking an experienced OCI L3 / Technical Lead to own the reliability, performance, security, and cost efficiency of our OCI workloads. You will serve as the highest technical escalation point for OCI operations, lead architecture and automation initiatives, mentor engineers, and collaborate cross-functionally to ensure resilient, compliant, and scalable solutions. This role combines hands-on engineering with leadership responsibilities across incident management, platform engineering, and cloud governance.
________________________________________
Key Responsibilities
1) L3 Operations & Escalation Management
• Act as the final technical escalation point for critical incidents, complex problems, and performance issues across OCI services (Compute, Networking, Storage, IAM, Load Balancer, WAF, OKE/Kubernetes, DBaaS/Autonomous DB, Exadata Cloud Service).
• Lead root cause analysis (RCA), produce corrective/preventive action plans, and drive problem management per ITIL.
• Own on-call rotations for priority incidents; coordinate across L2/L1 teams and vendors for swift resolution.
2) Architecture, Design & Governance
• Design and review high-availability, disaster recovery (HA/DR) architectures leveraging OCI regions, ADs, Fault Domains, Backup/Archive Storage, Data Guard (for Oracle DB), and multi-cloud patterns as needed.
• Define landing zone architectures, tenancy/subscription structure, compartment strategy, IAM policies, tagging, and cost governance.
• Establish standards for network segmentation (VCNs, subnets), routing, VPN/ FastConnect, NSGs/Security Lists, and WAF
3) Observability, Performance & Reliability
• Implement and optimize Monitoring, Logging, Alarms, APM, Tracing, and Log Analytics in OCI.
• Capacity plans, and performance baselines; drive performance tuning of compute, networking, databases, and storage.
4) Security, Compliance & Risk
• Enforce OCI security best practices: IAM least privilege, vaults/keys, secrets management, Cloud Guard, vulnerability scanning, CIS benchmarks, and Security Zones.
• Partner with GRC teams on audit readiness, regulatory compliance (e.g., ISO 27001, SOC 2, PCI DSS), data residency, and incident response tabletop exercises.
• Drive patching baselines, image hardening, and secure configuration drift detection.
5) Cost Management & FinOps
• Implement tagging, budgets, usage reports, and cost policies; recommend rightsizing, storage tiers, autoscaling, and reservations/committed use discounts.
• Run monthly cost reviews and produce optimization recommendations; integrate with FinOps dashboards/tools.
6) Migration & Modernization
• Lead migrations into OCI (re-host, re-platform, re-architect) for workloads including Oracle Databases, app servers, microservices, and data pipelines.
• Guide adoption of managed services (Autonomous DB, OKE, Streaming, Functions, API Gateway, Data Integration) and container strategies.
7) Stakeholder Leadership & Mentoring
• Serve as technical lead for cross-functional projects; translate business needs into robust cloud designs.
• Mentor L1/L2 engineers; deliver runbooks, playbooks, and capability uplift training.
• Collaborate with DBAs, App Owners, Security, Network, DevOps, and Product teams for end-to-end outcomes.
________________________________________
Required Qualifications
• 8–12+ years in cloud/infra engineering; 4+ years hands-on with OCI at scale.
• Deep expertise across OCI core services: Compute, VCN/Networking, Block/Object/Archive Storage, Load Balancer, WAF, IAM, Cloud Guard, Logging/Monitoring, OKE/Kubernetes, Autonomous DB/Exadata Cloud.
• Strong in scripting (Python/Bash/PowerShell)
• Solid understanding of ITIL and incident/problem/change processes.
• Proven experience with HA/DR architectures, performance tuning, and cost optimization.
• Hands-on with security hardening, compliance frameworks, and audit support.
________________________________________
Preferred Certifications (Nice to Have)
• OCI Architect Professional
• OCI Cloud Operations Associate / OCI Security Professional
• Oracle Autonomous Database / Exadata Cloud certifications
• CKA/CKAD (Kubernetes), Terraform Associate
• ITIL v4 Foundation/Managing Professional
________________________________________
Technical Stack (Representative)
• Cloud: OCI (Tenancy, Compartments, IAM, Policies, Tags, Budgets)
• Compute/Containers: Compute instances, OKE, OCI Registry, Functions
• Networking: VCN, Subnets, DRG, NAT, Service Gateway, VPN, FastConnect, NSG, WAF, Load Balancer
• Storage/DB: Block/Object/Archive, File Storage, Autonomous DB, Exadata Cloud Service, Data Guard
• Observability: OCI Monitoring, Alarms, Logging, Log Analytics, APM
• Security: Cloud Guard, Security Zones, Vault, KMS, IAM, Policies
• Automation: Terraform, Ansible, Python/Bash/PowerShell, OCI CLI/SDK
• ITSM: Remedy/ServiceNow/Jira (incidents, changes, CMDB), Confluence/Wiki
________________________________________
Soft Skills & Attributes
• Systems thinking, strong analytical and troubleshooting skills.
• Clear communication (can articulate trade-offs and risk).
• Ownership mindset; calm under pressure during Major Incidents.
• Collaborative leadership and mentorship; ability to influence without authority.
________________________________________
Typical Day / Week
• Morning: Review alarms, dashboards, capacity & cost trends; act on exceptions.
• Daytime: Lead solution designs, review plan of actions, mentor engineers, handle L3 escalations.
• Weekly: Architecture council, change advisory board (CAB), cost/security review, RCA review.
• Monthly/Quarterly: DR drills, failure mode analysis, audit support, roadmap updates.
________________________________________
Education
• Bachelor’s/Master’s in Computer Science, Information Technology, or equivalent experience.

About SSA Company
About
Company social profiles
Similar jobs (10)
Senior Cloud Site Reliability Engineer (CSRE) – Azure
About Searce:
Searce is an AI-native, engineering-led modern technology consultancy that empowers
clients to futurify their businesses by delivering real, intelligent business outcomes. As a
trusted partner for over 3,000 clients globally, Searce specializes in cloud modernization,
data engineering, applied AI, and robust cloud platform security. Driven by a "HAPPIER"
cultural mindset and our proprietary evlos problem-solving framework, we eliminate
bureaucratic fluff to build working prototypes fast and scale enterprise production
environments intelligently. We don't just fix systems; we leverage multi-cloud technologies
to transform client operations into distinct competitive advantages.
Position Overview:
We are looking for a high-caliber Senior or Lead Cloud Site Reliability Engineer (CSRE) to
architect, secure, and stabilize next-generation hybrid and multi-cloud environments.
Operating at the intersection of infrastructure design, security compliance, and production
operations, you will serve as the technical Subject Matter Expert (SME) across GCP, Azure,
and AWS.
Whether optimizing a microservice mesh on GKE, tuning autoscaling on AKS, or driving a
massive disaster recovery drill across AWS regions, your focus will be absolute reliability. For
the Lead path, you will couple this deep engineering toolkit with stakeholder management
and mentorship to drive an elite operational culture.
Experience & Level Expectation:
Years of Experience: 3 to 10 years of intensive, hands-on production operations
experience in a dedicated DevOps, Cloud Platform Engineering, or SRE role.
Associate level (3-5 Years): Expected to show flawless execution of IaC, advanced
triaging of infrastructure failures, and ownership of the CI/CD and deployment
lifecycles.
Intermediate level (5-10 Years): Expected to take architectural ownership, serve as
primary Incident Commander for complex outages, design cross-cloud governance
frameworks, and act as a reliable bridge between technical teams and client
leadership.
Key Responsibilities & Role Expectations:
Multi-Cloud Platforms & Orchestration: Design, configure, and maintain
production-grade Kubernetes clusters across major platforms (AKS).
Manage advanced network routing, service meshes (e.g., Istio), and multi-tenant
isolation.
Infrastructure as Code (IaC) & GitOps: Build declarative, enterprise-grade, reusable
infrastructure components using Terraform or Crossplane. Standardize automated
environment provisioning to eliminate configuration drift across multi-branch
environments.
Incident Management & Reliability (SRE): Own and optimize the production on-call
rotation. Lead rapid mitigation strategies for Sev-1/Sev-2 system outages, reducing
Mean Time to Recovery (MTTR) through centralized log and metric correlation.
Root Cause Analysis (RCA): Facilitate rigorous, blameless post-incident reviews to
identify core architectural vulnerabilities and establish long-term fixes preventing
recurrence.
Lifecycle, Patching & Upgrades: Plan and execute zero-downtime cluster upgrades,
operating system patching strategies (Linux/Windows), database lifecycle updates,
and multi-region Disaster Recovery (DR) failover drills.
Core Core Operations & Legacy Integration: Manage enterprise-level hybrid
networking architecture (VPCs, Firewalls, Load Balancers, DNS routing, and DHCP
configurations) while effectively connecting cloud native services to legacy
infrastructures like Active Directory.
Security & Governance: Embed Zero Trust policies, secure secrets management
(Secrets Manager/Key Vault), and continuous vulnerability patching into the
automated SDLC pipeline.
Required Technical Skills:
- Microsoft Azure: Azure Virtual Machines, Virtual Networks, Azure Active Directory, Azure Update Management.
- Containers & Orchestration
- Production-level management of GKE, AKS, and EKS.
- Advanced mastery of Docker, Helm, Kubernetes StatefulSets, Pod Disruption
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Cloud Expertise(Azure):
• Strong understanding of cloud services and resources like AI services, webapp, database, including monitoring tools like Azure Monitor and Log Analytics.
• Experience with Infrastructure as Code (IaC) tools such as Arm template / Bicep/Terraform.
• Deep understanding of Networking concepts(DNS, DHCP , Hub and Spoke).
• Understanding on policies and security aspects of cloud.
Kubernetes & Helm:
• In-depth knowledge of Kubernetes concepts such as pods, services, ingress, config maps, and secrets.
• Understand of Kubernetes templates and its deployment.
• Proficiency with Helm/ Kustomize or equivalent for Kubernetes package management and deployment automation.
• Implement Kubernetes best practices, including security, networking, and scaling.
• Concepts of Docker and Containers
CI/CD & Programming:
• Hands-on experience with YAML-based CI/CD pipelines (e.g., Azure DevOps, GitHub Actions).
• Familiarity with scripting and automation tools such as PowerShell, Azure CLI, or Bash.
• Proven skill in python programming and concepts.
Monitoring and Observability:
Expertise in creating and managing Grafana dashboards for visualizing metrics and logs.
• Knowledge of Log Analytics & Azure Application Insights for performance monitoring and tracing.
Platform Engineering Lead (For client company)
Location: Pune, India
Experience: 7+ years
What Success Looks Like
- Engineering teams ship faster with confidence and built-in guardrails.
- Cloud cost, security, and reliability are predictable, measurable, and well-managed.
- CI/CD pipelines are trusted, standardized, and production-ready.
- Platform decisions reduce cognitive load instead of introducing unnecessary process.
Scope & Expectations
This is a hands-on leadership role combining architecture and implementation.
You will:
- Build, not just review.
- Own the platform roadmap—not just infrastructure tickets.
- Act as a force multiplier for product engineering teams rather than becoming a bottleneck.
- Drive platform strategy while remaining deeply involved in execution.
Key Responsibilities
Platform & Cloud Architecture
- Own Zoop's platform and cloud architecture across GCP and AWS.
- Design reusable, opinionated platform patterns instead of one-off infrastructure.
- Build and evolve Zoop's Internal Developer Platform (IDP), including:
- Self-service environments
- Golden paths (paved roads)
- Standardized templates
- Built-in engineering guardrails
- Lead Kubernetes and cloud-native adoption at scale.
- Drive infrastructure automation using Terraform, Pulumi, or similar Infrastructure-as-Code (IaC) tools.
CI/CD, Reliability & Developer Experience
- Establish robust CI/CD practices with quality gates and production readiness.
- Improve deployment safety through automation and testing.
- Define and monitor:
- Golden Signals
- SLIs
- SLOs
- Incident response processes
- Reduce operational toil and improve developer productivity.
- Make observability a first-class capability using cost-efficient monitoring systems.
- Build an observability platform that multiple engineering teams can easily integrate into their applications.
Security, Privacy & Compliance
- Build security-by-default into infrastructure and deployment pipelines.
- Lead implementation and continuous compliance for:
- DPDP Act (India)
- ISO 27001:2022
- SOC 2 Type II
- Implement:
- Zero Trust architecture
- Least-privilege access
- Secure data isolation
FinOps & Cloud Optimization
- Make cloud costs transparent and accountable across engineering teams.
- Establish FinOps practices including:
- Budgets
- Cost alerts
- Optimization routines
- Drive build-vs-buy decisions using clear ROI analysis.
AI, Data & MLOps Foundations
- Build secure and scalable foundations for AI and MLOps workloads.
- Define guardrails for AI systems and sensitive data handling.
Leadership & Collaboration
- Partner closely with engineering teams to align infrastructure strategy with product goals.
- Mentor engineers and guide teams through technical change.
- Balance long-term platform initiatives with practical execution.
What We're Looking For
Experience
- 7+ years of experience building and operating production infrastructure.
- Experience scaling engineering platforms in high-growth or regulated companies.
- Strong hands-on expertise in:
- Kubernetes and the cloud-native ecosystem
- Service Mesh technologies
- Policy Engines
- GCP, AWS (Azure exposure is a plus)
- Terraform and Infrastructure as Code
Engineering & Operations
- Strong understanding of SDLC and modern CI/CD systems (Jenkins, GitOps, etc.).
- Experience with observability tools such as:
- Grafana
- Prometheus
- New Relic
- Comfortable reading and contributing to production systems written in:
- Go
- Python
- Node.js
Security & Compliance
- Practical experience implementing ISO 27001 and SOC 2 controls.
- Strong understanding of:
- Data protection
- Privacy
- Identity and access management
- Security best practices
Mindset
We're looking for someone who is:
- Action-oriented with sound engineering judgment.
- Analytical, cost-conscious, and reliability-focused.
- Collaborative, calm under pressure, and open to feedback.
- Comfortable challenging decisions and explaining trade-offs when necessary.
Nice to Have
- Experience in fintech, identity, or other regulated industries.
- Built Internal Developer Platforms (IDPs) or shared infrastructure tooling.
- Contributions to open-source projects.
As a DevOps Engineer at YOYO, you'll own the infrastructure and delivery backbone that keeps our platform running as we grow. You'll build the CI/CD, cloud infrastructure, and observability that let a small, fast-moving team ship confidently and you'll keep our AI and data workloads reliable and affordable at scale. This is a hands-on role with real ownership: you won't be maintaining someone else's setup, you'll be shaping ours. You'll work closely with the backend, AI/ML, and data teams to make deployment boring, incidents rare, and scaling a non-event.
If you are Interested DM me on LinkedIn - Saquib Mundagnur
What You'll Own
CI/CD & developer experience - Build and maintain fast, reliable CI/CD pipelines so engineers ship multiple times a day with confidence. - Make the path from commit to production simple, safe, and repeatable, with sensible automated testing, rollbacks, and release controls.Cloud infrastructure & IaC - Own our cloud infrastructure (AWS/GCP) end to end, managed as code (Terraform or similar) — no click-ops. - Design for scale and cost-efficiency as store and conversation volumes grow.
Containers & orchestration - Run our services on containers/Kubernetes: deployments, autoscaling, networking, and resource management. - Support the specific needs of AI/ML workloads, including GPU-backed inference and batch processing for the speech pipeline.
Reliability & observability (SRE) - Own uptime, performance, and incident response — monitoring, logging, tracing, alerting, on-call, and blameless postmortems. - Define and defend SLOs; keep the platform dependable as it scales across clients.
Data & pipeline infrastructure - Support the infrastructure behind large-scale, edge-to-cloud data movement and processing (audio ingestion, ASR/AI pipelines, analytics). - Keep data workloads reliable, performant, and cost-aware.
Security & compliance - Bake security into the platform: secrets management, IAM/least-privilege, encryption in transit and at rest, network hardening, and vulnerability management. - Support compliance readiness (including India's DPDP Act and enterprise-client security requirements) for a product that handles sensitive customer conversations.
Cost & scale - Own cloud cost visibility and optimization; make scaling decisions that balance reliability and spend.
What You'll Bring - 6+ years in DevOps, SRE, platform, or infrastructure engineering, running production systems at meaningful scale.
- Strong hands-on experience with a major cloud provider (**AWS or Azure or GCP**) and Infrastructure-as-Code (**Terraform** or equivalent).
- Solid experience with **containers and Kubernetes** in production. - Experience building and owning **CI/CD** pipelines (e.g. GitHub Actions, GitLab CI, Jenkins, Argo, or similar).
- Comfort with a scripting/automation language (Python, Go, or Bash) and a strong automation-first mindset.
- Real experience with **observability** (Prometheus/Grafana, ELK, Datadog, OpenTelemetry, or similar) and running incident response / on-call.
- A security-conscious approach — secrets, IAM, encryption, and least-privilege as defaults. - Startup temperament: ownership, pragmatism, and a bias to automate and ship. - Based in or willing to relocate to Bangalore, and up for an onsite/hybrid, in-person team
About Searce
Searce is a global, AI-native, engineering-led technology consultancy and a Premier Google
Cloud Partner — recognized as the Google Cloud Workplace AI Transformation Partner of the
Year, APAC (2026). With 20+ years of experience and 3,000+ clients across 10+ countries, we
help businesses stay ahead of the cloud curve.
The Role
We're looking for a Lead Cloud Security & Reliability Engineer with deep GCP expertise to own
end-to-end cloud reliability and security forAPAC enterprise clients. As Lead, you'll set the architectural direction, mentor your squad, and drive measurable client outcomes across multi-
cloud environments.
What You'll Do
Own Client Delivery — Lead 24x7 GCP cloud operations forAPAC clients. Define SLO frameworks and ensure adherence.
Architect Solutions — Design scalable, secure GCP-primary architectures with multi-cloud awareness.
Drive Reliability — Lead incident response, RCA, and long-term remediation across production systems.
Mentor & Elevate — Coach and grow a squad of Senior CSREs.
Drive FinOps — Own cloud cost governance and optimization with quantified impact.
Be the Expert — Represent Searce's technical depth in global client conversations.
What We're Looking For
Experience
7–12 years total with 5+ years on GCP cloud infrastructure
Strong background in Cloud Managed Services / MSP environments
Proven experience leading a team in client-facing delivery
Multi-cloud exposure (AWS/Azure secondary) preferred
Technical Skills (Must-Have)
- GCP: GKE, IAM, VPC, Cloud Monitoring, Stackdriver, KMS — demonstrated in work
- experience
- Kubernetes: GKE — production cluster management, Helm
- IaC: Terraform — module-level, reusable frameworks
- Observability: Prometheus, Grafana, Thanos or equivalent
- Security: IAM, Zero-trust, DevSecOps, CSPM tools
- Scripting: Python or Go
- FinOps: GCP cost governance demonstrated
Nice to Have
- GCP Professional Cloud Architect / Pro DevOps Engineer certification
- AWS / Azure secondary experience
- CKA (Certified Kubernetes Administrator)
- ITIL / change management awareness
- APAC client delivery experience
Why Searce?
🏆 Google Cloud Partner of the Year — APAC 2026
🌍 Work with APAC enterprise clients across multiple industries
🤖 AI-first, engineering-led culture
📈 Lead-level ownership with real career growth
🤝 HAPPIER values — Humble, Adaptable, Positive, Passionate, Innovative, Excellence,
Responsible
Job Description: Lead - Cloud Engineering (AWS / Azure)
Role Title: Lead - Cloud Engineering
Experience Level: 10+ Years
Domain Focus: Healthcare AI & Cloud Infrastructure
Location: Remote
Job Overview
We are seeking an experienced Lead - Cloud Engineering with over 10 years of IT experience to lead our cloud strategy, architecture, and infrastructure teams. In this role, you will oversee end-to-end cloud deployment, multi-cloud migration, and scalable architecture designed to support cutting-edge Generative AI applications in the healthcare technology domain.
The ideal candidate brings deep technical expertise in both AWS and Azure, strong hands-on capability in cloud infrastructure, and proven leadership experience driving security, compliance, and team growth.
Key Responsibilities
Cloud Architecture & Migration
- Lead the architecture, design, and execution of cloud migrations, deployments, and modernizations across AWS and Azure environments.
- Drive Infrastructure as Code (IaC) standards using Terraform, CloudFormation, or Bicep to ensure scalable, automated infrastructure provisioning.
- Build high-availability, low-latency architectures optimized for data-intensive Generative AI and Machine Learning workloads.
Security & Healthcare Compliance
- Enforce healthcare security standards including HIPAA, HITRUST, SOC 2, and data governance best practices across all cloud assets.
- Implement Zero-Trust security, Identity Access Management (IAM), data encryption key management, and continuous vulnerability monitoring.
Leadership & Team Management
- Manage, mentor, and scale a high-performing team of DevOps, Cloud, and SRE Engineers.
- Drive Agile workflows, sprint planning, incident response frameworks, and SLA compliance.
- Collaborate closely with Data Engineering, AI/ML, and Software Product teams to align infrastructure with business roadmaps.
Operations & FinOps
- Establish cloud cost optimization strategies (FinOps) to manage computing costs associated with AI models and large-scale data processing.
- Manage monitoring, alerting, and telemetry frameworks (e.g., Prometheus, Datadog, CloudWatch) to ensure 99.99% uptime.
Key Requirements
- Experience: 10+ years of overall IT experience with at least 5+ years in a cloud leadership or lead architect role.
- Cloud Platforms: Advanced hands-on expertise with both AWS (e.g., EC2, S3, EKS, Bedrock, SageMaker) and Azure (e.g., AKS, Azure OpenAI, Blob, Virtual Machines).
- DevOps & IaC: Strong background in Terraform, Docker, Kubernetes, CI/CD pipelines (GitHub Actions, GitLab CI, or Jenkins).
- Domain Knowledge: Prior experience building or managing cloud environments within Healthcare, Life Sciences, or HealthTech is strongly preferred.
- AI/ML Familiarity: Experience supporting cloud infrastructure for machine learning pipelines, LLM deployments, or GPU compute management.
- Certifications (Preferred): AWS Certified Solutions Architect – Professional, Azure Solutions Architect Expert, or Certified Kubernetes Administrator (CKA).
Bachelor's degree in Computer Science, Information Technology, or a related field (or equivalent experience).
3+ years of hands-on experience with Microsoft Azure, including IaaS, PaaS, networking, identity (Entra ID), and governance services.
2+ years of experience with DevOps tooling and practices, including CI/CD pipelines (Azure DevOps or GitHub Actions), Infrastructure as Code (Terraform, Bicep, or ARM), and version control (Git).
2+ years' experience in infrastructure design, platform engineering, or architecture.
Strong proficiency with automation and scripting — PowerShell, Azure CLI, Python, or Bash.
Solid understanding of containerisation (Docker) and familiarity with orchestration platforms (Azure Kubernetes Service).
Deep understanding of cloud architecture, deployment patterns, and management best practices, including the Azure Well-Architected Framework.
Experience defining standards, governance frameworks, and reference architectures for cloud platform consumption.
Strong knowledge of networking concepts, storage configurations, virtualisation technologies, and disaster recovery/backup design for hybrid environments.
Experience with Azure DevOps (Repos, Pipelines, Boards, Artifacts) or equivalent platforms.
Familiarity with identity and access management (IAM), including Entra ID, RBAC, PIM.
Strong understanding of security, compliance requirements, and IT governance frameworks (ITIL, COBIT, or similar).
Preferred
We are looking for a hands-on Senior AWS Cloud Engineer to lead the infrastructure build, optimization, automation, and production deployment of a Multi-Agent AI Chatbot Platform hosted on AWS. The development environment is already in place, and the successful candidate will drive the solution through testing, integrations, and production go-live.
Key Responsibilities
- Review, validate, and optimize existing Terraform code and AWS infrastructure.
- Establish and manage integrations with enterprise platforms such as ServiceNow, Workday, and other third-party systems.
- Design, build, and support secure, scalable, and highly available AWS environments.
- Implement and automate CI/CD pipelines and Infrastructure-as-Code practices.
- Lead infrastructure testing, performance tuning, and production readiness activities.
- Drive deployment and operationalization of the platform in the Production environment.
- Implement cloud governance, security, monitoring, and FinOps best practices.
- Troubleshoot and resolve complex cloud infrastructure issues.
Required Skills & Experience
- 10+ years of IT experience with strong expertise in AWS Cloud Engineering.
- Proven experience designing, deploying, and managing AWS production environments.
- Strong hands-on experience with Terraform and Infrastructure-as-Code.
- Experience with CI/CD pipeline automation and DevOps practices.
- Expertise in AWS services including VPC, IAM, EC2, S3, Lambda, CloudWatch, and networking.
- Experience in performance optimization, reliability, and cloud cost management (FinOps).
- Strong scripting and automation skills.
- Experience integrating enterprise applications through APIs and secure connectivity patterns.
We are looking for a hands-on Senior AWS Cloud Engineer to lead the infrastructure build, optimization, automation, and production deployment of a Multi-Agent AI Chatbot Platform hosted on AWS. The development environment is already in place, and the successful candidate will drive the solution through testing, integrations, and production go-live.
Key Responsibilities
- Review, validate, and optimize existing Terraform code and AWS infrastructure.
- Establish and manage integrations with enterprise platforms such as ServiceNow, Workday, and other third-party systems.
- Design, build, and support secure, scalable, and highly available AWS environments.
- Implement and automate CI/CD pipelines and Infrastructure-as-Code practices.
- Lead infrastructure testing, performance tuning, and production readiness activities.
- Drive deployment and operationalization of the platform in the Production environment.
- Implement cloud governance, security, monitoring, and FinOps best practices.
- Troubleshoot and resolve complex cloud infrastructure issues.
Required Skills & Experience
- 10+ years of IT experience with strong expertise in AWS Cloud Engineering.
- Proven experience designing, deploying, and managing AWS production environments.
- Strong hands-on experience with Terraform and Infrastructure-as-Code.
- Experience with CI/CD pipeline automation and DevOps practices.
- Expertise in AWS services including VPC, IAM, EC2, S3, Lambda, CloudWatch, and networking.
- Experience in performance optimization, reliability, and cloud cost management (FinOps).
- Strong scripting and automation skills.
- Experience integrating enterprise applications through APIs and secure connectivity patterns.
Preferred Skills
- Experience supporting AI/GenAI, chatbot, or multi-agent platforms.
- AWS Certifications (Associate/Professional).
- Banking or Financial Services domain experience.

This is a senior role in our Application & Database Modernization pillar, on a specific mission: leading a large-scale VMware-to-AWS migration and landing it in production. VMware estates are exactly where modernization programs go to stall — sprawling dependency graphs, undocumented workloads, and a hypervisor bill that grows while the migration deck gathers dust. Our customers need that estate moved, cut over, and retired, on a date in the contract.
You'll lead the design and implementation of the target infrastructure: Landing Zones built to AWS best practices, EKS-based container platforms, Infrastructure as Code across the stack with Terraform and CloudFormation, and HA, DR, security, and backup strategies that hold up in production. Aedeon absorbs the discovery, dependency-mapping, and validation grind that would otherwise consume the project's first two quarters. You own the judgment calls, architecture tradeoffs, cutover sequencing, risk decisions and you carry them through to production. You'll also mentor engineers and shape long-term infrastructure strategy with customer stakeholders. If you want your migration experience to end in retired VMware clusters rather than revised project plans, this is the role.
What you will do?
- Lead the design and implementation of infrastructure for a large-scale VMware-to-AWS migration — from discovery through production cutover.
- Architect and build secure, scalable, and highly available AWS environments, including Landing Zones that follow AWS best practices.
- Design and implement containerized application platforms on Amazon EKS; exposure to ECS is a plus.
- Implement Infrastructure as Code using Terraform and AWS CloudFormation.
- Define and enforce best practices across security, backups, high availability (HA), disaster recovery (DR), monitoring, and operations.
- Enable configuration management using tools such as Ansible, Chef, or similar.
- Build CI/CD pipelines and manage the complete build and release lifecycle for customer applications.
- Drive automation across provisioning, deployment, and operational workflows.
- Retire legacy infrastructure and land cloud-native architectures in its place.
- Improve and maintain DevOps platforms: Jenkins, Git repositories, monitoring, and observability stacks.
- Provide L3-level support for complex infrastructure and platform issues.
- Work with stakeholders on technical strategy and long-term architecture.
- Mentor engineers, contribute to talent evaluation, and support team development.
What are we looking for?
- 6+ years of hands-on DevOps experience, with strong expertise in designing and managing cloud infrastructure.
- Strong experience in VMware-to-AWS migration projects or large-scale infrastructure migrations.
- 4+ years of Terraform and CloudFormation for Infrastructure as Code (IaC).
- 4+ years in configuration management, systems engineering, and managing production-grade infrastructure.
- Solid Linux and/or Windows administration background.
- Deep understanding of AWS services, including VPC, EC2, IAM, EKS, ECS, RDS, S3, Backup, CloudWatch, etc.
- Experience with Kubernetes (EKS), ECS, and Docker in production environments.
- Hands-on experience designing HA, DR, security, and backup strategies.
- Experience with Landing Zone setup, multi-account strategy, and AWS governance frameworks.
- Proficiency in Git with a strong understanding of branching and merging strategies.
- Experience with CI/CD pipelines, automation, and operational tooling.
- A problem-solving mindset that's proactive and oriented toward reliability, performance, and scalability.
You will be preferred if
- Experience across multiple cloud platforms (AWS, Azure, GCP).
- AWS certifications such as Solutions Architect Associate/Professional or DevOps Engineer Professional.
- Exposure to data platforms such as Amazon EMR, Redshift, Lake Formation, and SageMaker.
- Experience designing cloud architectures at an L3/Architect level.









