System Engineer - Infrastructure Automation at Deltek · Remote only · 3 - 8 years · Profitable · Remote only · Posted 30 Jun 2026

Position Responsibilities :
Job Description
We are seeking a highly motivated c to join our Infrastructure team. This is a hands-on role for someone who enjoys day-to-day infrastructure operations (servers, storage, tickets) and is passionate about transforming manual processes into intelligent, automated workflows using GenAI and modern automation platforms.
The ideal candidate has a strong engineering and scripting background, thrives in collaborative environments, and is excited to work closely with infrastructure engineers and team leads to understand how work gets done today—and then automate it.
This role blends traditional infrastructure operations with next‑generation automation, making it ideal for someone who wants to shape the future of infrastructure operations using platforms such as Power Platform, Tines, and GenAI tools (e.g., Claude).
Key Responsibilities
Infrastructure Operations (Core)
- Perform day-to-day infrastructure operations, including:
- Server provisioning and builds (Windows/Linux)
- Storage administration and lifecycle tasks
- Incident, request, and problem tickets via ITSM platforms
- Support operational hygiene such as patching, compliance checks, renewals, and standard infra maintenance
- Participate in on-call rotations and operational escalations as required
Automation & GenAI Enablement (Primary Focus)
- Identify repetitive, manual, or error-prone infrastructure processes and automate them end-to-end
- Work directly with infrastructure engineers, SMEs, and team leads to:
- Understand current-state processes
- Map workflows and decision points
- Translate them into automated and AI-assisted solutions
- Design and build automation using:
- Microsoft Power Platform (Power Automate, Logic Apps, Copilot Studio)
- Tines for event-driven and security/infrastructure workflows
- GenAI tools (e.g., Claude, Copilot) for decision support, ticket enrichment, summarization, and remediation logic
- Integrate automation with infrastructure and ITSM tools (ServiceNow, scripting engines, APIs, schedulers)
Engineering & Scripting
- Develop and maintain scripts and automation using:
- PowerShell (primary)
- Python, Bash, or similar languages (as applicable)
- Build reusable automation components, libraries, and templates
- Implement guardrails, logging, error handling, and rollback mechanisms
- Create documentation, runbooks, and knowledge articles for automated workflows
Collaboration & Continuous Improvement
- Partner with cross-functional teams (server, storage, network, security, endpoint) to standardize and automate shared workflows
- Continuously look for opportunities to apply AI for:
- Ticket triage and enrichment
- Root cause analysis
- Compliance drift detection
- Infrastructure lifecycle insights
- Track and communicate automation outcomes (time saved, risk reduced, ticket volume decreased)
Required Qualifications
- 5+ years of experience in infrastructure operations (servers, storage, virtualization, or cloud)
- Strong scripting and automation skills (PowerShell required; Python preferred)
- Hands-on experience with infrastructure workflows such as provisioning, patching, compliance, and lifecycle management
- Demonstrated ability to analyze processes and convert them into automated solutions
- Experience working directly with engineers and technical stakeholders to improve operational processes
- Comfort using APIs, webhooks, and integrating multiple systems
Preferred Qualifications
- Experience with Power Platform, Logic Apps, or similar low-code automation tools
- Experience with Tines or comparable orchestration platforms
- Hands-on exposure to GenAI tools (Claude, Copilot, or equivalent) for operational automation or engineering productivity
- Familiarity with ITSM tools (ServiceNow preferred)
- Experience in hybrid environments (on‑prem + cloud)
- Strong documentation and communication skills
What Success Looks Like
- Manual infrastructure tasks are replaced with reliable, observable automation
- Engineers spend less time on repetitive work and more time on high-value engineering
- Incidents and tickets are enriched, routed, or resolved faster through AI-assisted workflows
- Infrastructure processes are consistent, scalable, and easier to operate

About Deltek
About
Project-based businesses transform the world we live in. Deltek innovates and delivers software and solutions that power them to achieve their purpose. Our industry-specific software and information solutions maximize our customers' performance at every stage of the project lifecycle by enabling superior levels of project intelligence, management and collaboration.
Deltek is the recognized global standard for project-based businesses across government contracting and professional services industries, helping more than 30,000 organizations of all sizes deliver on their mission.
With over 4,200 employees worldwide, our team of industry experts is passionately committed to creating exceptional customer experiences.
Similar jobs (10)
Job Description:
Position: Infrastructure / Automation Engineer
Location: Chennai / Pune / Hyderabad / Bangalore
Shift: UK Shift
Experience : 6-8 years overall
Interviews: 2 Interview rounds.
Work Mode: Onsite
Notice Period: Immediate Joiner/Serving Notice Period only
Project Scope:
This engagement automates end-to-end infrastructure build provisioning across on-prem and Azure. The team will build a shared ServiceNow-led intake, approval, orchestration, and closed-loop status model, using Jenkins/Ansible for on-prem builds and Azure DevOps/Terraform for cloud builds. Automation includes standards, security/compliance controls, scan gates, CMDB/change updates, and handover.
Build automated on-prem infrastructure provisioning and configuration workflows.
Automate server builds from approved request through VM/host deployment, OS install, domain join, and baseline configuration.
Develop Jenkins, Ansible, PowerShell, and integrations for network, storage, DNS, IP, agents, patching, monitoring, backup, and security.
Integrate SCVMM/Hyper-V, VMware, OneView, SCCM, Satellite, and related tools.
Encode infrastructure standards as reusable, version-controlled automation artifacts.
Support testing, UAT, defect resolution, documentation, and hypercare.
Required Skills:
6–8 years in infrastructure engineering, systems administration, or automation.
Bachelor’s degree or equivalent experience.
Jenkins, Ansible
PowerShell, Python/Bash
Windows/Linux
VMware, Hyper-V/SCVMM
OneView, SCCM
Satellite, Git, REST APIs
We are seeking a highly skilled Senior DevOps Engineer with 8+ years of professional experience to join our team. In this role, you will design, implement, and optimize cloud infrastructure and CI/CD processes.
You will collaborate closely with development, QA, and operations teams to deliver scalable, secure, automated, and reliable solutions on AWS.
The ideal candidate will have strong hands-on experience with AWS, Terraform, Git/GitHub, Jenkins, PowerShell, Python, AWS Systems Manager (SSM) Documents, and Active Directory (AD) administration, along with a passion for automation, efficiency, and operational excellence.
Key Responsibilities
- Design, build, and maintain scalable cloud infrastructure on AWS.
- Develop and manage Infrastructure as Code (IaC) using Terraform.
- Build, maintain, and optimize CI/CD pipelines using Jenkins and GitHub.
- Use AWS Systems Manager (SSM) to support operational automation and system administration.
- Automate system tasks and administrative workflows using PowerShell and other scripting languages.
- Create, maintain, and execute custom SSM Documents for configuration management, patching, automation, and troubleshooting.
- Manage and administer Active Directory (AD), including users, groups, permissions, policies, authentication, and integration with AWS services.
- Implement and manage version-control workflows in Git and GitHub.
- Ensure infrastructure and deployment processes follow best practices for security, reliability, scalability, and cost optimization.
- Monitor and troubleshoot production systems to ensure high availability, performance, and reliability.
- Collaborate with development teams to improve software delivery processes and release management.
- Mentor junior engineers and contribute to the development of DevOps standards and best practices.
Required Skills and Experience
- 8+ years of professional experience, including at least 4 years in DevOps or Site Reliability Engineering (SRE) roles.
- Strong hands-on experience with AWS services, including EC2, VPC, IAM, S3, EKS, Lambda, and related services.
- Proven expertise in Terraform for Infrastructure as Code.
- Experience administering both Linux and Windows operating systems.
- Experience designing and managing CI/CD pipelines using Jenkins and GitHub.
- Proficiency with Git workflows and source-code management best practices.
- Strong PowerShell scripting skills; familiarity with Python and/or Bash is a plus.
- Experience with AWS Systems Manager (SSM), including creating and managing SSM Documents.
- Hands-on experience managing Active Directory, including users, groups, policies, authentication, permissions, and AWS integration.
- Solid understanding of cloud networking, security, monitoring, and troubleshooting.
- Excellent problem-solving, communication, collaboration, and decision-making skills.
Nice-to-Have Skills
- Experience with containerization and orchestration technologies, such as Docker, Kubernetes, and Amazon EKS.
- Knowledge of monitoring and observability tools, such as Amazon CloudWatch, New Relic, and Sumo Logic.
- An AWS certification, such as AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect.
Who You Are
- You are eager to learn new technologies and continuously improve your skills.
- You make sound decisions and take ownership of your work.
- You are proactive, self-motivated, and comfortable taking initiative.
- You are a strong communicator who enjoys collaborating with cross-functional teams.
- You are committed to improving processes, automation, reliability, and operational efficiency.
- Strong hands-on experience in Microsoft Azure Cloud.
- Good understanding of Azure services such as Compute, Storage, Event Hub, Event Subscription, Storage Queue, and PaaS services.
- Basic understanding of Azure AI Foundry and AI-related Azure service setup.
- Good Azure networking basics: VNet, subnet, routing, and basic troubleshooting.
- Strong knowledge of Terraform, especially:
- Terraform state
- plan / apply
- troubleshooting failures
- migration risks
- Terraform Enterprise concepts
- Strong Python coding capability, not just basic scripting.
- Experience using Python for API integration, automation, JSON/YAML handling, and internal tooling.
- Good understanding of CI/CD pipelines.
- Ability to troubleshoot pipeline failures.
- Comfortable with YAML and JSON.
- Ability to troubleshoot Azure infrastructure/platform issues.
- Ability to collect logs/evidence and coordinate with network/app/Microsoft support teams.
- Basic awareness of agentic AI / LLM concepts.
- Awareness of security and cost best practices.
Good to Have Skills
- Hands-on experience with Harness.
- Hands-on experience with Terraform Enterprise.
- Exposure to LangGraph / LangChain.
- Exposure to agentic AI workflows or skill creation.
- Exposure to Claude or enterprise LLM integrations.
- Knowledge of Azure ML Workspace, model registry, and managed endpoints.
- MLOps / LLMOps knowledge.
- FinOps / Azure cost optimization experience.
- Azure certifications: AZ-104, AZ-305, AZ-400, AZ-500.
Screening Priority:
Azure Cloud + Terraform + Python Coding + CI/CD Troubleshooting + YAML/JSON + Basic Agentic AI Awareness
Job Description:
- Job title : Senior DevOps Engineer (On-Premises Automation)
- Experience: 8+ years
- Location: Chennai / Pune / Hyderabad / Bangalore
- Shift: UK Shift
- Work Mode: Onsite (WFO)
Roles and Responsibilities:
- This engagement automates end-to-end infrastructure build provisioning across on-prem and Azure. The team will build a shared ServiceNow-led intake, approval, orchestration, and closed-loop status model, using Jenkins/Ansible for on-prem builds and Azure DevOps/Terraform for cloud builds. Automation includes standards, security/compliance controls, scan gates, CMDB/change updates, and handover.
Key Responsibilities:
- Lead secure and repeatable on-prem build provisioning pipelines
- Lead Jenkins CI/CD for automated on-prem provisioning and configuration.
- Create pipeline templates for VM/physical host, OS deployment, configuration, certification, and rollback.
- Embed hardened-image validation, scan gates, secrets handling, approvals, and audit evidence.
- Lead Ansible roles, PowerShell modules, test automation, and code reviews.
- Implement logging, monitoring, and root-cause analysis; support UAT and hypercare.
Required Qualifications:
- 8–10 years in DevOps, platform engineering, or infrastructure automation.
- Bachelor’s degree or equivalent experience.
- Jenkins
- Git, CI/CD
- Ansible, PowerShell
- Python, IaC
- security scanning, secrets management
- Windows/Linux.
Essential Duties and Responsibilities:
• Build new automations in Python: API integrations, data pipelines, scheduled jobs, and process replacements scoped with operating partners.
• Maintain the existing Power Automate estate, both unattended cloud and desktop flows. Triage failures, repair flows, and keep unattended runs healthy on the bot-server farm. Operate the Power Platform space around them: environments, solutions, connection references, and pipeline-managed deployments.
• Migrate Power Automate flows to Python where the economics favor it. Retire flows rather than patching them indefinitely.
• Integrate systems over REST APIs. Handle JSON and XML transformation, authentication (OAuth, service principals), and error handling that survives flaky endpoints.
• Author SQL queries, tables, and stored procedures that support automations.
• Operate what you build. Instrument jobs with logging, monitoring, and alerting so failures surface before the business notices them. Write runbooks.
• Improve how automations run. Today they run as scheduled jobs on VMs. Help evaluate and move toward containerized or Azure-native execution (Functions, Container Apps) where it reduces operational load.
• Use AI coding tools as a core part of daily development, within company governance and review standards.
• Document what you build so the next engineer, or an operating partner, can understand and extend it.
Knowledge, Skills and Abilities:
• Python proficiency: clean scripting, packaging, error handling, structured logging, and enough testing to trust a job running unattended at 2 a.m.
• Power Automate strength across cloud and desktop flows: able to read, debug, and repair complex unattended flows built by someone else, plus the platform administration around them. You do not need to love the platform. You do need to support it capably, including solo coverage when other developers are out.
• REST API integration experience, including authentication patterns and rate-limit handling.
• SQL Server competence: comfortable writing T-SQL and authoring queries, tables, and stored procedures through a reviewed, versioned release process.
• Working knowledge of Azure: DevOps pipelines at minimum; Functions, Container Apps, or AKS exposure a plus.
• Daily fluency with AI-assisted development. You should be able to describe, in concrete detail, how you structure work with an agentic coding tool: what you delegate, what you review, where it fails, and how you catch it.
• PowerShell and shell scripting for glue work on Windows and Linux hosts.
• Production instincts: idempotent jobs, retries with backoff, alerting thresholds that page on real problems and stay quiet otherwise.
• Plain written and verbal communication. You will work directly with non-technical process owners who need to understand what an automation does and what to do when it stops.
Training and Experience:
• 3 to 5 years in automation engineering, RPA, or software engineering roles with automations shipped to production and operated afterward.
• A track record you can walk through: what you built, what broke, and what you changed.
• Demonstrated, current use of AI coding tools in real work. Candidates will be asked to describe their workflow in specifics; vague answers end the conversation.
Job Overview
We are looking for a technically strong IT Infrastructure & Automation Engineer to manage, implement, migrate, and optimize enterprise IT environments across Microsoft infrastructure, endpoint management, identity, collaboration, security, and AI-powered automation.
The ideal candidate should have strong hands-on experience with Microsoft technologies, excellent communication skills, and the ability to deliver infrastructure and digital transformation projects end-to-end. Experience with non-Microsoft technologies and multi-vendor environments will be an added advantage.
Key Responsibilities
Microsoft Intune & Endpoint Management
- Implement and manage Microsoft Intune environments.
- Configure device enrollment, application deployment, policies, compliance, and security baselines.
- Manage Windows, macOS, iOS, and Android devices.
- Configure Conditional Access, endpoint security, and compliance controls.
- Manage Windows Autopilot and device lifecycle activities.
- Monitor and resolve endpoint compliance and configuration issues.
SCCM / Configuration Manager
- Deploy, configure, and maintain SCCM infrastructure.
- Manage OS deployment, imaging, software distribution, and patch management.
- Support SCCM and Intune co-management.
- Generate reports and troubleshoot SCCM issues.
- Assist with migration to modern endpoint management.
Microsoft Exchange
- Administer Exchange Online and Hybrid Exchange environments.
- Support mailbox and tenant migration activities.
- Manage mail flow, connectors, transport rules, and email security.
- Configure SPF, DKIM, DMARC, anti-spam, and anti-phishing controls.
- Troubleshoot messaging and mail delivery issues.
Active Directory & Identity
- Manage Active Directory and Group Policy environments.
- Configure Microsoft Entra ID integration and synchronization.
- Implement SSO, MFA, Conditional Access, and identity security.
- Handle user provisioning, onboarding, offboarding, and access management.
- Support hybrid identity environments.
SharePoint & Microsoft 365
- Implement and administer SharePoint Online.
- Configure sites, permissions, document management, and governance.
- Support file-share and legacy platform migrations.
- Integrate SharePoint with Teams, Power Platform, and business applications.
- Support knowledge and collaboration solutions.
- AI & Automation
- Implement Microsoft 365 Copilot and Copilot Studio solutions.
- Develop AI-powered assistants and agents for business functions.
- Work with Azure OpenAI, Azure AI Search, and Document Intelligence.
- Design RAG and knowledge management solutions.
- Integrate AI solutions with SharePoint, Teams, Dataverse, and enterprise applications.
- Implement workflow automation using Power Automate and Power Platform.
- Support AI governance, security, compliance, and responsible AI adoption.
- Security & Compliance
- Work with Microsoft Defender, Conditional Access, Identity Protection, and endpoint security.
- Implement Microsoft Purview, DLP, information protection, and governance solutions.
- Support security and compliance initiatives across Microsoft environments.
- Identify infrastructure and security gaps and recommend improvements.
Required Technical Skills
- Microsoft Intune
- SCCM / Microsoft Configuration Manager
- Exchange Online & Hybrid Exchange
- Active Directory
- Microsoft Entra ID
- SharePoint Online
- Microsoft Teams Administration
- Windows Server Administration
- Microsoft 365 Administration
- PowerShell
- Microsoft Defender Suite
- Conditional Access
- Identity Protection
- Microsoft Purview
- Endpoint Security
- Device Compliance
- Data Loss Prevention
- Microsoft 365 Copilot
- Copilot Studio
- Azure OpenAI
- Azure AI Services
- Power Automate
- Power Platform Integration
- Preferred Experience
Experience with AWS, GCP, VMware, Citrix, Cisco, Fortinet, Palo Alto, Linux, Veeam, ServiceNow, Trend Micro, CrowdStrike, or similar enterprise technologies will be an advantage.
Knowledge of enterprise architecture, cybersecurity, governance, and compliance frameworks is also preferred.
Candidate Profile
- 5+ years of relevant implementation and administration experience.
- Strong troubleshooting and analytical skills.
- Excellent communication and stakeholder management abilities.
- Strong documentation and reporting skills.
- Ability to manage multiple projects independently.
- Comfortable working with business stakeholders and technical teams.
- Bachelor's degree in IT, Computer Science, Engineering, or a related field.
Preferred Certifications
- Microsoft Certified: Endpoint Administrator Associate
- Microsoft 365 Certified: Administrator Expert
- Microsoft Certified: Identity and Access Administrator Associate
- Microsoft Certified: Azure Administrator Associate
- Microsoft Certified: Azure AI Engineer Associate
- Microsoft Certified: Information Protection Administrator Associate
- ITIL Foundation
Key Deliverables
- Intune implementation and endpoint management.
- SCCM deployment, migration, and optimization.
- Exchange Online and Hybrid Exchange implementation.
- Active Directory and Entra ID integration.
- SharePoint Online deployment and governance.
- Microsoft 365 security and compliance implementation.
- AI and Copilot deployment and governance.
- Technical documentation, knowledge transfer, and operational handover.
- Integration of Microsoft and third-party enterprise technologies.
Job Description: AI Engineer – GenAI Platform Automation
Experience: 10+ Years
Location: Remote – Pan India
Employment Type: Haparz Payroll
Work Mode: Remote
Notice Period: Immediate / Short Notice Preferred
About the Role
We are looking for a senior AI Engineer – GenAI Platform Automation to lead automation initiatives across enterprise Generative AI, Data Science, Data Engineering, and Analytics platforms.
The role focuses on building scalable, secure, and self-service automation capabilities across infrastructure provisioning, CI/CD, cloud environments, AI workload deployment, governance, observability, and operational excellence. The ideal candidate will have strong hands-on experience in platform engineering, cloud automation, DevOps, Infrastructure-as-Code, Python, and enterprise GenAI ecosystems.
Key Responsibilities
- Lead end-to-end automation initiatives for enterprise GenAI, Data Science, Data Engineering, Metadata, Data Quality, Event Streaming, and Analytics platforms.
- Design self-service automation for infrastructure provisioning, environment onboarding, deployment, governance, monitoring, and operational workflows.
- Build automation capabilities supporting the AI lifecycle, including experimentation, model training, deployment, inference, observability, and lifecycle management.
- Develop scalable Infrastructure-as-Code solutions using Terraform and cloud-native automation frameworks.
- Design and maintain enterprise CI/CD pipelines, automated testing, deployment, and release processes using modern DevOps toolchains.
- Automate Kubernetes, containers, serverless, and distributed computing environments in collaboration with cloud and platform engineering teams.
- Develop automation solutions for GenAI and Agentic AI applications, including MCP-enabled services, API integrations, workflow automation, and event-driven architectures.
- Implement observability, monitoring, logging, tracing, alerting, automated remediation, and reliability engineering practices.
- Work with architecture, security, governance, engineering, and business teams to ensure enterprise standards and compliance requirements are met.
- Conduct technical design reviews, automation assessments, code reviews, and establish engineering best practices.
- Provide technical leadership and mentorship to engineering teams adopting automation-first and platform engineering practices.
What We’re Looking For
- 10+ years of hands-on experience in platform engineering, automation engineering, cloud engineering, DevOps, or distributed systems.
- Strong experience building enterprise self-service platforms supporting AI/ML, Data Science, Data Engineering, or Advanced Analytics workloads.
- Strong expertise in automation frameworks, CI/CD, DevOps, Infrastructure-as-Code, and software delivery lifecycle automation.
- Hands-on experience with Terraform and cloud-native infrastructure automation.
- Strong experience with Python for automation, orchestration, scripting, tooling, and platform engineering.
- Experience with Bitbucket, Bamboo, Jira, Confluence, or similar enterprise DevOps toolchains.
- Experience working with Kubernetes, containers, serverless platforms, YARN, and distributed processing environments.
- Knowledge of Generative AI and Agentic AI architectures, MCP frameworks, APIs, workflow automation, and enterprise AI platforms.
- Experience with event-driven architectures and technologies such as Kafka and streaming platforms.
- Strong understanding of cloud engineering, networking, security, scalability, resilience, and cost optimization.
- Experience implementing observability solutions covering monitoring, logging, tracing, alerting, and operational dashboards.
- Understanding of metadata management, data lineage, data governance, and semantic-layer concepts is highly valuable.
Good to Have
- Experience supporting enterprise GenAI platforms, AI governance, model management, and AI operationalization.
- Experience with GitOps, DevSecOps, Platform Engineering, and Reliability Engineering practices.
- Exposure to data governance, data quality, metadata management, and model lifecycle automation.
- Experience creating reusable internal developer platforms and self-service engineering tools at enterprise scale.
- Banking, AML, fraud detection, financial crime, or risk analytics domain experience is an advantage.
We are looking for a hands-on Azure DevOps & Infrastructure Engineer to manage enterprise Azure environments, automate infrastructure, build CI/CD pipelines, and support cloud modernization initiatives.
🔑 Key Responsibilities
• Design, implement & support CI/CD pipelines and deployment automation
• Perform TFS / Azure DevOps Server → Azure DevOps Services migration
• Build & manage Azure infrastructure using Infrastructure as Code (IaC)
• Support Azure cloud infrastructure, networking, security, monitoring & production environments
• Drive cloud migration, application modernization & platform transformation initiatives
• Work with Terraform, containerization, DevSecOps & cloud automation
• Troubleshoot infrastructure/deployment issues and support incident & on-call operations
• Create technical documentation, runbooks and deployment standards
🎯 Mandatory Requirements
✅ AZ-204 – Azure Developer Associate – Mandatory
✅ 5–8 years total IT experience
✅ 3+ years hands-on Azure DevOps / Azure Infrastructure experience
✅ Strong experience in CI/CD & Azure DevOps
✅ Hands-on experience with TFS / Azure DevOps Server to Azure DevOps Services migration
✅ Strong understanding of Azure cloud infrastructure & cloud operations
✅ Experience with IaC, provisioning & infrastructure automation
✅ Exposure to cloud/data migration and application modernization
✅ Experience supporting Azure networking, security & monitoring
✅ Experience in production support / managed services / cloud operations
✅ Willingness to work UK/US shifts and participate in on-call support
⭐ Preferred / Optional
• AZ-400 – Microsoft Certified DevOps Engineer Expert
• AZ-104 – Azure Administrator Associate
• HashiCorp Terraform Associate
• Container platforms, orchestration & DevSecOps experience
🎓 Education: Bachelor's degree in Computer Science / IT / Engineering or related field.
Amura’s Vision
We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.
Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.
Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.
These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.
We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence.
Role Overview
We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.
This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.
You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability.
Key Responsibilities
Cloud Infrastructure & Platform Engineering (AWS)
- Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
- Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
- Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
- Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
- Build reusable platform templates and shared infrastructure modules.
AI/ML Infrastructure & MLOps
- Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
- Support GPU-based workloads and optimize compute/storage usage.
- Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
- Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines.
CI/CD, Automation & Developer Productivity
- Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
- Automate deployments, environment provisioning, and release workflows.
- Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
- Implement automated patching, scaling, backups, cleanup workflows, and drift detection.
Containers, Kubernetes & Platform Reliability
- Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
- Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
- Optimize infrastructure for performance, resilience, and cost-efficiency.
- Implement progressive deployment strategies including blue/green, canary, and rolling deployments.
Observability, Incident Response & SRE Practices
- Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
- Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
- Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.
FinOps, Cost Governance & Security
- Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
- Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
- Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.
Collaboration, Leadership & Platform Culture
- Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
- Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.
Skills & Qualifications
Must-Have:
- 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
- Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
- Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Here are answers to some questions you may have
Where is your office?
Chennai (Velachery)
Work Model
Work from Office – because great stories are built in person!
Do you have an online presence?
https://amura.ai (we are @AmuraHealth on all social media)
Company Overview:
Planview is hiring a DevOps Engineer in Bengaluru, India to support Planview SaaS applications across the product line. You will work in a global, collaborative team — owning CI/CD pipelines, cloud infrastructure, and automation to keep deployments fast and systems reliable.
Responsibilities
- Build and maintain CI/CD pipelines in Jenkins for reliable, fast delivery.
- Manage containerized workloads on Docker and ECS — task definitions, services, and clusters.
- Provision and manage AWS infrastructure using Terraform (CloudFormation a plus).
- Automate configuration and deployment tasks using Python (Ansible a plus).
- Set up and maintain monitoring and alerting via New Relic (CloudWatch, Datadog, or Prometheus/Grafana a plus).
- Write Shell and Python scripts to automate operations and reduce manual work.
- Manage Git workflows — branching, merge strategies, and pull request reviews.
- Administer and support MSSQL databases underpinning the product line — backups, restores, and basic performance troubleshooting.
- Troubleshoot deployment, performance, and infrastructure issues with development teams.
- Participate in on-call rotations and drive incident response.
- Continuously improve infrastructure resilience and deployment speed.
- Apply AI-assisted engineering tools (e.g., GitHub Copilot, Claude Code) to speed up IaC authoring, pipeline debugging, and day-to-day scripting.
Qualifications
Must-Have Skills
- Experience: 4–6 years of experience in DevOps, SRE, or Infrastructure Engineering.
- OS: Linux & Windows administration (systemd, package management, log analysis).
- Cloud: AWS (Active Directory, ECS, EC2, CloudFront, S3, VPC, IAM, RDS, Lambda basics).
- Source Control: Git — branching, merge/rebase, PR reviews.
- CI/CD (Jenkins): Pipeline creation and basic Groovy scripting.
- Containerization (Docker + ECS): Task definitions, services, and clusters.
- IaC: Terraform.
- Monitoring: New Relic.
- Scripting: Bash and Python scripting for automation.
- Networking Basics: DNS, load balancers, security groups, VPNs.
- Logging: ELK stack / CloudWatch Logs.
- Database: MSSQL administration — backups, restores, basic performance troubleshooting.
- Infrastructure Automation: Hands-on experience automating infrastructure provisioning, configuration, and deployment end-to-end.
- AI-Assisted Engineering: Comfortable working with AI coding/DevOps assistants (e.g., GitHub Copilot, Claude Code) for IaC generation, scripting, and troubleshooting — verified via a mandatory AI proficiency assessment during interviews.
Nice-to-Have Skills
• Configuration Management: Ansible.
• Additional IaC: CloudFormation.
• Architecture: Knowledge of microservices architecture.
• Cloudflare: DNS, CDN, WAF.
• Artifact Repositories: Nexus, JFrog Artifactory, ECR.
• Other CI/CD Tools: GitHub Actions.
• AWS cost optimization / FinOps awareness.
• Datadog, CloudWatch, or Prometheus/Grafana.
• AIOps: Exposure to AI-driven anomaly detection, root-cause analysis, or incident triage (e.g., Dynatrace Davis AI, Datadog Bits AI, Harness AIDA).
• Database Basics: RDS backups, restores, performance tuning.






