
DevOps Engineer (Cloud & Infrastructure)
š Noida | š Full-Time | š§ Experience: 2ā4 years
About TestMu AI
TestMu AI (formerly LambdaTest) is an AI-native platform designed to move software testing beyond simple automation into the era of agentic intelligence. It provides end-to-end AI agents that manage the entire Quality Engineering lifecycle.
- Full-Stack AI Agents:Ā Autonomously plan, author, execute, and analyze tests across the SDLC.
- Comprehensive Coverage:Ā Supports web, mobile, and enterprise applications.
- Real-World Testing:Ā Scale execution across real devices, browsers, and custom environments.
About the Role
This isn't a role for someone who just wants to "maintain" systems. As aĀ DevOps EngineerĀ at TestMu AI, you are the architect of the automated highways that power our AI agents. You will step into a fast-paced environment where you bridge the gap between cloud-native automation and core infrastructure.
You will manage complex CI/CD pipelines, troubleshoot deep-seated Linux issues, and ensure our hybrid-cloud environment (AWS/Azure) is as resilient as the code it runs.
Key Responsibilities: The Pillars of Growth
A. DevOps & Automation (50% Focus)
- Platform Orchestration: Lead the migration to modular, self-healingĀ TerraformĀ andĀ HelmĀ templates.
- Agentic CI/CD: ArchitectĀ GitHub ActionsĀ workflows that treat AI agents as first-class citizens, automating environment promotion and risk scoring.
- Kubernetes Mastery: Advanced management ofĀ DockerĀ andĀ K8sĀ clusters to support scalable production workloads.
- Predictive Observability: UseĀ Prometheus, Grafana, and ELKĀ to move from reactive alerts to autonomous anomaly detection.
B. Networking & Data Center Mastery (30% Focus)
- Hybrid Networking: Design and troubleshoot VPCs and subnets inĀ Azure/AWS/GCP, paired with physicalĀ VLANsĀ and switches in our data centers.
- Bare-Metal Lifecycle: Automate hardware provisioning, RAID setup, and firmware updates for our real-device cloud.
- Remote Admin: Master out-of-band management (iDRAC, iLO, IPMI) to ensure 100% remote operational capability.
- Core Protocols: Own the lifecycle of DNS, DHCP, Load Balancing, and IPAM across distributed environments.
C. Development & Scripting (20% Focus)
- Backend Integration: Debug and optimizeĀ Python or GoĀ code; understanding how logic interacts with system-level resources.
- Advanced Scripting: Write idempotentĀ Bash/PythonĀ scripts to automate complex, multi-server operational tasks.
- Agentic Tooling: Support the integration of LLM-based developer tools into DevOps workflows to eliminate "toil".
The Interview Journey
We value your ability to solve problems under pressure more than your ability to memorize documentation.
- Technical Round 1 (DevOps Leads): A live session focused on real-world debugging scenarios and Linux fundamentals.
- Technical Round 2 (Hiring Manager / Pod Lead): An assessment of your architectural thinking, automation strategy, and team alignment.
- Technical Round 3 (SVP Engineering / VP DevOps): Strategic discussion on scalability, infrastructure vision, and technical leadership.
- Final Round (CEO): Mission alignment, cultural fit, and the "big picture" at TestMu AI.
Growth Timeline
This is a high-visibility role. You will receive direct mentorship from our senior engineering leadership. As you master our production environment, you will have a clear path to move intoĀ Senior DevOps EngineerĀ orĀ Infrastructure ArchitectĀ roles as our pods scale.
Perks That Matter
Health Cover: Comprehensive insurance for you and your family.
Fresh Meals: Daily catered meals at the office.
Transport: Safe cab facilities for eligible shifts.
Pod Budgets: Dedicated engagement budgets for team building and offsites.

About TestMu AI (Formely LambdaTest)
About
TestMu AI (formerly LambdaTest) is the worldās first full-stack Agentic AI Quality Engineering Platform.
We built TestMu AI for a reality where software is written by AI and must be shipped at machine speed.
Tech stack
Connect with the team
Similar jobs
Location: Bangalore preferred / Hybrid as applicable
Experience: 3+ years
Education: B.E/B.Tech in Computer Science, Engineering or a related technical discipline
Salary: Above market standards, flexible for the right candidate
Career growth: Long-term opportunity with potential to lead DevOps architecture and cloud platform operations
About FrontM
FrontM builds software platforms for frontline workforces operating in remote and low-connectivity environments, with a strong focus on the maritime industry. The platform supports communication, collaboration, healthcare, learning, welfare and operational workflows across mobile, web, kiosk and connected device environments.
The platform runs across cloud infrastructure, constrained networks and specialised customer environments, requiring reliable DevOps practices, strong observability, secure architecture and careful operational discipline.
Role Summary
As a Senior DevOps Engineer, you will take ownership of FrontMās AWS cloud infrastructure, CI/CD pipelines, platform reliability and technical operations. You will work closely with the VP of Delivery, CTO and CEO to maintain secure, scalable and high-availability infrastructure for FrontMās production systems.
This role requires strong hands-on DevOps experience, broad AWS knowledge, Kubernetes experience and the ability to troubleshoot complex networking and production issues across multi-domain SaaS environments.
Key Responsibilities
Cloud Infrastructure & DevOps Architecture (ā45%)
Ā·Ā Own, maintain and improve AWS cloud infrastructure for FrontM platforms
Ā·Ā Create and maintain Terraform scripts for infrastructure deployment and management
Ā·Ā Manage Kubernetes workloads deployed within AWS EKS
Ā·Ā Support multi-zone AWS infrastructure design for availability, resilience and scale
Ā·Ā Maintain AWS services including Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
Ā·Ā Contribute to DevOps architecture planning in line with FrontMās platform roadmap
CI/CD, Operations & Platform Reliability (ā35%)
Ā·Ā Build, maintain and improve CI/CD pipelines for backend and platform services
Ā·Ā Oversee technical operations with hands-on administration, monitoring and release support
Ā·Ā Ensure continuous server uptime, stability, performance and maintainability
Ā·Ā Debug, respond to and restore system outages in production and staging environments
Ā·Ā Improve observability across infrastructure and applications, including migration from Elastic stack to logz.io
Ā·Ā Support backend stability, scale and performance across Node.js, Java and related services
Security, Networking & Production Support (ā20%)
Ā·Ā Maintain AWS security configurations, access controls and monitoring practices
Ā·Ā Support complex networking requirements across multi-domain SaaS implementations
Ā·Ā Troubleshoot network, infrastructure and access issues with internal teams and customer-side users
Ā·Ā Work with backend teams to support API integrations and infrastructure abstractions for complex requirements
Ā·Ā Document operational procedures, incident findings and technical support steps clearly
Required Technical Skills
Cloud Infrastructure & AWS
Ā·Ā Strong hands-on experience with AWS infrastructure and cloud operations
Ā·Ā Experience with Route 53, EC2, API Gateway, VPC, VPN, AWS Cognito, ElastiCache, DynamoDB and Lambda
Ā·Ā Experience with AWS security setup, monitoring and multi-zone infrastructure
Ā·Ā Ability to manage infrastructure using Terraform
Kubernetes, CI/CD & Observability
Ā·Ā Strong experience with Kubernetes, preferably AWS EKS
Ā·Ā Extensive CI/CD and DevOps experience
Ā·Ā Experience with infrastructure observability and application monitoring tools
Ā·Ā Ability to diagnose production bottlenecks, server failures and performance issues
Backend, Networking & SaaS Operations
Ā·Ā Experience supporting Node.js, Java and backend system procedures for stability and scale
Ā·Ā Good understanding of APIs, integrations and backend service dependencies
Ā·Ā Experience with complex networking and multi-domain SaaS implementations
Ā·Ā Ability to troubleshoot technical issues with non-technical end users
Nice to Have
Ā·Ā Experience with MongoDB clusters in MongoDB Atlas
Personal Attributes
Ā·Ā Strong ownership mindset for uptime, reliability and production stability
Ā·Ā Practical problem-solving approach with the ability to act quickly during incidents
Ā·Ā Clear written and spoken communication in English
Ā·Ā Ability to work independently and coordinate with senior management when required
Ā·Ā Comfortable working in fast-moving engineering teams
Ā·Ā Attention to detail in security, monitoring, documentation and operational processes
Why join FrontM?
Long-Term Career Growth
Opportunity to work on cloud infrastructure used by global maritime and remote workforce customers, with scope to grow into DevOps architecture and platform leadership roles.
Engineering Challenges That Matter
Work on infrastructure that supports applications used in remote, low-bandwidth and operationally demanding environments.
Broad Technical Ownership
Take responsibility across cloud infrastructure, Kubernetes, CI/CD, observability, networking, security and production reliability.
Apply now
Join a team focused on building reliable software infrastructure for real-world use cases and contribute to systems used across the global maritime workforce.
Role : Principal Devops Engineer
About the Client
It is a Product base company that has to build a platform using AI and ML technology for their transportation and logiticsThey also have a presence in the global market
Responsibilities and Requirements
⢠Experience in designing and maintaining high volume and scalable micro-services architecture on cloud infrastructure
⢠Knowledge in Linux/Unix Administration and Python/Shell Scripting
⢠Experience working with cloud platforms like AWS (EC2, ELB, S3, Auto-scaling, VPC, Lambda), GCP, Azure
⢠Knowledge in deployment automation, Continuous Integration and Continuous Deployment (Jenkins, Maven, Puppet, Chef, GitLab) and monitoring tools like Zabbix, Cloud Watch Monitoring, Nagios
⢠Knowledge of Java Virtual Machines, Apache Tomcat, Nginx, Apache Kafka, Microservices architecture, Caching mechanisms
⢠Experience in enterprise application development, maintenance and operations
⢠Knowledge of best practices and IT operations in an always-up, always-available service
⢠Excellent written and oral communication skills, judgment and decision-making skill
About Kiru:
Kiru is a forward-thinking payments startup on a mission to revolutionise the digital payments landscape in Africa and beyond. Our innovative solutions will reshape how people transact, making payments safer, faster, and more accessible. Join us on our journey to redefine the future of payments.
Position Overview:
We are searching for a highly skilled and motivated DevOps Engineer to join our dynamic team in Pune, India. As a DevOps Engineer at Kiru, you will play a critical role in ensuring our payment infrastructure's reliability, scalability, and security.
Key Responsibilities:
- Utilize your expertise in technology infrastructure configuration to manage and automate infrastructure effectively.
- Collaborate with cross-functional teams, including Software Developers and technology management, to design and implement robust and efficient DevOps solutions.
- Configure and maintain a secure backend environment focusing on network isolation and VPN access.
- Implement and manage monitoring solutions like ZipKin, Jaeger, New Relic, or DataDog and visualisation and alerting solutions like Prometheus and Grafana.
- Work closely with developers to instrument code for visualisation and alerts, ensuring system performance and stability.
- Contribute to the continuous improvement of development and deployment pipelines.
- Collaborate on the selection and implementation of appropriate DevOps tools and technologies.
- Troubleshoot and resolve infrastructure and deployment issues promptly to minimize downtime.
- Stay up-to-date with emerging DevOps trends and best practices.
- Create and maintain comprehensive documentation related to DevOps processes and configurations.
Qualifications:
- Bachelor's degree in Computer Science, Information Technology, or a related field (or equivalent work experience).
- Proven experience as a DevOps Engineer or in a similar role.
- Experience configuring infrastructure on Microsoft Azure
- Experience with Kubernetes as a container orchestration technology
- Experience with Terraform and Azure ARM or Bicep templates for infrastructure provisioning and management.
- Experience configuring and maintaining secure backend environments, including network isolation and VPN access.
- Proficiency in setting up and managing monitoring and visualization tools such as ZipKin, Jaeger, New Relic, DataDog, Prometheus, and Grafana.
- Ability to collaborate effectively with developers to instrument code for visualization and alerts.
- Strong problem-solving and troubleshooting skills.
- Excellent communication and teamwork skills.
- A proactive and self-motivated approach to work.
Desired Skills:
- Experience with Azure Kubernetes Services and managing identities across Azure services.
- Previous experience in a financial or payment systems environment.
About Kiru:
At Kiru, we believe that success is achieved through collaboration. We recognise that every team member has a vital role to play, and it's the partnerships we build within our organisation that drive our customers' success and our growth as a business.
We are more than just a team; we are a close-knit partnership. By bringing together diverse talents and fostering powerful collaborations, we innovate, share knowledge, and continually learn from one another. We take pride in our daily achievements but never stop challenging ourselves and supporting each other. Together, we reach new heights and envision a brighter future.
Regardless of your career journey, we provide the guidance and resources you need to thrive. You will have everything required to excel through training programs, mentorship, and ongoing support. At Kiru, your success is our success, and that success matters because we are the essential partners for the world's most critical businesses. These companies manufacture, transport, and supply the world's essential goods.
Equal Opportunities and Accommodations Statement:
Kiru is committed to fostering a workplace and global community where inclusion is celebrated and where you can bring your authentic selfbecause that's who we're interested in. If you are interested in this role but don't meet every qualification in the job description, don't hesitate to apply. We are an equal opportunity employer.
About Hive
Hive is the leading provider of cloud-based AI solutions for content understanding,
trusted by the worldās largest, fastest growing, and most innovative organizations. The
company empowers developers with a portfolio of best-in-class, pre-trained AI models, serving billions of customer API requests every month. Hive also offers turnkey software applications powered by proprietary AI models and datasets, enabling breakthrough use cases across industries. Together, Hiveās solutions are transforming content moderation, brand protection, sponsorship measurement, context-based ad targeting, and more.
Hive has raised over $120M in capital from leading investors, including General Catalyst, 8VC, Glynn Capital, Bain & Company, Visa Ventures, and others. We have over 250 employees globally in our San Francisco, Seattle, and Delhi offices. Please reach out if you are interested in joining the future of AI!
About Role
Our unique machine learning needs led us to open our own data centers, with an
emphasis on distributed high performance computing integrating GPUs. Even with these data centers, we maintain a hybrid infrastructure with public clouds when the right fit. As we continue to commercialize our machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering for our customers. Our ideal candidate is someone who is
able to thrive in an unstructured environment and takes automation seriously. You believe there is no task that canāt be automated and no server scale too large. You take pride in optimizing performance at scale in every part of the stack and never manually performing the same task twice.
Responsibilities
ā Create tools and processes for deploying and managing hardware for Private Cloud Infrastructure.
ā Improve workflows of developer, data, and machine learning teams
ā Manage integration and deployment tooling
ā Create and maintain monitoring and alerting tools and dashboards for various services, and audit infrastructure
ā Manage a diverse array of technology platforms, following best practices and
procedures
ā Participate in on-call rotation and root cause analysis
Requirements
ā Minimum 5 - 10 years of previous experience working directly with Software
Engineering teams as a developer, DevOps Engineer, or Site Reliability
Engineer.
ā Experience with infrastructure as a service, distributed systems, and software design at a high-level.
ā Comfortable working on Linux infrastructures (Debian) via the CLIAble to learn quickly in a fast-paced environment.
ā Able to debug, optimize, and automate routine tasks
ā Able to multitask, prioritize, and manage time efficiently independently
ā Can communicate effectively across teams and management levels
ā Degree in computer science, or similar, is an added plus!
Technology Stack
ā Operating Systems - Linux/Debian Family/Ubuntu
ā Configuration Management - Chef
ā Containerization - Docker
ā Container Orchestrators - Mesosphere/Kubernetes
ā Scripting Languages - Python/Ruby/Node/Bash
ā CI/CD Tools - Jenkins
ā Network hardware - Arista/Cisco/Fortinet
ā Hardware - HP/SuperMicro
ā Storage - Ceph, S3
ā Database - Scylla, Postgres, Pivotal GreenPlum
ā Message Brokers: RabbitMQ
ā Logging/Search - ELK Stack
ā AWS: VPC/EC2/IAM/S3
ā Networking: TCP / IP, ICMP, SSH, DNS, HTTP, SSL / TLS, Storage systems,
RAID, distributed file systems, NFS / iSCSI / CIFS
Who we are
We are a group of ambitious individuals who are passionate about creating a revolutionary AI company. At Hive, you will have a steep learning curve and an opportunity to contribute to one of the fastest growing AI start-ups in San Francisco. The work you do here will have a noticeable and direct impact on the
development of the company.
Thank you for your interest in Hive and we hope to meet you soon
- Good knowledge of at least one language (C#, Java, Python, Go, PHP, Node.js)
- Have enough experience on application and infrastructure architectures
- Design and plan cloud solution architecture
- Design for security, network, and compliances
- Analyze and optimize technical and business processes
- Ensure solution and operational reliability
- Manage and provision cloud infrastructure
- Manage IaaS, PaaS, and SaaS solutions
- Design strategies around cloud governance, migration, Cloud operations and DevOps
- Design highly scalable, available, and reliable cloud applications
- Build and test applications
- Deploy applications on cloud
- Integration with cloud services
Certification:
- Architect level certificate of any cloud (AWS, GCP, Azure)

ļ· A Strong Devops experience of at least 4+ years
ļ· Strong Experience in Unix/Linux/Python scripting
ļ· Strong networking knowledge,vSphere networking stack knowledge desired.
ļ· Experience on Docker and Kubernetes
ļ· Experience with cloud technologies (AWS/Azure)
ļ· Exposure to Continuous Development Tools such as Jenkins or Spinnaker
ļ· Exposure to configuration management systems such as Ansible
ļ· Knowledge of resource monitoring systems
ļ· Ability to scope and estimate
ļ· Strong verbal and communication skills
ļ· Advanced knowledge of Docker and Kubernetes.
ļ· Exposure to Blockchain as a Service (BaaS) like - Chainstack/IBM blockchain platform/Oracle Blockchain Cloud/Rubix/VMWare etc.
ļ· Capable of provisioning and maintaining local enterprise blockchain platforms for Development and QA (Hyperledger fabric/Baas/Corda/ETH).
Job Summary
You'd be meticulously analyzing project requirements and carry forward the development of highly robust, scalable and easily maintainable backend applications, work independently, and you'll have the support & opportunity to thrive in a fast-paced environment.
Ā
Responsibilities and Duties:
Ā
- building and setting up new development tools and infrastructure
- understanding the needs of stakeholders and conveying this to developers
- working on ways to automate and improve development and release processes
- testing and examining code written by others and analysing results
- ensuring that systems are safe and secure against cybersecurity threats
- identifying technical problems and developing software updates and āfixesā
- working with software developers and software engineers to ensure that development follows established processes and works as intended
- planning out projects and being involved in project management decisions
Ā
Skill Requirements:
Ā
- Managing GitHub (example: - creating branches for test, QA, development and production, creating Release tags, resolve merge conflict)
- Setting up of the servers based on the projects in either AWS or Azure (test, development, QA, staging and production)
- AWS S3 configuring and s3 web hosting, Archiving data from s3 to s3-glacier
- Deploying the build(application) to the servers using AWS CI/CD and Jenkins (Automated and manual)
- AWS Networking and Content delivery (VPC, Route 53 and CloudFront)
- Managing databases like RDS, Snowflake, Athena, Redis and Elasticsearch
- Managing IAM roles and policies for the functions like Lambda, SNS, aws cognito, secret manager, certificate manager, Guard Duty, Inspector EC2 and S3.
- AWS Analytics (Elasticsearch, Athena, Glue and kinesis).
- AWS containers (elastic container registry, elastic container service, elastic Kubernetes service, Docker Hub and Docker compose
- AWS Auto scaling group (launch configuration, launch template) and load balancer
- EBS (snapshots, volumes and AMI.)
- AWS CI/CD build spec scripting, Jenkins groovy scripting, shell scripting and python scripting.
- Sagemaker, Textract, forecast, LightSail
- Android and IOS automation building
- Monitoring tools like cloudwatch, cloudwatch log group, Alarm, metric dashboard, SNS(simple notification service), SES(simple email service)
- Amazon MQ
- Operating system Linux and windows
- X-Ray, Cloud9, Codestar
- Fluent Shell Scripting
- Soft Skills
- Scripting Skills , Good to have knowledge (Python, Javascript, Java,Node.js)
- Knowledge On Various DevOps Tools And Technologies
Ā
Ā
Qualifications and Skills
Job Type: Full-time
Experience: 4 - 7 yrs
Qualification: BE/ BTech/MCA.
Location: Bengaluru, Karnataka
- Proficient in Java, Node or Python
- Experience with NewRelic, Splunk, SignalFx, DataDog etc.
- Monitoring and alerting experience
- Full stack development experience
- Hands-on with building and deploying micro services in Cloud (AWS/Azure)
- Experience with terraform w.r.t Infrastructure As Code
- Should have experience troubleshooting live production systems using monitoring/log analytics tools
- Should have experience leading a team (2 or more engineers)
- Experienced using Jenkins or similar deployment pipeline tools
- Understanding of distributed architectures
Ā








