Principal Architect / Scalability Lead (AWS) at MNC with 5000+ employees ¡ Gurugram ¡ 9 - 18 years ¡ âš30L - âš70L / yr ¡ Posted 19 Mar 2026

Principal Architect / Scalability Lead (AWS)
at MNC with 5000+ employees
Job Title: Principal Architect / Scalability Lead (AWS)
đ Location: Gurgaon (Hybrid)
đ Employment Type: Full-Time
Role Overview
We are seeking a Principal Architect / Scalability Lead with deep expertise in AWS and large-scale distributed systems to architect and scale cloud-native products from MVP to enterprise scale.
This role demands a senior technical leader who has proven experience designing systems that handle high throughput, large concurrent workloads, and enterprise-grade reliability, while ensuring exceptional end-user experience.
You will work closely with Product, Data Engineering, AI/ML, and Backend teams to define architecture standards, scalability roadmaps, and engineering best practices.
Key Responsibilities
đ Architecture & Scalability Leadership
- Architect highly scalable, resilient, and high-performance cloud-native systems on AWS.
- Design distributed systems capable of supporting 100K+ concurrent users.
- Lead architecture evolution from MVP to enterprise-grade deployment.
- Translate business and consumer requirements into robust technical architecture.
- Drive scalability planning, capacity modeling, and performance engineering.
đ End-to-End Ownership
- Own full SDLC visibility from discovery and design to release, monitoring, and optimization.
- Establish best practices for:
- Microservices architecture
- Distributed systems design
- Observability & monitoring
- DevSecOps & CI/CD
- Ensure system uptime, fault tolerance, and cost efficiency.
â AWS Cloud & Infrastructure
- Design and implement scalable systems using AWS services.
- Lead containerization and orchestration using Docker and Kubernetes (EKS).
- Architect secure, automated CI/CD pipelines.
- Drive cloud cost optimization and infrastructure efficiency.
đ Performance & Reliability Engineering
- Define and enforce SLAs, SLOs, and reliability metrics.
- Lead performance testing, load testing, and scalability validation.
- Implement monitoring, alerting, and observability frameworks.
- Design fault-tolerant and highly available systems.
đ§ Backend, Data & AI Collaboration
- Provide architectural guidance for:
- Backend services using Node.js and Python
- Frontend platforms using React / Next.js
- Data platforms using Snowflake
- Collaborate with Data Engineering and AI/ML teams on data-intensive and AI-driven systems.
- Design architectures supporting asynchronous processing, caching, and event-driven workflows.
đĽ Leadership & Governance
- Mentor senior engineers and guide architecture best practices.
- Lead architecture governance and design reviews.
- Influence senior stakeholders with data-driven technical decisions.
- Drive cross-functional alignment across Product, Engineering, Data, and AI teams.
Required Qualifications
- 8â15 years of experience in software engineering.
- Proven experience scaling distributed systems handling 100K+ users or high-throughput workloads.
- Deep hands-on expertise in AWS cloud architecture.
- Strong experience with Docker, Kubernetes, and container orchestration.
- Expertise in microservices, caching strategies, asynchronous processing, and distributed systems.
- Strong understanding of performance engineering and reliability frameworks.
- Experience building enterprise-grade systems for large-scale organizations.
Preferred Skills
- Experience with event-driven architectures (Kafka, SQS, SNS, etc.).
- Knowledge of database scalability and data warehousing (Snowflake).
- Exposure to Data Engineering and AI/ML platforms.
- Strong stakeholder communication and strategic thinking skil

Similar jobs (10)
Skill: Java Architect
Location: Pune Baner Office
Note: Client F2F is mandate
Budget:33 LPA
Experience: 10+ Years
Mandate Skills- Cloud+ Java Architect
Â
Who are we looking for?
We are looking for a Technical Architect, having strong software design and development experience of 9+ years on Core Java, Spring Boot, REST API, AWS & Microservices.
Technical Skills:Â
Proficient in software Design and development and familiar with technologies - Java, Java-J2EE, Spring Boot, Hibernate, Ajax, REST API, Microservices etc.
Working knowledge of JVM internals
Working knowledge in Mongo DB
Working knowledge of any database (MySQL or HSQLDB)
Working knowledge of No-SQL database (Mongo or Dynamo DB)
Working experience with messaging (JMS/RabbitMQ)Â
R & D on new advanced cloud-based technologies in a test-driven agile development.
Experience in designing and architecting systems with high scalability and performance requirements
Ability to design infrastructure for performance evaluation and reporting of cloud-based services, namely AWS
In depth knowledge of key AWS services like EC2, S3, Lambda, CloudWatch etc.
Certification on AWS architecture desirable
Roles and Responsibilities:Â
Participate and contribute in platform requirements/story development.
Contribute to the solutioning, design and design alternatives to the requirements/stories and also participate in design reviews.
Involve in Platform Sprint activities.
Development of assigned stories in appropriate languages defined for each module.
Participate in peer code reviewsÂ
Develop use cases and do unit test cases and execute them part of continuous integration pipeline.
Understand stack deployment and resolve issues with components in stacks across the ecommerce platform.
Process Skills:Â
Agile â Scrum and Test-Driven Development
Qualification:
Bachelor of Engineering (Computer background preferred)
We are hiring a Technical Architect to design robust, scalable and secure systems.
Responsibilities
- Define the architecture for new products and platforms
- Review designs and guide engineering teams
- Set coding, security and performance standards
- Evaluate tools and technologies
Requirements
- 8+ years in software development
- Strong system design and cloud experience
- Hands-on experience with microservices
About the Role
We are looking for a Solution Architect with 10â15 years of experience to design scalable, secure, and cloud-ready solutions. The role involves working with customers, Senior Solution Architects, business analysts, project managers, and engineering teams across cloud, modernization, microservices, integration, and digital transformation initiatives.
Must Have Skills
- 10â15 years of experience in software development, architecture, and enterprise application delivery.
- Strong experience with Azure and/or AWS and cloud-native architecture.
- Hands-on experience with Microservices, REST APIs, SOA, and system integration.
- Experience in application modernization and cloud migration.
- Knowledge of Docker, Kubernetes, CI/CD, and Infrastructure as Code.
- Understanding of cloud security, IAM, networking, monitoring, and disaster recovery.
- Ability to create architecture diagrams, solution designs, and technical documentation.
- Strong analytical, communication, and customer-facing skills.
Good To Have Skills
- AWS/Azure certifications.
- Experience with Serverless, Event-Driven Architecture, Messaging, and API Management.
- Exposure to SaaS / Multi-Tenant Architecture.
- Knowledge of Generative AI and AI-enabled solutions.
- Experience with pre-sales, POCs, estimations, and technical proposals.
- Exposure to application security and compliance.
- Experience working with global customers and distributed teams.
Responsibilities
- Design end-to-end solutions and independently own defined architecture workstreams.
- Translate business and technical requirements into application, cloud, integration, and deployment designs.
- Define microservice, API, integration, and data-flow architectures.
- Guide engineering teams on architecture implementation and conduct design/code reviews.
- Contribute to cloud migration, modernization, and digital transformation initiatives.
- Participate in customer workshops, technical discussions, POCs, and solution demonstrations.
- Identify and address architectural, performance, security, and technical risks.
- Ensure solutions follow defined architecture, security, and engineering standards.
- Work closely with the Senior Solution Architect on complex and strategic engagements.
- Develop reusable architecture patterns, templates, and technical accelerators
Next
We are looking for an AWS Solutions Architect to design secure, scalable cloud architectures.
Responsibilities
- Design AWS solutions following the Well-Architected Framework
- Build serverless architectures with AWS Lambda
- Define infrastructure with CloudFormation
- Design networking and security with VPC and IAM
- Guide engineering teams on cloud best practices and cost
Requirements
- 6+ years in cloud or software engineering, including 3+ on AWS
- AWS Solutions Architect certification preferred
- Strong networking and security knowledge
We're hiring a Cloud Architect (Contract) to work with our Equity Partners who builds profitable growth by acquiring and operating enterprise software companies. Refining a proprietary operating model across 40+ acquisitions and two decades of hands-on experience, now supercharged by our patented agentic AI platform . In this role, you'll take full architectural control of our CI/CD, observability, and event streaming infrastructure, build the standards every new acquisition plugs into, and use AI-assisted automation to keep 20+ products reliable without proportionally scaling headcount.
Job title: Cloud/Platform Architect (SRE)
Type: Global Remote | Contract
What You Bring
- 8â12 years in platform engineering, DevOps, or SRE, with growing ownership over time
- Deep Terraform experience across multi-account, multi-env setups
- Real production experience with event streaming at scale
- Hands-on Grafana, Prometheus, Loki, and strong AWS depth (ECS, EKS, IAM, VPC, RDS)
- SRE fundamentals: SLOs, error budgets, on-call design, post-mortems
- Bonus: acquisition or greenfield platform-building experience
Roles and Responsibilities
- Own everything outside core AWS infra: CI/CD, observability, event streaming, deployment, incidents
- Define the standards every future acquisition will plug into
- Keep 20+ enterprise products running at serious scale (millionsâbillions of requests)
- Build self-service tooling so product teams never wait on you
- Use AI/automation to kill toil â not to replace engineering judgement
Ready to build the platform that scales an entire portfolio? â let's connect.
Technical Architect â Product Engineering
Experience: 15+ Years
Location: Pune, India
Employment Type: Full-time
Desired Skills: Python, Technical Architecture, AWS, Microservices, SaaS / Multi-tenant Architecture, Kubernetes, System Design
About the Role
We are looking for a Senior Technical Architect to lead the architecture, design, and technical evolution of an enterprise SaaS product. This is a hands-on leadership role requiring deep technical expertise, strong product engineering experience, and the ability to build scalable, secure, and high-performance platforms.
The ideal candidate should be passionate about solving complex engineering problems, driving innovation, mentoring development teams, and effectively leveraging AI to accelerate software development.
Key Responsibilities
- Own the overall product architecture and technical roadmap.
- Design and build scalable, secure, and highly available enterprise applications.
- Lead the design and implementation of new product features from concept to production.
- Remain hands-on with coding and contribute to critical product components.
- Drive architecture reviews, code quality, performance optimization, and engineering best practices.
- Lead cloud architecture, security, scalability, and DevOps initiatives.
- Evaluate and adopt modern technologies to improve product capabilities and engineering efficiency.
- Leverage AI tools (ChatGPT, GitHub Copilot, Cursor, Claude, etc.) to accelerate software development, code reviews, testing, documentation, debugging, and productivity.
Required Skills & Qualifications
- 15+ years of software product engineering experience with at least 5 years in a Technical Architect role.
- Strong hands-on expertise in Python and modern backend frameworks.
- Deep experience with AWS services and cloud-native application architecture.
- Strong understanding of DevOps, CI/CD pipelines, Infrastructure as Code (Terraform/CloudFormation), Docker, Kubernetes, and container orchestration.
- Experience designing microservices, REST APIs, event-driven architectures, and distributed systems.
- Strong knowledge of SQL and NoSQL databases.
- Experience with scalable SaaS platforms, multi-tenant architectures, and secure application design.
- Excellent understanding of software design patterns, performance tuning, observability, and system reliability.
- Strong analytical, problem-solving, and decision-making skills.
Read This Before Anything Else
We have 6 developers who can ship. What we don't have is someone who turns that into a real engineering function: real architecture, real leverage, real AI-driven advantage. If that gap sounds like an opportunity rather than a headache, you're in the right place. If it sounds like a lot of undefined work with no playbook handed to you, this one probably isn't for you. That's completely okay. There are plenty of great roles that fit differently.
About CraftMyPlate
CraftMyPlate is Hyderabad's go-to platform for food experiences for micro-events: house parties, birthdays, office celebrations, festive gatherings, and more. We're building the operating system for how India discovers, customises, and orders food for smaller events. We're backed by established founders and investors, and we're funded and growing fast. The next phase of that growth runs through engineering.
Where We Stand
Some numbers, because they matter more than adjectives. Order volume has grown 50x in two years, and we're compounding at roughly 3x year over year, without giving up equity to fund it. That means the business runs on its own economics. The growth is real demand, not runway bought with dilution, and every efficient architectural decision this role makes directly protects that.
Most people size up an opportunity by asking what's going to change in ten years. The more useful question, and the one this company is built around, is what won't change. People will keep gathering. They'll keep celebrating, hosting, and marking festivals, in 10 years and in 20. That permanence is the bet. You're not building infrastructure for a trend cycle. You're building for a category that outlasts the current AI wave, the next funding round, and probably us too.
The Technical Reality
Here's an honest read of the engineering problem, not a sanitized version of it.
Event-driven commerce doesn't scale like typical e-commerce. Demand isn't smooth, it's spiky: weekends, festival calendars, and event dates create real load concentration, and each order is tied to a hard deadline that can't slip the way a shipped package can. That has direct architectural consequences: systems need to handle bursty, unpredictable traffic without paying for idle capacity the rest of the time, which is exactly why we're serverless-first on AWS rather than running a fixed fleet sized for peak.
Underneath that, every order touches multiple systems that have to stay consistent: kitchen and vendor fulfillment status, inventory across partners, payment gateway settlement, and refunds, often in real time and often across more than one vendor for a single event. Getting that consistency right across SQL and NoSQL stores, without it becoming a source of support tickets and manual reconciliation, is a real architecture problem, not a CRUD problem.
The AI-agent layer is the next lever, and it's a business lever as much as a technical one. Every workflow we can hand to a well-orchestrated agent instead of a new hire is a workflow that scales without adding headcount, which is exactly how a company grows 3x a year without diluting equity to fund the team behind it. That's why agent orchestration across multiple LLMs, using LangGraph, sits in the "go deep" tier of this role rather than being a nice-to-have.
You'll likely find some of this framing right and some of it worth challenging once you're actually in the codebase. That's expected, and honestly preferred over someone who just nods along.
Why This Role Exists
You'll be the most senior technical person in the company, reporting directly to the founder. Not a manager brought in to run standups. An owner. You set the architecture, you write code yourself, and you make the team materially better. You also own where AI and automation take this company next, starting with our first in-house AI agent product (details shared in the interview), and expanding from there into how the company runs, department by department: HR, finance, marketing, design, development, all sitting on an engineering layer that you design.
If you've outgrown a role where you plan but don't build, or where good ideas die in a committee, this is built to be the opposite of that.
What You'll Own
- Architecture, end to end. Scalable, cost-efficient systems from day one, not "fix it later" engineering. You own the decisions and their long-term consequences.
- Hands-on building. You are still writing code and shipping. This isn't a seat where you review other people's work all day. You lead by building.
- The engineering team. Directly manage, mentor, and level up our 6 developers. Build the technical bar, the review culture, and the calibration that lets the team ship independently.
- The AI-agent roadmap. Own the architecture behind our first AI agent product, then the broader strategy for AI agents and automation across every function in the company, with engineering as the layer underneath all of it.
- Team scaling. Build the next layer of leads under you so execution quality scales without you being the bottleneck.
- Technical accountability. When something breaks, you fix it. You don't escalate and wait.
Our Stack, and the Depth We Expect
Not everything on this list needs the same level of mastery. Some of it you need to own at an architectural level. The rest you need to be strong enough to build yourself, direct the team on, or delegate to AI agents with confidence.
Go deep here. This is where the real architecture decisions live, and where the business impact is highest:
- AWS, serverless first. You should be genuinely well versed in AWS application development, not just "have used AWS." You should be able to design and guide serverless architecture (Lambda, API Gateway, DynamoDB, Step Functions, and similar) as our default way of building, because our demand curve is spiky by nature and fixed infrastructure is money left on the table.
- TypeScript, our primary language across backend and frontend.
- Agent orchestration across multiple LLMs, using LangGraph. This is core to our AI roadmap and our path to scaling operations without scaling headcount. You own how it's architected, not just how it's used.
Working proficiency. Build it yourself, direct the team, or hand it to an AI agent and know if the output is right.
This Is You If
- You've built and shipped real production systems yourself, not just reviewed other people's architecture from a distance.
- You go deep wherever the problem is, and you're comfortable owning the exact stack described above, not just "full-stack" in the abstract.
- You've made engineers around you measurably better, whether or not you've held the title for it yet.
- You're already using AI coding tools and agents seriously, like Claude, Cursor, or similar tools, as part of how you build, not as something you tried once. We'll likely explore this together in the interview.
- You have a bias toward leverage over hours. You'd rather automate or systematize a problem than grind through it. But when something's live and needs to be done right, you see it through completely, with no half-finished work.
- You want to build something for years, not land somewhere comfortable. We'll know the difference from how you talk about your last three years.
This Might Not Be the Right Fit If
- You'd prefer a stable, well-defined role with clear boundaries and someone else making the calls. That's a fair thing to want, just not what this is.
- You'd rather receive direction than bring us architecture and AI strategy yourself.
- You haven't yet gotten hands-on with AI coding tools in your daily work.
- You're drawn more to the title than the work behind it.
If none of that sounds like you, we'd love to hear from you.
Requirements
- 5 to 7 years of experience in software engineering, with real ownership of architecture-level decisions, not just feature delivery.
- Prior experience leading or mentoring engineers, formally or informally.
- Tier-1 or Tier-1+ engineering college strongly preferred (IIT, BITS, top NIT tier, or equivalent). We'll consider other institutions only with clearly commendable, verifiable work: real systems you can walk us through in depth, strong open-source contributions, or a track record that speaks for itself. Pedigree is a proxy for speed, not a checkbox. We test for the underlying ability regardless.
- Comfortable in an early-stage environment: undefined problems, few processes, and the expectation that you help define both.
Compensation
Competitive, with equity. We're formalizing a structured ESOP program alongside this hire. Specific numbers are discussed directly in later interview rounds.
If reading this got you a little excited about what you'd build here, we'd genuinely love to talk. If it didn't quite land, no hard feelings. We just want the right fit for both sides.
Amuraâs VisionÂ
We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.
Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.
Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.
These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.
We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence.Â
Role OverviewÂ
We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.
This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.
You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability.Â
Key ResponsibilitiesÂ
Cloud Infrastructure & Platform Engineering (AWS)Â
- Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
- Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
- Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
- Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
- Build reusable platform templates and shared infrastructure modules.Â
AI/ML Infrastructure & MLOpsÂ
- Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
- Support GPU-based workloads and optimize compute/storage usage.
- Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
- Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines.Â
CI/CD, Automation & Developer ProductivityÂ
- Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
- Automate deployments, environment provisioning, and release workflows.
- Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
- Implement automated patching, scaling, backups, cleanup workflows, and drift detection.Â
Containers, Kubernetes & Platform Reliability
- Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
- Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
- Optimize infrastructure for performance, resilience, and cost-efficiency.
- Implement progressive deployment strategies including blue/green, canary, and rolling deployments.Â
Observability, Incident Response & SRE Practices
- Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
- Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
- Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.
FinOps, Cost Governance & Security
- Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
- Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
- Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.
Collaboration, Leadership & Platform Culture
- Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
- Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.
Skills & Qualifications
Must-Have:
- 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
- Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
- Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.Â
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.Â
Here are answers to some questions you may have
Where is your office?
Chennai (Velachery)
Work Model
Work from Office â because great stories are built in person!
Do you have an online presence?
https://amura.ai (we are @AmuraHealth on all social media)
About the Role
We are hiring Staff / Principal Engineers to take full, hands-on ownership of Blitzy's most critical production-grade systems and to deliver high-leverage features that materially improve customer outcomes and engineering velocity. This is the most senior individual contributor role at the company today.
This is not a Senior-plus role, an architecture-only role, or a promotion-track role. We are looking for someone who has already operated at Principal / Staff+ scope in a highly technical environment and expects to spend their time writing, reviewing, and shipping production code.
This role is 100% hands-on. Leverage comes from system ownership, execution quality, and durable technical decisions â not people management or process.
Responsibilities
- Own mission-critical production systems end-to-end, ensuring correctness, scalability, performance, reliability, and operational excellence.
- Design, build, and ship high-impact backend systems and features that improve product reliability, performance, and customer value.
- Architect scalable services and cloud infrastructure using technologies such as Python, REST, gRPC, Kubernetes, and Terraform.
- Identify and resolve complex technical bottlenecks that limit engineering quality, system performance, or organizational velocity.
- Build and operate LLM-powered systems and validation loops that evaluate correctness, consistency, durability, and production performance.
- Design and evolve data architectures incorporating relational, NoSQL, graph, and vector databases to support complex enterprise applications and semantic retrieval.
- Modernize and improve complex enterprise systems while balancing reliability, maintainability, scalability, and delivery speed.
- Set and uphold engineering quality standards through hands-on technical leadership, sound technical judgment, and ownership of long-term technical decisions.
Qualifications
- Direct experience with Python as a primary programming language, backend frameworks, and microservices architectures.
- Expertise in REST and gRPC, with proficiency in Node.js and JavaScript.
- Proficiency in GCP, along with experience using at least one additional cloud platform such as AWS or Azure.
- Advanced knowledge of Kubernetes and Terraform in production environments.
- Experience operating highly available production systems, including monitoring, scalability, reliability, performance optimization, and operational tooling.
- Strong knowledge of SQL and NoSQL databases, including PostgreSQL, MySQL, MongoDB, Cassandra, or DynamoDB.
- Familiarity with graph databases such as Neo4j and vector databases or embedding infrastructure for semantic search and retrieval.
- Hands-on experience building and operating LLM-powered systems in production, including evaluation, validation, regression testing, tracing, and failure analysis.
- Working knowledge of LangSmith or comparable LLM observability and evaluation tools; familiarity with OpenAI, Anthropic, or similar model providers is a plus.
- Ability to contribute across the full stack, with a strong understanding of frontend architecture and the ability to debug, design, and ship across frontend, backend, infrastructure, and AI systems.
- Understanding of large-scale enterprise software systems, including architecture, integration, deployment, modernization, and long-term maintainability.
- Proven track record of operating at Staff+, Principal Engineer, or equivalent level, independently driving complex technical initiatives and delivering high-impact outcomes with minimal supervision.
Blitzy is a Cambridge, MA based AI software development platform on a mission to revolutionize the software development life cycle by autonomously building custom software to unlock the next industrial revolution. We're transforming how enterprises build software, turning enterprise requirements into enterprise grade code with an agentic software development platform that can autonomously execute 80% of the quantum of software development work. We're backed by multiple tier 1 investors, and have proven success as founders of previous start-ups.
Our Culture
Who we are:
Led by two pioneering co-founders we are one of the fastest growing companies in the U.S., creating our own category of enterprise autonomous software development. We automate thousands of hours of software development for our customers, which includes strong representation within the Fortune 500.
How we work:
- We move Blitzy Fast: Time is both our companyâs and our clientsâ most precious asset. We move quickly and decisively to innovate internally and deliver exceptional software externally.
- Championship Mindset: We operate like a professional sports team. We win as a team by holding ourselves and each other to high standards, collaborating in-person, and remaining focused on the mission.
- Passion for Invention: Weâre pushing the frontier of whatâs possible, requiring constant innovation and iteration.
- We Work for the Customer: We focus on delivering outsized value to the customers we work with and expanding those relationships into deep, meaningful partnerships.
- We believe in being âeveryday athletesâ: taking care of ourselves so we can bring our best minds to work. We promote great sleep, movement, and restorative activities forÂ
Blitzy is an equal opportunity employer committed to building a diverse and inclusive team. We believe different perspectives make us stronger.
Job Description: Lead - Cloud Engineering (AWS / Azure)
Role Title: Lead - Cloud Engineering
Experience Level: 10+ Years
Domain Focus: Healthcare AI & Cloud Infrastructure
Location: Remote
Job Overview
We are seeking an experienced Lead - Cloud Engineering with over 10 years of IT experience to lead our cloud strategy, architecture, and infrastructure teams. In this role, you will oversee end-to-end cloud deployment, multi-cloud migration, and scalable architecture designed to support cutting-edge Generative AI applications in the healthcare technology domain.
The ideal candidate brings deep technical expertise in both AWS and Azure, strong hands-on capability in cloud infrastructure, and proven leadership experience driving security, compliance, and team growth.
Key Responsibilities
Cloud Architecture & Migration
- Lead the architecture, design, and execution of cloud migrations, deployments, and modernizations across AWS and Azure environments.
- Drive Infrastructure as Code (IaC) standards using Terraform, CloudFormation, or Bicep to ensure scalable, automated infrastructure provisioning.
- Build high-availability, low-latency architectures optimized for data-intensive Generative AI and Machine Learning workloads.
Security & Healthcare Compliance
- Enforce healthcare security standards including HIPAA, HITRUST, SOC 2, and data governance best practices across all cloud assets.
- Implement Zero-Trust security, Identity Access Management (IAM), data encryption key management, and continuous vulnerability monitoring.
Leadership & Team Management
- Manage, mentor, and scale a high-performing team of DevOps, Cloud, and SRE Engineers.
- Drive Agile workflows, sprint planning, incident response frameworks, and SLA compliance.
- Collaborate closely with Data Engineering, AI/ML, and Software Product teams to align infrastructure with business roadmaps.
Operations & FinOps
- Establish cloud cost optimization strategies (FinOps) to manage computing costs associated with AI models and large-scale data processing.
- Manage monitoring, alerting, and telemetry frameworks (e.g., Prometheus, Datadog, CloudWatch) to ensure 99.99% uptime.
Key Requirements
- Experience: 10+ years of overall IT experience with at least 5+ years in a cloud leadership or lead architect role.
- Cloud Platforms: Advanced hands-on expertise with both AWS (e.g., EC2, S3, EKS, Bedrock, SageMaker) and Azure (e.g., AKS, Azure OpenAI, Blob, Virtual Machines).
- DevOps & IaC: Strong background in Terraform, Docker, Kubernetes, CI/CD pipelines (GitHub Actions, GitLab CI, or Jenkins).
- Domain Knowledge: Prior experience building or managing cloud environments within Healthcare, Life Sciences, or HealthTech is strongly preferred.
- AI/ML Familiarity: Experience supporting cloud infrastructure for machine learning pipelines, LLM deployments, or GPU compute management.
- Certifications (Preferred): AWS Certified Solutions Architect â Professional, Azure Solutions Architect Expert, or Certified Kubernetes Administrator (CKA).







