Senior Solution Architect at AI tech startup · Remote only · 5 - 10 years · ₹15L - ₹23L / yr · Remote only · Posted 7 Jan 2026

Key Responsibilities
1. Platform & Application Architecture
- Architect and oversee end-to-end platform design using GCP services (Vertex AI, Cloud Run, GKE, Pub/Sub, Firestore, BigQuery, and IAM).
- Define the architecture for multi-tenant SaaS deployment—ensuring secure data boundaries across customers while maintaining a single codebase.
- Collaborate closely with backend developers to optimize Node.js APIs, data pipelines, and event-driven microservices.
- Design front-end integration flows and ensure UI/UX alignment with backend AI and agentic processes.
2. AI / LLM / Agentic System Architecture
- Lead architecture for all AI initiatives involving RAG pipelines, AI agents, NLP, and conversational AI.
- Evaluate and integrate LLMs and APIs from OpenAI, Google Gemini, Anthropic (Claude), Mistral, and others.
- Define frameworks for AI agent orchestration, prompt engineering, and function/tool calling workflows using frameworks like LangChain or LlamaIndex.
- Establish standards for vector databases (Pinecone, FAISS, Chroma), embeddings, chunking strategies, and retrieval optimization.
- Architect an AgentOps layer for monitoring, safety filtering, and AI model governance.
3. MLOps, PromptOps & DevOps Integration
- Implement CI/CD pipelines for AI models, prompts, and agents using Cloud Build, Terraform, and GitHub Actions.
- Design PromptOps and MLOps pipelines to manage model lifecycle, versioning, rollback, and production promotion.
- Define observability and monitoring mechanisms (OpenTelemetry, Cloud Monitoring, custom logs for LLM outputs).
- Collaborate with DevOps engineers to define blue/green and canary deployments across environments.
4. Security, Compliance & Scalability
- Design tenant-isolated data layers using IAM, VPC-SC, and KMS encryption policies to meet regulatory standards (HIPAA, GxP, GDPR).
- Implement data governance, access control, and threat detection mechanisms for AI-driven workflows.
- Create performance optimization strategies: GPU/TPU allocation, autoscaling policies, token budgets, response caching, and FinOps visibility.
5. Cross-Functional Leadership
- Partner with Product Managers and Business Analysts to translate business requirements into architectural blueprints.
- Work with UX/UI designers to ensure front-end experiences align with AI agent capabilities and backend APIs.
- Provide architectural guidance and technical mentorship to developers, ensuring adherence to best practices across projects.
- Contribute to platform evolution, reference architecture libraries, and reusable design patterns.
Key Skills & Qualifications
Technical Expertise
- Strong background in cloud architecture, specifically Google Cloud Platform (GCP)—Vertex AI, Cloud Run, Cloud Storage, IAM, Pub/Sub, BigQuery, VPC, and Cloud Functions.
- Proficiency with Node.js, TypeScript, and API development for scalable microservices.
- Deep understanding of LLMs, AI agent frameworks, retrieval-augmented generation (RAG), and vector search.
- Experience with LangChain, LlamaIndex, CrewAI, and Agentic AI design patterns.
- Strong understanding of data architecture, including NoSQL/SQL databases (MongoDB, Firestore, PostgreSQL).
- Hands-on experience with container orchestration (Docker, Kubernetes) and Infrastructure-as-Code (Terraform).
Preferred Skills
- Prior experience designing multi-tenant SaaS platforms.
- Exposure to frontend integration (Angular, React, or similar) and UI/UX flow understanding.
- Knowledge of AI safety, risk frameworks, and prompt governance.
- Strong analytical and communication skills for working across business and technical stakeholders.
- Awareness of FinOps and cost optimization in token-based AI ecosystems.
Qualifications
- Bachelor’s or Master’s degree in Computer Science, AI, or Cloud Architecture.
- 8–12 years of professional experience, with at least 3 years focused on AI platform architecture or enterprise-grade SaaS system design.
Certifications in GCP Professional Cloud Architect or Machine Learning Engineer are a plus.

Similar jobs (10)
Job Description: Lead - Cloud Engineering (AWS / Azure)
Role Title: Lead - Cloud Engineering
Experience Level: 10+ Years
Domain Focus: Healthcare AI & Cloud Infrastructure
Location: Remote
Job Overview
We are seeking an experienced Lead - Cloud Engineering with over 10 years of IT experience to lead our cloud strategy, architecture, and infrastructure teams. In this role, you will oversee end-to-end cloud deployment, multi-cloud migration, and scalable architecture designed to support cutting-edge Generative AI applications in the healthcare technology domain.
The ideal candidate brings deep technical expertise in both AWS and Azure, strong hands-on capability in cloud infrastructure, and proven leadership experience driving security, compliance, and team growth.
Key Responsibilities
Cloud Architecture & Migration
- Lead the architecture, design, and execution of cloud migrations, deployments, and modernizations across AWS and Azure environments.
- Drive Infrastructure as Code (IaC) standards using Terraform, CloudFormation, or Bicep to ensure scalable, automated infrastructure provisioning.
- Build high-availability, low-latency architectures optimized for data-intensive Generative AI and Machine Learning workloads.
Security & Healthcare Compliance
- Enforce healthcare security standards including HIPAA, HITRUST, SOC 2, and data governance best practices across all cloud assets.
- Implement Zero-Trust security, Identity Access Management (IAM), data encryption key management, and continuous vulnerability monitoring.
Leadership & Team Management
- Manage, mentor, and scale a high-performing team of DevOps, Cloud, and SRE Engineers.
- Drive Agile workflows, sprint planning, incident response frameworks, and SLA compliance.
- Collaborate closely with Data Engineering, AI/ML, and Software Product teams to align infrastructure with business roadmaps.
Operations & FinOps
- Establish cloud cost optimization strategies (FinOps) to manage computing costs associated with AI models and large-scale data processing.
- Manage monitoring, alerting, and telemetry frameworks (e.g., Prometheus, Datadog, CloudWatch) to ensure 99.99% uptime.
Key Requirements
- Experience: 10+ years of overall IT experience with at least 5+ years in a cloud leadership or lead architect role.
- Cloud Platforms: Advanced hands-on expertise with both AWS (e.g., EC2, S3, EKS, Bedrock, SageMaker) and Azure (e.g., AKS, Azure OpenAI, Blob, Virtual Machines).
- DevOps & IaC: Strong background in Terraform, Docker, Kubernetes, CI/CD pipelines (GitHub Actions, GitLab CI, or Jenkins).
- Domain Knowledge: Prior experience building or managing cloud environments within Healthcare, Life Sciences, or HealthTech is strongly preferred.
- AI/ML Familiarity: Experience supporting cloud infrastructure for machine learning pipelines, LLM deployments, or GPU compute management.
- Certifications (Preferred): AWS Certified Solutions Architect – Professional, Azure Solutions Architect Expert, or Certified Kubernetes Administrator (CKA).
Greetings from zealous sevices!
GoLang (Soln Architect) -(9 AM to 6PM)
Experience: 10-12 Yrs
Need a senior Solution Architect - No BGV - ok with consultants
Own end-to-end solution architecture for a cloud-based cybersecurity platform operating at massive scale (millions of assets), spanning multiple product modules and a broad third-party integration ecosystem.
What They'll Do
Define and own the overall solution architecture across multiple product modules (e.g., detection, response, asset inventory, reporting), ensuring they work as a coherent, scalable system rather than disconnected parts
Design for scale: architecture that reliably handles millions of assets/endpoints with predictable performance, cost, and reliability as volume grows
Lead cloud infrastructure architecture decisions, compute, storage, networking, multi-region/multi-tenant strategy, disaster recovery
Architect and govern third-party integrations (SIEM/SOAR, ticketing, identity providers, cloud provider APIs, threat intel feeds, etc.), define integration patterns, data contracts, and reusable frameworks rather than one-off connections
Create and maintain architecture roadmaps, capacity plans, and technical strategy aligned with product/business roadmap
Set architectural standards and guardrails (scalability, security, data flow, API design) across engineering teams
Act as the technical bridge between engineering, product, security, and infrastructure/DevOps teams
Conduct architecture reviews, threat-model system designs, and identify scaling bottlenecks or single points of failure before they hurt production
Own key non-functional requirements: performance, availability, security, cost-efficiency, compliance
Support presales/enterprise customer conversations on architecture, especially for large or complex deployments
Evaluate and select core technologies (data stores, messaging systems, orchestration, IaC tooling)
What They Should Have
8–12+ years in software/infra engineering, with 4+ years specifically in a solution/enterprise architect role
Proven experience architecting systems at scale (millions of records/assets/events), not just designing on paper, but having operated such systems in production
Deep cloud architecture expertise (AWS/GCP/Azure), multi-account/multi-tenant design, cost optimization, networking, IAM
Strong grasp of distributed systems patterns: data partitioning, eventual consistency, event-driven architecture, caching strategies
Experience designing integration architectures (APIs, webhooks, message queues) across many heterogeneous third-party systems
Security architecture fluency,this is a security product, so the architect must think adversarially about their own system
Ability to produce architecture documentation (C4 diagrams, ADRs, data flow diagrams) that both engineers and executives can use
Track record of cross-functional influence, able to align multiple engineering teams around a shared architecture without direct authority
Cybersecurity domain knowledge (SIEM/EDR/CNAPP/CSPM concepts, threat detection pipelines), or equivalent adjacent domain with fast ramp-up ability
Nice to Have
Experience with specific integration ecosystems relevant to your product (e.g., Splunk, ServiceNow, Okta, AWS Security Hub)
Familiarity with IaC (Terraform), service mesh, and Kubernetes at scale Enterprise/regulated-industry experience (SOC 2, FedRAMP, ISO 27001) shaping architecture decisions
Prior experience scaling a product from ""hundreds of thousands"" to ""tens of millions"" of managed assets
Read This Before Anything Else
We have 6 developers who can ship. What we don't have is someone who turns that into a real engineering function: real architecture, real leverage, real AI-driven advantage. If that gap sounds like an opportunity rather than a headache, you're in the right place. If it sounds like a lot of undefined work with no playbook handed to you, this one probably isn't for you. That's completely okay. There are plenty of great roles that fit differently.
About CraftMyPlate
CraftMyPlate is Hyderabad's go-to platform for food experiences for micro-events: house parties, birthdays, office celebrations, festive gatherings, and more. We're building the operating system for how India discovers, customises, and orders food for smaller events. We're backed by established founders and investors, and we're funded and growing fast. The next phase of that growth runs through engineering.
Where We Stand
Some numbers, because they matter more than adjectives. Order volume has grown 50x in two years, and we're compounding at roughly 3x year over year, without giving up equity to fund it. That means the business runs on its own economics. The growth is real demand, not runway bought with dilution, and every efficient architectural decision this role makes directly protects that.
Most people size up an opportunity by asking what's going to change in ten years. The more useful question, and the one this company is built around, is what won't change. People will keep gathering. They'll keep celebrating, hosting, and marking festivals, in 10 years and in 20. That permanence is the bet. You're not building infrastructure for a trend cycle. You're building for a category that outlasts the current AI wave, the next funding round, and probably us too.
The Technical Reality
Here's an honest read of the engineering problem, not a sanitized version of it.
Event-driven commerce doesn't scale like typical e-commerce. Demand isn't smooth, it's spiky: weekends, festival calendars, and event dates create real load concentration, and each order is tied to a hard deadline that can't slip the way a shipped package can. That has direct architectural consequences: systems need to handle bursty, unpredictable traffic without paying for idle capacity the rest of the time, which is exactly why we're serverless-first on AWS rather than running a fixed fleet sized for peak.
Underneath that, every order touches multiple systems that have to stay consistent: kitchen and vendor fulfillment status, inventory across partners, payment gateway settlement, and refunds, often in real time and often across more than one vendor for a single event. Getting that consistency right across SQL and NoSQL stores, without it becoming a source of support tickets and manual reconciliation, is a real architecture problem, not a CRUD problem.
The AI-agent layer is the next lever, and it's a business lever as much as a technical one. Every workflow we can hand to a well-orchestrated agent instead of a new hire is a workflow that scales without adding headcount, which is exactly how a company grows 3x a year without diluting equity to fund the team behind it. That's why agent orchestration across multiple LLMs, using LangGraph, sits in the "go deep" tier of this role rather than being a nice-to-have.
You'll likely find some of this framing right and some of it worth challenging once you're actually in the codebase. That's expected, and honestly preferred over someone who just nods along.
Why This Role Exists
You'll be the most senior technical person in the company, reporting directly to the founder. Not a manager brought in to run standups. An owner. You set the architecture, you write code yourself, and you make the team materially better. You also own where AI and automation take this company next, starting with our first in-house AI agent product (details shared in the interview), and expanding from there into how the company runs, department by department: HR, finance, marketing, design, development, all sitting on an engineering layer that you design.
If you've outgrown a role where you plan but don't build, or where good ideas die in a committee, this is built to be the opposite of that.
What You'll Own
- Architecture, end to end. Scalable, cost-efficient systems from day one, not "fix it later" engineering. You own the decisions and their long-term consequences.
- Hands-on building. You are still writing code and shipping. This isn't a seat where you review other people's work all day. You lead by building.
- The engineering team. Directly manage, mentor, and level up our 6 developers. Build the technical bar, the review culture, and the calibration that lets the team ship independently.
- The AI-agent roadmap. Own the architecture behind our first AI agent product, then the broader strategy for AI agents and automation across every function in the company, with engineering as the layer underneath all of it.
- Team scaling. Build the next layer of leads under you so execution quality scales without you being the bottleneck.
- Technical accountability. When something breaks, you fix it. You don't escalate and wait.
Our Stack, and the Depth We Expect
Not everything on this list needs the same level of mastery. Some of it you need to own at an architectural level. The rest you need to be strong enough to build yourself, direct the team on, or delegate to AI agents with confidence.
Go deep here. This is where the real architecture decisions live, and where the business impact is highest:
- AWS, serverless first. You should be genuinely well versed in AWS application development, not just "have used AWS." You should be able to design and guide serverless architecture (Lambda, API Gateway, DynamoDB, Step Functions, and similar) as our default way of building, because our demand curve is spiky by nature and fixed infrastructure is money left on the table.
- TypeScript, our primary language across backend and frontend.
- Agent orchestration across multiple LLMs, using LangGraph. This is core to our AI roadmap and our path to scaling operations without scaling headcount. You own how it's architected, not just how it's used.
Working proficiency. Build it yourself, direct the team, or hand it to an AI agent and know if the output is right.
This Is You If
- You've built and shipped real production systems yourself, not just reviewed other people's architecture from a distance.
- You go deep wherever the problem is, and you're comfortable owning the exact stack described above, not just "full-stack" in the abstract.
- You've made engineers around you measurably better, whether or not you've held the title for it yet.
- You're already using AI coding tools and agents seriously, like Claude, Cursor, or similar tools, as part of how you build, not as something you tried once. We'll likely explore this together in the interview.
- You have a bias toward leverage over hours. You'd rather automate or systematize a problem than grind through it. But when something's live and needs to be done right, you see it through completely, with no half-finished work.
- You want to build something for years, not land somewhere comfortable. We'll know the difference from how you talk about your last three years.
This Might Not Be the Right Fit If
- You'd prefer a stable, well-defined role with clear boundaries and someone else making the calls. That's a fair thing to want, just not what this is.
- You'd rather receive direction than bring us architecture and AI strategy yourself.
- You haven't yet gotten hands-on with AI coding tools in your daily work.
- You're drawn more to the title than the work behind it.
If none of that sounds like you, we'd love to hear from you.
Requirements
- 5 to 7 years of experience in software engineering, with real ownership of architecture-level decisions, not just feature delivery.
- Prior experience leading or mentoring engineers, formally or informally.
- Tier-1 or Tier-1+ engineering college strongly preferred (IIT, BITS, top NIT tier, or equivalent). We'll consider other institutions only with clearly commendable, verifiable work: real systems you can walk us through in depth, strong open-source contributions, or a track record that speaks for itself. Pedigree is a proxy for speed, not a checkbox. We test for the underlying ability regardless.
- Comfortable in an early-stage environment: undefined problems, few processes, and the expectation that you help define both.
Compensation
Competitive, with equity. We're formalizing a structured ESOP program alongside this hire. Specific numbers are discussed directly in later interview rounds.
If reading this got you a little excited about what you'd build here, we'd genuinely love to talk. If it didn't quite land, no hard feelings. We just want the right fit for both sides.
We're hiring a Cloud Architect (Contract) to work with our Equity Partners who builds profitable growth by acquiring and operating enterprise software companies. Refining a proprietary operating model across 40+ acquisitions and two decades of hands-on experience, now supercharged by our patented agentic AI platform . In this role, you'll take full architectural control of our CI/CD, observability, and event streaming infrastructure, build the standards every new acquisition plugs into, and use AI-assisted automation to keep 20+ products reliable without proportionally scaling headcount.
Job title: Cloud/Platform Architect (SRE)
Type: Global Remote | Contract
What You Bring
- 8–12 years in platform engineering, DevOps, or SRE, with growing ownership over time
- Deep Terraform experience across multi-account, multi-env setups
- Real production experience with event streaming at scale
- Hands-on Grafana, Prometheus, Loki, and strong AWS depth (ECS, EKS, IAM, VPC, RDS)
- SRE fundamentals: SLOs, error budgets, on-call design, post-mortems
- Bonus: acquisition or greenfield platform-building experience
Roles and Responsibilities
- Own everything outside core AWS infra: CI/CD, observability, event streaming, deployment, incidents
- Define the standards every future acquisition will plug into
- Keep 20+ enterprise products running at serious scale (millions–billions of requests)
- Build self-service tooling so product teams never wait on you
- Use AI/automation to kill toil — not to replace engineering judgement
Ready to build the platform that scales an entire portfolio? — let's connect.
Technical Architect – Product Engineering
Experience: 15+ Years
Location: Pune, India
Employment Type: Full-time
Desired Skills: Python, Technical Architecture, AWS, Microservices, SaaS / Multi-tenant Architecture, Kubernetes, System Design
About the Role
We are looking for a Senior Technical Architect to lead the architecture, design, and technical evolution of an enterprise SaaS product. This is a hands-on leadership role requiring deep technical expertise, strong product engineering experience, and the ability to build scalable, secure, and high-performance platforms.
The ideal candidate should be passionate about solving complex engineering problems, driving innovation, mentoring development teams, and effectively leveraging AI to accelerate software development.
Key Responsibilities
- Own the overall product architecture and technical roadmap.
- Design and build scalable, secure, and highly available enterprise applications.
- Lead the design and implementation of new product features from concept to production.
- Remain hands-on with coding and contribute to critical product components.
- Drive architecture reviews, code quality, performance optimization, and engineering best practices.
- Lead cloud architecture, security, scalability, and DevOps initiatives.
- Evaluate and adopt modern technologies to improve product capabilities and engineering efficiency.
- Leverage AI tools (ChatGPT, GitHub Copilot, Cursor, Claude, etc.) to accelerate software development, code reviews, testing, documentation, debugging, and productivity.
Required Skills & Qualifications
- 15+ years of software product engineering experience with at least 5 years in a Technical Architect role.
- Strong hands-on expertise in Python and modern backend frameworks.
- Deep experience with AWS services and cloud-native application architecture.
- Strong understanding of DevOps, CI/CD pipelines, Infrastructure as Code (Terraform/CloudFormation), Docker, Kubernetes, and container orchestration.
- Experience designing microservices, REST APIs, event-driven architectures, and distributed systems.
- Strong knowledge of SQL and NoSQL databases.
- Experience with scalable SaaS platforms, multi-tenant architectures, and secure application design.
- Excellent understanding of software design patterns, performance tuning, observability, and system reliability.
- Strong analytical, problem-solving, and decision-making skills.
Platform Engineering Lead (For client company)
Location: Pune, India
Experience: 7+ years
What Success Looks Like
- Engineering teams ship faster with confidence and built-in guardrails.
- Cloud cost, security, and reliability are predictable, measurable, and well-managed.
- CI/CD pipelines are trusted, standardized, and production-ready.
- Platform decisions reduce cognitive load instead of introducing unnecessary process.
Scope & Expectations
This is a hands-on leadership role combining architecture and implementation.
You will:
- Build, not just review.
- Own the platform roadmap—not just infrastructure tickets.
- Act as a force multiplier for product engineering teams rather than becoming a bottleneck.
- Drive platform strategy while remaining deeply involved in execution.
Key Responsibilities
Platform & Cloud Architecture
- Own Zoop's platform and cloud architecture across GCP and AWS.
- Design reusable, opinionated platform patterns instead of one-off infrastructure.
- Build and evolve Zoop's Internal Developer Platform (IDP), including:
- Self-service environments
- Golden paths (paved roads)
- Standardized templates
- Built-in engineering guardrails
- Lead Kubernetes and cloud-native adoption at scale.
- Drive infrastructure automation using Terraform, Pulumi, or similar Infrastructure-as-Code (IaC) tools.
CI/CD, Reliability & Developer Experience
- Establish robust CI/CD practices with quality gates and production readiness.
- Improve deployment safety through automation and testing.
- Define and monitor:
- Golden Signals
- SLIs
- SLOs
- Incident response processes
- Reduce operational toil and improve developer productivity.
- Make observability a first-class capability using cost-efficient monitoring systems.
- Build an observability platform that multiple engineering teams can easily integrate into their applications.
Security, Privacy & Compliance
- Build security-by-default into infrastructure and deployment pipelines.
- Lead implementation and continuous compliance for:
- DPDP Act (India)
- ISO 27001:2022
- SOC 2 Type II
- Implement:
- Zero Trust architecture
- Least-privilege access
- Secure data isolation
FinOps & Cloud Optimization
- Make cloud costs transparent and accountable across engineering teams.
- Establish FinOps practices including:
- Budgets
- Cost alerts
- Optimization routines
- Drive build-vs-buy decisions using clear ROI analysis.
AI, Data & MLOps Foundations
- Build secure and scalable foundations for AI and MLOps workloads.
- Define guardrails for AI systems and sensitive data handling.
Leadership & Collaboration
- Partner closely with engineering teams to align infrastructure strategy with product goals.
- Mentor engineers and guide teams through technical change.
- Balance long-term platform initiatives with practical execution.
What We're Looking For
Experience
- 7+ years of experience building and operating production infrastructure.
- Experience scaling engineering platforms in high-growth or regulated companies.
- Strong hands-on expertise in:
- Kubernetes and the cloud-native ecosystem
- Service Mesh technologies
- Policy Engines
- GCP, AWS (Azure exposure is a plus)
- Terraform and Infrastructure as Code
Engineering & Operations
- Strong understanding of SDLC and modern CI/CD systems (Jenkins, GitOps, etc.).
- Experience with observability tools such as:
- Grafana
- Prometheus
- New Relic
- Comfortable reading and contributing to production systems written in:
- Go
- Python
- Node.js
Security & Compliance
- Practical experience implementing ISO 27001 and SOC 2 controls.
- Strong understanding of:
- Data protection
- Privacy
- Identity and access management
- Security best practices
Mindset
We're looking for someone who is:
- Action-oriented with sound engineering judgment.
- Analytical, cost-conscious, and reliability-focused.
- Collaborative, calm under pressure, and open to feedback.
- Comfortable challenging decisions and explaining trade-offs when necessary.
Nice to Have
- Experience in fintech, identity, or other regulated industries.
- Built Internal Developer Platforms (IDPs) or shared infrastructure tooling.
- Contributions to open-source projects.
We’re on hunt for AI Architect
Responsibilities:
- 10–15+ years overall experience, with recent hands-on AI/GenAI architecture ownership.
- Must have architected enterprise AI platforms/solutions end-to-end, not just individual ML models or PoCs.
- Strong GenAI/LLM production experience: RAG, embeddings, vector DBs, hybrid search, reranking, evaluation, guardrails.
- Strong Agentic AI understanding: agents, tool calling, workflows, orchestration, human-in-the-loop.
- Experience taking AI solutions from architecture → production → scale, ideally across multiple business teams/use cases.
- Strong cloud architecture — Azure/AWS preferred; hybrid/on-prem experience is a plus.
- Must understand enterprise security, governance, Responsible AI, observability and LLMOps/MLOps.
- Should be able to articulate build-vs-buy, MVP-vs-target architecture, cost/performance/security tradeoffs.
- Strong stakeholder-facing / consulting ability — can work with business leaders, engineering, security and data teams and influence without authority.
There is scope to move to the US for this role if you are aligned for the same, else this will be a WFO role from Hyderabad location
About the role
You will be a pivotal member of the leadership team, responsible for driving the company's product and platform strategy while overseeing engineering excellence across AI, data, platform, DevOps and infrastructure. This is a mission-critical role for an engineering leader who is both strategic and hands-on, and who can build scalable technologies for global enterprise clients in a fast-paced innovation ecosystem.
Reports to: CEO · Location: Bengaluru, India — hybrid, 3 days a week in office
What you will do
Technology & product leadership
- Define and drive the technology vision, architecture and long-term platform roadmap.
- Oversee the architecture, design and delivery of highly scalable enterprise systems.
- Ensure engineering excellence, velocity and reliability across the product lifecycle.
Engineering & platform management
- Lead platform engineering, product technology, DevOps, infrastructure and Quality Engineering.
- Build robust cloud-native systems using the Azure, AWS and GCP ecosystems.
- Oversee operational effectiveness, including uptime, production reliability and cost optimisation.
Innovation & AI strategy
- Spearhead Mission AI by developing scalable, production-grade AI/ML and GenAI capabilities.
- Own the GenAI/LLM solutions architecture.
- Core specialisation in AI agents and autonomous workflows, data extraction and intelligent automation.
- Direct hands-on architectural oversight of large language models, applied AI and multi-layered deep learning products.
- Extensive expertise in building intelligent document automation systems — similar in complexity to patent parsing, legal workflow automation and structured decision intelligence tools.
Technical leadership
- A proven track record overseeing product architecture, tech strategy and cross-functional engineering execution across core full-stack and AI platforms.
- Collaborate with executive leadership on business strategy, client requirements and product delivery.
- Build, mentor and scale high-performing engineering teams with a growth mindset.
- Establish a strong technology culture grounded in ownership, innovation and continuous learning.
What success looks like
- Mission AI: build a world-class AI system leveraging GenAI, ML and enterprise-grade data engineering.
- Hyper-scaling: architect and scale the platform to match global industry leaders in the IP space.
- Culture building: develop a strong engineering organisation with high ownership, performance and innovation DNA.
Qualifications & experience
- Preferably under 15 years of enterprise software engineering experience across B2B SaaS or technology-first companies; fewer is fine for an exceptional engineer.
- A proven track record of taking early-stage AI/ML prototypes and scaling them into robust enterprise SaaS platforms featuring multi-agent orchestration, complex workflows and decision intelligence.
- Experienced in building and mentoring agile, lean startup teams of full-stack and machine learning engineers from the ground up.
- Proven leadership in defining and executing technology strategy and platform roadmaps.
- Extensive cloud-native engineering experience with Azure, AWS and GCP.
Technical expertise
- Strong full-stack engineering background (Java, Python, JavaScript frameworks).
- Expertise with JS frameworks such as React, Angular and Node.js.
- Experience building and scaling distributed systems and microservices.
- Strong knowledge of databases (SQL, NoSQL), data modelling and unstructured data management.
- Strong understanding of Agile methodologies and tools (Atlassian, Git, CI/CD pipelines).
Behavioural & leadership competencies
- Product and delivery management expertise, end to end, including delivery and customer support.
- Excellent communication, with the ability to influence executive stakeholders.
- High technical proficiency combined with strong business acumen.
- Strong analytical and decision-making skills.
Job Description:
- Infrastructure Management: Design, implement, and manage scalable, reliable, and secure cloud infrastructure using AWS, GCP, and/or Azure.
- CI/CD Pipelines: Develop and maintain continuous integration and continuous deployment (CI/CD) pipelines to streamline the development lifecycle.
- Automation: Automate infrastructure provisioning, configuration management, and application deployment processes.
- Monitoring and Performance: Implement monitoring, logging, and alerting solutions to ensure system health, performance, and reliability.
- Security: Ensure the security of cloud infrastructure and applications, including identity management and compliance with industry standards.
- Collaboration: Work closely with client and development teams to integrate DevOps practices and deliver high-quality software.
- Documentation: Maintain comprehensive documentation of infrastructure, configurations, and processes.
- Innovation: Stay current with emerging technologies and industry trends, integrating them into the DevOps strategy as appropriate.
Qualifications:
- Education: Bachelor's degree in Computer Science, Information Technology, or a related field.
- Experience: 7 - 10 years of overall experience with relevant experience of at least 7 years in DevOps and served as a lead or senior engineer.
Senior Cloud Site Reliability Engineer (CSRE) – Azure
About Searce:
Searce is an AI-native, engineering-led modern technology consultancy that empowers
clients to futurify their businesses by delivering real, intelligent business outcomes. As a
trusted partner for over 3,000 clients globally, Searce specializes in cloud modernization,
data engineering, applied AI, and robust cloud platform security. Driven by a "HAPPIER"
cultural mindset and our proprietary evlos problem-solving framework, we eliminate
bureaucratic fluff to build working prototypes fast and scale enterprise production
environments intelligently. We don't just fix systems; we leverage multi-cloud technologies
to transform client operations into distinct competitive advantages.
Position Overview:
We are looking for a high-caliber Senior or Lead Cloud Site Reliability Engineer (CSRE) to
architect, secure, and stabilize next-generation hybrid and multi-cloud environments.
Operating at the intersection of infrastructure design, security compliance, and production
operations, you will serve as the technical Subject Matter Expert (SME) across GCP, Azure,
and AWS.
Whether optimizing a microservice mesh on GKE, tuning autoscaling on AKS, or driving a
massive disaster recovery drill across AWS regions, your focus will be absolute reliability. For
the Lead path, you will couple this deep engineering toolkit with stakeholder management
and mentorship to drive an elite operational culture.
Experience & Level Expectation:
Years of Experience: 3 to 10 years of intensive, hands-on production operations
experience in a dedicated DevOps, Cloud Platform Engineering, or SRE role.
Associate level (3-5 Years): Expected to show flawless execution of IaC, advanced
triaging of infrastructure failures, and ownership of the CI/CD and deployment
lifecycles.
Intermediate level (5-10 Years): Expected to take architectural ownership, serve as
primary Incident Commander for complex outages, design cross-cloud governance
frameworks, and act as a reliable bridge between technical teams and client
leadership.
Key Responsibilities & Role Expectations:
Multi-Cloud Platforms & Orchestration: Design, configure, and maintain
production-grade Kubernetes clusters across major platforms (AKS).
Manage advanced network routing, service meshes (e.g., Istio), and multi-tenant
isolation.
Infrastructure as Code (IaC) & GitOps: Build declarative, enterprise-grade, reusable
infrastructure components using Terraform or Crossplane. Standardize automated
environment provisioning to eliminate configuration drift across multi-branch
environments.
Incident Management & Reliability (SRE): Own and optimize the production on-call
rotation. Lead rapid mitigation strategies for Sev-1/Sev-2 system outages, reducing
Mean Time to Recovery (MTTR) through centralized log and metric correlation.
Root Cause Analysis (RCA): Facilitate rigorous, blameless post-incident reviews to
identify core architectural vulnerabilities and establish long-term fixes preventing
recurrence.
Lifecycle, Patching & Upgrades: Plan and execute zero-downtime cluster upgrades,
operating system patching strategies (Linux/Windows), database lifecycle updates,
and multi-region Disaster Recovery (DR) failover drills.
Core Core Operations & Legacy Integration: Manage enterprise-level hybrid
networking architecture (VPCs, Firewalls, Load Balancers, DNS routing, and DHCP
configurations) while effectively connecting cloud native services to legacy
infrastructures like Active Directory.
Security & Governance: Embed Zero Trust policies, secure secrets management
(Secrets Manager/Key Vault), and continuous vulnerability patching into the
automated SDLC pipeline.
Required Technical Skills:
- Microsoft Azure: Azure Virtual Machines, Virtual Networks, Azure Active Directory, Azure Update Management.
- Containers & Orchestration
- Production-level management of GKE, AKS, and EKS.
- Advanced mastery of Docker, Helm, Kubernetes StatefulSets, Pod Disruption







