Staff Software Engineer, AI Agents at Asha Health (YC F24) · Bengaluru (Bangalore) · 3 - 7 years · ₹100L - ₹100L / yr · Raised funding · Posted 21 Jun 2026

About Asha Health
Asha Health helps medical practices launch their own AI clinics. We're backed by Y Combinator, General Catalyst, 186 Ventures, Reach Capital and many more. We recently raised an oversubscribed seed round from some of the best investors in Silicon Valley. Our team includes AI product leaders from companies like Google, physician executives from major health systems, and more.
About the Role
We're looking for a top 0.01% Software Engineer to join our engineering team in our Bangalore office.
4.6 fundamentally changed the game, which means that high intelligence, high agency engineers can now do the work of 10+ good engineers. It doesn't make sense to have anyone but the best on the team.
Since low level coding has become easier, what we expect from engineers on our team has expanded. Engineers on our team are expected to:
- Ship features end to end at a rapid pace
- Deeply research the domain and be their own product managers
- Ensure reliability and quality is best-in-class
- Design robust eng architecture, and develop testing and observability tools for each feature pre-launch
- Build each feature with deep customer empathy, meaning planning out and building stellar UX yourself
This means to thrive in a startup environment like ours, you not only need to be a stellar engineer, but you need to:
- Be super adept with AI development tools and building the AI systems that build your features for you (Conductor, Browser agents, QA agents, Ralph loops, adverserial agents, and more).
- Have exceptional product and UX taste, meaning you can ship features that are more effective than those historically designed by teams of product managers and designers.
- Take the highest level of ownership around feature outcomes, reliability, and observability.
Other Points to Note
- We are growing rapidly, our work has impact on tens of thousands of patients if not more.
- On our team, everything you do is on the bleeding edge of applied AI.
- We expect a high level of commitment from everyone on the team, most folks work 6 days and lead every project with intensity. It's a high ask, and we only bring on the best people. We compensate significantly above market, accordingly.

About Asha Health (YC F24)
About
Asha Health is a Y Combinator backed AI healthcare startup. We help medical practices spin up their own AI clinic. We've raised an oversubscribed seed round backed by top Silicon Valley investors, and are growing rapidly. Our team consists of AI product experts from companies like Google, as well as senior physician executives from major health systems.
Tech stack
Candid answers by the company
We help medical practices spin up their own AI clinic.
Similar jobs (10)
About Us
We believe the future of software development is AI-native — where engineers operate at a higher level of abstraction and quality remains non-negotiable.
Incubyte is a software craft consultancy where the “how” of building software matters as much as the “what”.
We partner with companies of all sizes, from helping enterprises build, scale, and modernize to early-stage founders bring their ideas to life.
Our engineers operate in an AI-native development model, using AI as a collaborator across the SDLC to accelerate development while upholding the discipline of software craftsmanship. Guided by Software Craftsmanship and Extreme Programming practices, we build reliable, maintainable, and scalable systems with speed, without compromising quality. If this way of building software resonates with you, we’d like to talk.
Our Guiding Principles
These principles define how we work at Incubyte. They are non-negotiable.
Relentless Pursuit of Quality with Pragmatism
We build high-quality systems without losing sight of delivery.
Extreme Ownership
We take responsibility end-to-end for decisions, execution, and outcomes.
Proactive Collaboration
We collaborate closely, challenge each other, and solve problems together.
Active Pursuit of Mastery
We continuously improve our craft and raise our bar.
Invite, Give, and Act on Feedback
We seek, give, and act on feedback to get better every day.
Ensuring Client Success
We act as trusted partners and focus on real outcomes, not just output.
Experience Level
This role is ideal for engineers with total 3+ years of experience with a proven track record of shipping complex projects successfully.
An experienced individual contributor and leader who thrives in large, complex projects with widespread impact.
What You’ll Do as a Software Craftsperson
- Design and build high-quality, maintainable systems using disciplined engineering practices such as TDD, continuous refactoring, and pair programming
- Operate in an AI-native development model, using AI as a collaborator to explore architecture and design, accelerate development, and continuously improve systems while applying strong judgment to ensure that speed never compromises quality
- Take end-to-end ownership of outcomes from problem understanding and system design to implementation, deployment, and operation in production
- Make thoughtful design decisions that balance simplicity, scalability, and long-term maintainability in real-world systems
- Maintain a high bar for engineering quality through rigorous testing, code reviews, and continuous feedback
- Investigate and resolve production issues, and implement systemic improvements to prevent recurrence
- Work directly with clients, navigate ambiguity, and translate business problems into well-designed technical solutions
- Contribute to improving team practices, tooling, and systems to raise the overall quality and effectiveness of engineering
Requirements
What You’ll Bring
- 3+ years of experience building high-quality, production systems (flexible based on demonstrated capability)
- Strong fundamentals in software engineering, including object-oriented design, system design, and testing practices such as TDD
- Demonstrated ability to build simple, maintainable, and scalable systems with a focus on long-term reliability
- Proficiency in one or more modern technologies, Python, PHP, JavaScript, or TypeScript, with the ability to learn new technologies quickly
- Deep experience working with Git in collaborative environments, including managing shared codebases, conducting code reviews, and maintaining a high bar for quality
- Ability to operate effectively in an AI-native workflow using AI as a collaborator to explore solutions and accelerate development, while applying strong judgment to ensure correctness, quality, and maintainability
- Clear thinking and strong problem-solving ability, with the capacity to break down complex problems into simple, well-structured solutions
- A strong sense of ownership — you take responsibility for outcomes, care deeply about quality, and are not comfortable shipping work that does not meet your standards.
Benefits
Life at Incubyte
We are a remote-first company with structured flexibility. Teams commit to shared rhythms during core hours, ensuring smooth collaboration while maintaining autonomy. Twice a year, we come together in person for a co-working sprint and once a year for a retreat - with all travel expenses covered.
Our environment is built for crafters: pairing, refactoring, experimenting with AI, and pushing the boundaries of software excellence. We are all lifelong learners, and our work is our passion.
Perks
- Dedicated learning & development budget.
- Sponsorship for conference talks.
- Comprehensive medical & term insurance.
- Employee-friendly leave policies.
- Home Office fund
- Medical Insurance
AI Engineer
LLMs, Agents & AI Services
📍 Mumbai (On-site) | Full-time | 2-4 years
About the Role:
Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.
AI is core to how we design, deliver, and scale software for our customers.
We are hiring an AI Engineer for a dedicated client engagement building a complex production AI platform, working on the AI capabilities and agentic features at the core of the product.
The mandatory requirement for this role is at least one AI feature personally shipped to production for real users, with operational ownership.
The role suits someone who thinks quickly on solutioning, can take an ambiguous problem to a working prototype in days, and has the discipline to carry it through to production with predictable economics.
You will work alongside the Senior AI Engineer and the wider pod, with ownership of parts of the AI surface area of the product.
Responsibilities:
Solutioning and POCs
Translate ambiguous customer problems into working POCs at speed.
Pick the right model, framework, and architecture, and demonstrate value early before scaling investment.
LLM Application Development
Build AI features and services using LLM APIs from OpenAI, Anthropic, Google, and self-hosted open-weight models (Llama, Qwen, Mistral).
Choose the right model per use case based on cost, latency, capability, and context-window trade-offs.
Agentic System Design
Design and implement agentic workflows using LangGraph, CrewAI, AutoGen, LlamaIndex Agents, or custom orchestration.
Cover tool use, planning, memory, and multi-step reasoning appropriate to the problem.
API and Service Development
Build production AI services and APIs using Python and FastAPI.
Handle streaming responses, async processing, structured outputs, retries, and graceful degradation when models or tools fail.
Retrieval and Tool Integration
Implement RAG pipelines with vector databases (Pinecone, Weaviate, Qdrant, pgvector, Chroma), embeddings, chunking strategies, hybrid search, and reranking.
Integrate external tools, internal APIs, and document sources through tool-calling and MCP-style patterns.
Cost Analysis and Unit Economics
Model the per-request and per-user cost of every AI feature before it ships.
Track token usage, prompt caching, batching, and model-routing strategies.
Drive measurable improvements in unit economics.
Production Hardening
Add observability and tracing (LangSmith, Langfuse, OpenTelemetry), guardrails, content safety checks, prompt injection defences, and fallback behaviour.
Prompt Engineering and Evaluation
Design, test, and iterate prompts with measured outcomes.
Build evaluation harnesses for accuracy, hallucination, latency, and cost.
Run benchmarks across models and prompt variants before locking in a design.
Requirements:
AI Feature Shipped to Production (Mandatory)
Must have personally built and shipped at least one AI feature that runs in production for real users, with operational ownership.
POCs, internal demos, and one-off scripts do not qualify.
2 to 4 Years of Professional Software or AI Engineering Experience
With at least one production AI feature owned end to end.
Strong Python Proficiency and API Development with FastAPI
Comfort with type hints, async, packaging, testing, streaming responses, and authentication.
Production-grade Python, not notebook-only code.
Hands-on Depth Across the LLM and Agent Stack
Working experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or self-hosted open-weight models (vLLM, Ollama, Together, Replicate).
Working familiarity with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.
Working knowledge of RAG, embeddings, and vector databases (Pinecone, Weaviate, Qdrant, pgvector, Chroma).
Solutioning Speed and POC Velocity
Demonstrated ability to move from a fuzzy problem to a working prototype in days.
Strong instinct for what to build first, what to defer, and what to throw away.
Cost Discipline for Production AI
Ability to calculate, monitor, and optimise the cost of LLM APIs, tokens, embeddings, vector store usage, and infrastructure.
Treats unit economics as a first-class concern.
AWS Familiarity
Working knowledge of EC2, S3, IAM, and at least one of Bedrock, SageMaker, or equivalent.
Comfortable in a Fast-Moving Environment
Self-directed, comfortable with ambiguity, takes ownership without being asked, and ships under shifting priorities.
Strong Written and Spoken English Communication
Able to explain trade-offs to non-AI engineers, designers, product managers, and clients in plain language.
Nice to Have
- fine-tuning or LoRA, QLoRA, PEFT exposure
- MCP server authoring
- eval framework experience (LangSmith, Promptfoo, Ragas, DeepEval)
- open-source AI contributions
- multi-modal models (vision, audio)
About Us
We believe the future of software development is AI-native — where engineers operate at a higher level of abstraction and quality remains non-negotiable.
Incubyte is a software craft consultancy where the “how” of building software matters as much as the “what”.
We partner with companies of all sizes, from helping enterprises build, scale, and modernize to early-stage founders bring their ideas to life.
Our engineers operate in an AI-native development model, using AI as a collaborator across the SDLC to accelerate development while upholding the discipline of software craftsmanship. Guided by Software Craftsmanship and Extreme Programming practices, we build reliable, maintainable, and scalable systems with speed, without compromising quality. If this way of building software resonates with you, we’d like to talk.
Our Guiding Principles
These principles define how we work at Incubyte. They are non-negotiable.
Relentless Pursuit of Quality with Pragmatism
We build high-quality systems without losing sight of delivery.
Extreme Ownership
We take responsibility end-to-end for decisions, execution, and outcomes.
Proactive Collaboration
We collaborate closely, challenge each other, and solve problems together.
Active Pursuit of Mastery
We continuously improve our craft and raise our bar.
Invite, Give, and Act on Feedback
We seek, give, and act on feedback to get better every day.
Ensuring Client Success
We act as trusted partners and focus on real outcomes, not just output.
Job Description
This is a remote position.
This is a remote position.
Experience Level
This role is ideal for engineers with 3–6 years of experience and a strong background in building scalable, production-grade software systems.
We are looking for hands-on Software Engineers with deep expertise in Python and TypeScript, with exposure to AI/LLM systems and modern infrastructure tooling.
What You’ll Do as a Software Craftsperson
• Take full ownership of the software development lifecycle for complex, cross-functional initiatives — from design through production readiness.
• Build and maintain robust, scalable backend systems and APIs using Python and TypeScript, following clean code and software craftsmanship principles.
• Design and deliver features end-to-end, balancing scope, quality, and long-term maintainability.
• Identify technical and product issues beyond the immediate scope of work, proactively raising risks and driving solutions.
• Shape team practices around code quality, testing, tooling, and continuous improvement using DevEx and DORA principles.
• Collaborate closely with clients and internal teams to understand requirements, clarify priorities, and align on outcomes.
• Mentor and guide fellow engineers to raise overall team performance and promote a culture of learning.
• Leverage AI tools (LLMs, agentic frameworks, etc.) to accelerate design, development, testing, and delivery where applicable.
Requirements
What You’ll Bring
3–6 years of overall software engineering experience with a strong track record of owning and delivering complex production systems.
Must-Have Skills
• Python (must-have): Deep expertise in writing idiomatic, testable, production-grade Python — including advanced OOP, data structures, algorithms, and software engineering best practices.
• TypeScript (must-have): Strong proficiency in building and maintaining type-safe, scalable applications across frontend and/or backend TypeScript codebases.
• Strong system design skills: ability to architect scalable, maintainable, and observable systems with a focus on reliability and long-term operability.
• Solid engineering practices: experience with TDD, CI/CD, code reviews, refactoring, and continuous deployment in Agile or eXtreme Programming environments.
• Working knowledge of relational databases, web server ecosystems, REST/gRPC APIs, and performance optimisation.
• Experience with source control, bug tracking, user story writing, and maintaining clear technical documentation.
Good-to-Have Skills
• AI / LLM experience: hands-on exposure to building applications using LLMs (GPT, Claude, Gemini, or similar) for tasks such as document classification, entity extraction, or structured data generation.
• LLM orchestration frameworks: familiarity with LangChain, LlamaIndex, Mastra, Agno, or similar agentic frameworks.
• Prompt engineering: experience refining prompts and orchestration patterns to improve response accuracy, consistency, and structured outputs.
• Vector databases and observability tooling: exposure to tools like Pinecone, PgVector, Qdrant, LangSmith, or DeepEval.
• Container orchestration and infrastructure-as-code: experience with Kubernetes, Terraform, and Docker for deploying and managing production workloads.
Benefits
Life at Incubyte
We are a remote-first company with structured flexibility. Teams commit to shared rhythms during core hours, ensuring smooth collaboration while maintaining autonomy. Twice a year, we come together in person for a co-working sprint and once a year for a retreat — with all travel expenses covered.
Our environment is built for crafters: experimenting with real-world systems, solving complex infrastructure challenges, and contributing to cutting-edge AI initiatives. We are all lifelong learners, and our work is our passion.
Perks
• Dedicated learning & development budget
• Sponsorship for conference talks
• Comprehensive medical & term insurance
• Employee-friendly leave policies
• Home Office fund
• Medical Insurance
We are looking for an Engineering Lead to own the entire technology stack — from onboarding and underwriting to disbursals, repayments, and collections — and to build the engineering function into something genuinely AI-native.
What You'll Own
● Full tech stack: backend, frontend, infrastructure, integrations, and data pipelines
● Real-time underwriting and decisioning systems
● LOS/LMS architecture — onboarding, disbursals, repayments, and collections
● Integrations with bureaus, KYC providers, account aggregators, and payment gateways
● Reconciliation systems — disbursement, repayment, and NACH reconciliation end-to-end
● AWS infrastructure: scaling, reliability, uptime, and cloud cost ownership ● Data infrastructure for the credit and risk team — feature pipelines, model serving, experiment infrastructure
● Engineering leadership: hiring, sprint planning, code reviews, and execution standards
● Compliance systems: RBI guidelines, DPDP, KYC/AML, e-NACH, e-sign
AI-Native Engineering
This is a core part of the role, not a bonus. You will build a machine-readable knowledge base of the entire codebase — architecture, data models, service contracts, coding standards, decision history — so that AI agents working on code have the context to produce accurate, consistent output. You will build skills for code review, developer onboarding, and recurring engineering workflows. You will build a code review pipeline where agents do the first pass on every pull request. The knowledge base and the skills improve over time as the team grows and the product evolves.
What We're Looking For
● 7+ years in software engineering, with at least 2 years leading teams or architecture
● Strong hands-on experience with Python, Django, and React Native
● Deep expertise in AWS and cloud-native architecture
● Experience with both SQL and NoSQL databases
● Strong understanding of distributed systems, microservices, and API design
● Experience owning reconciliation or payment flow infrastructure in a lending or payments context
● Prior experience in fintech / NBFC / digital lending — mandatory
● Strong understanding of the full loan lifecycle — mandatory
● You have used LLMs seriously as engineering tools and have strong opinions about what makes AI-assisted development produce good output versus mediocre output
Bonus: Kubernetes / Kafka, AI/ML-driven underwriting, Account Aggregator framework, e-NACH / e-Sign / Video KYC integrations
What Success Looks Like
● scales with strong uptime, performance, and reliability
● Reconciliation runs cleanly — no financial discrepancies surface late ● A new engineer joins and is writing standard, correct code within their first week
● The credit team is never blocked on an engineering dependency
● Engineering health metrics are tracked and visibly improving
● AI agents are doing the structured first pass on code reviews, and the system gets smarter over time
The Role
You own AI systems end to end. From the speech-to-text models that turn audio into text, to the diarization that separates and identifies speakers, to the agentic layer that turns conversation into memory and action, to the observability and evaluation that keep all of it honest in production. This is a wide role by design. You will own model selection, serving, and production reliability. If you want to tune one model and ignore the system around it, this is not the role.
What You Will Own
• Speech-to-text. Evaluate, integrate, and optimize STT models across cloud and self-hosted. Drive accuracy and cost trade-offs with ground-truth metrics.
• Speaker diarization and identification. Push accuracy on hard, real-world, multi-speaker audio.
• Agentic AI. Build the memory and retrieval pipeline, LLM orchestration, and the agent workflows that sit on top of captured conversation.
• Model serving and infrastructure. Stand up and optimize self-hosted serving (vLLM, Triton class). Own latency, throughput, and cost per user.
Observability
An always-on wearable means models run in production every second, on messy real-world audio. You own the visibility into that.
• Instrument the full audio-to-memory pipeline: STT, diarization, retrieval, and LLM calls.
• Define and track model-quality SLOs in production: transcription drift, diarization error over time, retrieval relevance, latency, throughput, and cost per user.
• Build dashboards and alerting so model degradation is caught before users feel it.
• Trace failures across a distributed, always-on system using metrics, logs, and traces.
• Close the loop. Production signals feed back into evaluation and model selection.
Evaluation
We do not ship what we cannot measure. You own the systems that prove a model is actually better, not just newer.
• Build and own ground-truth evaluation harnesses for every model in the stack.
• Measure with real metrics: WER for transcription, DER for diarization, Recall and F1 for retrieval and speaker identification.
• Build and maintain labeled benchmark datasets that reflect real, messy, multi-speaker audio.
• Run regression and A/B evaluations on every model swap, prompt change, or pipeline update. Nothing ships on a vibe.
• Reject anecdotal proxies, single confidence scores, and cherry-picked examples as evidence of quality.
What We Are Looking For
• 3 to 5 years as an AI/ML engineer with production systems behind you. Engineering and production experience is non-negotiable.
• Depth across the modern AI stack: LLMs, speech models, vector retrieval, model serving.
• Strong software engineering. You write code that ships and survives contact with real users.
• Fluency in Python and the production ML ecosystem.
• Comfort with cloud infrastructure (GCP a plus) and containerized deployment on Kubernetes.
• A working command of observability and evaluation. You measure first and trust metrics over intuition.
• First-principles reasoning and metric discipline.
Nice to Have
• Research background or publications. A strong signal, not a substitute for production work.
• Audio and speech ML experience (STT, diarization, voice).
• Experience self-hosting and optimizing open models.
• Experience with LLM gateway and agent orchestration patterns.
• Experience building eval harnesses or production model-monitoring systems.
Requirements
Agentic work is must. Audio is good to have
. Self hosting models is a must
Experience with LLM gateway and agent orchestration is a must have
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Job Description – AI Engineer (End-to-End Development & Deployment)
Role Summary
We are looking for an AI Engineer with hands-on experience in designing, developing, deploying, and maintaining Generative/Agentic AI solutions in production. The ideal candidate should have end-to-end ownership of AI applications, from development to deployment, monitoring, and optimization.
Key Responsibilities
● Design, build, and deploy Generative/Agentic AI solutions.
● Develop applications using LLMs, RAG, AI agents, and vector databases.
● Build scalable APIs and integrate AI solutions with enterprise applications.
● Implement CI/CD pipelines, containerization, and MLOps best practices.
● Monitor, optimize, and maintain production AI systems.
● Collaborate with cross-functional teams to deliver business-driven AI solutions.
Required Skills
● Strong programming skills in Python.
● Experience with vector databases (e.g., Pinecone, FAISS, ChromaDB) and graph memory systems
● Knowledge of atleast one agent development framework: Google ADK (preferred), LangChain/LangGraph/LlamaIndex, CrewAI
● Experience with LLMs, RAG, GenAI, AgenticAI Agents
● Hands-on experience with FastAPI, and REST APIs.
● Knowledge of Docker, Kubernetes, Git, CI/CD.
● Experience with AWS, Azure, or GCP.
● Experience with security compliance, monitoring and observability tools such as AWS CloudWatch, Azure Monitor, Google Cloud Monitoring.
Job Title: Full Stack AI Engineer
Location: Remote/Hyderabad
Experience Level: 3-5
Salary Range: 12-18LPA
Application Link:https://beyond.ciltriq.com/apply/BUILD
Description:
Join a team building AI-powered systems that solve complex business problems and automate operational workflows across document processing, voice agents, enterprise integrations, workflow automation, and multi-agent systems.
Strong full-stack foundations: frontend state management, asynchronous user experiences and performance; backend API design, authentication, data modelling, databases, queues and distributed systems.
Strong coding ability in Python and JavaScript or TypeScript, with practical experience in modern frontend frameworks and backend services.
Requirements:
- Design and build complete systems: frontend applications, backend services, APIs, databases, data pipelines and integrations with customer systems.
- Build multi-agent workflows with clear agent responsibilities, tool access, shared state, context management, routing, handoffs and coordination across sequential and parallel tasks.
- Make agent execution dependable through durable state, checkpoints, retries, timeouts, idempotency, recovery and human approval or review where needed.
- Deliver document-processing pipelines, voice agents and retrieval-based AI applications, connecting model outputs to useful actions in real business workflows.
- Own quality in production: automated tests, AI evaluations, guardrails, observability, access controls, deployments, incident response and clear documentation.
- Choose where AI adds value and where deterministic software is the better fit. Balance accuracy, latency, cost, security and maintainability.
- Improve reusable engineering foundations, review code and help other engineers grow as the team expands.
- A solid understanding of tool calling, structured outputs, retrieval, context and memory management, model selection and evaluation.
- Practical cloud and deployment experience, including containers, CI/CD, secrets management, logging, monitoring and production debugging.
- Ability to reason from first principles, investigate failures across system boundaries and communicate technical decisions clearly to customers and teammates.
- Useful additional experience: Document AI and OCR, real-time voice systems, enterprise integrations, agent protocols such as MCP, and orchestration frameworks.
- Useful additional experience: Mentoring engineers or building reusable platforms.
Location: Hyderabad, India (home base), deployed at client sites in India. Occasional Middle East exposure possible.
About the Role
You will work as a senior AI engineer who embeds inside a customer's business. Your job is to learn how the business makes money, find the highest value problem, and build a working system that solves it.
Four behaviors define this role:
- Go where the work happens. You work onsite with the customer, in the room where decisions are made.
- Show working software early. You build a prototype in days, not a document in weeks.
- One person owns the outcome. You are the single point of accountability for the result.
- Stay after go-live. You keep running and improving the system after launch.
You are the single point of accountability. You are not a solo builder. A full KnackLabs engineering team in Hyderabad builds and runs the production systems behind you.
This role involves extended onsite deployments at client locations in other cities, sometimes up to six months at a stretch. Please apply only if you are ready for this way of working.
What you'll own
- Discovery - Learn how the customer makes money. Find the highest value problem to solve first.
- The prototype - Build a working prototype fast, using real or sample data, to prove the idea.
- The roadmap - Decide what to build, in what order, and set clear success measures tied to business outcomes.
- The build - Design and ship the production system with the Hyderabad engineering team. This includes data integration, agents, retrieval, and evaluations.
- The client relationship - Be the trusted technical contact for the customer, from engineers to senior leaders.
- Go live and after - Deploy the system, watch how it performs, fix problems, and improve it over time.
- Feedback to the product - Share what you learn in the field so the vendor's product and our internal tools get better.
What we are looking for
- Around 4 or more years of software engineering experience, including customer-facing or client delivery work.
- Strong programming skills in Python. Working knowledge of TypeScript or JavaScript.
- A full-stack development experience with strength in backend technologies.
- Production experience with large language models, including prompt engineering and agent development.
- You build with AI coding tools like Claude Code or Codex as your default way of working, and you have shipped real apps or agents this way.
- Experience building retrieval-augmented generation (RAG) systems: chunking, embeddings, vector databases, retrieval, and reranking.
- Experience building and deploying AI systems.
- Experience integrating with APIs and enterprise systems.
- Experience with at least one cloud platform (AWS, Azure, or GCP).
- Clear communication. You can explain a technical choice to an engineer and to a business leader.
- High ownership and comfort with ambiguity. You can take an unclear problem and turn it into a plan.
- Willingness to work onsite at client locations in India for extended periods, and to travel as the work needs.
Nice to have
- Experience with on-premises or private cloud (VPC) deployments.
- Experience with observability and tracing tools such as LangSmith or Braintrust.
- Experience with data engineering and pipelines.
- A history of side projects, open source contributions, or products you shipped end-to-end.
- Experience in embedded or forward-deployed roles before.
- Experience working at a consulting or professional services firm in a client-facing delivery role.
Stack and tools
- Languages: Python and TypeScript.
- Models: Claude and other frontier or open-source models, chosen to fit the customer.
- AI patterns: RAG, agents, prompt engineering, and evaluations.
- Vector and retrieval: vector databases and retrieval pipelines.
- Cloud: AWS, Azure, or GCP, on public or private cloud.
- Integration: REST APIs and enterprise system connectors.
Location: Hyderabad, India. Based at the KnackLabs headquarters, with occasional travel to client locations for workshops and reviews. This role does not involve extended onsite deployments.
About the Role
You will work as an AI Architect who designs the systems behind our client engagements: AI agents, RAG systems, automation platforms, and the conventional backend systems around them.
This is a hands-on design role, not a slideware role. You will scope architectures with clients, make the hard technical decisions, defend them in review, and stay accountable for how the systems perform in production.
You will work directly with clients. Everyone at KnackLabs does. You will sit in design discussions with client engineering teams, present architecture decisions to technical and business stakeholders, and answer for the choices you make.
A full KnackLabs engineering team in Hyderabad builds with you. You own the technical design and the quality of what ships.
What you'll own
- Architecture - Design AI agents, RAG systems, integrations, and the scalable backend systems around them, for multiple client engagements.
- Technical scoping - Work directly with clients to turn a business problem into a system design, with clear trade-offs and clear reasons.
- Scale and reliability - Make sure what we build handles real load: data stores, queues, caching, horizontal scaling, and fault tolerance.
- Design reviews - Review designs and builds across engagements. Set the technical bar and hold it.
- Evaluation strategy - Define how we measure accuracy, safety, latency, and cost for the AI systems we ship.
- Guiding engineers - Raise the level of the engineers building with you, through reviews and direct pairing.
- Feedback to the platform - Feed what you learn across engagements back into our platform and internal tools.
What we are looking for
- Around 7 or more years of software engineering experience, including direct work with customers on design or delivery.
- Full-stack development experience with strength in backend technologies.
- Experience designing and building scalable applications. You understand how large-scale distributed systems work: data partitioning, queues, caching, horizontal scaling, and fault tolerance.
- At least 2 years of strong, hands-on AI experience with large language models in production.
- You build with AI coding tools like Claude Code or Codex as your default way of working. You understand Claude Skills, have written skills yourself, use them actively, and have contributed to them.
- Hands-on experience building retrieval-augmented generation (RAG) systems: chunking, embeddings, vector databases, retrieval, and reranking.
- Hands-on experience building AI agents.
- Strong programming skills in Python. Working knowledge of TypeScript or JavaScript.
- Experience with at least one cloud platform (AWS, Azure, or GCP).
- Clear communication. You can explain an architecture decision to an engineer and to a business leader, and defend it under questioning.
- High ownership and comfort with ambiguity. You can take an unclear problem and turn it into a design.
Nice to have
- Experience building evaluations to measure accuracy, safety, latency, and cost.
- Experience with observability and tracing tools such as LangSmith or Braintrust.
- Experience with on-premises or private cloud (VPC) deployments.
- Experience deploying AI systems in regulated industries such as insurance, banking, or the public sector.
- Experience with data engineering and pipelines.
- A history of side projects, open source contributions, or products you shipped end-to-end.
- Experience working at a consulting or professional services firm in a client-facing delivery role.
Stack and tools
- Languages: Python and TypeScript.
- Models: Claude and other frontier or open-source models, chosen to fit the customer.
- AI patterns: RAG, agents, prompt engineering, skills, and evaluations.
- Vector and retrieval: vector databases and retrieval pipelines.
- Cloud: AWS, Azure, or GCP, on public or private cloud.
- Integration: REST APIs and enterprise system connectors.
Strong AI Engineer / Machine Learning Engineer profiles.
2
Mandatory (Experience 1) – Must have minimum 3+ years of hands-on experience in Data Science, Machine Learning, Applied AI, NLP, Deep Learning, or Generative AI solutions.
3
Mandatory (Experience 2) – Must have strong hands-on experience in Python programming, SQL, data analysis, feature engineering, model development, and production-grade ML applications.
4
Mandatory (Experience 3) – Must have experience working with Machine Learning and Deep Learning frameworks such as PyTorch, TensorFlow, Keras, Scikit-learn, or equivalent.
5
Mandatory (Experience 4) – Must have hands-on experience working on NLP, embeddings, semantic search, text classification, document understanding, recommendation systems, or similar AI/ML use cases.
6
Mandatory (Experience 5) – Must have experience working with Large Language Models (LLMs) such as GPT, Llama, Mistral, Claude, Gemini, Phi, or similar foundation models.
7
Mandatory (Experience 6) – Must have hands-on experience building or implementing RAG (Retrieval Augmented Generation) systems, vector search, knowledge retrieval, embeddings, chunking, indexing, or semantic retrieval solutions.
8
Mandatory (Experience 7) – Must have experience working with Git, CI/CD practices, production environments, and scalable AI/ML systems.
9
Mandatory (CTC) – The CTC breakup offered will be 75% fixed + 25% variable, as per company policy.
10
Mandatory (Age) - Candidate's Age should be below 30 Years
11
Preferred (Experience 1) – Experience with MLFlow, Kubeflow, Airflow, Prefect, Feature Stores, Model Registry, or MLOps/LLMOps frameworks.
12
Preferred (Experience 2) – Experience working with Vector Databases, Spark, PySpark, distributed ML pipelines, large-scale data processing, or real-time ML systems..
13
Preferred (Experience 3) – Familiarity with Docker, Kubernetes, Azure, AWS, GCP, cloud-native AI deployments, and scalable ML architecture.
14
Preferred (Company) – Candidates from AI-first startups, Fintech, Banking, Lending, Fraud Analytics, Risk Analytics, Product Companies, SaaS organizations, or data-driven technology companies
15
Mandatory ( Pedigree) - B.TECH / M.TECH from Tier 1 Colleges (IIT's, NIT's, BITS) are Considered.






