Cutshort logo
For Employers
Google Accelerator Startup Building AI-Marketing Platform. logo
Founding AI Engineer
Google Accelerator Startup Building AI-Marketing Platform.

Founding AI Engineer at Google Accelerator Startup Building AI-Marketing Platform. · Mumbai · 3 - 8 years · ₹30L - ₹60L / yr (ESOP available) · Posted 13 Jul 2026

Orangemint Technologies Private Limited's logo

Founding AI Engineer

at Google Accelerator Startup Building AI-Marketing Platform.

3 - 8 yrs
₹30L - ₹60L / yr (ESOP available)
Mumbai
Skills
skill iconPython
Large Language Models (LLM)
Fullstack Developer
API
Agentic AI
Model Context Protocol (MCP)

We're building in a market that hit $47B in 2026 and is on track to hit $107.5B by 2028 and we're not chasing that opportunity from the sidelines. Our founding team brings marketing expertise from Kotak and Amazon paired with deep AI chops, we were selected as one of just 20 startups by Google, and we already have paying customers like Kotak and Myntra using the product in production. Join us early and help build the infrastructure that turns AI marketing adoption (already at 88%) into real, scaled results for customers who are already paying for it, not case studies.


Founding Product Engineer, Applied AI

Location: Mumbai • Hybrid • Reports to: CTO


WHAT THIS ROLE ACTUALLY IS

Weve built a multi-agent AI system for marketing. It works well out of the box but every client (their data, workflows, brand rules, compliance constraints, output preferences) is different. Generic doesnt ship results. Customised does.

This role is the bridge: you sit between our agent architecture and the client’s reality. You understand both deeply, and you make the agents do what each specific client needs them to do.

Think Palantir Forward Deployed Engineer, but for AI agents. An engineer who owns product outcomes end-to-end.


WHAT YOU'LL DO

Day-to-day split: roughly 60% in Codex and Claude Code shipping customisations, 40% in client rooms understanding what to ship next.

  • Customise the agents. Tune prompts, build new agent capabilities, configure workflows, wire up integrations. Codex and Claude Code are your primary IDEs. You ship working code, not specs for someone else to build.
  • Own client outcomes end-to-end. From discovery to live deployment, you are the person responsible for whether the agents actually solve the client’s problem. No PM-engineer handoff to hide behind.
  • Problem-solve in real time. Client says “the output tone is off for our brand” or “the research agent is missing X data source” — you diagnose, fix, and redeploy, usually within days.
  • Hold your own on creative. When a CMO says “this doesn’t feel right,” you can question intelligently — is it the prompt, the input brief, the brand guidelines themselves, or the client’s own confusion? Most engineers reach for the prompt. The good ones diagnose first.
  • Feed learnings back to core product. Anything that gets customised 3+ times becomes a platform feature. You decide what graduates.
  • Build the playbook. As we scale from 10 to 50 to 100 clients, the customisation patterns you create become the foundation others work from.

WHAT YOU NEED

  • Real engineering chops. You can read a codebase, write production code, debug agent failures, and ship to live clients. AI tools are your accelerant, not your crutch.
  • Comfort with agents as a system. You understand how LLMs fail, why prompts drift, what context engineering means, when to use tools vs. fine-tuning, why eval harnesses matter. If “MCP,” “tool use,” and “agent harness” feel like jargon, this isn’t the role.
  • Client-facing instincts. Comfortable in a room with a CMO, an enterprise IT head, and a junior brand manager — and you can switch register for each in the same meeting. If client calls feel like tax to you, don’t apply.
  • Some creative literacy. You don’t need to be a designer. You do need to be able to tell the difference between “this output is technically broken” and “this output is technically fine but creatively wrong” — and chase the right thread.
  • Bias toward shipping. A working hack today beats a clean design next quarter.

YOU'LL THRIVE HERE IF

  • You’re an engineer who realised products live or die on customisation, not architecture.
  • You’ve been frustrated by being kept away from customers.
  • You like building with AI tools and want a role where that’s the job, not a side experiment.
  • You like marketing as a domain — or you’re willing to fall in love with it.

YOU WON'T THRIVE HERE IF

  • You want to optimise systems in a corner with headphones on.
  • You think client calls are below your pay grade.
  • You think AI is overhyped (we don’t).
  • You need a fully scoped ticket queue and weekly sprint rituals.


Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

Similar jobs (10)

company logo
Umama Sayed
Posted by Umama Sayed
Remote, Mumbai
2 - 4 yrs
Best in industry
skill iconPython
Large Language Models (LLM)
Generative AI
LangGraph
FastAPI
+7 more

AI Engineer

LLMs, Agents & AI Services

📍 Mumbai (On-site) | Full-time | 2-4 years


About the Role:

Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.

AI is core to how we design, deliver, and scale software for our customers.

We are hiring an AI Engineer for a dedicated client engagement building a complex production AI platform, working on the AI capabilities and agentic features at the core of the product.

The mandatory requirement for this role is at least one AI feature personally shipped to production for real users, with operational ownership.

The role suits someone who thinks quickly on solutioning, can take an ambiguous problem to a working prototype in days, and has the discipline to carry it through to production with predictable economics.

You will work alongside the Senior AI Engineer and the wider pod, with ownership of parts of the AI surface area of the product.


Responsibilities:

Solutioning and POCs

Translate ambiguous customer problems into working POCs at speed.

Pick the right model, framework, and architecture, and demonstrate value early before scaling investment.


LLM Application Development

Build AI features and services using LLM APIs from OpenAI, Anthropic, Google, and self-hosted open-weight models (Llama, Qwen, Mistral).

Choose the right model per use case based on cost, latency, capability, and context-window trade-offs.


Agentic System Design

Design and implement agentic workflows using LangGraph, CrewAI, AutoGen, LlamaIndex Agents, or custom orchestration.

Cover tool use, planning, memory, and multi-step reasoning appropriate to the problem.


API and Service Development

Build production AI services and APIs using Python and FastAPI.

Handle streaming responses, async processing, structured outputs, retries, and graceful degradation when models or tools fail.


Retrieval and Tool Integration

Implement RAG pipelines with vector databases (Pinecone, Weaviate, Qdrant, pgvector, Chroma), embeddings, chunking strategies, hybrid search, and reranking.

Integrate external tools, internal APIs, and document sources through tool-calling and MCP-style patterns.


Cost Analysis and Unit Economics

Model the per-request and per-user cost of every AI feature before it ships.

Track token usage, prompt caching, batching, and model-routing strategies.

Drive measurable improvements in unit economics.


Production Hardening

Add observability and tracing (LangSmith, Langfuse, OpenTelemetry), guardrails, content safety checks, prompt injection defences, and fallback behaviour.


Prompt Engineering and Evaluation

Design, test, and iterate prompts with measured outcomes.

Build evaluation harnesses for accuracy, hallucination, latency, and cost.

Run benchmarks across models and prompt variants before locking in a design.


Requirements:

AI Feature Shipped to Production (Mandatory)

Must have personally built and shipped at least one AI feature that runs in production for real users, with operational ownership.

POCs, internal demos, and one-off scripts do not qualify.


2 to 4 Years of Professional Software or AI Engineering Experience

With at least one production AI feature owned end to end.


Strong Python Proficiency and API Development with FastAPI

Comfort with type hints, async, packaging, testing, streaming responses, and authentication.

Production-grade Python, not notebook-only code.


Hands-on Depth Across the LLM and Agent Stack

Working experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or self-hosted open-weight models (vLLM, Ollama, Together, Replicate).

Working familiarity with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.

Working knowledge of RAG, embeddings, and vector databases (Pinecone, Weaviate, Qdrant, pgvector, Chroma).


Solutioning Speed and POC Velocity

Demonstrated ability to move from a fuzzy problem to a working prototype in days.

Strong instinct for what to build first, what to defer, and what to throw away.


Cost Discipline for Production AI

Ability to calculate, monitor, and optimise the cost of LLM APIs, tokens, embeddings, vector store usage, and infrastructure.

Treats unit economics as a first-class concern.


AWS Familiarity

Working knowledge of EC2, S3, IAM, and at least one of Bedrock, SageMaker, or equivalent.


Comfortable in a Fast-Moving Environment

Self-directed, comfortable with ambiguity, takes ownership without being asked, and ships under shifting priorities.


Strong Written and Spoken English Communication

Able to explain trade-offs to non-AI engineers, designers, product managers, and clients in plain language.


Nice to Have

  • fine-tuning or LoRA, QLoRA, PEFT exposure
  • MCP server authoring
  • eval framework experience (LangSmith, Promptfoo, Ragas, DeepEval)
  • open-source AI contributions
  • multi-modal models (vision, audio)
Read more
company logo
icon

The recruiter has not been active on this job recently. You may apply but please expect a delayed response.

Remote only
8 - 12 yrs
₹25L - ₹40L / yr
Agentic AI
skill iconReact.js
skill iconNodeJS (Node.js)
skill iconMongoDB
Generative AI (GenAI)
+6 more

Lead Product Engineer - Pod Captain 


A Little Bit About the Role


There is a version of this role that attracts the wrong person: a senior engineer who wants a title upgrade and a slightly bigger scope. That person will struggle here. The Lead PE is not a promotion for being a great individual engineer. It is a fundamentally different job.


Your success at this level is measured entirely by what your Island delivers — not by the quality of your own code. You will still write code, often the hardest and most risk-laden on the project: auth systems, payment integrations, data migrations, the things nobody else should touch. But that is a fraction of the role. The majority is managing delivery, owning the client relationship on the technical side, unblocking engineers before problems compound, and making scope calls under pressure with incomplete information.


You are also the technical and commercial bridge between your Island and the client. You own every piece of technical communication that goes to them — what is being built, what risks exist, what the honest timeline is. When a client pushes for something technically unsound, you push back directly and professionally, always with an alternative. When something breaks in production, you own the post-mortem process and the client receives it within 24 hours. You protect your team from unfiltered client pressure; you translate it into clear, actionable direction instead.


On top of delivery and client management, you are the AI operating standard for your Island. Every engineer on your pod uses Cursor and Claude within the rules you define. You set those rules before the first line of code is written, govern how they evolve, and you are accountable when AI-generated code ships without being understood.


If you are energised by all of that — by the accountability, by owning outcomes rather than tasks, by the combination of technical depth and people management — this role will suit you well. If you want to spend most of your time in code, this is not the right level.


AI Governance Expectation


You set the AI standard for your entire Island. If an engineer on your pod is shipping AI-generated code they do not understand, or underusing AI tools entirely, both are your problem to fix — not theirs alone.


At Lead level, AI governance means:

  • Authoring the .cursorrules file before the first PR is opened. This defines how every engineer on the team writes code, uses Cursor, and structures prompts. You update it as the team learns, and you submit it to the Fountane Mainland at the close of every engagement.
  • Evaluating new AI tools before they reach client code. You are the gate. A tool that looks useful in a demo can create untraceable technical debt in production. You decide what enters the workflow.
  • Governing prompt discipline across the team: scoped prompts, codebase context hygiene, no AI-generated code that the author cannot explain line by line.
  • Using Claude for architecture reasoning, sprint planning, risk documentation, and client-facing technical communication — not just Cursor for code generation. AI works at every layer of the Lead role, not just the coding layer.
  • Staying ahead of the Golden Stack. You should know what Cursor and Claude can and cannot do at the codebase level better than anyone on your Island. When the Mainland updates the standard toolchain, you integrate the change into your Island’s workflow first.


Key Responsibilities


Technical Architecture & Standards Ownership

Design the full-stack architecture for every engagement before the first sprint starts: service boundaries, API contracts, database schema, state management strategy across web and mobile surfaces, and the security baseline. Author the .cursorrules file that governs how the entire Island writes code and uses AI. Own the highest-risk code personally — auth systems, payment integrations, live data migrations — these are not delegated regardless of workload. Enforce Mainland security standards across the entire codebase: input validation, correct auth scoping, safe data handling. If a security issue reaches production that was visible in a PR, that is a failure you own.


CE Prototype Gatekeeping

Review every Concept Engineer prototype before any engineer touches it. This is not a courtesy review — you have veto authority. If the scope is technically unrealistic for the sprint, the architecture is unbuildable as specified, or the prototype makes promises the codebase cannot keep, the sprint does not start until those are resolved. The time to surface a structural problem is before it is built, not after it is demoed. Produce a written checklist verdict for every prototype review — approved, approved with conditions, or blocked with reasons.


Client Technical Communication & Scope Ownership

Own all technical communication with the client directly. This means: what is being built each sprint, what risks exist and what is being done about them, honest timelines when estimates change, and clear explanations of technical decisions in plain language. Push back on risky or scope-expanding requests professionally and always with an alternative — “we cannot do X because Y; here is what achieves the same outcome.” Manage scope creep actively: every request that expands the engagement goes through you, gets assessed for impact, and is either absorbed, priced, or declined with a rationale. Renewal conversations are yours to lead in coordination with the CE.


Island Integration Audit

When a prospective client wants fractional augmentation — one or two roles rather than a full Island — you conduct the Island Integration Audit. This means evaluating their existing engineering team against Fountane Mainland SOPs: coding standards, AI tool usage, deployment practices, and documentation quality. If they pass, fractional deployment proceeds. If they do not, you produce a written audit report that either recommends full Island deployment or a Mainland upskilling programme. You are protecting your team’s velocity — a fractional deployment into a dysfunctional team destroys it.


Delivery & Velocity Management

Review the work cycle board every morning. Flag anything stuck, stale, or at risk before it becomes a client problem — not after. Communicate risks to the CE proactively with options, not just problems. Track every engineer’s billing hours: a shortfall spotted on Wednesday is a Thursday conversation, not a Friday surprise. If an engineer is consistently under-billing, that is either a workload problem, a capability problem, or a motivation problem — each has a different resolution and you are responsible for diagnosing which one it is.


Production Incident Ownership

When something breaks in production, you own the resolution process end to end: take technical lead, make the call on rollback vs. hotfix, communicate status to the client directly if the incident affects them, and own the post-mortem. The client receives the post-mortem within 24 hours of resolution — this is a hard commitment, not a target. The post-mortem covers: what happened, why it happened, what was done to fix it, and what structural change prevents recurrence. You produce it; it is not delegated to the engineer who caused the issue.


Automation & Infrastructure Boundary Arbitration

The Automation Engineer owns the test gates within the CI/CD pipeline. The Infrastructure Engineer owns the pipeline configuration itself. When these overlap and conflict — which they will — you arbitrate. The resolution must be documented in the Island’s .cursorrules and SOP onboarding document so the boundary does not need to be relitigated every sprint. This is not optional housekeeping; an unresolved boundary between these two roles is one of the most common sources of deployment failures on Islands.


Full Stack Ownership — Web & Mobile

You review and own work across the entire application layer: React / Next.js web frontends, React Native or Flutter mobile apps, Node.js / Python backends, and database schema. You cannot delegate a review of mobile code because you do not know mobile, or backend code because your background is frontend. If there is a gap in your full-stack depth, that is something to resolve before taking this role, not on the job.


People Management & Mentoring

Hold bi-weekly 1:1s with every engineer on the Island — focused on growth, blockers, and career direction, not status updates. Status belongs on the board; 1:1s are for the person. Address performance gaps early, directly, and specifically: a pattern spotted in week two is a conversation in week two, not month three. Protect the team from unfiltered client anxiety — absorb it, process it, and give the team clear direction instead of noise. If an engineer is struggling, that is your problem to solve before it becomes the client’s problem to notice.


Network Contribution

At the close of every engagement, submit to the Fountane Mainland: the .cursorrules file, all prompt templates used, a short retrospective on what worked and what did not, and any reusable patterns or scripts developed during the engagement. This is a Lead-level obligation, not a nice-to-have. The network’s collective capability grows through what Leads bring back from the field. An engagement that ends without a Mainland submission has left the network no smarter than it was before.


Qualifications


  • 8+ years of software engineering experience, with at least two years in a role where you were accountable for delivery outcomes — not just your own code. You should be able to describe an engagement where something went wrong on your watch, what you did about it, and what you would do differently. If your experience is entirely individual-contributor, this role will be a difficult fit.
  • Genuinely full stack across web and mobile: React / Next.js, React Native or Flutter, Node.js or Python, and database architecture including schema design and migrations. You review and own work across all of these surfaces — not just the ones you are most comfortable in.
  • Hands-on, daily Cursor and Claude usage with the ability to set and enforce standards for how a team uses them — not just personal proficiency. You should be able to describe the .cursorrules file you would write for a standard SMB engagement and explain the decisions behind it.
  • Proven ability to communicate technical risk and decisions directly to non-technical clients — not through a PM, directly, in plain language. You should be able to describe a client pushback you navigated and how you handled it.
  • Experience managing scope on a real engagement: a client who wanted more than was in the contract, and how you resolved it without damaging the relationship or the team’s velocity.
  • Has conducted security-aware architecture design: auth systems, input validation, safe data handling. You should be able to describe a security decision you made and why.
  • Has led or coordinated a small engineering team — run 1:1s, addressed performance gaps, made scope calls under pressure. Management experience at a staffing agency or outsourcing firm counts if the accountability was real.
  • Comfortable having direct, specific performance conversations early — not after the situation has compounded. You should be able to describe a time you addressed a performance issue at the first sign rather than waiting.
  • Bachelor’s or Master’s in Computer Science or a related field, or equivalent experience.


About Fountane


Fountane is a technology ventures lab - one part product studio, one part startup engine. We build high-quality software and AI products for clients ranging from fast-moving startups to large enterprises, and we co-build and invest in new companies when we see the right opportunity.


Founded in 2017 and headquartered in Minneapolis, we have grown to 60+ people across four continents and were recognised as one of America’s fastest-growing companies, ranking No. 699 on the Inc. 5000 with 595% three-year growth.


We are serious about craft, direct about expectations, and operate at the frontier of AI-assisted engineering. Lead PEs here carry real authority and real accountability — over delivery, over people, and over client outcomes. If that combination is what you are looking for, this is the place.


Read more
Neosapien
Neosapien
Agency job
via by Nehlata Pandey
Bengaluru (Bangalore)
3 - 7 yrs
₹15L - ₹40L / yr (ESOP available)
skill iconPython
Large Language Models (LLM) tuning
Agentic AI
Retrieval Augmented Generation (RAG)
skill iconMachine Learning (ML)
+2 more

The Role

You own AI systems end to end. From the speech-to-text models that turn audio into text, to the diarization that separates and identifies speakers, to the agentic layer that turns conversation into memory and action, to the observability and evaluation that keep all of it honest in production. This is a wide role by design. You will own model selection, serving, and production reliability. If you want to tune one model and ignore the system around it, this is not the role.

What You Will Own

•     Speech-to-text. Evaluate, integrate, and optimize STT models across cloud and self-hosted. Drive accuracy and cost trade-offs with ground-truth metrics.

•     Speaker diarization and identification. Push accuracy on hard, real-world, multi-speaker audio.

•     Agentic AI. Build the memory and retrieval pipeline, LLM orchestration, and the agent workflows that sit on top of captured conversation.

•     Model serving and infrastructure. Stand up and optimize self-hosted serving (vLLM, Triton class). Own latency, throughput, and cost per user.

Observability

An always-on wearable means models run in production every second, on messy real-world audio. You own the visibility into that.

•     Instrument the full audio-to-memory pipeline: STT, diarization, retrieval, and LLM calls.

•     Define and track model-quality SLOs in production: transcription drift, diarization error over time, retrieval relevance, latency, throughput, and cost per user.

•     Build dashboards and alerting so model degradation is caught before users feel it.

•     Trace failures across a distributed, always-on system using metrics, logs, and traces.

•     Close the loop. Production signals feed back into evaluation and model selection.

Evaluation

We do not ship what we cannot measure. You own the systems that prove a model is actually better, not just newer.

•     Build and own ground-truth evaluation harnesses for every model in the stack.

•     Measure with real metrics: WER for transcription, DER for diarization, Recall and F1 for retrieval and speaker identification.

•     Build and maintain labeled benchmark datasets that reflect real, messy, multi-speaker audio.

•     Run regression and A/B evaluations on every model swap, prompt change, or pipeline update. Nothing ships on a vibe.

•     Reject anecdotal proxies, single confidence scores, and cherry-picked examples as evidence of quality.

What We Are Looking For

•     3 to 5 years as an AI/ML engineer with production systems behind you. Engineering and production experience is non-negotiable.

•     Depth across the modern AI stack: LLMs, speech models, vector retrieval, model serving.

•     Strong software engineering. You write code that ships and survives contact with real users.

•     Fluency in Python and the production ML ecosystem.

•     Comfort with cloud infrastructure (GCP a plus) and containerized deployment on Kubernetes.

•     A working command of observability and evaluation. You measure first and trust metrics over intuition.

•     First-principles reasoning and metric discipline.

Nice to Have

•     Research background or publications. A strong signal, not a substitute for production work.

•     Audio and speech ML experience (STT, diarization, voice).

•     Experience self-hosting and optimizing open models.

•     Experience with LLM gateway and agent orchestration patterns.

•     Experience building eval harnesses or production model-monitoring systems.


Requirements

Agentic work is must. Audio is good to have

. Self hosting models is a must

 Experience with LLM gateway and agent orchestration is a must have

Read more
company logo
Umama Sayed
Posted by Umama Sayed
Mumbai
5 - 8 yrs
Best in industry
skill iconPython
Large Language Models (LLM)
Artificial Intelligence (AI)
Prompt engineering
LangGraph
+6 more

Senior AI Engineer

Code Generation, Agent Architecture & LLM Systems

📍 Mumbai (On-site) | Full-time | 5+ years


About the Role:

Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.

We are hiring a Senior AI Engineer for a dedicated client engagement focused on building an AI-powered application builder platform - a product where users describe software in plain English and the system generates, previews, and iteratively refines working code.

The mandatory requirement for this role is hands-on production experience shipping LLM-powered systems with agent architectures, with experience in code generation or developer tooling contexts a strong advantage.


The role is product-focused and deeply hands-on. You will own everything between the user's prompt and correct code landing in the project: the agentic loop, code generation pipeline, context management, evaluation suite, and model cost strategy.

You will work alongside the Senior MLOps Engineer who operationalises the infrastructure around your system, and collaborate closely with backend, frontend, and DevOps engineers.


Responsibilities:


Agent Architecture

Design and own the agentic loop for the platform - request interpretation, planning, tool-calling sequence (read file, edit file, run build, search code, install package), and stop conditions.

Make and revisit architectural decisions on single-agent vs. multi-agent designs, including planner/executor splits and dedicated build-repair sub-agents.


Code Generation Pipeline

Own the end-to-end generation flow: task classification, context gathering, planning, targeted edits, verification, and commit.

Implement diff/search-replace-based file editing with fuzzy matching and fallback strategies.

Enforce scope discipline so the agent makes minimal diffs and does not modify code it was not asked to touch.


Self-Repair Loop

Build and tune the automated repair loop that pipes compiler, lint, build, and runtime errors back to the model with retry budgets and model escalation.

This loop is the primary quality lever - the difference between 60-70% and 90%+ build success rates.


Context Management

Build file-relevance retrieval so the agent sees the right files, not the whole codebase: dependency graphs, AST/tree-sitter-based chunking, embeddings, recency signals, and hybrid retrieval.

Implement conversation summarisation and memory for long sessions, and address long-project degradation through codebase summaries and periodic consistency passes.

Own token budgeting and prompt caching strategy.


Prompt Engineering as a Discipline

Own the system prompt and per-task prompt variants (new feature, bug fix, styling change).

Maintain few-shot examples and enforce coding conventions, stack rules, and prohibited behaviours such as no hardcoded secrets and no whole-file rewrites.

Version prompts like code with changelogs and rollback capability.


Evaluation and Quality Measurement

Design and own the evaluation suite: representative test prompts run on every prompt and model change, scored on build success rate, instruction adherence, and output quality including LLM-as-judge and visual/screenshot checks where relevant.

Define regression gates that block quality-degrading changes from shipping.

Treat evals the way engineers treat automated testing: versioned, automated, and tracked over time.

This responsibility is non-negotiable at this level.


Model Strategy and Cost

Design model routing - cheap and fast models for classification and small edits, frontier models for complex generation.

Drive cost optimisation through prompt caching, diff-based edits over full-file rewrites, and tighter context selection.

Track cost per agent run and tokens per task; evaluate new model releases against the eval suite and lead migrations when results justify it.


Safety and Reliability of Agent Behaviour

Defend against prompt injection from user content and fetched web content.

Ensure secrets never appear in generated client code.

Define what the agent's tools may and may not do in collaboration with the platform team.

Contribute to output moderation and abuse-pattern awareness.


Mentorship and Engineering Standards

Run code reviews, define engineering conventions for AI work, and raise the engineering bar across the AI team.

Work closely with the Senior MLOps Engineer on handoff of eval design, prompt configurations, and model routing logic.


Requirements:


Hands-on Production Ownership of LLM-Powered Systems with Agent Architectures (Mandatory)

Must have personally shipped and operated at least one complex production AI system - agentic, multi-step, or code generation - with end-to-end ownership of architecture, evaluation, and cost.

POCs, internal demos, and tutorial-grade work do not qualify.


5+ Years of Professional Software or AI Engineering Experience

With at least 3 years focused on LLM applications, AI engineering, or production AI systems.

Candidates with strong backend backgrounds and a clear, substantive pivot into LLM systems qualify.


Strong Python Proficiency and Service Development

Production-grade Python with FastAPI or equivalent: type hints, async patterns, streaming responses, testing, and packaging.

Not notebook-only.


Depth Across LLM APIs and Agent Systems

Production experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or open-weight models (vLLM, Ollama, Together).

Production experience with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.

Hands-on with tool calling, structured outputs, and multi-step reasoning.


Demonstrated, Systematic Evaluation Practice - Non-Negotiable

Must have built evaluation harnesses that gate production releases, not ad-hoc testing.

Hands-on with at least one of LangSmith, Langfuse, Promptfoo, Ragas, or DeepEval.

Candidates with no systematic answer to evaluation should not be considered at senior level regardless of other strengths.


Cost Discipline for Production AI

Track record of measurable cost optimisation on production AI features.

Able to speak in specifics: cost per request, savings achieved through caching or model routing, context reduction decisions.


AWS Working Knowledge

Hands-on with EC2, S3, IAM, and Docker.

Comfort with CI/CD workflows and deploying AI services.


Awareness of LLM Security Failure Modes

Familiar with prompt injection patterns, understands that system prompt rules alone are insufficient, and has experience with output validation and content safety in production.


Nice to Have

  • Experience with AST/tree-sitter tooling, diff-based editing systems, or compiler-adjacent work
  • MCP server authoring
  • Open-source AI contributions
  • Published technical writing on LLM systems
  • Multi-modal model experience
  • Fine-tuning exposure (LoRA, QLoRA, PEFT)
Read more
Service Co
Service Co
Agency job
via by Rishika Teja
Pune
2 - 4 yrs
₹6L - ₹16L / yr
Artificial Intelligence (AI)
skill iconMachine Learning (ML)
skill iconPython
skill iconJavascript
Large Language Models (LLM)

Mandatory skills : Artificial Intelligence,Machine Learning,Large Language Model


 2–4 years of experience in software engineering or AI/ML roles

- Proficiency in Python (preferred) or JavaScript/TypeScript

- Basic understanding of LLMs, RAG, and AI application development

- Familiarity with APIs, microservices, and databases (SQL/NoSQL)

- Exposure to cloud platforms (AWS/Azure/GCP)

- Strong problem-solving and learning mindset

Read more
company logo
Stuti Jain
Posted by Stuti Jain
Hyderabad
7 - 10 yrs
₹30L - ₹45L / yr
skill iconPython
AI Agents
Large Language Models (LLM)
Prompt engineering
Retrieval Augmented Generation (RAG)
+7 more

Location: Hyderabad, India (home base), deployed at client sites in India. Occasional Middle East exposure possible.

About the Role

You will work as a senior AI consultant who embeds inside a customer's business. Your job is to learn how the business makes money, find the highest value problem, and build a working system that solves it.


You will not hand over a document and walk away. You will show working software early, own the roadmap, own the client relationship, and stay after go-live to run and improve the system.


Four behaviors define this role:

  1. Go where the work happens. You work onsite with the customer, in the room where decisions are made.
  2. Show working software early. You build a prototype in days, not a document in weeks.
  3. One person owns the outcome. You are the single point of accountability for the result.
  4. Stay after go-live. You keep running and improving the system after launch.


You are the single point of accountability. You are not a solo builder. A full KnackLabs engineering team in Hyderabad builds and runs the production systems behind you.


This role involves extended onsite deployments at client locations in other cities, sometimes up to six months at a stretch. Please apply only if you are ready for this way of working.

What you'll own

  1. Discovery - Learn how the customer makes money. Find the highest value problem to solve first.
  2. The prototype - Build a working prototype fast, using real or sample data, to prove the idea.
  3. The roadmap - Decide what to build, in what order, and set clear success measures tied to business outcomes.
  4. The build - Design and ship the production system with the Hyderabad engineering team. This includes data integration, agents, retrieval, and evaluations.
  5. The client relationship - Be the trusted technical contact for the customer, from engineers to senior leaders.
  6. Go live and after - Deploy the system, watch how it performs, fix problems, and improve it over time.
  7. Feedback to the product - Share what you learn in the field so our platform and internal tools get better.

What we are looking for

  1. Around 7 or more years of software engineering experience, including customer-facing or client delivery work.
  2. Experience working at a consulting or professional services firm in a client-facing delivery role.
  3. A full-stack development experience with strength in backend technologies.
  4. Strong programming skills in Python. Working knowledge of TypeScript or JavaScript.
  5. Production experience with large language models, including prompt engineering and agent development.
  6. You build with AI coding tools like Claude Code as your default way of working. You have built real apps and agents this way, not just used it for document generation or review.
  7. Experience building retrieval-augmented generation (RAG) systems: chunking, embeddings, vector databases, retrieval, and reranking.
  8. Experience building and deploying AI systems.
  9. Experience integrating with APIs and enterprise systems.
  10. Experience with at least one cloud platform (AWS, Azure, or GCP).
  11. Experience building evaluations to measure accuracy, safety, latency, and cost.
  12. Clear communication. You can explain a technical choice to an engineer and to a business leader.
  13. High ownership and comfort with ambiguity. You can take an unclear problem and turn it into a plan.
  14. Willingness to work onsite at client locations in India for extended periods, and to travel as the work needs.

Nice to have

  1. Experience deploying AI systems in regulated industries such as insurance, banking, or the public sector.
  2. Experience with on-premises or private cloud (VPC) deployments.
  3. Experience with observability and tracing tools such as LangSmith or Braintrust.
  4. Experience with data engineering and pipelines.
  5. A history of side projects, open source contributions, or products you shipped end-to-end.
  6. Experience in embedded or forward-deployed roles before.

Stack and tools

  1. Languages: Python and TypeScript.
  2. Models: Claude and other frontier or open source models, chosen to fit the customer.
  3. AI patterns: RAG, agents, prompt engineering, and evaluations.
  4. Vector and retrieval: vector databases and retrieval pipelines.
  5. Cloud: AWS, Azure, or GCP, on public or private cloud.
  6. Integration: REST APIs and enterprise system connectors.


Read more
company logo
Ranitha Nair
Posted by Ranitha Nair
Chennai, Pune, Gurugram
4 - 10 yrs
₹25L - ₹60L / yr
Agentic AI
Generative AI
Retrieval Augmented Generation (RAG)
LangGraph
LangChain
+6 more

Staff Software Engineer — AI-native, high agency


ZoomRx Technology Team · Chennai / Pune / Gurugram (hybrid)


The bet

The last 12 months redefined what “senior engineer” means. We’re hiring the people who already operate the new way.

If you’ve spent this year rewiring how you ship — turning ambiguous PRDs into shipped features with Claude Code and Codex in the loop, replacing rote work with agents, treating AI as leverage instead of assistance — read on.


About ZoomRx

We work on hard problems at the intersection of data, healthcare, and technology, for the world’s largest biopharma companies. Our products — Ferma.AI, PERxCEPT, HCP-Pt Conversations — shape how they understand markets and make decisions. We’re flat; engineers own outcomes, not tickets.

Apply Here:-Hey! I came across an amazing job opportunity at ZoomRx that I think could be a perfect match for you. Check out the details and apply here: Staff Software Engineer


The role

You’ll own engineering work end-to-end across our products and platform — from ambiguous problem to shipped feature.

•    Translate fuzzy product and engineering asks into shipped systems; you don’t wait for a manager to break the work down for you.

•    Build with Claude Code, Codex, and the broader agentic stack as your default workflow — not novelty.

•    Design and ship backend systems, APIs, and data platforms that hold up at scale.

•    Partner directly with product owners and stakeholders. Product instinct matters as much as code.


How we engineer

We write evals, not just tests. We engineer the harness: model routing, context and prompt design, tokenomics — first-class craft, not afterthoughts. We author skills and MCP plugins to specialize our toolchain — we don’t just consume it. Open-weights models get explored when they unlock cost, sovereignty, or capability that frontier APIs can’t. We treat the engineering loop itself as something to instrument and improve, week over week.


How we operate

Closer to a Forward-Deployed Engineer than a heads-down coder. Twice-daily written updates are the default, not an ask. Decisions, trade-offs, and approaches live where everyone can see them — silence stays comfortable because the writing is already flowing. We respond fast on chat and email and treat documents as how we think and align.


What we look for

•    4–10 years building real software. Stack is secondary; strong fundamentals and learning agility are not.

•    You’ve already rewired how you work in the last 6–12 months around AI tooling. You can tell us specifically what changed, and why.

•    High agency. You spot the problem, frame it, and ship the fix without waiting to be asked.

•    Communicates in writing by default. Surfaces progress, decisions, and blockers fast — not weekly, not when asked.

•    No cognitive surrender. You think with AI, not through it — you spot slop, push back on the model, and ship higher-quality work because of the loop, not despite it.

•    Sharp judgment on trade-offs, prioritization, and when to push back.

•    Comfort with ambiguity and a bias to action.


Python, FastAPI, LangGraph, and Postgres are common in our stack — but we’d rather hire a generalist who thinks AI-native than a specialist who doesn’t.


Why now

We’re not retrofitting AI onto how we used to work. We’re rebuilding the tech function around it — operating model, tooling, hiring bar. Engineers joining now shape what that looks like.

If “what changed in the last six months of how you work” is a question you have a real answer to, we’d love to talk.


What’s in our stack today

Tooling our engineers run on every day: Claude Code, Codex, MCP servers, and the broader agentic toolchain.


Product stack: Python, FastAPI, LangGraph, Temporal, Anthropic Claude, OpenAI, Google Gemini (Vertex AI), Milvus, PostgreSQL, Redis, Kubernetes (GKE) on GCP, LangFuse, SigNoz, Sentry.

We don’t expect everyone to have touched all of these — most of our engineers picked up half on the job.

Read more
company logo
Rohan Jain
Posted by Rohan Jain
Pune
1 - 3 yrs
₹20L - ₹25L / yr
Software engineering
Fullstack Developer

Summary

We are looking for an exceptional AI-native full-stack product engineer who can take product specs, prototypes, or PRDs and turn them into clean, tested, QA’d, merge-ready branches.


We are building a fast-moving software company and currently use AI coding tools heavily in our development workflow. Right now, the delivery process is: prototype review, spec clarification, implementation planning, AI-assisted coding, QA, checking what was actually built, and preparing features for final merge.


We want to have this entire workflow managed by the right engineer.


Your job will be to take a feature brief, PRD, prototype, or Loom walkthrough and own the process from:


Input → clarification → implementation plan → AI-assisted build → self-review → QA → PR → merge-ready branch.


We are looking for someone who can take ownership, not create more management work.


What You’ll Do


You will be responsible for:

  • Reviewing PRDs, prototypes, tickets, or Loom walkthroughs
  • Asking clear, batched clarification questions
  • Creating concise implementation plans before coding
  • Using AI coding tools to accelerate development
  • Reviewing and improving AI-generated code
  • Building full-stack product features
  • Writing or updating tests where appropriate
  • Running QA before handing work over
  • Creating clean, reviewable pull requests
  • Documenting what was built, what changed, and what risks remain
  • Preparing branches that are actually ready to merge
  • Communicating clearly and asynchronously


What “Done” Means


A feature is not done when the code is written. A feature is done when:

  • The implementation matches the PRD/prototype
  • Acceptance criteria are completed
  • Code is clean and understandable
  • Build/lint/type checks pass
  • Relevant tests are added or updated
  • Manual QA has been completed
  • Screenshots or Loom walkthrough are provided when relevant
  • PR description is clear
  • Risks, assumptions, and deviations are documented
  • The branch is ready for review/merge


Ideal Candidate


You are probably a strong fit if you are:

  • An experienced full-stack engineer with strong product judgment
  • Experienced with AI coding tools such as Claude Code
  • Comfortable taking a vague product idea and turning it into a working feature
  • Able to review AI-generated code critically
  • Strong with GitHub PR workflows
  • Detail-oriented with good QA discipline
  • Clear in written communication
  • Comfortable working in a fast-moving startup environment
  • Able to work independently without needing constant hand-holding
  • Good at asking the right questions early


You Are Not a Fit, please do not apply if:

  • You blindly trust AI-generated code
  • You need every task broken down into tiny instructions
  • You say “done” without testing your work
  • You create large messy PRs
  • You cannot explain your own code
  • You are only comfortable with narrow frontend or backend tasks
  • You are not comfortable with ambiguity
  • You dislike writing implementation plans or QA notes
  • You require constant synchronous direction


Required Skills


Please include your experience with:


  • Full-stack web development
  • Frontend frameworks
  • Backend/API development
  • Git/GitHub
  • Pull request workflow
  • Testing and QA
  • AI-assisted coding tools
  • Debugging and self-review
  • Product-focused software development
  • Nice to Have
  • Startup experience
  • Experience as a founding engineer or product engineer
  • Strong UX/product taste
  • Experience with automated testing
  • Experience with CI/CD
  • Experience with preview deployments
  • Experience turning Figma/prototypes into production features
  • Experience working directly with founders


How We Work


We will usually provide one or more of:

  • PRD
  • Prototype
  • Figma
  • Loom walkthrough
  • Linear/Jira/Notion ticket
  • Product brief
  • Screenshots
  • Existing codebase


You will be expected to return:

  • Clarifying questions, if needed
  • Implementation plan
  • Branch/PR
  • QA notes
  • Test results
  • Screenshots/Loom if relevant
  • Clear summary of what was built
  • Any risks, tradeoffs, or unresolved questions


Trial Task


We will start with a paid trial task.


The trial will involve a small real feature or product improvement. You will be asked to:


  • Review the brief/prototype
  • Write a short implementation plan
  • Build the feature
  • QA it
  • Open a clean PR
  • Provide a Loom or written walkthrough
  • Document assumptions, risks, and anything not completed


We are not just evaluating whether the code works. We are evaluating whether you can own the delivery process end-to-end and how well you leverage AI.


What Success Looks Like


Success means the team can give you a PRD or prototype and receive back a merge-ready branch without needing to manage every step.


Great output looks like:


“I reviewed the feature brief, found two ambiguities, made reasonable assumptions, implemented the core flow, added tests, QA’d happy path and edge cases, documented one tradeoff, and the branch is ready for review. Management's input is only needed on this one product decision.”


Preference will be given to candidates who can work with high ownership and potentially become a long-term product engineering partner.


To Apply, please consider answering the following questions:

  • What AI coding tools do you use regularly, and how do you use them?
  • Walk us through your process for turning a PRD or prototype into a merge-ready branch.
  • How do you QA your own work before asking for review?
  • What are the biggest risks of AI-generated code?
  • Share an example of a PR, project, or product feature you built that you are proud of.
  • What stack are you strongest in?
  • Are you comfortable doing a paid trial task?
  • What does “merge-ready” mean to you?
Read more
company logo
Vijay Muthu
Posted by Vijay Muthu
Delhi, Gurugram, Noida, Ghaziabad, Faridabad
2 - 5 yrs
₹8L - ₹12L / yr
Generative AI
Chatbot
Large Language Models (LLM) tuning
Agentic AI
Prompt engineering
+4 more

About MyOperator


MyOperator is a Business AI Operator, a category-leader that unifies WhatsApp, Calls, and AI-powered chat & voice bots into one intelligent business communication platform. Unlike fragmented communication tools, MyOperator combines automation, intelligence, and workflow integration to help businesses run WhatsApp campaigns, manage calls, deploy AI chatbots, and track performance — all from a single, no-code platform. Trusted by 12,000+ brands including Amazon, Domino's, Apollo, and Razorpay, MyOperator enables faster responses, higher resolution rates, and scalable customer engagement — without fragmented tools or increased headcount


Role Summary

We’re hiring a Front Deployed Engineer (FDE)—a customer-facing, field-deployed engineer who owns the end-to-end delivery of AI bots/agents.

This role is “frontline”: you’ll work directly with customers (often onsite), translate business reality into bot workflows, do prompt engineering + knowledge grounding, ship deployments, and iterate until it works reliably in production.

Think: solutions engineer + implementation engineer + prompt engineer, with a strong bias for execution.


Responsibilities-


Requirement Discovery & Stakeholder Interaction

  • Join customer calls alongside Sales and Revenue teams.
  • Ask targeted questions to understand business objectives, user journeys, automation expectations, and edge cases.
  • Identify data sources (CRM, APIs, Excel, SharePoint, etc.) required for the solution.
  • Act as the AI subject-matter expert during client discussions.


Use Case & Solution Documentation

  • Convert discussions into clear, structured use case documents, including:
  • Problem statement & goals.
  • Current vs. proposed conversational flows.
  • Chatbot conversation logic, integrations, and dependencies.
  • Assumptions, limitations, and success criteria.


Customer Delivery Ownership

  • Own deployment of AI bots for customer use-cases (lead qualification, support, booking, etc.). Run workshops to capture processes, FAQs, edge cases, and success metrics. Drive the go-live process: requirements through monitoring and improvement.


Prompt Engineering & Conversation Design

  • Craft prompts, tool instructions, guardrails, fallbacks, and escalation policies for stable behavior. Build structured conversational flows: intents, entities, routing, handoff, and compliant responses. Create reusable prompt patterns and "prompt packs."


Testing, Debugging & Iteration

  • Analyze logs to find failure modes (misclassification, hallucination, poor handling). Create test sets ("golden conversations"), run regressions, and measure improvements. Coordinate with Product/Engineering for platform needs.


Integrations & Technical Coordination

  • Integrate bots with APIs/webhooks (CRM, ticketing, internal tools) to complete workflows. Troubleshoot production issues and coordinate fixes/root-cause analysis.



What Success Looks Like

  • Customer bots go live quickly and show high containment + high task completion with low escalation.
  • You can diagnose failures from transcripts/logs and fix them with prompt/workflow/knowledge changes.
  • Customers trust you as the “AI delivery owner”—clear communication, realistic timelines, crisp execution.


Requirements (Must Have)

  • 2–5 years in customer-facing delivery roles: implementation, solutions engineering, customer success engineering, or similar.
  • Hands-on comfort with LLMs and prompt engineering (structured outputs, guardrails, tool use, iteration).
  • Strong communication: workshops, requirement capture, crisp documentation, stakeholder management.
  • Technical fluency: APIs/webhooks concepts, JSON, debugging logs, basic integration troubleshooting.
  • Willingness to be front deployed (customer calls/visits as needed).


Good to Have (Nice to Have)

  • Experience with chatbots/voicebots, IVR, WhatsApp automation, conversational AI platforms with at least a couple of projects. 
  • Understanding of metrics like containment, resolution rate, response latency, CSAT drivers.
  • Prior SaaS onboarding/delivery experience in mid-market or enterprises.


Working Style & Traits We Value

  • High agency: you don’t wait for perfect specs—you create clarity and ship.
  • Customer empathy + engineering discipline.
  • Strong bias for iteration: deploy → learn → improve.
  • Calm under ambiguity (real customer environments are chaotic by default).


Read more
company logo
Archita Srivastava
Posted by Archita Srivastava
Hyderabad
4 - 8 yrs
₹15L - ₹25L / yr
skill iconPython
TypeScript
skill iconJavascript
Large Language Models (LLM)
Agentic AI
+1 more

Location: Hyderabad, India (home base), deployed at client sites in India. Occasional Middle East exposure possible.

About the Role

You will work as a senior AI engineer who embeds inside a customer's business. Your job is to learn how the business makes money, find the highest value problem, and build a working system that solves it.


Four behaviors define this role:

  1. Go where the work happens. You work onsite with the customer, in the room where decisions are made.
  2. Show working software early. You build a prototype in days, not a document in weeks.
  3. One person owns the outcome. You are the single point of accountability for the result.
  4. Stay after go-live. You keep running and improving the system after launch.


You are the single point of accountability. You are not a solo builder. A full KnackLabs engineering team in Hyderabad builds and runs the production systems behind you.


This role involves extended onsite deployments at client locations in other cities, sometimes up to six months at a stretch. Please apply only if you are ready for this way of working.

What you'll own

  1. Discovery - Learn how the customer makes money. Find the highest value problem to solve first.
  2. The prototype - Build a working prototype fast, using real or sample data, to prove the idea.
  3. The roadmap - Decide what to build, in what order, and set clear success measures tied to business outcomes.
  4. The build - Design and ship the production system with the Hyderabad engineering team. This includes data integration, agents, retrieval, and evaluations.
  5. The client relationship - Be the trusted technical contact for the customer, from engineers to senior leaders.
  6. Go live and after - Deploy the system, watch how it performs, fix problems, and improve it over time.
  7. Feedback to the product - Share what you learn in the field so the vendor's product and our internal tools get better.


What we are looking for

  1. Around 4 or more years of software engineering experience, including customer-facing or client delivery work.
  2. Strong programming skills in Python. Working knowledge of TypeScript or JavaScript.
  3. A full-stack development experience with strength in backend technologies.
  4. Production experience with large language models, including prompt engineering and agent development.
  5. You build with AI coding tools like Claude Code or Codex as your default way of working, and you have shipped real apps or agents this way.
  6. Experience building retrieval-augmented generation (RAG) systems: chunking, embeddings, vector databases, retrieval, and reranking.
  7. Experience building and deploying AI systems.
  8. Experience integrating with APIs and enterprise systems.
  9. Experience with at least one cloud platform (AWS, Azure, or GCP).
  10. Clear communication. You can explain a technical choice to an engineer and to a business leader.
  11. High ownership and comfort with ambiguity. You can take an unclear problem and turn it into a plan.
  12. Willingness to work onsite at client locations in India for extended periods, and to travel as the work needs.

Nice to have

  1. Experience with on-premises or private cloud (VPC) deployments.
  2. Experience with observability and tracing tools such as LangSmith or Braintrust.
  3. Experience with data engineering and pipelines.
  4. A history of side projects, open source contributions, or products you shipped end-to-end.
  5. Experience in embedded or forward-deployed roles before.
  6. Experience working at a consulting or professional services firm in a client-facing delivery role.

Stack and tools

  1. Languages: Python and TypeScript.
  2. Models: Claude and other frontier or open-source models, chosen to fit the customer.
  3. AI patterns: RAG, agents, prompt engineering, and evaluations.
  4. Vector and retrieval: vector databases and retrieval pipelines.
  5. Cloud: AWS, Azure, or GCP, on public or private cloud.
  6. Integration: REST APIs and enterprise system connectors.


Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos