Cutshort logo
For Employers
AIspire360 Inc logo
Senior Voice AI Engineer — HIPAA-Compliant AI Receptionist Platform
Senior Voice AI Engineer — HIPAA-Compliant AI Receptionist Platform

Senior Voice AI Engineer — HIPAA-Compliant AI Receptionist Platform at AIspire360 Inc · Remote only · 3 - 10 years · ₹10L - ₹30L / yr · Bootstrapped · Remote only · Posted 12 May 2026

AIspire360 Inc's logo

Senior Voice AI Engineer — HIPAA-Compliant AI Receptionist Platform

Amit Mohan's profile picture
Posted by Amit Mohan
3 - 10 yrs
₹10L - ₹30L / yr
Remote only
Skills
Artificial Intelligence (AI)
HIPAA
Backend testing
API management

Company: AISpire360 Role: Senior Voice AI Engineer (Independent Contractor, Full-Time Engagement) Location: Remote (India-based preferred; must overlap 4–6 hours daily with London + US Eastern) Engagement: Full-time independent contractor — long-term Compensation: Competitive, paid in USD, monthly. Senior staff-level rates.

About AISpire360

We are building the operating system for medical practices in the US. Our flagship product is ARIA, an AI receptionist that handles inbound and outbound calls, intake, scheduling, eligibility, prior authorization workflows, and after-hours coverage for medical practices. We are in active production deployment with our first practice — a spine surgery practice at the Hospital for Special Surgery in New York. Multiple specialty practices are in our deployment pipeline.

The platform is built to be specialty-agnostic. The voice agent stack is the universal foundation. Knowledge layers (spine surgery, primary care, behavioral health, dermatology, cardiology, pediatrics, etc.) plug in as configuration. One platform, many specialties.

We are venture-backed, founder-led, and shipping. The technical decisions you make here will be in production within weeks.

What You'll Own

You are the technical anchor for the voice agent. You will be responsible for:

Real-time voice system

  • Voice agent stack on ElevenLabs Conversational AI (and evaluation of Retell, Vapi, LiveKit alternatives as we scale)
  • STT and TTS configuration, voice tuning, language and accent handling
  • Turn-taking, barge-in, end-of-speech detection, latency budgets
  • Driving end-to-end voice latency below 700ms (we currently target 600ms)
  • Quality vs. cost tradeoffs across voice tier selection (Premium / Turbo / Standard)
  • Silence-detection optimization for outbound calls (95% silence-discount realization)

HIPAA + PHI infrastructure

  • BAA-compliant hosting architecture (AWS-based, BAA in place)
  • PHI detection and redaction pipeline (AWS Comprehend Medical + custom pre-filters)
  • Encryption in transit and at rest, audit logging, 7-year retention compliance
  • Access controls, RBAC, kill-switch infrastructure
  • HIPAA Security Rule compliance for technical safeguards

Agent orchestration and tools

  • LLM integration (GPT-4o, Claude Sonnet, with model routing for cost optimization)
  • Tool calling for scheduling, eligibility checks, prior auth lookups, calendar booking, task creation
  • Multi-turn conversation state management
  • Multi-topic call handling and structured task generation
  • Safety-critical keyword detection and escalation routing

Integrations

  • EHR integration (Epic, Athena, eClinicalWorks — at least one in V1)
  • Insurance eligibility APIs (Availity, Change Healthcare, payer-specific)
  • Calendar integration (Google Calendar, Microsoft 365, Cal.com)
  • SMS and email gateways (Twilio, SendGrid)
  • Practice-specific webhooks for task manager surfacing

Multi-tenant platform foundations

  • Per-practice configuration loading (specialty profile + practice profile)
  • Per-tenant isolation, data segregation
  • Specialty-specific intake schemas, triage rules, safety keywords
  • Production observability (tracing, metrics, latency monitoring per practice)

Required Experience

You should be able to demonstrate, in a live walkthrough with code in hand, that you have shipped production voice AI systems. Specifically:

  • 3+ years of production voice/conversational AI experience. Not chatbots — actual real-time voice agents handling live phone calls. ElevenLabs, Retell, Vapi, LiveKit, OpenAI Realtime, Deepgram, or equivalent.
  • Deep familiarity with at least 2 of: ElevenLabs Conversational AI, Retell AI, Vapi, LiveKit Agents, OpenAI Realtime API. Production deployment experience required, not demo experience.
  • HIPAA and PHI handling experience in production. Not theoretical. You have signed BAAs, deployed in HIPAA-compliant environments, and built PHI redaction pipelines.
  • AWS infrastructure expertise. Specifically: ECS or Fargate, Lambda, S3 with encryption, KMS, CloudWatch, IAM, VPC, BAA-eligible service selection.
  • AWS Comprehend Medical or equivalent PHI detection experience.
  • Twilio or equivalent telephony provider experience. SIP trunking, inbound/outbound call flows, recording and transcription pipelines.
  • Latency optimization in real-time systems. You have hit sub-700ms voice latency in production and can explain the trade-offs you made.
  • LLM production experience. GPT-4o, Claude, or equivalent. Prompt engineering, tool calling, structured outputs, model routing for cost.
  • Python or TypeScript fluency. Production code, not scripts.
  • API integration experience. REST, webhooks, OAuth, async patterns. Bonus if you've integrated with EHRs (Epic FHIR, Athena, eClinicalWorks).

Nice to Have

  • Healthcare domain experience (EHRs, clinical workflows, medical terminology)
  • Prior auth or eligibility verification workflow experience
  • Multi-tenant SaaS architecture experience
  • SOC 2 audit experience
  • Experience with knowledge graphs (PostgreSQL + pgvector, Neo4j, or similar)
  • Open-source contributions to LangChain, LiveKit, Pipecat, or similar voice/agent frameworks
  • Demonstrated ability to ship to production fast (not perfectionism — pragmatic engineering)

What This Engagement Looks Like

  • Independent contractor. You invoice us monthly. We pay in USD via wire transfer or Wise.
  • Full-time engagement. This is your primary work. We expect ~40 hours/week of focused output.
  • Long-term. We are not looking for a 3-month engagement. Expect this to run 12+ months.
  • You sign a BAA with us. This is non-negotiable due to the nature of the work.
  • You are eligible for future equity grants as we formalize the team structure.
  • Direct working relationship with our CTO (London-based) and founders (NYC and India).
  • Time-zone overlap required: at least 4–6 hours of daily overlap with both London and US Eastern. Indian Standard Time evenings (6 PM IST onward) work well.

How We'll Evaluate You

Our hiring process is fast and substantive. No take-homes. No multi-round bureaucracy.

  1. Initial screen (30 min): Background, what you've shipped, compensation expectations, BAA-readiness.
  2. Technical deep-dive (90 min): Walk us through a production voice AI system you've shipped. Code in hand. Architecture diagrams. Real production numbers — latency, cost per call, scale. We will ask follow-up questions about specific decisions you made.
  3. System design (90 min, live): We give you a design problem — a HIPAA-compliant voice agent for a medical specialty we don't yet support. You design end-to-end on a whiteboard / Excalidraw. We probe trade-offs.
  4. Founders + CTO conversation (60 min): Mutual fit. You ask us hard questions. We answer honestly.

Total: 4–5 hours of your time. We'll move from first conversation to offer in 7–10 days for the right candidate.

What We Will NOT Accept

  • Take-home assignments produced by other AI tools. We will ask you to walk through code live; if you can't explain decisions you didn't make, the interview ends.
  • Resume claims you can't substantiate. We verify employment via EPFO/UAN where applicable.
  • "I've done LLM chatbots" experience without production voice agent experience. They are not the same skill.
  • Engagement structures that conflict with this being your primary work.

How to Apply

Send a brief message with:

  1. A link to a production voice AI system you've shipped (or, if NDA'd, a description of the architecture and your specific contribution).
  2. Two or three sentences on why this role fits you specifically — not a generic cover letter.
  3. Your current compensation expectation in USD per month.
  4. Your earliest start date and current time-zone availability.

We read every message. We respond within 5 business days, including with "no" answers. We don't ghost.


Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About AIspire360 Inc

Founded :
2025
Type :
Products & Services
Size :
0-20
Stage :
Bootstrapped

About

N/A

Company social profiles

N/A

Similar jobs (10)

company logo
Aswathy Vimal
Posted by Aswathy Vimal
Remote only
6 - 10 yrs
₹24L - ₹36L / yr
skill iconPython
skill iconJavascript
TypeScript
Large Language Models (LLM) tuning

The Role

We’re looking for a Senior Applied AI & Data Engineer to become our first dedicated AI and data engineer.

You’ll build conversational AI experiences across web, mobile, and in-store channels while developing the data foundation behind them. You’ll make key technical decisions and own your work through to production.

What You’ll Do

• Build AI assistants using tool calling to work with real product, search, and order systems

• Design guardrails and evaluation sets to ensure AI responses are accurate and safe

• Build real-time and voice-enabled AI experiences

• Improve product data quality through AI-assisted enrichment and review workflows

• Build data pipelines, analytics, and personalisation systems

• Work closely with web and mobile developers and help guide technical implementation

What You’ll Need

• 6+ years of experience building and running production backend systems

• Strong Python skills, plus experience with JavaScript/TypeScript backends

• Experience shipping at least one LLM-powered feature to real users

• Experience with search and relevance

• Experience building data pipelines and analytics stores

• Comfortable deploying and monitoring services on a major cloud platform

• Strong communication skills and the ability to work independently

Nice to Have

• Experience with speech or voice AI

• E-commerce or retail technology experience

• Experience building multilingual products

Read more
A design and development studio
A design and development studio
Agency job
via by Ariba Khan
Gurugram
8 - 15 yrs
Best in industry
skill iconPython
Generative AI (GenAI)
Large Language Models (LLM)
Voice Agent
Langchain
+1 more

About the Company

The client is revolutionising the way businesses operate through cutting-edge technological solutions. Their focus is on developing intelligent agents and agentic workflows that automate processes and eliminate the need for human effort wherever possible. By leveraging

advanced AI and machine learning, they create systems that enhance productivity and drive efficiency.

Their expertise extends to the fintech, healthcare and medical technology sectors, where they develop innovative solutions that improve patient outcomes and streamline medical operations.


From medical devices to healthcare platforms, their work sits at the intersection of technology and medicine, pushing the boundaries of what's possible. The team is dedicated to continuous learning and growth, ensuring the team members are always at the forefront of the tech landscape.


About the Role

This is a senior, hands-on engineering role at the heart of our product team. You will be one of the most technical people in the room — setting the architecture for our real-time voice AI

agents and building the hardest parts of it yourself. From the systems that power live conversations to the interfaces our clients rely on, you will own how the product is engineered end to end.


We are looking for a genuine lead full-stack engineer with the depth to make architecture decisions that hold up as we scale, and the appetite to still be in the code every day. You should be as comfortable designing the backend services behind a live voice agent as you are shaping a clean interface on top of them — and comfortable being the person others turn to when something is hard.


You will work directly with the founder and product leadership on a fast-moving product, with real influence over technical direction. This is a role for someone who wants ownership at the level of "how the whole thing is built," not just individual features — and who raises the bar for

everyone around them.


What You'll Own

Set the technical direction

  • Own the architecture of our core systems — the real-time voice agents, backend
  • services, data and APIs — making the decisions that keep the product fast, reliable and scalable as it grows.
  • Lead the hardest engineering problems and solve them personally.
  • Establish engineering standards — code quality, review practices, testing and technical patterns that the team builds to.
  • Drive technical strategy with the founder and product leadership — shaping the roadmap, flagging risk early, and turning product ambition into a sound technical plan.

Build the product end to end

  • Design, build and ship features across the stack — backend services, APIs and front-ends — owning them from idea to production.
  • Build the client-facing surfaces — dashboards, review tools and configuration interfaces that let our clients run and trust the product.
  • Design and evolve the data models and APIs that hold up as we scale across clients.

Make it reliable and fast

  • Own production quality — put the monitoring and alerting in place so issues are caught before clients feel them, and performance stays within target.
  • Care about performance — find and fix bottlenecks across the stack.
  • Build for correctness — put the testing and evaluation in place that keeps the product behaving predictably as it changes.

Lead through the team

  • Mentor and grow engineers — through code review, pairing, and setting a technical example others learn from.
  • Multiply the team's output — unblock others and lift the overall quality of the codebase.
  • Take features from ambiguity to done — turn a rough product goal into a shipped, working capability with minimal hand-holding, and help others do the same.


What We're Looking For

  • 8+ years of professional software engineering experience, with significant depth across backend and a track record of owning systems, not just features.
  • Strong backend engineering, ideally in Python — building and scaling production services and APIs..
  • Proven architecture and system-design ability — you have designed systems that scaled, and can reason clearly about trade-offs.
  • Solid fundamentals across APIs, databases and cloud infrastructure.
  • Experience building real-time and/or AI-powered products — or clear, demonstrable ability to lead in this area.
  • A history of technical leadership — setting standards, mentoring engineers, and being trusted with the hardest problems — while remaining hands-on.
  • Excellent communication and a genuine ownership mindset — someone who can be handed an ambiguous, high-stakes problem and be trusted to see it through.

Nice to Have

  • Experience working with AI / large language models in production.
  • Experience with voice or other real-time products.
  • Exposure to healthcare, fintech, or other regulated / high-stakes domains.
  • Experience as an early or senior engineer in a startup, where you set direction and wore many hats.
Read more
company logo
Agency job
via by Vrishali Mishra
Bengaluru (Bangalore)
4 - 6 yrs
₹20L - ₹40L / yr
skill iconPython
Speech-to-Text (STT)
ASR
Text-to-Speech (TTS)
Large Language Models (LLM)
+3 more

About Us

Invorto is our Voice AI product, bringing intelligent voice agents to real-world customer and operational use cases. Our voice pipeline is built in Python, running an STT → LLM → TTS architecture on top of the Pipecat framework.

This is a chance to work on hard problems in voice AI — latency, accuracy, naturalness, and reliability — building zero-to-one, owning your area end-to-end, and shipping to production at scale.

Note: This is a customer-facing role, and strong communication skills are essential.

About the Role

We're looking for a Voice AI Research Engineer to join the Invorto team and help build and continuously improve the voice AI systems that power our intelligent voice agents. This role is focused on the specialized craft of voice AI — designing evaluation and automation frameworks that ensure our STT, LLM, and TTS pipeline performs reliably in real-world, production conditions.

 

What You'll Do

  • Design and build automated testing and quality frameworks for our STT → LLM → TTS voice pipeline, built on Pipecat
  • Evaluate and benchmark STT, LLM, and TTS/ASR components on accuracy, latency, naturalness, and robustness across accents, languages, and real-world audio conditions
  • Work hands-on with STT, TTS, and ASR models — fine-tuning, evaluating, and improving them for production use cases
  • Identify failure modes and edge cases across the pipeline (background noise, accents, interruptions, turn-taking, latency, pipeline-stage handoffs) and build systems to catch them before production
  • Collaborate closely with engineering to integrate quality checks and automation into the voice agent development lifecycle within the Pipecat-based architecture
  • Research and stay current with advances in voice AI, and bring in new techniques, models, and tools to improve pipeline performance
  • Work directly with customers to understand real-world voice use cases and translate them into evaluation criteria and quality benchmarks
  • Partner with product and engineering to define what "production-grade quality" means for voice agents and drive the team toward it

 

What We're Looking For

  • 4–6 years of experience, with a specialization in voice AI systems and automated quality evaluation
  • Hands-on experience with STT (Speech-to-Text), TTS (Text-to-Speech), and ASR (Automatic Speech Recognition) models
  • Experience designing and building automated testing/evaluation frameworks for voice or speech systems
  • Strong understanding of what drives voice AI quality — accuracy, latency, naturalness, and robustness to real-world variability
  • Strong programming skills in Python; familiarity with Pipecat or similar voice pipeline/orchestration frameworks is a plus
  • Understanding of STT → LLM → TTS pipeline architectures and the trade-offs involved at each stage
  • Research mindset — comfortable exploring new models, techniques, and tools and translating them into practical improvements
  • Excellent communication skills — this is a customer-facing role, and you'll regularly engage directly with customers to understand needs and validate quality expectations


Read more
company logo
Deepak Sharma
Posted by Deepak Sharma
Remote, Los Angeles California, Florida, New York, Virginia
1 - 5 yrs
$1.5K - $4.5K / yr
skill iconPython
skill iconNodeJS (Node.js)
skill iconGo Programming (Golang)
SIP
Web Realtime Communication (WebRTC)
+7 more

We’re building a powerful, AI-driven communication platform — a next-generation alternative to RingCentral or 8x8 — powered by OpenAI, LangChain, and SIP/WebRTC. We're looking for a Full-Stack Software Developer who’s passionate about building real-time, AI-enabled voice infrastructure and who’s excited to work in a fast-moving, founder-led environment.

This is an opportunity to build from scratch, take ownership of core systems, and innovate on the edge of VoIP + AI.

What You’ll Do

  • Design and build AI-driven voice and messaging features (e.g. smart IVRs, call transcription, virtual agents)
  • Develop backend services using Python, Node.js, or Golang
  • Integrate OpenAI, Whisper, and LangChain with real-time VoIP systems like Twilio, SIP, or WebRTC
  • Create scalable APIs, handle call logic, and build AI pipelines
  • Collaborate with the founder and early team on product strategy and infrastructure
  • Participate in occasional in-person strategy meetings (Delhi, Bangalore, or nearby)

Must-Have Skills

  • Strong programming experience in Python, Node.js, or Go
  • Hands-on experience with VoIP/SIP, WebRTC, or tools like Twilio, Asterisk, Plivo
  • Experience integrating with LLM APIs, OpenAI, or speech-to-text models
  • Solid understanding of backend design, Docker, Redis, PostgreSQL
  • Ability to work independently and deliver production-grade code

Nice to Have

  • Familiarity with LangChain or agent-based AI systems
  • Knowledge of call routing logic, STUN/TURN, or media servers (e.g. FreeSWITCH)
  • Interest in building scalable cloud-first SaaS products

Work Setup

  • 🏠 Remote work
  • 🕐 Full-time role
  • 💼 Direct collaboration with founder (technical)
  • 🧘‍♂️ Flexible hours, strong ownership culture
Read more
Neosapien
Neosapien
Agency job
via by Nehlata Pandey
Bengaluru (Bangalore)
3 - 7 yrs
₹15L - ₹40L / yr (ESOP available)
skill iconPython
Large Language Models (LLM) tuning
Agentic AI
Retrieval Augmented Generation (RAG)
skill iconMachine Learning (ML)
+2 more

The Role

You own AI systems end to end. From the speech-to-text models that turn audio into text, to the diarization that separates and identifies speakers, to the agentic layer that turns conversation into memory and action, to the observability and evaluation that keep all of it honest in production. This is a wide role by design. You will own model selection, serving, and production reliability. If you want to tune one model and ignore the system around it, this is not the role.

What You Will Own

•     Speech-to-text. Evaluate, integrate, and optimize STT models across cloud and self-hosted. Drive accuracy and cost trade-offs with ground-truth metrics.

•     Speaker diarization and identification. Push accuracy on hard, real-world, multi-speaker audio.

•     Agentic AI. Build the memory and retrieval pipeline, LLM orchestration, and the agent workflows that sit on top of captured conversation.

•     Model serving and infrastructure. Stand up and optimize self-hosted serving (vLLM, Triton class). Own latency, throughput, and cost per user.

Observability

An always-on wearable means models run in production every second, on messy real-world audio. You own the visibility into that.

•     Instrument the full audio-to-memory pipeline: STT, diarization, retrieval, and LLM calls.

•     Define and track model-quality SLOs in production: transcription drift, diarization error over time, retrieval relevance, latency, throughput, and cost per user.

•     Build dashboards and alerting so model degradation is caught before users feel it.

•     Trace failures across a distributed, always-on system using metrics, logs, and traces.

•     Close the loop. Production signals feed back into evaluation and model selection.

Evaluation

We do not ship what we cannot measure. You own the systems that prove a model is actually better, not just newer.

•     Build and own ground-truth evaluation harnesses for every model in the stack.

•     Measure with real metrics: WER for transcription, DER for diarization, Recall and F1 for retrieval and speaker identification.

•     Build and maintain labeled benchmark datasets that reflect real, messy, multi-speaker audio.

•     Run regression and A/B evaluations on every model swap, prompt change, or pipeline update. Nothing ships on a vibe.

•     Reject anecdotal proxies, single confidence scores, and cherry-picked examples as evidence of quality.

What We Are Looking For

•     3 to 5 years as an AI/ML engineer with production systems behind you. Engineering and production experience is non-negotiable.

•     Depth across the modern AI stack: LLMs, speech models, vector retrieval, model serving.

•     Strong software engineering. You write code that ships and survives contact with real users.

•     Fluency in Python and the production ML ecosystem.

•     Comfort with cloud infrastructure (GCP a plus) and containerized deployment on Kubernetes.

•     A working command of observability and evaluation. You measure first and trust metrics over intuition.

•     First-principles reasoning and metric discipline.

Nice to Have

•     Research background or publications. A strong signal, not a substitute for production work.

•     Audio and speech ML experience (STT, diarization, voice).

•     Experience self-hosting and optimizing open models.

•     Experience with LLM gateway and agent orchestration patterns.

•     Experience building eval harnesses or production model-monitoring systems.


Requirements

Agentic work is must. Audio is good to have

. Self hosting models is a must

 Experience with LLM gateway and agent orchestration is a must have

Read more
company logo
Umama Sayed
Posted by Umama Sayed
Mumbai
5 - 8 yrs
Best in industry
skill iconPython
Large Language Models (LLM)
Artificial Intelligence (AI)
Prompt engineering
LangGraph
+6 more

Senior AI Engineer

Code Generation, Agent Architecture & LLM Systems

📍 Mumbai (On-site) | Full-time | 5+ years


About the Role:

Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.

We are hiring a Senior AI Engineer for a dedicated client engagement focused on building an AI-powered application builder platform - a product where users describe software in plain English and the system generates, previews, and iteratively refines working code.

The mandatory requirement for this role is hands-on production experience shipping LLM-powered systems with agent architectures, with experience in code generation or developer tooling contexts a strong advantage.


The role is product-focused and deeply hands-on. You will own everything between the user's prompt and correct code landing in the project: the agentic loop, code generation pipeline, context management, evaluation suite, and model cost strategy.

You will work alongside the Senior MLOps Engineer who operationalises the infrastructure around your system, and collaborate closely with backend, frontend, and DevOps engineers.


Responsibilities:


Agent Architecture

Design and own the agentic loop for the platform - request interpretation, planning, tool-calling sequence (read file, edit file, run build, search code, install package), and stop conditions.

Make and revisit architectural decisions on single-agent vs. multi-agent designs, including planner/executor splits and dedicated build-repair sub-agents.


Code Generation Pipeline

Own the end-to-end generation flow: task classification, context gathering, planning, targeted edits, verification, and commit.

Implement diff/search-replace-based file editing with fuzzy matching and fallback strategies.

Enforce scope discipline so the agent makes minimal diffs and does not modify code it was not asked to touch.


Self-Repair Loop

Build and tune the automated repair loop that pipes compiler, lint, build, and runtime errors back to the model with retry budgets and model escalation.

This loop is the primary quality lever - the difference between 60-70% and 90%+ build success rates.


Context Management

Build file-relevance retrieval so the agent sees the right files, not the whole codebase: dependency graphs, AST/tree-sitter-based chunking, embeddings, recency signals, and hybrid retrieval.

Implement conversation summarisation and memory for long sessions, and address long-project degradation through codebase summaries and periodic consistency passes.

Own token budgeting and prompt caching strategy.


Prompt Engineering as a Discipline

Own the system prompt and per-task prompt variants (new feature, bug fix, styling change).

Maintain few-shot examples and enforce coding conventions, stack rules, and prohibited behaviours such as no hardcoded secrets and no whole-file rewrites.

Version prompts like code with changelogs and rollback capability.


Evaluation and Quality Measurement

Design and own the evaluation suite: representative test prompts run on every prompt and model change, scored on build success rate, instruction adherence, and output quality including LLM-as-judge and visual/screenshot checks where relevant.

Define regression gates that block quality-degrading changes from shipping.

Treat evals the way engineers treat automated testing: versioned, automated, and tracked over time.

This responsibility is non-negotiable at this level.


Model Strategy and Cost

Design model routing - cheap and fast models for classification and small edits, frontier models for complex generation.

Drive cost optimisation through prompt caching, diff-based edits over full-file rewrites, and tighter context selection.

Track cost per agent run and tokens per task; evaluate new model releases against the eval suite and lead migrations when results justify it.


Safety and Reliability of Agent Behaviour

Defend against prompt injection from user content and fetched web content.

Ensure secrets never appear in generated client code.

Define what the agent's tools may and may not do in collaboration with the platform team.

Contribute to output moderation and abuse-pattern awareness.


Mentorship and Engineering Standards

Run code reviews, define engineering conventions for AI work, and raise the engineering bar across the AI team.

Work closely with the Senior MLOps Engineer on handoff of eval design, prompt configurations, and model routing logic.


Requirements:


Hands-on Production Ownership of LLM-Powered Systems with Agent Architectures (Mandatory)

Must have personally shipped and operated at least one complex production AI system - agentic, multi-step, or code generation - with end-to-end ownership of architecture, evaluation, and cost.

POCs, internal demos, and tutorial-grade work do not qualify.


5+ Years of Professional Software or AI Engineering Experience

With at least 3 years focused on LLM applications, AI engineering, or production AI systems.

Candidates with strong backend backgrounds and a clear, substantive pivot into LLM systems qualify.


Strong Python Proficiency and Service Development

Production-grade Python with FastAPI or equivalent: type hints, async patterns, streaming responses, testing, and packaging.

Not notebook-only.


Depth Across LLM APIs and Agent Systems

Production experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or open-weight models (vLLM, Ollama, Together).

Production experience with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.

Hands-on with tool calling, structured outputs, and multi-step reasoning.


Demonstrated, Systematic Evaluation Practice - Non-Negotiable

Must have built evaluation harnesses that gate production releases, not ad-hoc testing.

Hands-on with at least one of LangSmith, Langfuse, Promptfoo, Ragas, or DeepEval.

Candidates with no systematic answer to evaluation should not be considered at senior level regardless of other strengths.


Cost Discipline for Production AI

Track record of measurable cost optimisation on production AI features.

Able to speak in specifics: cost per request, savings achieved through caching or model routing, context reduction decisions.


AWS Working Knowledge

Hands-on with EC2, S3, IAM, and Docker.

Comfort with CI/CD workflows and deploying AI services.


Awareness of LLM Security Failure Modes

Familiar with prompt injection patterns, understands that system prompt rules alone are insufficient, and has experience with output validation and content safety in production.


Nice to Have

  • Experience with AST/tree-sitter tooling, diff-based editing systems, or compiler-adjacent work
  • MCP server authoring
  • Open-source AI contributions
  • Published technical writing on LLM systems
  • Multi-modal model experience
  • Fine-tuning exposure (LoRA, QLoRA, PEFT)
Read more
Leadsquared
Leadsquared
Agency job
via by Vrishali Mishra
Bengaluru (Bangalore)
2 - 4 yrs
₹25L - ₹45L / yr
Large Language Models (LLM) tuning

About LeadSquared

LeadSquared is a leading sales execution and marketing automation platform trusted by 2,000+ businesses globally, including healthcare, education, financial services, and real estate. Headquartered in Bengaluru with offices across the US, UK, UAE, and Southeast Asia, we empower sales teams to close faster, smarter, and at scale.

Our AI team is at the forefront of integrating cutting-edge large language model capabilities into enterprise workflows — building intelligent agents, copilots, and automation systems that redefine how businesses operate.

Role Overview

We are looking for a Senior AI Engineer with hands-on experience building LLM-powered agents and agentic AI systems. You will design, develop, and deploy autonomous AI pipelines that solve complex, multi-step business problems — from lead qualification and follow-up automation to intelligent CRM workflows and beyond.

This role is ideal for someone who is deeply excited about the frontier of AI, can move fast, and wants their work to directly impact millions of sales professionals worldwide.

Key Responsibilities

•

Design and build LLM-powered agentic systems using frameworks such as LangChain, LlamaIndex, AutoGen, or CrewAI to automate complex, multi-step workflows.

•

Develop and maintain Retrieval-Augmented Generation (RAG) pipelines with vector databases (Pinecone, Weaviate, Chroma, pgvector) for domain-specific knowledge grounding.

•

Build and integrate tool-use and function-calling capabilities into AI agents, enabling dynamic interaction with internal APIs, databases, and third-party services.

•

Implement prompt engineering strategies including chain-of-thought, few-shot prompting, and structured output parsing to ensure reliable agent behavior.

•

Design evaluation frameworks and observability pipelines (LangSmith, Helicone, custom metrics) to monitor agent performance, accuracy, and cost.

•

Collaborate with product, sales, and domain teams to translate business requirements into AI-driven solutions and features.

•

Optimize LLM inference for latency and cost using techniques like caching, model distillation, quantization, and batching.

•

Stay current with the rapidly evolving LLM ecosystem and proactively propose improvements and new approaches.

•

Contribute to internal best practices, documentation, and knowledge-sharing across the engineering org.

Required Qualifications

Experience

•

2–4 years of professional software engineering experience, with at least 1–2 years focused on LLM/AI systems.

•

Proven experience shipping LLM-based products or agentic AI systems into production environments.

Technical Skills

•

Strong proficiency in Python and familiarity with async programming patterns for AI pipelines.

•

Hands-on experience with LLM APIs: OpenAI (GPT-4o), Anthropic (Claude), Google (Gemini), or open-source models (Llama, Mistral).

•

Experience with agentic frameworks: LangChain, LangGraph, LlamaIndex, AutoGen, CrewAI, or similar.

•

Solid understanding of RAG architectures, embedding models, and semantic search.

•

Experience with vector databases and similarity search infrastructure.

•

Knowledge of REST APIs, microservices architecture, and containerization (Docker/Kubernetes).

Problem-Solving & Mindset

•

Strong ability to decompose ambiguous, open-ended problems into structured AI system designs.

•

Experience with prompt debugging, LLM evaluation, and iterative refinement workflows.

•

Ability to balance research exploration with engineering pragmatism to ship reliable systems.

Preferred Qualifications

•

Experience with multi-agent orchestration and agent memory systems (short-term and long-term).

•

Familiarity with fine-tuning or RLHF workflows for domain adaptation.

•

Background in NLP, information retrieval, or conversational AI.

•

Prior experience in B2B SaaS or CRM domain is a plus.

•

Contributions to open-source AI/ML projects or published research/blogs.

•

Experience with cloud platforms: AWS, GCP, or Azure — particularly AI/ML services

Read more
company logo
Remote only
5 - 10 yrs
₹35L - ₹65L / yr
Systems design
APIs & Integration
Data Infrastructure
TypeScript
skill iconReact.js
+4 more

Procedure is hiring for WorkHero.


WorkHero is building the AI-powered back office for the skilled trades, starting with the $50B+ HVAC industry. Small contractors are great at their trade but lose 20+ hours a week to invoicing, permits, scheduling, and paperwork. WorkHero combines expert office managers with automation and AI tooling, enabling a small team to take real ownership of that back-office work


We’re hiring a senior engineer to own our real-time voice stack end to end—AI agents operating on live phone calls—and the data platform that turns those calls into insight: call → transcript → events → warehouse → dashboards. You’ll own meaningful systems end to end alongside a small, senior team with deep experience in AI, product, and the trades.


What you’ll build

  • New product screens and flows (jobs, customers, invoices, scheduling) in React and React Native, especially AI chat UI (chat & tool result rendering, streaming responses, human review and feedback loops)
  • AI workflows in production: tool-using agents, RAG/search, classification/extraction, and human-in-the-loop flows
  • Automations: Contribute new features and improvements to our AI-powered business automation platform


In addition, you’ll own our first investments into a realtime voice stack and the call-data platform behind it. For example:

  • Realtime voice agents on live phone calls: telephony/WebRTC integration, streaming speech-to-text and text-to-speech, turn-taking, interruption handling, and latency optimization
  • Voice pipeline reliability: backpressure, failover, graceful degradation, and monitoring for live calls
  • Call-data pipeline: transcripts, events, and structured extraction flowing from every call into the warehouse
  • Analytics & dashboards: data modeling and conversation-intelligence features on top of call data
  • Evals & monitoring for voice agents: quality metrics, drift detection, and cost/latency tracking
  • Cloud infrastructure: scaling our platform with infrastructure as code, queues and orchestration, and CI/CD


Responsibilities

  • Analyze requirements and propose innovative AI-native solutions to technical problems
  • Write clean scalable code
  • Own the voice and data stack end-to-end: design, build, test, deploy, and operate
  • Optimize the performance, latency, and cost of our real-time AI systems
  • Respond to critical system issues and ensure continuous system reliability
  • Mentor team members and collaborate across teams, especially with product and subject matter experts
  • Work to understand the needs of our users and think creatively about how to solve design challenges in your work
  • This is a Remote role. We expect a minimum 4 hours overlap with the WorkHero team (11 AM - 3 PM ET).


Qualifications

  • Senior-level backend experience (typically 5+ years) shipping production systems that you've owned
  • Hands-on experience with realtime voice or streaming systems: telephony (SIP/Twilio), WebRTC, streaming STT/TTS, or frameworks like LiveKit or Pipecat — or comparable experience with demanding realtime/streaming infrastructure
  • Data engineering fundamentals: event pipelines, data modeling, warehousing, and analytics on production data
  • Strong proficiency in a typed backend language (TypeScript preferred; comparable experience welcome)
  • Hands-on experience with LLM-powered features (usage, prompting, optimization, etc) and AI architectures
  • The ability to work with infrastructure as code (terraform), cloud, and CI/CD systems at scale. We're a small team, so we own the whole stack!
  • Excitement to leverage AI coding tools to their maximum benefit. We love Claude Code and Cursor and are constantly looking for better ways to leverage our time to build fast and build for scale.


Nice to have

  • experience with voice-AI platforms (Vapi, Retell, Bland, Deepgram, LiveKit) or conversation-intelligence products (e.g. Gong-style analytics)
  • experience scaling cloud infrastructure, especially AWS, and how to get the most out of key AWS services
  • experience with workflow automation tools like n8n or Lindy
  • experience with React for building internal dashboards
  • experience with HVAC or back-office business workflows


WorkHero is committed to building a diverse team. We encourage candidates from all backgrounds to apply.

Read more
Service Co
Service Co
Agency job
via by Rishika Teja
Pune, Mumbai
5 - 10 yrs
₹15L - ₹40L / yr
Artificial Intelligence (AI)
Generative AI
Large Language Models (LLM)
Retrieval Augmented Generation (RAG)

Hiring for AI Engineer


Exp: 5 - 10 yrs

Edu : BE/B.Tech/MCA

Work Location : Pune / Mumbai


Skill Set:


Total experience ranging from 5–10 years in software engineering/AI roles

Min 5 years strong programming experience in Python is a MUST

Min 3.5 years hands-on experience in AI with LLMs, RAG pipelines, and AI frameworks

2+ years shipping LLM systems in production

Experience with cloud platforms (AWS/Azure/GCP)

Read more
company logo
Shefali Gupta
Posted by Shefali Gupta
Remote, Delhi, Gurugram, Noida, Ghaziabad, Faridabad, Bengaluru (Bangalore)
2 - 10 yrs
₹5L - ₹15L / yr
skill iconAmazon Web Services (AWS)
Google Cloud Platform (GCP)
skill iconDocker
API
skill iconFlask
+4 more

Job Title: Senior AI/ML Engineer

Company: Timble Technologies Pvt. Ltd

Location: Gurugram (Hybrid)

Experience: 2 TO 5 Years


About Us

Timble Glance is a high-growth AI RegTech and B2B SaaS company catering to top-tier BFSI and enterprise clients. We build cutting-edge systems powering 30+ high-scale APIs for digital identity verification, fraud detection, document intelligence, and compliance automation.

Role Overview

We are looking for a hands-on Senior AI/ML Engineer to design, develop, and productionize high-throughput AI/ML and Generative AI systems. You will own the full lifecycle—from problem formulation and data pipelines to deep learning architectures, RAG systems, LLMOps, and model governance—delivering sub-second latency and high reliability across our enterprise products.


Key Responsibilities


·       Model Architecture & Deployment: Design, train, and deploy production-scale ML/Deep Learning and GenAI systems (computer vision, document intelligence, OCR, NLP, fraud risk classification, and LLM applications).

·       GenAI & LLM Solutions: Develop robust LLM workflows including prompt engineering, fine-tuning, RAG pipelines, semantic search, vector indexing (Pinecone/Milvus/Chroma), and safety guardrails.

·       Pipelines & Engineering: Build performant feature extraction and data pipelines; write modular, vectorized, production-grade Python (NumPy, Pandas) and advanced SQL.

·       MLOps & Monitoring: Establish end-to-end MLOps/LLMOps standards—model registries, CI/CD, experiment tracking, drift detection, A/B testing, latency optimization, and cost governance.

·       Responsible AI & Security: Ensure model decisions comply with enterprise data security, privacy standards, and auditability required by the BFSI sector.

·       Collaboration & Ownership: Translate complex business requirements into technical roadmaps, conduct rigorous code reviews, and mentor junior engineers.


Required Qualifications & Skills


·       Education: B.Tech / M.Tech in Computer Science, AI/ML, Mathematics, or a related field—Tier-1 institutes (IIT, IIIT, NIT) strongly preferred.

·       Experience: 2+ years of hands-on experience developing, deploying, and maintaining ML/Deep Learning or GenAI models in production environments.

·       GenAI & NLP Stack: Hands-on experience with LLMs, embeddings, RAG architectures, and frameworks such as LangChain, LlamaIndex, or Hugging Face.

·       Deep Learning Frameworks: Strong proficiency in PyTorch or TensorFlow, with deep knowledge of transformer architectures and modern NLP/CV models.

·       Software & Data Engineering: Expert-level Python skills (pytest, Git, OOP, asynchronous programming), solid SQL proficiency, and familiarity with data workflows.

·       Deployment & Cloud: Practical exposure to cloud platforms (AWS/GCP), containerization (Docker), API frameworks (FastAPI/Flask), and basic orchestration (Kubernetes).


Preferred Qualifications

·       Prior domain experience in Fintech, RegTech, Identity Verification (KYC/AML), Fraud Intelligence, or B2B SaaS.

·       Experience optimizing models for low latency and inference cost (e.g., ONNX, TensorRT, model quantization).

·       Familiarity with workflow orchestrators such as Airflow, Prefect, or Kubeflow.

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos