Cutshort logo
For Employers

AI Architect at VMax eSolutions India Pvt Ltd · Hyderabad · 10 - 15 years · ₹35L - ₹45L / yr · Profitable · Posted 25 Dec 2025

VMax eSolutions India Pvt Ltd's logo

AI Architect

Bachu Sai Nikheel's profile picture
Posted by Bachu Sai Nikheel
10 - 15 yrs
₹35L - ₹45L / yr
Hyderabad
Skills
Generative AI
PEFT (Parameter-Efficient Fine-Tuning)
Voice processing
Artificial Intelligence (AI)
GPU computing
Retrieval Augmented Generation (RAG)
skill iconMachine Learning (ML)
Large Language Models (LLM) tuning

We are seeking an experienced AI Architect to design, build, and scale production-ready AI voice conversation agents deployed locally (on-prem / edge / private cloud) and optimized for GPU-accelerated, high-throughput environments.

You will own the end-to-end architecture of real-time voice systems, including speech recognition, LLM orchestration, dialog management, speech synthesis, and low-latency streaming pipelines—designed for reliability, scalability, and cost efficiency.

This role is highly hands-on and strategic, bridging research, engineering, and production infrastructure.


Key Responsibilities

Architecture & System Design

  • Design low-latency, real-time voice agent architectures for local/on-prem deployment
  • Define scalable architectures for ASR → LLM → TTS pipelines
  • Optimize systems for GPU utilization, concurrency, and throughput
  • Architect fault-tolerant, production-grade voice systems (HA, monitoring, recovery)

Voice & Conversational AI

  • Design and integrate:
  • Automatic Speech Recognition (ASR)
  • Natural Language Understanding / LLMs
  • Dialogue management & conversation state
  • Text-to-Speech (TTS)
  • Build streaming voice pipelines with sub-second response times
  • Enable multi-turn, interruptible, natural conversations

Model & Inference Engineering

  • Deploy and optimize local LLMs and speech models (quantization, batching, caching)
  • Select and fine-tune open-source models for voice use cases
  • Implement efficient inference using TensorRT, ONNX, CUDA, vLLM, Triton, or similar

Infrastructure & Production

  • Design GPU-based inference clusters (bare metal or Kubernetes)
  • Implement autoscaling, load balancing, and GPU scheduling
  • Establish monitoring, logging, and performance metrics for voice agents
  • Ensure security, privacy, and data isolation for local deployments

Leadership & Collaboration

  • Set architectural standards and best practices
  • Mentor ML and platform engineers
  • Collaborate with product, infra, and applied research teams
  • Drive decisions from prototype → production → scale

Required Qualifications

Technical Skills

  • 7+ years in software / ML systems engineering
  • 3+ years designing production AI systems
  • Strong experience with real-time voice or conversational AI systems
  • Deep understanding of LLMs, ASR, and TTS pipelines
  • Hands-on experience with GPU inference optimization
  • Strong Python and/or C++ background
  • Experience with Linux, Docker, Kubernetes

AI & ML Expertise

  • Experience deploying open-source LLMs locally
  • Knowledge of model optimization:
  • Quantization
  • Batching
  • Streaming inference
  • Familiarity with voice models (e.g., Whisper-like ASR, neural TTS)

Systems & Scaling

  • Experience with high-QPS, low-latency systems
  • Knowledge of distributed systems and microservices
  • Understanding of edge or on-prem AI deployments

Preferred Qualifications

  • Experience building AI voice agents or call automation systems
  • Background in speech processing or audio ML
  • Experience with telephony, WebRTC, SIP, or streaming audio
  • Familiarity with Triton Inference Server / vLLM
  • Prior experience as Tech Lead or Principal Engineer

What We Offer

  • Opportunity to architect state-of-the-art AI voice systems
  • Work on real-world, high-scale production deployments
  • Competitive compensation and equity (if applicable)
  • High ownership and technical influence
  • Collaboration with top-tier AI and infrastructure talent
Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About VMax eSolutions India Pvt Ltd

Founded :
2008
Type :
Services
Size :
100-1000
Stage :
Profitable

About

VMAX is an ISO 90012015 and ISO 27001:2013 certified organisation based in Hyderabad, India. It builds custom, tailor-made, and scalable products that can be easily integrated with third-party systems. VMAX provides cutting-edge solutions to its clients and has helped organizations emerge as winners in their respective industries.
Read more

Company social profiles

twitter

Similar jobs (10)

company logo
Stuti Jain
Posted by Stuti Jain
Hyderabad
7 - 10 yrs
₹25L - ₹35L / yr
Retrieval Augmented Generation (RAG)
skill iconAmazon Web Services (AWS)

Location: Hyderabad, India. Based at the KnackLabs headquarters, with occasional travel to client locations for workshops and reviews. This role does not involve extended onsite deployments.

About the Role

You will work as an AI Architect who designs the systems behind our client engagements: AI agents, RAG systems, automation platforms, and the conventional backend systems around them.

This is a hands-on design role, not a slideware role. You will scope architectures with clients, make the hard technical decisions, defend them in review, and stay accountable for how the systems perform in production.


You will work directly with clients. Everyone at KnackLabs does. You will sit in design discussions with client engineering teams, present architecture decisions to technical and business stakeholders, and answer for the choices you make.


A full KnackLabs engineering team in Hyderabad builds with you. You own the technical design and the quality of what ships.

What you'll own

  1. Architecture - Design AI agents, RAG systems, integrations, and the scalable backend systems around them, for multiple client engagements.
  2. Technical scoping - Work directly with clients to turn a business problem into a system design, with clear trade-offs and clear reasons.
  3. Scale and reliability - Make sure what we build handles real load: data stores, queues, caching, horizontal scaling, and fault tolerance.
  4. Design reviews - Review designs and builds across engagements. Set the technical bar and hold it.
  5. Evaluation strategy - Define how we measure accuracy, safety, latency, and cost for the AI systems we ship.
  6. Guiding engineers - Raise the level of the engineers building with you, through reviews and direct pairing.
  7. Feedback to the platform - Feed what you learn across engagements back into our platform and internal tools.

What we are looking for

  1. Around 7 or more years of software engineering experience, including direct work with customers on design or delivery.
  2. Full-stack development experience with strength in backend technologies.
  3. Experience designing and building scalable applications. You understand how large-scale distributed systems work: data partitioning, queues, caching, horizontal scaling, and fault tolerance.
  4. At least 2 years of strong, hands-on AI experience with large language models in production.
  5. You build with AI coding tools like Claude Code or Codex as your default way of working. You understand Claude Skills, have written skills yourself, use them actively, and have contributed to them.
  6. Hands-on experience building retrieval-augmented generation (RAG) systems: chunking, embeddings, vector databases, retrieval, and reranking.
  7. Hands-on experience building AI agents.
  8. Strong programming skills in Python. Working knowledge of TypeScript or JavaScript.
  9. Experience with at least one cloud platform (AWS, Azure, or GCP).
  10. Clear communication. You can explain an architecture decision to an engineer and to a business leader, and defend it under questioning.
  11. High ownership and comfort with ambiguity. You can take an unclear problem and turn it into a design.

Nice to have

  1. Experience building evaluations to measure accuracy, safety, latency, and cost.
  2. Experience with observability and tracing tools such as LangSmith or Braintrust.
  3. Experience with on-premises or private cloud (VPC) deployments.
  4. Experience deploying AI systems in regulated industries such as insurance, banking, or the public sector.
  5. Experience with data engineering and pipelines.
  6. A history of side projects, open source contributions, or products you shipped end-to-end.
  7. Experience working at a consulting or professional services firm in a client-facing delivery role.

Stack and tools

  1. Languages: Python and TypeScript.
  2. Models: Claude and other frontier or open-source models, chosen to fit the customer.
  3. AI patterns: RAG, agents, prompt engineering, skills, and evaluations.
  4. Vector and retrieval: vector databases and retrieval pipelines.
  5. Cloud: AWS, Azure, or GCP, on public or private cloud.
  6. Integration: REST APIs and enterprise system connectors.


Read more
company logo
Umama Sayed
Posted by Umama Sayed
Mumbai
5 - 8 yrs
Best in industry
skill iconPython
Large Language Models (LLM)
Artificial Intelligence (AI)
Prompt engineering
LangGraph
+6 more

Senior AI Engineer

Code Generation, Agent Architecture & LLM Systems

📍 Mumbai (On-site) | Full-time | 5+ years


About the Role:

Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.

We are hiring a Senior AI Engineer for a dedicated client engagement focused on building an AI-powered application builder platform - a product where users describe software in plain English and the system generates, previews, and iteratively refines working code.

The mandatory requirement for this role is hands-on production experience shipping LLM-powered systems with agent architectures, with experience in code generation or developer tooling contexts a strong advantage.


The role is product-focused and deeply hands-on. You will own everything between the user's prompt and correct code landing in the project: the agentic loop, code generation pipeline, context management, evaluation suite, and model cost strategy.

You will work alongside the Senior MLOps Engineer who operationalises the infrastructure around your system, and collaborate closely with backend, frontend, and DevOps engineers.


Responsibilities:


Agent Architecture

Design and own the agentic loop for the platform - request interpretation, planning, tool-calling sequence (read file, edit file, run build, search code, install package), and stop conditions.

Make and revisit architectural decisions on single-agent vs. multi-agent designs, including planner/executor splits and dedicated build-repair sub-agents.


Code Generation Pipeline

Own the end-to-end generation flow: task classification, context gathering, planning, targeted edits, verification, and commit.

Implement diff/search-replace-based file editing with fuzzy matching and fallback strategies.

Enforce scope discipline so the agent makes minimal diffs and does not modify code it was not asked to touch.


Self-Repair Loop

Build and tune the automated repair loop that pipes compiler, lint, build, and runtime errors back to the model with retry budgets and model escalation.

This loop is the primary quality lever - the difference between 60-70% and 90%+ build success rates.


Context Management

Build file-relevance retrieval so the agent sees the right files, not the whole codebase: dependency graphs, AST/tree-sitter-based chunking, embeddings, recency signals, and hybrid retrieval.

Implement conversation summarisation and memory for long sessions, and address long-project degradation through codebase summaries and periodic consistency passes.

Own token budgeting and prompt caching strategy.


Prompt Engineering as a Discipline

Own the system prompt and per-task prompt variants (new feature, bug fix, styling change).

Maintain few-shot examples and enforce coding conventions, stack rules, and prohibited behaviours such as no hardcoded secrets and no whole-file rewrites.

Version prompts like code with changelogs and rollback capability.


Evaluation and Quality Measurement

Design and own the evaluation suite: representative test prompts run on every prompt and model change, scored on build success rate, instruction adherence, and output quality including LLM-as-judge and visual/screenshot checks where relevant.

Define regression gates that block quality-degrading changes from shipping.

Treat evals the way engineers treat automated testing: versioned, automated, and tracked over time.

This responsibility is non-negotiable at this level.


Model Strategy and Cost

Design model routing - cheap and fast models for classification and small edits, frontier models for complex generation.

Drive cost optimisation through prompt caching, diff-based edits over full-file rewrites, and tighter context selection.

Track cost per agent run and tokens per task; evaluate new model releases against the eval suite and lead migrations when results justify it.


Safety and Reliability of Agent Behaviour

Defend against prompt injection from user content and fetched web content.

Ensure secrets never appear in generated client code.

Define what the agent's tools may and may not do in collaboration with the platform team.

Contribute to output moderation and abuse-pattern awareness.


Mentorship and Engineering Standards

Run code reviews, define engineering conventions for AI work, and raise the engineering bar across the AI team.

Work closely with the Senior MLOps Engineer on handoff of eval design, prompt configurations, and model routing logic.


Requirements:


Hands-on Production Ownership of LLM-Powered Systems with Agent Architectures (Mandatory)

Must have personally shipped and operated at least one complex production AI system - agentic, multi-step, or code generation - with end-to-end ownership of architecture, evaluation, and cost.

POCs, internal demos, and tutorial-grade work do not qualify.


5+ Years of Professional Software or AI Engineering Experience

With at least 3 years focused on LLM applications, AI engineering, or production AI systems.

Candidates with strong backend backgrounds and a clear, substantive pivot into LLM systems qualify.


Strong Python Proficiency and Service Development

Production-grade Python with FastAPI or equivalent: type hints, async patterns, streaming responses, testing, and packaging.

Not notebook-only.


Depth Across LLM APIs and Agent Systems

Production experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or open-weight models (vLLM, Ollama, Together).

Production experience with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.

Hands-on with tool calling, structured outputs, and multi-step reasoning.


Demonstrated, Systematic Evaluation Practice - Non-Negotiable

Must have built evaluation harnesses that gate production releases, not ad-hoc testing.

Hands-on with at least one of LangSmith, Langfuse, Promptfoo, Ragas, or DeepEval.

Candidates with no systematic answer to evaluation should not be considered at senior level regardless of other strengths.


Cost Discipline for Production AI

Track record of measurable cost optimisation on production AI features.

Able to speak in specifics: cost per request, savings achieved through caching or model routing, context reduction decisions.


AWS Working Knowledge

Hands-on with EC2, S3, IAM, and Docker.

Comfort with CI/CD workflows and deploying AI services.


Awareness of LLM Security Failure Modes

Familiar with prompt injection patterns, understands that system prompt rules alone are insufficient, and has experience with output validation and content safety in production.


Nice to Have

  • Experience with AST/tree-sitter tooling, diff-based editing systems, or compiler-adjacent work
  • MCP server authoring
  • Open-source AI contributions
  • Published technical writing on LLM systems
  • Multi-modal model experience
  • Fine-tuning exposure (LoRA, QLoRA, PEFT)
Read more
A design and development studio
A design and development studio
Agency job
via by Ariba Khan
Gurugram
8 - 15 yrs
Best in industry
skill iconPython
Generative AI (GenAI)
Large Language Models (LLM)
Voice Agent
Langchain
+1 more

About the Company

The client is revolutionising the way businesses operate through cutting-edge technological solutions. Their focus is on developing intelligent agents and agentic workflows that automate processes and eliminate the need for human effort wherever possible. By leveraging

advanced AI and machine learning, they create systems that enhance productivity and drive efficiency.

Their expertise extends to the fintech, healthcare and medical technology sectors, where they develop innovative solutions that improve patient outcomes and streamline medical operations.


From medical devices to healthcare platforms, their work sits at the intersection of technology and medicine, pushing the boundaries of what's possible. The team is dedicated to continuous learning and growth, ensuring the team members are always at the forefront of the tech landscape.


About the Role

This is a senior, hands-on engineering role at the heart of our product team. You will be one of the most technical people in the room — setting the architecture for our real-time voice AI

agents and building the hardest parts of it yourself. From the systems that power live conversations to the interfaces our clients rely on, you will own how the product is engineered end to end.


We are looking for a genuine lead full-stack engineer with the depth to make architecture decisions that hold up as we scale, and the appetite to still be in the code every day. You should be as comfortable designing the backend services behind a live voice agent as you are shaping a clean interface on top of them — and comfortable being the person others turn to when something is hard.


You will work directly with the founder and product leadership on a fast-moving product, with real influence over technical direction. This is a role for someone who wants ownership at the level of "how the whole thing is built," not just individual features — and who raises the bar for

everyone around them.


What You'll Own

Set the technical direction

  • Own the architecture of our core systems — the real-time voice agents, backend
  • services, data and APIs — making the decisions that keep the product fast, reliable and scalable as it grows.
  • Lead the hardest engineering problems and solve them personally.
  • Establish engineering standards — code quality, review practices, testing and technical patterns that the team builds to.
  • Drive technical strategy with the founder and product leadership — shaping the roadmap, flagging risk early, and turning product ambition into a sound technical plan.

Build the product end to end

  • Design, build and ship features across the stack — backend services, APIs and front-ends — owning them from idea to production.
  • Build the client-facing surfaces — dashboards, review tools and configuration interfaces that let our clients run and trust the product.
  • Design and evolve the data models and APIs that hold up as we scale across clients.

Make it reliable and fast

  • Own production quality — put the monitoring and alerting in place so issues are caught before clients feel them, and performance stays within target.
  • Care about performance — find and fix bottlenecks across the stack.
  • Build for correctness — put the testing and evaluation in place that keeps the product behaving predictably as it changes.

Lead through the team

  • Mentor and grow engineers — through code review, pairing, and setting a technical example others learn from.
  • Multiply the team's output — unblock others and lift the overall quality of the codebase.
  • Take features from ambiguity to done — turn a rough product goal into a shipped, working capability with minimal hand-holding, and help others do the same.


What We're Looking For

  • 8+ years of professional software engineering experience, with significant depth across backend and a track record of owning systems, not just features.
  • Strong backend engineering, ideally in Python — building and scaling production services and APIs..
  • Proven architecture and system-design ability — you have designed systems that scaled, and can reason clearly about trade-offs.
  • Solid fundamentals across APIs, databases and cloud infrastructure.
  • Experience building real-time and/or AI-powered products — or clear, demonstrable ability to lead in this area.
  • A history of technical leadership — setting standards, mentoring engineers, and being trusted with the hardest problems — while remaining hands-on.
  • Excellent communication and a genuine ownership mindset — someone who can be handed an ambiguous, high-stakes problem and be trusted to see it through.

Nice to Have

  • Experience working with AI / large language models in production.
  • Experience with voice or other real-time products.
  • Exposure to healthcare, fintech, or other regulated / high-stakes domains.
  • Experience as an early or senior engineer in a startup, where you set direction and wore many hats.
Read more
The industry’s only Manufacturing Operating System
The industry’s only Manufacturing Operating System
Agency job
via by Ariba Khan
Hyderabad
10 - 15 yrs
Best in industry
Artificial Intelligence (AI)
Generative AI (GenAI)
Retrieval Augmented Generation (RAG)
Large Language Models (LLM)

We’re on hunt for AI Architect


Responsibilities:

  • 10–15+ years overall experience, with recent hands-on AI/GenAI architecture ownership.
  • Must have architected enterprise AI platforms/solutions end-to-end, not just individual ML models or PoCs.
  • Strong GenAI/LLM production experience: RAG, embeddings, vector DBs, hybrid search, reranking, evaluation, guardrails.
  • Strong Agentic AI understanding: agents, tool calling, workflows, orchestration, human-in-the-loop.
  • Experience taking AI solutions from architecture → production → scale, ideally across multiple business teams/use cases.
  • Strong cloud architecture — Azure/AWS preferred; hybrid/on-prem experience is a plus.
  • Must understand enterprise security, governance, Responsible AI, observability and LLMOps/MLOps.
  • Should be able to articulate build-vs-buy, MVP-vs-target architecture, cost/performance/security tradeoffs.
  • Strong stakeholder-facing / consulting ability — can work with business leaders, engineering, security and data teams and influence without authority.


There is scope to move to the US for this role if you are aligned for the same, else this will be a WFO role from Hyderabad location

Read more
Neosapien
Neosapien
Agency job
via by Nehlata Pandey
Bengaluru (Bangalore)
3 - 7 yrs
₹15L - ₹40L / yr (ESOP available)
skill iconPython
Large Language Models (LLM) tuning
Agentic AI
Retrieval Augmented Generation (RAG)
skill iconMachine Learning (ML)
+2 more

The Role

You own AI systems end to end. From the speech-to-text models that turn audio into text, to the diarization that separates and identifies speakers, to the agentic layer that turns conversation into memory and action, to the observability and evaluation that keep all of it honest in production. This is a wide role by design. You will own model selection, serving, and production reliability. If you want to tune one model and ignore the system around it, this is not the role.

What You Will Own

•     Speech-to-text. Evaluate, integrate, and optimize STT models across cloud and self-hosted. Drive accuracy and cost trade-offs with ground-truth metrics.

•     Speaker diarization and identification. Push accuracy on hard, real-world, multi-speaker audio.

•     Agentic AI. Build the memory and retrieval pipeline, LLM orchestration, and the agent workflows that sit on top of captured conversation.

•     Model serving and infrastructure. Stand up and optimize self-hosted serving (vLLM, Triton class). Own latency, throughput, and cost per user.

Observability

An always-on wearable means models run in production every second, on messy real-world audio. You own the visibility into that.

•     Instrument the full audio-to-memory pipeline: STT, diarization, retrieval, and LLM calls.

•     Define and track model-quality SLOs in production: transcription drift, diarization error over time, retrieval relevance, latency, throughput, and cost per user.

•     Build dashboards and alerting so model degradation is caught before users feel it.

•     Trace failures across a distributed, always-on system using metrics, logs, and traces.

•     Close the loop. Production signals feed back into evaluation and model selection.

Evaluation

We do not ship what we cannot measure. You own the systems that prove a model is actually better, not just newer.

•     Build and own ground-truth evaluation harnesses for every model in the stack.

•     Measure with real metrics: WER for transcription, DER for diarization, Recall and F1 for retrieval and speaker identification.

•     Build and maintain labeled benchmark datasets that reflect real, messy, multi-speaker audio.

•     Run regression and A/B evaluations on every model swap, prompt change, or pipeline update. Nothing ships on a vibe.

•     Reject anecdotal proxies, single confidence scores, and cherry-picked examples as evidence of quality.

What We Are Looking For

•     3 to 5 years as an AI/ML engineer with production systems behind you. Engineering and production experience is non-negotiable.

•     Depth across the modern AI stack: LLMs, speech models, vector retrieval, model serving.

•     Strong software engineering. You write code that ships and survives contact with real users.

•     Fluency in Python and the production ML ecosystem.

•     Comfort with cloud infrastructure (GCP a plus) and containerized deployment on Kubernetes.

•     A working command of observability and evaluation. You measure first and trust metrics over intuition.

•     First-principles reasoning and metric discipline.

Nice to Have

•     Research background or publications. A strong signal, not a substitute for production work.

•     Audio and speech ML experience (STT, diarization, voice).

•     Experience self-hosting and optimizing open models.

•     Experience with LLM gateway and agent orchestration patterns.

•     Experience building eval harnesses or production model-monitoring systems.


Requirements

Agentic work is must. Audio is good to have

. Self hosting models is a must

 Experience with LLM gateway and agent orchestration is a must have

Read more
company logo
Sandeep C
Posted by Sandeep C
Bengaluru (Bangalore)
8 - 16 yrs
₹1L - ₹2L / yr (ESOP available)
Large Language Models (LLM)
Agentic AI
Applied mathematics

Key Responsibilities:

·      Architectural Leadership: Design and lead the development of robust, scalable AI architectures, ensuring high performance, reliability, and security.

·      Applied Mathematics & Statistics: Apply statistical analysis, numerical computation, and mathematical modeling to derive insights from large-scale data and optimize model performance.

·      Deep Learning Development: Design, train, and deploy advanced Deep Learning (DL) models.

·      Technical Mentorship: Mentor engineering teams on best practices for AI/ML, coding standards, and architectural design.

·      Model Optimization: Optimize models for speed, efficiency, and accuracy using techniques like pruning, quantization, or GPU acceleration.

·      Strategy & Innovation: Evaluate and select appropriate AI frameworks, tools, and platforms, staying abreast of cutting-edge research and industry trends.

Qualifications:

Required:

·      Education: Master's or PhD in Computer Science, Applied Mathematics, Statistics, Physics, or a related quantitative field.

·      Experience: 10+ years of experience in software development, with at least 3-5 years in a Applied Mathematics and Deep learning.

·      AI/ML Expertise: Proven experience designing and deploying deep learning models in production using frameworks.

·      Mathematics/Statistics: Strong proficiency in linear algebra, calculus, probability, and statistical methods.

·      Programming Skills: Expert-level coding skills in Python (NumPy, Pandas, Scikit-learn) and experience with languages like Java or C++.

Key Competencies:

  • Strategic mindset with deep operational awareness.
  • Excellent communication and stakeholder management skills.
  • Ability to simplify complex technical concepts for executive reporting.
  • Strong leadership, people development, and cross-functional influencing skills.

Bias for action and a relentless focus on continuous improvement.

Read more
Service Co
Service Co
Agency job
via by Rishika Teja
Pune, Mumbai
5 - 10 yrs
₹15L - ₹40L / yr
Artificial Intelligence (AI)
Generative AI
Large Language Models (LLM)
Retrieval Augmented Generation (RAG)

Hiring for AI Engineer


Exp: 5 - 10 yrs

Edu : BE/B.Tech/MCA

Work Location : Pune / Mumbai


Skill Set:


Total experience ranging from 5–10 years in software engineering/AI roles

Min 5 years strong programming experience in Python is a MUST

Min 3.5 years hands-on experience in AI with LLMs, RAG pipelines, and AI frameworks

2+ years shipping LLM systems in production

Experience with cloud platforms (AWS/Azure/GCP)

Read more
Leadsquared
Leadsquared
Agency job
via by Vrishali Mishra
Bengaluru (Bangalore)
2 - 4 yrs
₹25L - ₹45L / yr
Large Language Models (LLM) tuning

About LeadSquared

LeadSquared is a leading sales execution and marketing automation platform trusted by 2,000+ businesses globally, including healthcare, education, financial services, and real estate. Headquartered in Bengaluru with offices across the US, UK, UAE, and Southeast Asia, we empower sales teams to close faster, smarter, and at scale.

Our AI team is at the forefront of integrating cutting-edge large language model capabilities into enterprise workflows — building intelligent agents, copilots, and automation systems that redefine how businesses operate.

Role Overview

We are looking for a Senior AI Engineer with hands-on experience building LLM-powered agents and agentic AI systems. You will design, develop, and deploy autonomous AI pipelines that solve complex, multi-step business problems — from lead qualification and follow-up automation to intelligent CRM workflows and beyond.

This role is ideal for someone who is deeply excited about the frontier of AI, can move fast, and wants their work to directly impact millions of sales professionals worldwide.

Key Responsibilities

•

Design and build LLM-powered agentic systems using frameworks such as LangChain, LlamaIndex, AutoGen, or CrewAI to automate complex, multi-step workflows.

•

Develop and maintain Retrieval-Augmented Generation (RAG) pipelines with vector databases (Pinecone, Weaviate, Chroma, pgvector) for domain-specific knowledge grounding.

•

Build and integrate tool-use and function-calling capabilities into AI agents, enabling dynamic interaction with internal APIs, databases, and third-party services.

•

Implement prompt engineering strategies including chain-of-thought, few-shot prompting, and structured output parsing to ensure reliable agent behavior.

•

Design evaluation frameworks and observability pipelines (LangSmith, Helicone, custom metrics) to monitor agent performance, accuracy, and cost.

•

Collaborate with product, sales, and domain teams to translate business requirements into AI-driven solutions and features.

•

Optimize LLM inference for latency and cost using techniques like caching, model distillation, quantization, and batching.

•

Stay current with the rapidly evolving LLM ecosystem and proactively propose improvements and new approaches.

•

Contribute to internal best practices, documentation, and knowledge-sharing across the engineering org.

Required Qualifications

Experience

•

2–4 years of professional software engineering experience, with at least 1–2 years focused on LLM/AI systems.

•

Proven experience shipping LLM-based products or agentic AI systems into production environments.

Technical Skills

•

Strong proficiency in Python and familiarity with async programming patterns for AI pipelines.

•

Hands-on experience with LLM APIs: OpenAI (GPT-4o), Anthropic (Claude), Google (Gemini), or open-source models (Llama, Mistral).

•

Experience with agentic frameworks: LangChain, LangGraph, LlamaIndex, AutoGen, CrewAI, or similar.

•

Solid understanding of RAG architectures, embedding models, and semantic search.

•

Experience with vector databases and similarity search infrastructure.

•

Knowledge of REST APIs, microservices architecture, and containerization (Docker/Kubernetes).

Problem-Solving & Mindset

•

Strong ability to decompose ambiguous, open-ended problems into structured AI system designs.

•

Experience with prompt debugging, LLM evaluation, and iterative refinement workflows.

•

Ability to balance research exploration with engineering pragmatism to ship reliable systems.

Preferred Qualifications

•

Experience with multi-agent orchestration and agent memory systems (short-term and long-term).

•

Familiarity with fine-tuning or RLHF workflows for domain adaptation.

•

Background in NLP, information retrieval, or conversational AI.

•

Prior experience in B2B SaaS or CRM domain is a plus.

•

Contributions to open-source AI/ML projects or published research/blogs.

•

Experience with cloud platforms: AWS, GCP, or Azure — particularly AI/ML services

Read more
company logo
Rishu Dutta
Posted by Rishu Dutta
Gurugram
7 - 12 yrs
₹20L - ₹50L / yr
Retrieval Augmented Generation (RAG)
Agentic AI
Multi-agent Systems

Role Overview 

We are looking for an AI Engineer to design, build, and ship production AI systems, including agentic AI applications, for enterprise clients. This is a hands-on engineering role: you will write production code, build and evaluate models and agents, and work closely with architects and product teams to take solutions from prototype to scale. 


Key Responsibilities 

Design and build agentic AI systems: agent workflows, tool/function-calling, memory, and human-in-the-loop patterns. Build and productionise RAG pipelines, prompt-based applications, and LLM integrations across providers. Develop and maintain data and ML pipelines: feature engineering, model training, evaluation, and monitoring. Integrate AI systems with enterprise applications (CRMs, ERPs, ITSM tools) via APIs, events, and MCP-based tool servers. Implement guardrails, prompt-injection defences, and evaluation frameworks to keep AI systems safe and reliable in production. 

Write clean, tested, production-grade code and participate actively in code and design reviews. 

Collaborate with architects, product managers, and delivery teams to translate requirements into working AI solutions. Troubleshoot and optimise AI systems for accuracy, latency, and cost in production. 


Required Qualifications 

8–12 years of hands-on software engineering experience, with a strong, unbroken technical track record. Hands-on experience building and shipping AI/ML systems in production, not just POCs. 

Practical experience with agentic AI systems and at least one major agent framework (LangGraph, CrewAI, AutoGen, OpenAI Agents SDK, Bedrock Agents/Strands, or Semantic Kernel). 

Experience with LLM/GenAI systems: RAG pipelines, prompt engineering, structured outputs, and tool calling across providers. 

Strong Python skills (TypeScript/Node.js a plus), with production-grade testing, CI/CD, and API design practices. Working knowledge of ML fundamentals: model evaluation, feature engineering, and experimentation. Cloud-native experience on AWS and/or Azure: containers, serverless, event backbones, and vector databases. Understanding of LLM safety and reliability practices: guardrails, prompt-injection defences, and observability. 



Read more
New York, Los Angeles California
3 - 5 yrs
$2.5K - $5.5K / yr
skill iconPython
Artificial Intelligence (AI)
skill iconMachine Learning (ML)
Multi-Agent System
Full Stack Development
+17 more

We are building an advanced, AI-driven multi-agent software system designed to revolutionize task automation and code generation. This is a futuristic AI platform capable of:


✅ Real-time self-coding based on tasks  

✅ Autonomous multi-agent collaboration  

✅ AI-powered decision-making  

✅ Cross-platform compatibility (Desktop, Web, Mobile)  


We are hiring a highly skilled **AI Engineer & Full-Stack Developer** based in India, with a strong background in AI/ML, multi-agent architecture, and scalable, production-grade software development.


### Responsibilities:


- Build and maintain a multi-agent AI system (AutoGPT, BabyAGI, MetaGPT concepts)  

- Integrate large language models (GPT-4o, Claude, open-source LLMs)  

- Develop full-stack components (Backend: Python, FastAPI/Flask, Frontend: React/Next.js)  

- Work on real-time task execution pipelines  

- Build cross-platform apps using Electron or Flutter  

- Implement Redis, Vector databases, scalable APIs  

- Guide the architecture of autonomous, self-coding AI systems  


### Must-Have Skills:


- Python (advanced, AI applications)  

- AI/ML experience, including multi-agent orchestration  

- LLM integration knowledge  

- Full-stack development: React or Next.js  

- Redis, Vector Databases (e.g., Pinecone, FAISS)  

- Real-time applications (websockets, event-driven)  

- Cloud deployment (AWS, GCP)  


### Good to Have:


- Experience with code-generation AI models (Codex, GPT-4o coding abilities)  

- Microservices and secure system design  

- Knowledge of AI for workflow automation and productivity tools  


Join us to work on cutting-edge AI technology that builds the future of autonomous software.

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos