Lead AI Engineer at Talent Pro · Remote only · 3 - 6 years · ₹30L - ₹40L / yr · Bootstrapped · Remote only · Posted 27 Mar 2025

Need Excellent Communication skills as the company is dealing with US Clients also
• 3+ years in AI development, with experience in multi-agent systems, logistics, or related fields.
• Proven experience in conducting A/B testing and beta testing for AI systems.
• Hands-on experience with CrewAI and LangChain tools.
• Should have hands-on experience working with end-to-end chatbot development, specifically with Agentic and RAG-based chatbots. It is essential that the candidate has been involved in the entire lifecycle of chatbot creation, from design to deployment.
• Should have practical experience with LLM application deployment.
• Proficiency in Python and machine learning frameworks (e.g., TensorFlow, PyTorch).
• Experience in setting up monitoring dashboards with tools like Grafana, Tableau, or similar.
• Proficiency with cloud platforms (AWS, Azure, GCP)

Similar jobs (10)
Job Description:
We are looking for a hands-on AI Engineer with experience in Generative AI and Agentic AI to build and deploy production-ready AI solutions.
Key Responsibilities:
- Develop and deploy GenAI and Agentic AI applications.
- Build RAG pipelines, LLM workflows, and AI agents.
- Develop solutions using Python, LangChain, LangGraph, LlamaIndex, or similar frameworks.
- Implement tool calling, context retrieval, and LLM orchestration.
- Integrate AI solutions with APIs and cloud platforms.
- Work with AWS/Azure/GCP, Docker, and CI/CD.
Required Skills:
- Strong Python programming skills.
- 3+ years of GenAI/Agentic AI experience.
- RAG and LLM orchestration.
- LangChain / LangGraph / LlamaIndex / AutoGen / CrewAI / Semantic Kernel.
- MCP and A2A knowledge.
- Cloud, APIs, Docker, and CI/CD experience.
Preferred Experience:
Hands-on experience building and deploying production-ready AI solutions.
Key Skills:
• Agentic AI / AI Agents
• Python or Java
• LLMs & Generative AI
• RAG & Vector Databases
• LangChain / LangGraph / AutoGen / CrewAI
• REST APIs & Microservices
• Prompt Engineering
• AI Workflow Automation
Roles & Responsibilities:
• Design and develop AI agents and agentic workflows
• Build scalable backend services using Python or Java
• Integrate LLMs, APIs, tools, and external systems into AI workflows
• Develop RAG-based solutions and intelligent automation
• Design multi-step AI workflows with tool/function calling
• Evaluate, monitor, and optimize AI agent performance
• Collaborate with engineering and product teams to deliver production-ready AI solutions
Experience - 4 to 6 year
Location – Ahmedabad/Pune/Indore
- Additional Job Description
Additional Job Description
Required Skills and Experience:
- Strong proficiency in Python and experience with ML/AI libraries (scikit-learn, TensorFlow, PyTorch, Hugging Face ecosystem).
- Hands-on experience with LLMs, RAG, vector databases, and retrieval pipelines.
- Practical experience deploying agentic workflows and building multi-step, tool-enabled agents.
- Experience using Garak (or similar LLM red-teaming/vulnerability scanners) to identify model weaknesses and harden deployments.
- Demonstrated experience implementing content filtering / moderation systems.
- Solid skills working with structured and unstructured data and advanced feature engineering.
- Familiarity with cloud GenAI platforms and services (Azure AI Services preferred; AWS/GCP acceptable).
- Experience building APIs/microservices; containerization (Docker), orchestration (Kubernetes).
- Strong understanding of model evaluation, performance profiling, inference cost optimization, and observability.
- Good knowledge of security, data governance, and privacy best practices for AI systems.
Role: AI Developer
Experience: 3–4 Years
Employment Type: Full-Time
Location: Goregaon, Mumbai
About the Role
We are looking for an experienced AI Developer with 3–4 years of software development experience and strong hands-on exposure to Generative AI, AI Agents, Copilots, and AI-powered application development.
The candidate will be responsible for building production-ready AI solutions, developing agentic workflows, modernizing legacy applications, and integrating LLM capabilities into enterprise applications.
Key Responsibilities
- Design, develop, and deploy AI Agents and agentic workflows for enterprise use cases.
- Build AI Copilots and LLM-powered applications using modern AI frameworks and APIs.
- Develop RAG-based applications using embeddings, vector databases, and enterprise data.
- Work on legacy application migration and modernization, leveraging AI-assisted development and code transformation techniques.
- Analyze legacy codebases and design strategies for AI-driven migration, refactoring, and modernization.
- Integrate LLMs with enterprise applications, APIs, databases, and third-party systems.
- Implement tool calling, function calling, multi-agent workflows, and workflow automation.
- Perform prompt engineering, context optimization, model evaluation, and AI application testing.
- Take ownership of AI solutions from POC and prototyping through production deployment.
- Collaborate with product managers, architects, and engineering teams to convert business requirements into scalable AI solutions.
- Stay updated with emerging technologies in Generative AI, Agentic AI, LLMs, and AI-assisted software development.
Required Skills
- 3–4 years of professional software development experience.
- Strong proficiency in Python and/or JavaScript/TypeScript.
- Hands-on experience developing Generative AI / LLM-based applications.
- Strong understanding of AI Agents, RAG, Prompt Engineering, LLM APIs, and embeddings.
- Experience with frameworks such as LangChain, LangGraph, Semantic Kernel, AutoGen, or equivalent.
- Experience working with REST APIs, databases, Git, and cloud environments.
- Hands-on experience with vector databases such as Pinecone, Weaviate, Chroma, FAISS, or equivalent.
- Good understanding of software architecture, debugging, testing, and deployment practices.
Good to Have
- Experience with Microsoft Copilot / Copilot Studio.
- Experience working with Claude, OpenAI, Gemini, Azure OpenAI, or open-source LLMs.
- Experience in legacy application migration, modernization, or code conversion.
- Knowledge of Azure AI / AWS / Google Cloud AI services.
- Experience with MCP, multi-agent systems, tool calling, and AI orchestration.
- Experience building enterprise-grade AI solutions with focus on security, scalability, and performance.
Hiring for AI Engineer
Exp: 6 - 12 yrs
Edu : BE/B.Tech/MCA
Work Location : Pune / Mumbai
Skill Set:
Total experience ranging from 5–10 years in software engineering/AI roles
Min 5 years strong programming experience in Python or Typescript is a MUST
Min 2.5 years hands-on experience in AI with LLMs, RAG pipelines, and AI frameworks
2+ years shipping LLM systems in production
Experience with cloud platforms (AWS/Azure/GCP)
Senior AI Engineer
Code Generation, Agent Architecture & LLM Systems
📍 Mumbai (On-site) | Full-time | 5+ years
About the Role:
Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.
We are hiring a Senior AI Engineer for a dedicated client engagement focused on building an AI-powered application builder platform - a product where users describe software in plain English and the system generates, previews, and iteratively refines working code.
The mandatory requirement for this role is hands-on production experience shipping LLM-powered systems with agent architectures, with experience in code generation or developer tooling contexts a strong advantage.
The role is product-focused and deeply hands-on. You will own everything between the user's prompt and correct code landing in the project: the agentic loop, code generation pipeline, context management, evaluation suite, and model cost strategy.
You will work alongside the Senior MLOps Engineer who operationalises the infrastructure around your system, and collaborate closely with backend, frontend, and DevOps engineers.
Responsibilities:
Agent Architecture
Design and own the agentic loop for the platform - request interpretation, planning, tool-calling sequence (read file, edit file, run build, search code, install package), and stop conditions.
Make and revisit architectural decisions on single-agent vs. multi-agent designs, including planner/executor splits and dedicated build-repair sub-agents.
Code Generation Pipeline
Own the end-to-end generation flow: task classification, context gathering, planning, targeted edits, verification, and commit.
Implement diff/search-replace-based file editing with fuzzy matching and fallback strategies.
Enforce scope discipline so the agent makes minimal diffs and does not modify code it was not asked to touch.
Self-Repair Loop
Build and tune the automated repair loop that pipes compiler, lint, build, and runtime errors back to the model with retry budgets and model escalation.
This loop is the primary quality lever - the difference between 60-70% and 90%+ build success rates.
Context Management
Build file-relevance retrieval so the agent sees the right files, not the whole codebase: dependency graphs, AST/tree-sitter-based chunking, embeddings, recency signals, and hybrid retrieval.
Implement conversation summarisation and memory for long sessions, and address long-project degradation through codebase summaries and periodic consistency passes.
Own token budgeting and prompt caching strategy.
Prompt Engineering as a Discipline
Own the system prompt and per-task prompt variants (new feature, bug fix, styling change).
Maintain few-shot examples and enforce coding conventions, stack rules, and prohibited behaviours such as no hardcoded secrets and no whole-file rewrites.
Version prompts like code with changelogs and rollback capability.
Evaluation and Quality Measurement
Design and own the evaluation suite: representative test prompts run on every prompt and model change, scored on build success rate, instruction adherence, and output quality including LLM-as-judge and visual/screenshot checks where relevant.
Define regression gates that block quality-degrading changes from shipping.
Treat evals the way engineers treat automated testing: versioned, automated, and tracked over time.
This responsibility is non-negotiable at this level.
Model Strategy and Cost
Design model routing - cheap and fast models for classification and small edits, frontier models for complex generation.
Drive cost optimisation through prompt caching, diff-based edits over full-file rewrites, and tighter context selection.
Track cost per agent run and tokens per task; evaluate new model releases against the eval suite and lead migrations when results justify it.
Safety and Reliability of Agent Behaviour
Defend against prompt injection from user content and fetched web content.
Ensure secrets never appear in generated client code.
Define what the agent's tools may and may not do in collaboration with the platform team.
Contribute to output moderation and abuse-pattern awareness.
Mentorship and Engineering Standards
Run code reviews, define engineering conventions for AI work, and raise the engineering bar across the AI team.
Work closely with the Senior MLOps Engineer on handoff of eval design, prompt configurations, and model routing logic.
Requirements:
Hands-on Production Ownership of LLM-Powered Systems with Agent Architectures (Mandatory)
Must have personally shipped and operated at least one complex production AI system - agentic, multi-step, or code generation - with end-to-end ownership of architecture, evaluation, and cost.
POCs, internal demos, and tutorial-grade work do not qualify.
5+ Years of Professional Software or AI Engineering Experience
With at least 3 years focused on LLM applications, AI engineering, or production AI systems.
Candidates with strong backend backgrounds and a clear, substantive pivot into LLM systems qualify.
Strong Python Proficiency and Service Development
Production-grade Python with FastAPI or equivalent: type hints, async patterns, streaming responses, testing, and packaging.
Not notebook-only.
Depth Across LLM APIs and Agent Systems
Production experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or open-weight models (vLLM, Ollama, Together).
Production experience with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.
Hands-on with tool calling, structured outputs, and multi-step reasoning.
Demonstrated, Systematic Evaluation Practice - Non-Negotiable
Must have built evaluation harnesses that gate production releases, not ad-hoc testing.
Hands-on with at least one of LangSmith, Langfuse, Promptfoo, Ragas, or DeepEval.
Candidates with no systematic answer to evaluation should not be considered at senior level regardless of other strengths.
Cost Discipline for Production AI
Track record of measurable cost optimisation on production AI features.
Able to speak in specifics: cost per request, savings achieved through caching or model routing, context reduction decisions.
AWS Working Knowledge
Hands-on with EC2, S3, IAM, and Docker.
Comfort with CI/CD workflows and deploying AI services.
Awareness of LLM Security Failure Modes
Familiar with prompt injection patterns, understands that system prompt rules alone are insufficient, and has experience with output validation and content safety in production.
Nice to Have
- Experience with AST/tree-sitter tooling, diff-based editing systems, or compiler-adjacent work
- MCP server authoring
- Open-source AI contributions
- Published technical writing on LLM systems
- Multi-modal model experience
- Fine-tuning exposure (LoRA, QLoRA, PEFT)
AI Engineer
LLMs, Agents & AI Services
📍 Mumbai (On-site) | Full-time | 2-4 years
About the Role:
Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.
AI is core to how we design, deliver, and scale software for our customers.
We are hiring an AI Engineer for a dedicated client engagement building a complex production AI platform, working on the AI capabilities and agentic features at the core of the product.
The mandatory requirement for this role is at least one AI feature personally shipped to production for real users, with operational ownership.
The role suits someone who thinks quickly on solutioning, can take an ambiguous problem to a working prototype in days, and has the discipline to carry it through to production with predictable economics.
You will work alongside the Senior AI Engineer and the wider pod, with ownership of parts of the AI surface area of the product.
Responsibilities:
Solutioning and POCs
Translate ambiguous customer problems into working POCs at speed.
Pick the right model, framework, and architecture, and demonstrate value early before scaling investment.
LLM Application Development
Build AI features and services using LLM APIs from OpenAI, Anthropic, Google, and self-hosted open-weight models (Llama, Qwen, Mistral).
Choose the right model per use case based on cost, latency, capability, and context-window trade-offs.
Agentic System Design
Design and implement agentic workflows using LangGraph, CrewAI, AutoGen, LlamaIndex Agents, or custom orchestration.
Cover tool use, planning, memory, and multi-step reasoning appropriate to the problem.
API and Service Development
Build production AI services and APIs using Python and FastAPI.
Handle streaming responses, async processing, structured outputs, retries, and graceful degradation when models or tools fail.
Retrieval and Tool Integration
Implement RAG pipelines with vector databases (Pinecone, Weaviate, Qdrant, pgvector, Chroma), embeddings, chunking strategies, hybrid search, and reranking.
Integrate external tools, internal APIs, and document sources through tool-calling and MCP-style patterns.
Cost Analysis and Unit Economics
Model the per-request and per-user cost of every AI feature before it ships.
Track token usage, prompt caching, batching, and model-routing strategies.
Drive measurable improvements in unit economics.
Production Hardening
Add observability and tracing (LangSmith, Langfuse, OpenTelemetry), guardrails, content safety checks, prompt injection defences, and fallback behaviour.
Prompt Engineering and Evaluation
Design, test, and iterate prompts with measured outcomes.
Build evaluation harnesses for accuracy, hallucination, latency, and cost.
Run benchmarks across models and prompt variants before locking in a design.
Requirements:
AI Feature Shipped to Production (Mandatory)
Must have personally built and shipped at least one AI feature that runs in production for real users, with operational ownership.
POCs, internal demos, and one-off scripts do not qualify.
2 to 4 Years of Professional Software or AI Engineering Experience
With at least one production AI feature owned end to end.
Strong Python Proficiency and API Development with FastAPI
Comfort with type hints, async, packaging, testing, streaming responses, and authentication.
Production-grade Python, not notebook-only code.
Hands-on Depth Across the LLM and Agent Stack
Working experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or self-hosted open-weight models (vLLM, Ollama, Together, Replicate).
Working familiarity with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.
Working knowledge of RAG, embeddings, and vector databases (Pinecone, Weaviate, Qdrant, pgvector, Chroma).
Solutioning Speed and POC Velocity
Demonstrated ability to move from a fuzzy problem to a working prototype in days.
Strong instinct for what to build first, what to defer, and what to throw away.
Cost Discipline for Production AI
Ability to calculate, monitor, and optimise the cost of LLM APIs, tokens, embeddings, vector store usage, and infrastructure.
Treats unit economics as a first-class concern.
AWS Familiarity
Working knowledge of EC2, S3, IAM, and at least one of Bedrock, SageMaker, or equivalent.
Comfortable in a Fast-Moving Environment
Self-directed, comfortable with ambiguity, takes ownership without being asked, and ships under shifting priorities.
Strong Written and Spoken English Communication
Able to explain trade-offs to non-AI engineers, designers, product managers, and clients in plain language.
Nice to Have
- fine-tuning or LoRA, QLoRA, PEFT exposure
- MCP server authoring
- eval framework experience (LangSmith, Promptfoo, Ragas, DeepEval)
- open-source AI contributions
- multi-modal models (vision, audio)
The Role
You own AI systems end to end. From the speech-to-text models that turn audio into text, to the diarization that separates and identifies speakers, to the agentic layer that turns conversation into memory and action, to the observability and evaluation that keep all of it honest in production. This is a wide role by design. You will own model selection, serving, and production reliability. If you want to tune one model and ignore the system around it, this is not the role.
What You Will Own
• Speech-to-text. Evaluate, integrate, and optimize STT models across cloud and self-hosted. Drive accuracy and cost trade-offs with ground-truth metrics.
• Speaker diarization and identification. Push accuracy on hard, real-world, multi-speaker audio.
• Agentic AI. Build the memory and retrieval pipeline, LLM orchestration, and the agent workflows that sit on top of captured conversation.
• Model serving and infrastructure. Stand up and optimize self-hosted serving (vLLM, Triton class). Own latency, throughput, and cost per user.
Observability
An always-on wearable means models run in production every second, on messy real-world audio. You own the visibility into that.
• Instrument the full audio-to-memory pipeline: STT, diarization, retrieval, and LLM calls.
• Define and track model-quality SLOs in production: transcription drift, diarization error over time, retrieval relevance, latency, throughput, and cost per user.
• Build dashboards and alerting so model degradation is caught before users feel it.
• Trace failures across a distributed, always-on system using metrics, logs, and traces.
• Close the loop. Production signals feed back into evaluation and model selection.
Evaluation
We do not ship what we cannot measure. You own the systems that prove a model is actually better, not just newer.
• Build and own ground-truth evaluation harnesses for every model in the stack.
• Measure with real metrics: WER for transcription, DER for diarization, Recall and F1 for retrieval and speaker identification.
• Build and maintain labeled benchmark datasets that reflect real, messy, multi-speaker audio.
• Run regression and A/B evaluations on every model swap, prompt change, or pipeline update. Nothing ships on a vibe.
• Reject anecdotal proxies, single confidence scores, and cherry-picked examples as evidence of quality.
What We Are Looking For
• 3 to 5 years as an AI/ML engineer with production systems behind you. Engineering and production experience is non-negotiable.
• Depth across the modern AI stack: LLMs, speech models, vector retrieval, model serving.
• Strong software engineering. You write code that ships and survives contact with real users.
• Fluency in Python and the production ML ecosystem.
• Comfort with cloud infrastructure (GCP a plus) and containerized deployment on Kubernetes.
• A working command of observability and evaluation. You measure first and trust metrics over intuition.
• First-principles reasoning and metric discipline.
Nice to Have
• Research background or publications. A strong signal, not a substitute for production work.
• Audio and speech ML experience (STT, diarization, voice).
• Experience self-hosting and optimizing open models.
• Experience with LLM gateway and agent orchestration patterns.
• Experience building eval harnesses or production model-monitoring systems.
Requirements
Agentic work is must. Audio is good to have
. Self hosting models is a must
Experience with LLM gateway and agent orchestration is a must have
Location: Hyderabad, India. Based at the KnackLabs headquarters, with occasional travel to client locations for workshops and reviews. This role does not involve extended onsite deployments.
About the Role
You will work as an AI Architect who designs the systems behind our client engagements: AI agents, RAG systems, automation platforms, and the conventional backend systems around them.
This is a hands-on design role, not a slideware role. You will scope architectures with clients, make the hard technical decisions, defend them in review, and stay accountable for how the systems perform in production.
You will work directly with clients. Everyone at KnackLabs does. You will sit in design discussions with client engineering teams, present architecture decisions to technical and business stakeholders, and answer for the choices you make.
A full KnackLabs engineering team in Hyderabad builds with you. You own the technical design and the quality of what ships.
What you'll own
- Architecture - Design AI agents, RAG systems, integrations, and the scalable backend systems around them, for multiple client engagements.
- Technical scoping - Work directly with clients to turn a business problem into a system design, with clear trade-offs and clear reasons.
- Scale and reliability - Make sure what we build handles real load: data stores, queues, caching, horizontal scaling, and fault tolerance.
- Design reviews - Review designs and builds across engagements. Set the technical bar and hold it.
- Evaluation strategy - Define how we measure accuracy, safety, latency, and cost for the AI systems we ship.
- Guiding engineers - Raise the level of the engineers building with you, through reviews and direct pairing.
- Feedback to the platform - Feed what you learn across engagements back into our platform and internal tools.
What we are looking for
- Around 7 or more years of software engineering experience, including direct work with customers on design or delivery.
- Full-stack development experience with strength in backend technologies.
- Experience designing and building scalable applications. You understand how large-scale distributed systems work: data partitioning, queues, caching, horizontal scaling, and fault tolerance.
- At least 2 years of strong, hands-on AI experience with large language models in production.
- You build with AI coding tools like Claude Code or Codex as your default way of working. You understand Claude Skills, have written skills yourself, use them actively, and have contributed to them.
- Hands-on experience building retrieval-augmented generation (RAG) systems: chunking, embeddings, vector databases, retrieval, and reranking.
- Hands-on experience building AI agents.
- Strong programming skills in Python. Working knowledge of TypeScript or JavaScript.
- Experience with at least one cloud platform (AWS, Azure, or GCP).
- Clear communication. You can explain an architecture decision to an engineer and to a business leader, and defend it under questioning.
- High ownership and comfort with ambiguity. You can take an unclear problem and turn it into a design.
Nice to have
- Experience building evaluations to measure accuracy, safety, latency, and cost.
- Experience with observability and tracing tools such as LangSmith or Braintrust.
- Experience with on-premises or private cloud (VPC) deployments.
- Experience deploying AI systems in regulated industries such as insurance, banking, or the public sector.
- Experience with data engineering and pipelines.
- A history of side projects, open source contributions, or products you shipped end-to-end.
- Experience working at a consulting or professional services firm in a client-facing delivery role.
Stack and tools
- Languages: Python and TypeScript.
- Models: Claude and other frontier or open-source models, chosen to fit the customer.
- AI patterns: RAG, agents, prompt engineering, skills, and evaluations.
- Vector and retrieval: vector databases and retrieval pipelines.
- Cloud: AWS, Azure, or GCP, on public or private cloud.
- Integration: REST APIs and enterprise system connectors.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
KEY RESPONSIBILITIES:
•Build agents with persistent context & memory
•Design self-learning feedback loops
•Implement RAG pipelines for domain knowledge
•Manage conversation state & orchestration
•Integrate with LLM APIs (OpenAI, Claude, open-source)
Iterate fast — ship daily, measure weekly
MUST-HAVE SKILLS
•Python / TypeScript proficiency
•LangChain, CrewAI, AutoGen or custom frameworks
•Experience with vector DBs (Pinecone, Weaviate, Qdrant)
•Prompt engineering & evaluation pipelines
•Understanding of agent architectures (ReAct, tool-use)
Git, CI/CD, containerization basics












