Cutshort logo
For Employers
Haparz Pvt Ltd logo
Senior/Staff AI Evaluation Engineer
Senior/Staff AI Evaluation Engineer

Senior/Staff AI Evaluation Engineer at Haparz Pvt Ltd · Remote only · 6 - 10 years · ₹20L - ₹22L / yr · Profitable · Remote only · Posted 24 Sep 2026

Haparz Pvt Ltd's logo

Senior/Staff AI Evaluation Engineer

Jansi Rani's profile picture
Posted by Jansi Rani
6 - 10 yrs
₹20L - ₹22L / yr
Remote only
Skills
Large Language Models (LLM)
Retrieval Augmented Generation (RAG)
AI Evaluation
Agentic AI
AWS Bedrock

Job Description


Role: Senior/Staff AI Evaluation Engineer

Experience: 6+ Years

Location: Remote (Pan India)

Mode: Hyparz Payroll


About the Role

We are looking for a Senior/Staff AI Evaluation Engineer to lead the benchmarking, validation, reliability, and safety evaluation of next-generation Agentic AI platforms.

This role sits at the intersection of Quality Engineering, Software Engineering, and AI/ML. You will build automated evaluation frameworks, benchmark AI behavior, validate LLM/RAG/agentic workflows, and establish measurable quality standards for production AI systems.

What You'll Do

  • Design and implement automated evaluation frameworks for LLM, RAG, and agent-based applications.
  • Build and maintain golden datasets, benchmark datasets, test datasets, and regression suites for AI evaluation.
  • Develop measurable evaluation criteria for accuracy, relevance, consistency, safety, reliability, and agent behavior.
  • Perform adversarial testing to identify hallucinations, prompt vulnerabilities, unsafe behavior, and edge cases.
  • Evaluate LangGraph-based and multi-agent systems at the node, state-transition, routing, and workflow levels.
  • Use LangSmith for tracing, debugging, experiments, datasets, and evaluation workflows.
  • Work extensively with LangChain and LangGraph, including sub-graphs and conditional routing.
  • Build Python-based automation for functional, regression, integration, and end-to-end AI testing.
  • Evaluate prompts, embeddings, vector search, RAG pipelines, tool calling, and agentic workflows.
  • Validate AI systems against safety, security, privacy, and Responsible AI expectations.
  • Integrate evaluation and regression testing into GitHub-based CI/CD workflows.
  • Work with AWS and Amazon Bedrock to validate AI-powered application architectures.
  • Perform API and backend validation using REST APIs, JSON, SQL, and modern application architectures.
  • Collaborate with Engineering, Product, Data, and Quality teams to investigate failures and drive improvements.
  • Take ownership of ambiguous AI quality problems and convert them into repeatable, measurable evaluation approaches.

What We're Looking For

  • 6+ years of experience in software engineering, quality engineering, test automation, AI engineering, ML engineering, or a related technical discipline.
  • Demonstrable hands-on experience testing or evaluating LLM, NLP, ML, RAG, or agentic AI applications.
  • Strong Python development and automation experience.
  • Deep hands-on knowledge of LangChain, LangGraph, and LangSmith.
  • Experience evaluating LangGraph node execution, state transitions, conditional routing, sub-graphs, or multi-agent workflows.
  • Strong understanding of LLMs, prompt engineering, embeddings, vector databases/search, RAG, and AI agents.
  • Experience building golden datasets, benchmark datasets, or structured AI test datasets.
  • Experience with adversarial testing and AI safety evaluation.
  • Practical understanding of security, privacy, and Responsible AI evaluation.
  • Strong GitHub experience covering repositories, branching, pull requests, code reviews, and CI/CD.
  • Experience with Claude Code or comparable AI-assisted development tools.
  • Experience with AWS and Amazon Bedrock.
  • Strong REST API, JSON, and SQL knowledge.
  • Strong experience in automated, regression, integration, and end-to-end testing.
  • Excellent analytical, troubleshooting, communication, and problem-solving abilities


Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About Haparz Pvt Ltd

Founded :
2015
Type :
Services
Size :
20-100
Stage :
Profitable

About

At Haparz, we help businesses turn ideas into powerful digital products. We specialize in building custom mobile applications, websites, software solutions, AR/VR experiences, games, and blockchain applications that are secure, scalable, and designed to deliver real business impact. By combining technical expertise with a client-first approach, we create solutions tailored to each organization's unique goals.


We believe great products are built by great teams. Our experienced designers, developers, and technology specialists work closely with clients at every stage of the development journey—from ideation and strategy to deployment and ongoing support. Whether you're a startup launching your first product or an enterprise scaling your digital presence, we're committed to delivering high-quality solutions that help you innovate, grow, and stay ahead of the competition.

Read more

Company social profiles

bloginstagramlinkedinfacebook

Similar jobs (10)

company logo
Leena Lahari
Posted by Leena Lahari
Mumbai
2 - 4 yrs
₹6L - ₹15L / yr
Generative AI
skill iconPython
Test Automation (QA)
Manual testing
Functional testing
+7 more

Title: AI/ML Test Engineer – GenAI

Location - Mumbai

Technical Skills


• Strong experience in Generative AI, LLMs, and Agentic AI systems

• Hands-on expertise with AI evaluation frameworks (RAGAS, DeepEval, TruLens, LangSmith, Promptfoo, etc.)

• Proficiency in Python and AI/ML development libraries

• Knowledge of Prompt Engineering, prompt testing, and optimization

Ability to define and track evaluation metrics such as accuracy, relevance, groundedness, hallucination rate, latency, and user satisfaction

• Experience in creating automated evaluation pipelines and benchmarking frameworks

• Strong understanding of AI safety, guardrails, bias testing, and responsible AI practices

• Familiarity with REST APIs, JSON, vector databases, and knowledge retrieval systems

• Experience in A/B testing, human-in-the-loop evaluation, and red teaming

• Strong experience in Manual Testing of AI/GenAI applications, including functional, exploratory, UAT, regression, and end-to-end testing

• Expertise in validating Agent Reasoning, Tool Calling, Workflow Execution, and Response Quality

• Hands-on experience in Automation Testing using Python frameworks


Key Responsibilities

  • Design, execute, and automate evaluation strategies for Agentic AI applications.
  • Develop evaluation datasets, test cases, and benchmark suites.
  • Measure and improve agent performance, reasoning quality, tool usage, and workflow effectiveness.
  • Analyze model outputs and identify hallucinations, biases, safety risks, and failure patterns.
  • Collaborate with AI Engineers, Product Teams, and Domain Experts to improve agent quality and reliability.
  • Generate evaluation reports, dashboards, and actionable recommendations.


Read more
company logo
Umama Sayed
Posted by Umama Sayed
Mumbai
5 - 8 yrs
Best in industry
skill iconPython
Large Language Models (LLM)
Artificial Intelligence (AI)
Prompt engineering
LangGraph
+6 more

Senior AI Engineer

Code Generation, Agent Architecture & LLM Systems

📍 Mumbai (On-site) | Full-time | 5+ years


About the Role:

Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.

We are hiring a Senior AI Engineer for a dedicated client engagement focused on building an AI-powered application builder platform - a product where users describe software in plain English and the system generates, previews, and iteratively refines working code.

The mandatory requirement for this role is hands-on production experience shipping LLM-powered systems with agent architectures, with experience in code generation or developer tooling contexts a strong advantage.


The role is product-focused and deeply hands-on. You will own everything between the user's prompt and correct code landing in the project: the agentic loop, code generation pipeline, context management, evaluation suite, and model cost strategy.

You will work alongside the Senior MLOps Engineer who operationalises the infrastructure around your system, and collaborate closely with backend, frontend, and DevOps engineers.


Responsibilities:


Agent Architecture

Design and own the agentic loop for the platform - request interpretation, planning, tool-calling sequence (read file, edit file, run build, search code, install package), and stop conditions.

Make and revisit architectural decisions on single-agent vs. multi-agent designs, including planner/executor splits and dedicated build-repair sub-agents.


Code Generation Pipeline

Own the end-to-end generation flow: task classification, context gathering, planning, targeted edits, verification, and commit.

Implement diff/search-replace-based file editing with fuzzy matching and fallback strategies.

Enforce scope discipline so the agent makes minimal diffs and does not modify code it was not asked to touch.


Self-Repair Loop

Build and tune the automated repair loop that pipes compiler, lint, build, and runtime errors back to the model with retry budgets and model escalation.

This loop is the primary quality lever - the difference between 60-70% and 90%+ build success rates.


Context Management

Build file-relevance retrieval so the agent sees the right files, not the whole codebase: dependency graphs, AST/tree-sitter-based chunking, embeddings, recency signals, and hybrid retrieval.

Implement conversation summarisation and memory for long sessions, and address long-project degradation through codebase summaries and periodic consistency passes.

Own token budgeting and prompt caching strategy.


Prompt Engineering as a Discipline

Own the system prompt and per-task prompt variants (new feature, bug fix, styling change).

Maintain few-shot examples and enforce coding conventions, stack rules, and prohibited behaviours such as no hardcoded secrets and no whole-file rewrites.

Version prompts like code with changelogs and rollback capability.


Evaluation and Quality Measurement

Design and own the evaluation suite: representative test prompts run on every prompt and model change, scored on build success rate, instruction adherence, and output quality including LLM-as-judge and visual/screenshot checks where relevant.

Define regression gates that block quality-degrading changes from shipping.

Treat evals the way engineers treat automated testing: versioned, automated, and tracked over time.

This responsibility is non-negotiable at this level.


Model Strategy and Cost

Design model routing - cheap and fast models for classification and small edits, frontier models for complex generation.

Drive cost optimisation through prompt caching, diff-based edits over full-file rewrites, and tighter context selection.

Track cost per agent run and tokens per task; evaluate new model releases against the eval suite and lead migrations when results justify it.


Safety and Reliability of Agent Behaviour

Defend against prompt injection from user content and fetched web content.

Ensure secrets never appear in generated client code.

Define what the agent's tools may and may not do in collaboration with the platform team.

Contribute to output moderation and abuse-pattern awareness.


Mentorship and Engineering Standards

Run code reviews, define engineering conventions for AI work, and raise the engineering bar across the AI team.

Work closely with the Senior MLOps Engineer on handoff of eval design, prompt configurations, and model routing logic.


Requirements:


Hands-on Production Ownership of LLM-Powered Systems with Agent Architectures (Mandatory)

Must have personally shipped and operated at least one complex production AI system - agentic, multi-step, or code generation - with end-to-end ownership of architecture, evaluation, and cost.

POCs, internal demos, and tutorial-grade work do not qualify.


5+ Years of Professional Software or AI Engineering Experience

With at least 3 years focused on LLM applications, AI engineering, or production AI systems.

Candidates with strong backend backgrounds and a clear, substantive pivot into LLM systems qualify.


Strong Python Proficiency and Service Development

Production-grade Python with FastAPI or equivalent: type hints, async patterns, streaming responses, testing, and packaging.

Not notebook-only.


Depth Across LLM APIs and Agent Systems

Production experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or open-weight models (vLLM, Ollama, Together).

Production experience with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.

Hands-on with tool calling, structured outputs, and multi-step reasoning.


Demonstrated, Systematic Evaluation Practice - Non-Negotiable

Must have built evaluation harnesses that gate production releases, not ad-hoc testing.

Hands-on with at least one of LangSmith, Langfuse, Promptfoo, Ragas, or DeepEval.

Candidates with no systematic answer to evaluation should not be considered at senior level regardless of other strengths.


Cost Discipline for Production AI

Track record of measurable cost optimisation on production AI features.

Able to speak in specifics: cost per request, savings achieved through caching or model routing, context reduction decisions.


AWS Working Knowledge

Hands-on with EC2, S3, IAM, and Docker.

Comfort with CI/CD workflows and deploying AI services.


Awareness of LLM Security Failure Modes

Familiar with prompt injection patterns, understands that system prompt rules alone are insufficient, and has experience with output validation and content safety in production.


Nice to Have

  • Experience with AST/tree-sitter tooling, diff-based editing systems, or compiler-adjacent work
  • MCP server authoring
  • Open-source AI contributions
  • Published technical writing on LLM systems
  • Multi-modal model experience
  • Fine-tuning exposure (LoRA, QLoRA, PEFT)
Read more
NA
NA
Agency job
via by Ramya Munirathnam
Hyderabad, Pune
10 - 14 yrs
₹6L - ₹14L / yr
Generative AI
skill iconC#
skill icon.NET
AI Agents

AI Engineer


We are seeking an AI Engineering specialist focused on AI evaluation, and continuous quality improvement for Ezra MetLife's employee-facing AI platform. This role will establish and scale the testing strategy for enterprise AI agents, ensuring high response quality, reliability, and production readiness. The engineer will build automated regression testing framework (preferred Playwright ), define AI evaluation methodologies, analyze AI performance metrics, and partner with engineering teams to continuously improve answer quality, grounding accuracy, and customer experience. This position is critical to enabling confidence as Ezra expands its AI agent portfolio and employee-facing capabilities.


Required Skills & Experience


• C# and .NET development experience


• Experience with at least one AI evaluation framework (e.g., prompt evaluation, RAG evaluation, LLM quality assessment)


• Microsoft Agent Framework (preferred) or similar enterprise agent frameworks


• Experience with Azure OpenAI / Azure AI Foundry


• Microsoft 365 Agent SDK


• Azure AI Search, RAG pipelines, and retrieval quality testing


• Infrastructure as Code using Terraform


• Experience building automated testing and AI quality validation processes


• Familiarity with telemetry analysis, AI observability, and performance measurement


• Strong analytical skills with a passion for improving AI response quality and reliability


Skills: AI Agents~Core .NET Technologies~C# 5.0

Experience Required: 10 & Above

Location: Hyderabad :5+ relevant exp in AI + .NET

Read more
company logo
Banu S
Posted by Banu S
Bengaluru (Bangalore), Hyderabad
5 - 12 yrs
₹4L - ₹25L / yr
skill iconData Science
skill iconPython
Large Language Models (LLM) tuning
RAG
Langchain

Support with design and build to prove out agentic AI solution flow by working with other data 

scientists and engineers to build, train Large Language Model (LLM) architectures, RAG 

systems, and autonomous agentic workflows 

Key qualifications: 

  

>> AI solution design & Development: Design Agentic AI solutions using RAG (Retrieval-

Augmented Generation) and orchestration frameworks like LangGraph or LangChain. 

  

>> Model Fine-Tuning: Solid understanding and experience with Pre-train, fine-tune, and 

optimize open-source like BERT, LLama, and other proprietary foundation models for domain-

specific tasks 

  

>> Solid Stats and ML foundations and (vibe) coding skills with Python, PySpark 

  

>>  Implement validation frameworks and tracing practices (using tools like Arize) to monitor 

agent behavior, guard against model drift, and ensure compliance 

  

>> Collaborate with Engineering to deploy models securely on cloud and on-prem ecosystems 

 

Read more
company logo
Faisal AshrafNomani
Posted by Faisal AshrafNomani
Remote only
4 - 8 yrs
Best in industry
Agentic AI

Senior Agentic AI Engineer - (Freelance)

Positions: 2

Experience: Ideally 4(J–(J8 years with strong software-engineering fundamentals and recent hands-on Agentic AI experience.

Mission

Build UC2's governed AI agents capable of reasoning across and interacting safely with enterprise IT systems.

Mandatory capabilities

  • Python
  • LangGraph
  • Agentic AI
  • Tool/function calling
  • Stateful workflows
  • Structured outputs
  • Human-in-the-loop
  • Guardrails
  • Agent state/checkpointing
  • Agent evaluation
  • FastAPI
  • REST APIs
  • Async Python

Retry/timeout/error handling

Highly desirable

MCP, LangChain, Semantic Kernel, agent observability, event-driven architecture and experience integrating AI agents with ServiceNow/Splunk/Confluence or similar enterprise platforms.

The candidate should understand how to engineer:

Read more
Leadsquared
Leadsquared
Agency job
via by Vrishali Mishra
Bengaluru (Bangalore)
2 - 4 yrs
₹15L - ₹35L / yr
Test Automation (QA)
API
skill iconPython
skill iconJavascript
Regression Testing
+2 more

About Us

We’re building the next generation of AI-powered business software, and we’re looking for people who want to shape that future with us. With Lumen, we’re reimagining how users interact with CRM — moving beyond screens, menus and dashboards to an intelligent interface where users can simply ask AI to take actions, retrieve knowledge, generate insights and get work done. With Agent Studio, we’re enabling businesses to build, test and deploy their own AI agents for real-world workflows. And with Invorto, we’re bringing AI to voice, allowing businesses to create intelligent voice agents tailored to their customer and operational use cases.

What makes this especially exciting is the stage and scale of the opportunity. You’ll get to work on genuinely hard problems across LLMs, agents, reasoning, orchestration, voice AI, evaluation, reliability and enterprise security — not as isolated experiments, but as products used in real business workflows. You’ll have the opportunity to build zero-to-one, own meaningful parts of the product end-to-end, work closely with customers, experiment rapidly, and see your work reach production at scale.

Why join now? Because the playbook for enterprise AI is still being written. You won’t just be implementing someone else’s roadmap — you’ll help define the product, architecture and experiences that become that playbook. Expect high ownership, fast iteration, hard technical and product problems, direct customer impact, and the chance to build AI systems that have to work reliably in the real world — not just in a demo.

About the Role

We are looking for a QA Engineer who specializes in testing agentic AI platforms. You will design and automate quality processes for systems that involve LLMs, autonomous agents, tool use and orchestration across Lumen and Agent Studio — ensuring that AI-driven workflows behave reliably, safely and predictably in production, not just in a demo.

What You’ll Do

  • Design and build automated test suites and evaluation frameworks for agentic AI workflows, including multi-step and tool-calling behaviors.
  • Use AI/LLM-based QA tools and evaluation frameworks to test model outputs, agent decisions and end-to-end task completion at scale.
  • Define quality metrics and benchmarks for agent reliability, correctness, latency and safety, and track them over releases.
  • Identify edge cases, failure modes and regressions specific to non-deterministic AI systems, and build automated checks to catch them early.
  • Integrate automated agent/LLM testing into CI/CD pipelines to support fast, reliable iteration.
  • Partner closely with AI/ML and backend engineers to reproduce issues, root-cause failures and validate fixes.
  • Work with customers and customer-facing teams to understand real-world usage patterns and translate them into test scenarios.

What We’re Looking For

  • 2–4 years of QA/test automation experience, including hands-on work testing agentic AI or LLM-based platforms.
  • Practical experience using AI-focused QA/evaluation tools to test agent behavior, prompts and model outputs.
  • Strong scripting/automation skills (Python preferred) to build and maintain test frameworks.
  • Understanding of how LLM-based agents work — tool calling, orchestration, memory, reasoning chains — well enough to design meaningful test cases.
  • Comfort working with non-deterministic systems and designing evaluation approaches beyond traditional pass/fail testing.
  • Strong communication skills — this is a customer-facing role, and you will be expected to clearly articulate technical concepts, decisions and trade-offs to both technical and non-technical stakeholders, including customers.

Good to Have

  • Experience testing voice AI or real-time conversational systems.
  • Familiarity with CRM or enterprise SaaS platforms.
  • Exposure to enterprise security or compliance testing for AI systems.


Read more
company logo
Agency job
via by Vrishali Mishra
Bengaluru (Bangalore)
4 - 6 yrs
₹20L - ₹40L / yr
skill iconPython
Speech-to-Text (STT)
ASR
Text-to-Speech (TTS)
Large Language Models (LLM)
+3 more

About Us

Invorto is our Voice AI product, bringing intelligent voice agents to real-world customer and operational use cases. Our voice pipeline is built in Python, running an STT → LLM → TTS architecture on top of the Pipecat framework.

This is a chance to work on hard problems in voice AI — latency, accuracy, naturalness, and reliability — building zero-to-one, owning your area end-to-end, and shipping to production at scale.

Note: This is a customer-facing role, and strong communication skills are essential.

About the Role

We're looking for a Voice AI Research Engineer to join the Invorto team and help build and continuously improve the voice AI systems that power our intelligent voice agents. This role is focused on the specialized craft of voice AI — designing evaluation and automation frameworks that ensure our STT, LLM, and TTS pipeline performs reliably in real-world, production conditions.

 

What You'll Do

  • Design and build automated testing and quality frameworks for our STT → LLM → TTS voice pipeline, built on Pipecat
  • Evaluate and benchmark STT, LLM, and TTS/ASR components on accuracy, latency, naturalness, and robustness across accents, languages, and real-world audio conditions
  • Work hands-on with STT, TTS, and ASR models — fine-tuning, evaluating, and improving them for production use cases
  • Identify failure modes and edge cases across the pipeline (background noise, accents, interruptions, turn-taking, latency, pipeline-stage handoffs) and build systems to catch them before production
  • Collaborate closely with engineering to integrate quality checks and automation into the voice agent development lifecycle within the Pipecat-based architecture
  • Research and stay current with advances in voice AI, and bring in new techniques, models, and tools to improve pipeline performance
  • Work directly with customers to understand real-world voice use cases and translate them into evaluation criteria and quality benchmarks
  • Partner with product and engineering to define what "production-grade quality" means for voice agents and drive the team toward it

 

What We're Looking For

  • 4–6 years of experience, with a specialization in voice AI systems and automated quality evaluation
  • Hands-on experience with STT (Speech-to-Text), TTS (Text-to-Speech), and ASR (Automatic Speech Recognition) models
  • Experience designing and building automated testing/evaluation frameworks for voice or speech systems
  • Strong understanding of what drives voice AI quality — accuracy, latency, naturalness, and robustness to real-world variability
  • Strong programming skills in Python; familiarity with Pipecat or similar voice pipeline/orchestration frameworks is a plus
  • Understanding of STT → LLM → TTS pipeline architectures and the trade-offs involved at each stage
  • Research mindset — comfortable exploring new models, techniques, and tools and translating them into practical improvements
  • Excellent communication skills — this is a customer-facing role, and you'll regularly engage directly with customers to understand needs and validate quality expectations


Read more
company logo
Anish N
Posted by Anish N
Bengaluru (Bangalore)
3 - 5 yrs
₹10L - ₹20L / yr
skill iconPython
Generative AI
Agentic AI
LangChain
LlamaIndex
+3 more

Job Description:

We are looking for a hands-on AI Engineer with experience in Generative AI and Agentic AI to build and deploy production-ready AI solutions.

Key Responsibilities:

  • Develop and deploy GenAI and Agentic AI applications.
  • Build RAG pipelines, LLM workflows, and AI agents.
  • Develop solutions using Python, LangChain, LangGraph, LlamaIndex, or similar frameworks.
  • Implement tool calling, context retrieval, and LLM orchestration.
  • Integrate AI solutions with APIs and cloud platforms.
  • Work with AWS/Azure/GCP, Docker, and CI/CD.

Required Skills:

  • Strong Python programming skills.
  • 3+ years of GenAI/Agentic AI experience.
  • RAG and LLM orchestration.
  • LangChain / LangGraph / LlamaIndex / AutoGen / CrewAI / Semantic Kernel.
  • MCP and A2A knowledge.
  • Cloud, APIs, Docker, and CI/CD experience.

Preferred Experience:

Hands-on experience building and deploying production-ready AI solutions.

Read more
company logo
Bhawna Khemani
Posted by Bhawna Khemani
Bengaluru (Bangalore), Delhi, Gurugram, Noida, Ghaziabad, Faridabad
4 - 13 yrs
₹11L - ₹35L / yr
Generative AI
Large Language Models (LLM)
Retrieval Augmented Generation (RAG)
skill iconPython
LlamaIndex
+4 more

Generative AI Engineer 

Role Overview:

You will be responsible for the hands-on development, coding, and deployment of AI-powered features. Your focus is on writing clean, efficient code to integrate LLMs into our existing tech stack, building robust data pipelines for RAG, and ensuring the reliability of model outputs through rigorous testing and optimization.

Key Responsibilities

  • Application Implementation: Code and integrate LLM APIs (OpenAI, Anthropic, etc.) or local models into backend services using Python, FastAPI, etc.,
  • MCP Server Development: Design and implement custom MCP servers using the official SDKs (Python/TypeScript) to expose internal databases, APIs, and file systems to AI agents.
  • RAG Implementation: Build and maintain the "plumbing" for Retrieval-Augmented Generation—specifically coding the data ingestion scripts, text chunking logic, and metadata filtering.
  • Vector DB Management: Perform day-to-day operations on vector databases (Pinecone, Milvus, etc.), including indexing, querying, and optimizing search retrieval.
  • Prompt Programming: Develop, version-control, and refine complex prompt templates (using Jinja2 or similar) to ensure consistent structured outputs (JSON/YAML).
  • Agent Development: Implement multi-step workflows using LangChain, LangGraph, CrewAI etc.,, focusing on tool-calling logic and error handling.
  • Evaluation & Testing: Build automated test suites to detect "hallucinations" and measure accuracy using frameworks.
  • Performance Tuning: Implement caching layers and streaming responses to reduce latency and improve the end-user experience; Token optimization.
  • Data Pre-processing: Clean and tokenize datasets for model fine-tuning or high-quality context retrieval.

Technical Skills (The "Execution" Stack)

  • Language: Advanced Python (Asyncio, Pydantic) and optional TypeScript/Node.js (for full-stack integration).
  • AI Frameworks: Hands-on experience with any of LangChain, LlamaIndex, and Hugging Face Transformers. RAG and Vector search concepts.
  • Data Handling: Proficiency in SQL and handling unstructured data formats (PDFs, Markdown, JSON).
  • Deployment: Practical experience with Docker, GitHub Actions (CI/CD), and experience with OpenTelemetry, LangSmith, Weights & Biases etc., Understanding of evaluation/guardrails.
  • MCP/API Proficiency: Deep understanding of RESTful APIs, Streaming HTTP, MCP server vs client, JSONRPC
Read more
company logo
Lakshit Bagga
Posted by Lakshit Bagga
Remote only
10 - 50 yrs
₹1L - ₹50L / yr
Open-source LLMs
Playwright
cypress
Selenium
skill iconAmazon Web Services (AWS)
+1 more

Job title: Chief Agentic Quality Architect

Type: Full-Time | Contract

Location: Remote


Role Overview

Equity Partners builds profitable growth by acquiring and operating enterprise software companies — refining a proprietary operating model across 40+ acquisitions and two decades of hands-on experience, now supercharged by our patented agentic AI platform. We're hiring a Chief Agentic Quality Architect to lead the transition from traditional scripted testing to an AI-augmented quality ecosystem. In this role, you'll audit and rebuild our quality engineering foundations, deploy agentic automation across critical business flows, and build the guardrails needed to keep AI-generated code production-ready.


Key Skills

  • 10+ years in QA automation engineering, SDET, or test architecture roles
  • Expert-level proficiency in Playwright, Cypress, or Selenium
  • Hands-on experience using LLMs (Claude, GPT-4, etc.) and agentic frameworks to generate code or automate workflows
  • Deep understanding of integrating quality gates into AWS-based CI/CD pipelines or similar environments
  • Architectural mindset, with the ability to design "Behavioral Snapshots" to safeguard critical business logic during rapid transformation


Responsibilities

  • Conduct a comprehensive audit of the existing test estate across unit, integration, API, UI, sanity, and regression layers, producing a Current State & Gap Coverage Report
  • Architect and own a phased quality engineering roadmap across two-week, one-month, and three-month delivery horizons
  • Deploy agentic test generation — transforming business requirements into executable Playwright/Cypress scripts, generating synthetic test data, and implementing self-healing automation
  • Design regression strategies and quality gates specifically tuned to catch hallucinations and logic errors in AI-generated code
  • Build, mentor, and upskill a specialist QA team fluent in AI-assisted testing and agentic automation frameworks


If you're ready to rewrite the rules of quality for an AI-native world, we want you leading the charge.

Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos