Cutshort logo
For Employers
TGS The Global Skills logo
Sr Agentic Harness Engineer
Sr Agentic Harness Engineer

Sr Agentic Harness Engineer at TGS The Global Skills · Mumbai · 8 - 12 years · ₹15L - ₹30L / yr · Profitable · Posted 26 Aug 2026

TGS The Global Skills's logo

Sr Agentic Harness Engineer

Sakshi Bhardwaj's profile picture
Posted by Sakshi Bhardwaj
8 - 12 yrs
₹15L - ₹30L / yr
Mumbai
Skills
AWS Bedrock
Agentic AI
Harness
skill iconPython
JIRA
LLM Evaluation Frameworks
IT governance
AgentCore,

Responsibilities

·        Build and operate the agentic loop: trigger → orchestration → agent execution → output to JIRA → human accept/reject → next agent, across design, coding, review, and testing agents.

·        Implement model routing and retry logic across a provider-agnostic model layer (e.g., Claude via AWS Bedrock, self-hosted or alternative models as cost/sovereignty hedges), including business-continuity fallback if a given provider becomes unavailable.

·        Own token cost control and context window management — per-agent and per-run budgets, circuit breakers that halt runaway execution, and cost observability tied back to JIRA.

·        Stand up and maintain observability, alerting, and monitoring across the agent fleet (e.g., Langfuse or equivalent), so agent health, cost, and quality are visible in real time.

·        Implement agent governance and safety guardrails: deterministic pre/post hooks gating every LLM call, kill switches, prompt injection prevention and mitigation, and audit logging.

·        Integrate the harness with JIRA as the system of record and other business systems as needed, ensuring every agent action, decision, and human override is tracked with no side channels.

·        Pair directly with client engineers throughout — this is capability transfer, not black-box delivery. You'll document, demo, and hand over as you build.

·        Work in outcome-based delivery stages (spike → architecture sign-off → build → pilot) with gated milestones tied to working software demos, not fixed artifact checklists.

·        Participate actively in team discussion and design decisions — this team expects engineers to challenge ideas constructively and speak up, not defer silently.

Must-Have Experience

·        Hands-on production experience building agentic systems(not tutorial-level or personal-project experience.) Candidates should be able to speak concretely about systems they've shipped.

·        Practical experience with agentic frameworks such as LangChain, LangGraph, or equivalent orchestration frameworks.

·        Experience with LLM orchestration and model routing across multiple providers/models, including fallback and retry design.

·        Working knowledge of agent governance: guardrails, human-in-the-loop approval flows, kill switches, and audit trails.

·        Practical understanding of prompt injection risks and mitigation techniques.

·        Experience with token cost management and context window/memory handling at production scale — this is a named governance requirement for the engagement, not a nice-to-have.

·        Strong Python (or equivalent) engineering background, comfortable working in AWS environments (Bedrock/AgentCore exposure a strong plus).

·        Experience with observability/monitoring tooling for distributed or agentic systems (e.g., Langfuse, Datadog, or equivalent).

·        Comfortable working with JIRA/Atlassian APIs or similar ticketing-system-of-record integrations.

Nice to Have

·        Direct experience with AWS Bedrock AgentCore, Temporal (or similar workflow orchestration), or LiteLLM-style model gateways.

·        Exposure to Cursor or other AI-native IDEs in a production engineering context.

·        Experience with self-hosted open-weight models (e.g., DeepSeek, GLM) as cost or sovereignty hedges alongside commercial APIs.

·        Financial services or other regulated-industry background.

·        Familiarity with Claude Code, Claude Cowork, or Claude Skills.

Qualifications

·        Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.

·        3-5+ years in software/platform engineering, with a meaningful portion of that time specifically on agentic or LLM-orchestration systems (not general ML or data engineering alone).

Relevant Experience

·        Already built this kind of system and can talk through the trade-offs from experience, not theory.

·        Comfortable operating with ambiguity - technology choices (frameworks, specific models, tooling) are expected to evolve during the engagementand milestones are tied to outcomes rather than fixed deliverables.

·        Will contribute opinions - quiet execution without a point of view is not a fit for this team.

Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About TGS The Global Skills

Founded :
2007
Type :
Services
Size :
0-20
Stage :
Profitable

About

TGS The Global Skills is a recruitment and staffing solutions company helping businesses hire the right talent across India.

We specialize in IT Recruitment and specialized technology hiring, including AI/ML, Data, Cloud, SAP, Software, Engineering, Sales, and Operations.

Our goal: Connecting the right talent with the right opportunity.

Read more

Candid answers by the company

What is the location preference of jobs?

Delhi

Company social profiles

linkedin

Similar jobs (10)

company logo
Remote only
4 - 8 yrs
₹10L - ₹30L / yr
Generative AI (GenAI)
Large Language Models (LLM) tuning
Fine-tuning LLMs
Retrieval Augmented Generation (RAG)
skill iconPython
+2 more

Forward-deployed engineers (FDEs) are Mactores' services layer. You embed with the customer's team, own outcomes from discovery through the production cutover, and personally carry the delivery commitment.

The agent platform we deploy absorbs 60–70% of engagement work, discovery, assessment, design, and testing. You absorb the judgment: target architecture, refactoring trade-offs, model selection, cutover strategy, and the decisions an agent platform cannot make. The agent absorbs scale. You absorb judgment. 

This is not a staff-augmentation seat and not an advisory role. You ship.

 

What you will do?

  • Deliver production agentic AI systems and AWS modernization engagements on committed dates across three pillars: Data Platform Modernization, Application & Database Modernization, and AI Agents for Apps.
  • Build and productionize AI agents, orchestration, retrieval pipelines, evaluation harnesses, observability running against real customer data, not demo data.
  • Convert existing products into agents: expose product functionality as callable tools for agent-to-agent composition, or replace form-and-click UX with agent-native, intent-driven interfaces.
  • Convert existing Business processes into agents: expose process functionality as callable tools for agent-to-agent composition, or replace form-and-click UX with agent-native, intent-driven interfaces.
  • Embed directly with customer engineering teams. Run architecture sessions, defend design decisions, and align stakeholders from VP Engineering to CTO.
  • Make agent decisions traceable and defensible, validation runs in parallel with live workloads, and outputs hold up to internal audit and regulators (HIPAA, PCI-DSS, FSI-grade governance where the vertical demands it).
  • Feed field experience back into the platform and practice: your deployment patterns, integration playbooks, and edge cases shape how we deliver.


What are we looking for?

  • Excellent communication skills (English) — verbal and written. Non-negotiable. You will present architecture to customer CTOs, write documents that hold up in audit, and defend judgment calls in the room. If you can build but not explain, this role is not a fit.
  • You have shipped production agentic AI systems on AWS. Not POCs, not notebooks — systems running in production for real users. This is the primary qualification. Be prepared to walk through what you shipped, the decisions you made, and what broke.
  • Deep understanding of agentic architecture — you can design an agent system from first principles and explain why each component exists:
  • Agent design patterns: single-agent vs. multi-agent systems, supervisor/orchestrator patterns, hierarchical agent topologies, planner–executor separation, and when each applies.
  • Orchestration: building and operating orchestrator agents that decompose tasks, route work to specialist agents or tools, and manage state across multi-step workflows (LangGraph, Strands Agents, CrewAI, or equivalent).
  • Memory: short-term/working memory (context management, conversation state) and long-term memory (episodic and semantic stores, vector- and graph-backed retrieval), and the production trade-offs of each.
  • Reflection and self-correction: critique loops, self-evaluation, retry-with-feedback patterns, and evaluation harnesses that catch agent failures before customers do.
  • Tool use and function calling: schema design, tool-selection reliability, error handling, and agent-to-agent composition.
  • RAG and retrieval pipelines: chunking, embedding, hybrid retrieval, reranking, and grounding agent decisions in customer data.
  • Strong AWS production experience: Amazon Bedrock and AWS AI services, plus core platform services (Lambda, API Gateway, DynamoDB, RDS/Aurora, Glue, EMR, Redshift, Kinesis, or similar depending on specialization).
  • Solid software engineering fundamentals Python, TypeScript, CI/CD, infrastructure-as-code, testing-driven development discipline.
  • Experience with data or application modernization (database migration, legacy refactoring, data platform builds) is a strong plus, since agents run against these workloads.
  • Indicative experience: roughly 3–10 years in engineering roles, with agentic AI / GenAI as your current day job. We have demonstrated agent-native expertise over tenure — an engineer with 3–4 years of hands-on agentic AI work typically outperforms a 12-year generalist on this work.


You'll be preferred if you've:

  • US English verbal and written fluency 
  • Delivery experience in one or more of our verticals: Financial Services, Healthcare & Life Sciences, Internet & Software, Manufacturing, or Telco/Media/Entertainment/Gaming/Sports.
  • Model tuning and fine-tuning: systematic prompt engineering and optimization; parameter-efficient fine-tuning (LoRA/QLoRA or similar); instruction tuning; working knowledge of RLHF/DPO; sound judgment on when to fine-tune vs. prompt vs. RAG; and evaluation of tuned models against baselines. Fine-tuning experience on Amazon Bedrock or SageMaker is a plus.
  • Experience with compliance-sensitive AI systems (HIPAA, PCI-DSS, SOC 2, data residency).
  • Knowledge graph, code-analysis (AST), or CDC/streaming experience (Debezium, Kafka/MSK).
  • Solid software engineering fundamentals — Java, C++, Go Lang, .Net, Rust
  • Prior customer-facing consulting or forward-deployed experience.
  • AWS certifications (Solutions Architect Professional, Machine Learning Specialty, or Data Analytics).


Why This Role?

  • You own outcomes, not tickets. FDEs carry the delivery commitment personally — architecture, judgment, and cutover are yours.
  • You work agent-native from day one. Our delivery model would not function without agents. You build with the platform, not around it.
  • You ship. Engagements measured in weeks to production, legacy retired, outcomes named. No archived pilots.
  • You compound. Field delivery informs the Aedeon platform roadmap; the platform's growth expands what you can deliver. Few engineering roles sit in that loop.


Read more
Remote only
7 - 12 yrs
₹40L - ₹70L / yr (ESOP available)
Agentic AI
skill iconPython
API management
Anthropic Claude

Experience: 8+ years, senior candidates only | Type: Full-time | Location: Remote (India)


 ---

 WHAT WE'RE BUILDING


 See http://www.juliet.space


 We're building Juliet, an AI that runs marketing end to end. Our users are marketers, founders, CEOs, growth leads, agencies, and SMBs — not developers. They

 tell Juliet the goal. She plans, writes production code, and ships real marketing: conversion-optimized websites, launch assets, campaigns, audits, autonomously.


 That's the engineering problem in one line: the humans in the loop can't read code, so the agent has to get it right on her own — plan, build, self-correct,

 recover, ship.


 Under the hood: a browser-based studio backed by cloud sandboxes, a real-time SSE streaming pipeline, and a LangGraph agent working across 83 tools and 63 skill

 modules. The agent isn't bolted onto the product. She is the product.


 Small team, big ambitions. You'll ship things users touch daily, not write tickets about them.


 ---

 THE ROLE


 We're hiring one architect-level backend engineer to own Juliet's agentic infrastructure end to end. That means the agent graph, the execution environment, the

 streaming pipeline, the state and memory systems — and setting technical direction for the engineers working alongside you.


 This is a player-coach seat. You'll still write code every day, and your architectural calls become the product. You'll work directly with the founder. No PMs in

 between.


 Frontend is part of the system. You won't be leading it, but you'll need to understand how the agent's output reaches the browser and be able to ship full-stack

 features when needed.


 ---

 THE STACK


 AI agent (primary): Python 3.11, LangGraph 1.x + LangChain, Anthropic / Google / OpenAI model providers


 API (primary): NestJS 11, Supabase, Redis, PostgreSQL, Server-Sent Events


 Infra (primary): Modal cloud sandboxes, Docker, Netlify deployments


 Frontend (secondary): Next.js 15, React 19, TypeScript, Zustand, CodeMirror 6, XTerm.js


 Monorepo: Turborepo, pnpm


 ---

 WHAT YOU'LL WORK ON


 The majority of your time is here:


 Agentic AI workflows — Design, extend, and harden the LangGraph agent graph: multi-step planning, code generation, tool dispatch, self-correction, and recovery

 across 83 tools and 63 skill modules. This is the core of the product.


 Real-time streaming architecture — The SSE pipeline that carries every agent action from the Python backend through NestJS to the browser: event framing,

 reconnection, health monitoring, interrupt handling for plan approvals and clarifying questions.


 Agent execution environments — Sandbox lifecycle on Modal: container spin-up, file sync, terminal I/O, command execution, and live preview with per-asset esbuild

 bundling. The agent lives here.


 State and memory systems — LangGraph Postgres checkpointers, middleware-injected context (goals, design docs, memory anchors), conversation summarization. How

 the agent knows what it knows.


 Backend API and data layer — NestJS services, Supabase schema, Redis caching, quota enforcement, webhook handling. The plumbing the agent depends on.


 Marketing intelligence pipelines — AEO, CRO, and brand-perception audit engines: multi-LLM probing, parallel inference, streamed structured reports, result

 caching. Audit-at-scale infrastructure.


 The remaining ~25% of your time:


 Full-stack product features — Collaboration (roles and permissions), the Netlify deployment pipeline, subscription and quota flows, onboarding. You'll ship these

 end to end — backend first, frontend to close the loop.


 ---

 WHAT WE'RE LOOKING FOR


 Must-have:


 - 8+ years of professional software engineering, including meaningful time as a tech lead or systems architect who owned something end to end. Closer to ten is

 the norm for people who thrive here.

 - Both worlds on your resume: engineering rigor inside a large company and 0-to-1 ownership at an early-stage startup.

 - Production agentic systems experience. You've built and operated LLM agent systems in production with LangGraph, LangChain, or equivalent — agent graphs, tool

 use, state management, prompt engineering, evals. This means well beyond calling a chat endpoint.

 - Strong Python. You design and ship production Python daily. The agent codebase is yours to own.

 - Architect-level system design. You can own how data flows across four services, make tradeoffs under uncertainty, and defend every call.

 - AI-native development workflow. You drive Claude Code, Codex, or similar agentic tools as everyday instruments — not occasionally. You have opinions about

 working with coding agents because you do it constantly.

 - Real-time backend systems. You've built SSE, WebSocket, or streaming API infrastructure in production — not just consumed it.

 - Strong TypeScript. The API layer and most product features are in TypeScript. You're productive in it.


 Strong plus:


 - Background in developer tools, IDEs, or coding/execution platforms

 - Container runtimes and sandboxed execution (Modal, E2B, Firecracker, or similar)

 - Depth in PostgreSQL, Redis, and Supabase

 - LLM observability and evals tooling (LangSmith or similar)

 - NestJS or equivalent Node.js API framework experience

 - React/Next.js — enough to ship a full-stack feature without handoff

 - Exposure to marketing, growth, or publisher-facing products


 ---

 WHY THIS ROLE IS DIFFERENT

  

 You own the architecture. Not a feature factory. Not someone else's design doc. The technical execution of an AI product is yours to lead.


 The agent is the product. You're not adding AI to an existing system. You're building and operating the system that is the AI. Every architectural decision

 touches what Juliet can and can't do.


 Hard problems, always. The system spans cloud sandboxes, streaming infrastructure, multi-step agent graphs, and a full-stack web product — for non-technical

 users who can't course-correct a broken output. The bar is high.


 Small team, real leverage. Your code ships to users the same week. No layers of approval.


 ---

 HOW TO APPLY


 Send us:


 1. A short note on the most complex agentic system you've shipped: what broke, and what you'd redo. A link to something you've built that involves agent graphs, tool use, or autonomous multi-step execution


 2. What is one thing you would improve about Juliet? It could be a feature or a bug.

Read more
company logo
Ritesh Kalvellu
Posted by Ritesh Kalvellu
Bengaluru (Bangalore)
7 - 15 yrs
₹50L - ₹50L / yr (ESOP available)
skill iconPython
Large Language Models (LLM)

About us

MyRico builds personal AI agents for enterprise, the copilots and digital employees that make humans more productive. The MyRico agents sit at the intersection of enterprise memory, high-end security, and an ever-expanding set of capabilities. We're a small team shipping fast, and the product is live with real customers today.


The role

Full-time · Bangalore

You'll own the systems that make an autonomous agent trustworthy in production. This is not just prompt engineering, and it's not model training - it's that and all the engineering layers in between: model steering, memory architecture, orchestration design, latency optimization, deterministic vs non-deterministic systems tradeoff, deployment, and operator tooling.

Concretely, the kind of work you'd have done here last week would include enhancing agent memory systems, runtime and scheduling reliability and predictability, adding new capability and tools to deployed agents, operator tools, live system debugging and benchmarking various models for price, latency and quality. And that is just last week. We are a small nimble startup rapidly working to address customer needs in this growing space, so things change rapidly.


What we're looking for

  • More than 7 years of software engineering, with real production ownership of distributed or stateful systems - you've been paged for something you built and made it not happen again.
  • Strong understanding of LLM based native app building, combining classic and model driven applications to get the best of both. You've built on LLMs beyond demos: agent frameworks, tool use, context management, eval fixtures, and you know why "it worked in the transcript" isn't evidence.
  • Python and shell in production settings; comfortable in TypeScript/Node. You write boring, testable code and prefer the standard library to a new dependency.
  • Systems taste: append-only logs, idempotent reconciliation, fold-the-events state machines, and read-only debugging surfaces feel like home.
  • Evidence discipline: tests before features, claims backed by quoted observations, decisions written down.


Nice to have

  • Experience running the combination of multi-tenant and single-tenant / on-prem-style fleets with ability to handle per-customer isolation, upgrade paths, migration compatibility in both setups.
  • Security instincts for products that touch highly sensitive data and systems, including things like executives' email, calendars, and messages: least privilege, loopback-only services, secrets that never hit a log.
  • You already orchestrate AI coding agents in your own workflow and have opinions about where they break.


How we work

Small team, high trust, written decisions. Designs get adversarial review before code; PRs get automated review driven to zero open findings; features aren't done until verified on a live system. AI agents do a large share of the implementation, your leverage is judgment: framing the problem, freezing the right design, and knowing when the machine is wrong.

Read more
company logo
Rishu Dutta
Posted by Rishu Dutta
Gurugram
7 - 12 yrs
₹20L - ₹50L / yr
Retrieval Augmented Generation (RAG)
Agentic AI
Multi-agent Systems

Role Overview 

We are looking for an AI Engineer to design, build, and ship production AI systems, including agentic AI applications, for enterprise clients. This is a hands-on engineering role: you will write production code, build and evaluate models and agents, and work closely with architects and product teams to take solutions from prototype to scale. 


Key Responsibilities 

Design and build agentic AI systems: agent workflows, tool/function-calling, memory, and human-in-the-loop patterns. Build and productionise RAG pipelines, prompt-based applications, and LLM integrations across providers. Develop and maintain data and ML pipelines: feature engineering, model training, evaluation, and monitoring. Integrate AI systems with enterprise applications (CRMs, ERPs, ITSM tools) via APIs, events, and MCP-based tool servers. Implement guardrails, prompt-injection defences, and evaluation frameworks to keep AI systems safe and reliable in production. 

Write clean, tested, production-grade code and participate actively in code and design reviews. 

Collaborate with architects, product managers, and delivery teams to translate requirements into working AI solutions. Troubleshoot and optimise AI systems for accuracy, latency, and cost in production. 


Required Qualifications 

8–12 years of hands-on software engineering experience, with a strong, unbroken technical track record. Hands-on experience building and shipping AI/ML systems in production, not just POCs. 

Practical experience with agentic AI systems and at least one major agent framework (LangGraph, CrewAI, AutoGen, OpenAI Agents SDK, Bedrock Agents/Strands, or Semantic Kernel). 

Experience with LLM/GenAI systems: RAG pipelines, prompt engineering, structured outputs, and tool calling across providers. 

Strong Python skills (TypeScript/Node.js a plus), with production-grade testing, CI/CD, and API design practices. Working knowledge of ML fundamentals: model evaluation, feature engineering, and experimentation. Cloud-native experience on AWS and/or Azure: containers, serverless, event backbones, and vector databases. Understanding of LLM safety and reliability practices: guardrails, prompt-injection defences, and observability. 



Read more
company logo
Lakshit Bagga
Posted by Lakshit Bagga
Remote only
10 - 25 yrs
₹1L - ₹50L / yr
Odoo (OpenERP)


We're Hiring: Agentic Tools Architect

Contract | Remote-first | Global


We're currently recruiting for one of our clients, A company that specializes in acquiring enterprise software businesses and transforming them into AI-native, profitable, scalable operations. They run remote-first, globally distributed teams, and their internal operations platform (Odoo, Redmine, GitHub) currently runs on human-shaped workflows. They're looking for someone to rebuild it for agents.


The Role

Redesign the client's operational tooling so AI agents can create, route, and resolve work alongside people — with minimal oversight. You'll own the roadmap, build the plugins (Rails-first), and set the guardrails.


You'll:

  • Own feature roadmap across Odoo, Redmine, GitHub & adjacent systems
  • Design agent-operable APIs, MCP servers, and automation
  • Build Rails-based Redmine plugins, Odoo modules, GitHub Apps that survive version upgrades
  • Review AI-generated PRs and agent-authored config; set standards for access, integrity, observability

You bring:

  • 10+ years in software/platform engineering with deep SME knowledge of Odoo, Redmine, GitHub, or similar
  • Expert Ruby on Rails, production Redmine plugin experience
  • Real-world agentic dev experience (MCP, LLM tooling, agent workflows)
  • Strong grip on delivery/support/ops business processes


Nice to have: Python + Odoo modules, PostgreSQL, legacy modernization experience, PE-backed/SaaS scale-up background

Read more
company logo
Sandli Srivastava
Posted by Sandli Srivastava
Remote only
3 - 6 yrs
Best in industry
skill iconPython
Artificial Intelligence (AI)
skill iconReact.js
TypeScript
skill iconJavascript

About Us

We believe the future of software development is AI-native — where engineers operate at a higher level of abstraction and quality remains non-negotiable. 

Incubyte is a software craft consultancy where the “how” of building software matters as much as the “what”.  

We partner with companies of all sizes, from helping enterprises build, scale, and modernize to early-stage founders bring their ideas to life. 

Our engineers operate in an AI-native development model, using AI as a collaborator across the SDLC to accelerate development while upholding the discipline of software craftsmanship. Guided by Software Craftsmanship and Extreme Programming practices, we build reliable, maintainable, and scalable systems with speed, without compromising quality. If this way of building software resonates with you, we’d like to talk. 


Our Guiding Principles 

These principles define how we work at Incubyte. They are non-negotiable. 


Relentless Pursuit of Quality with Pragmatism 

  We build high-quality systems without losing sight of delivery. 

Extreme Ownership 

  We take responsibility end-to-end for decisions, execution, and outcomes. 

Proactive Collaboration 

  We collaborate closely, challenge each other, and solve problems together. 

Active Pursuit of Mastery 

  We continuously improve our craft and raise our bar. 

Invite, Give, and Act on Feedback 

We seek, give, and act on feedback to get better every day. 

Ensuring Client Success 

We act as trusted partners and focus on real outcomes, not just output. 


Experience Level


This role is ideal for engineers with total 3+ years of experience with a proven track record of shipping complex projects successfully.

An experienced individual contributor and leader who thrives in large, complex projects with widespread impact.


What You’ll Do as a Software Craftsperson 


  • Design and build high-quality, maintainable systems using disciplined engineering practices such as TDD, continuous refactoring, and pair programming 
  • Operate in an AI-native development model, using AI as a collaborator to explore architecture and design, accelerate development, and continuously improve systems while applying strong judgment to ensure that speed never compromises quality 
  • Take end-to-end ownership of outcomes from problem understanding and system design to implementation, deployment, and operation in production 
  • Make thoughtful design decisions that balance simplicity, scalability, and long-term maintainability in real-world systems 
  • Maintain a high bar for engineering quality through rigorous testing, code reviews, and continuous feedback 
  • Investigate and resolve production issues, and implement systemic improvements to prevent recurrence 
  • Work directly with clients, navigate ambiguity, and translate business problems into well-designed technical solutions 
  • Contribute to improving team practices, tooling, and systems to raise the overall quality and effectiveness of engineering 


Requirements


What You’ll Bring 


  • 3+ years of experience building high-quality, production systems (flexible based on demonstrated capability) 
  • Strong fundamentals in software engineering, including object-oriented design, system design, and testing practices such as TDD 
  • Demonstrated ability to build simple, maintainable, and scalable systems with a focus on long-term reliability 
  • Proficiency in one or more modern technologies, Python, PHP, JavaScript, or TypeScript, with the ability to learn new technologies quickly 
  • Deep experience working with Git in collaborative environments, including managing shared codebases, conducting code reviews, and maintaining a high bar for quality 
  • Ability to operate effectively in an AI-native workflow using AI as a collaborator to explore solutions and accelerate development, while applying strong judgment to ensure correctness, quality, and maintainability 
  • Clear thinking and strong problem-solving ability, with the capacity to break down complex problems into simple, well-structured solutions 
  • A strong sense of ownership — you take responsibility for outcomes, care deeply about quality, and are not comfortable shipping work that does not meet your standards.



Benefits


Life at Incubyte 


We are a remote-first company with structured flexibility. Teams commit to shared rhythms during core hours, ensuring smooth collaboration while maintaining autonomy. Twice a year, we come together in person for a co-working sprint and once a year for a retreat - with all travel expenses covered. 

Our environment is built for crafters: pairing, refactoring, experimenting with AI, and pushing the boundaries of software excellence. We are all lifelong learners, and our work is our passion. 


Perks

  • Dedicated learning & development budget. 
  • Sponsorship for conference talks. 
  • Comprehensive medical & term insurance. 
  • Employee-friendly leave policies. 
  • Home Office fund 
  • Medical Insurance
Read more
Remote only
3 - 6 yrs
₹20L - ₹30L / yr
Fullstack Developer
Fine-tuning LLMs
Model Context Protocol (MCP)
Artificial Intelligence (AI)
TypeScript
+2 more

About Us:


CLOUDSUFI, a Google Cloud Premier Partner, is a global leading provider of data-driven digital transformation across cloud-based enterprises. With a global presence and focus on Software & Platforms, Life sciences and Healthcare, Retail, CPG, financial services and supply chain, CLOUDSUFI is positioned to meet customers where they are in their data monetization journey.


Our Values


We are a passionate and empathetic team that prioritizes human values. Our purpose is to elevate the quality of lives for our family, customers, partners and the community.


Equal Opportunity Statement


CLOUDSUFI is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified candidates receive consideration for employment without regard to race, colour, religion, gender, gender identity or expression, sexual orientation and national origin status. We provide equal opportunities in employment, advancement, and all other areas of our workplace. Please explore more at https://www.cloudsufi.com/


Role :


A Software Engineer who builds the tools this company runs on. You build agent loops, and the loops build the solutions. You work towards a Company Brain that anyone here can ask.3–5 years’ experience · Reports to the CFO · 


THE KEY SKILL


You build the agent loops that build the solutions. You will not write every automation by hand. You build the loops that

write them. Ship a prototype every week. You ship something every day.


You’ll be building a Company Brain with access control. One system that holds what the company knows about finance,delivery and people. Anyone can ask it a question. Each person sees only what they are cleared to see.

One hard filter. If you cannot write and debug production code, and have not done it before, please do not apply.


CORE RESPONSIBILITIES


• Work the backlog: You pick items off a live, ranked backlog. You learn each function by building inside it. There is no discovery phase. What you learn goes back into the backlog and changes what comes next.


• Build the product: You design, build and ship tools that people use every day. Reconciliation, MIS, the deal desk,quote to cash, or whatever the real bottleneck turns out to be. You choose the tools and frameworks.


• Wire the data: Connect the systems each team already uses, so that the same number means the same thing everywhere.


• Make it visible: You build live dashboards and alerts that leaders read on their own, instead of asking someone for a report.


• Keep it running: You own uptime and accuracy for everything you build. Anything that touches money or people needs a person in the loop.


THE STACK


• Build with: Python and TypeScript. You write production code. Frontier model APIs from Anthropic, OpenAI or Google, with tool calling and structured output. At least one agent framework. MCP to connect agents to internal systems.Postgres and pgvector, or something similar, for retrieval. You deploy on GCP, and you debug your own work.


• Work in agents daily: Claude Code, Cursor or something like them, as the way you write code every day. You should have a clear view on how to run the loop, and on when a person has to step in.


• Connect to: The systems we already run on for accounting, CRM, hiring and IT support, along with Google Workspace.Most of the work is getting them to agree with each other.


• Check what you ship: Anything that produces a number needs a way to catch it going quietly wrong. Test sets, regression checks, and alerts on the output as well as on the job.


WHAT GOOD LOOKS LIKE


• Something you built is running by week two, and someone is using it.

• By day 90, time spent on reconciliation or reporting is down by a number you can defend to the CFO.

• Every tool you ship has a named owner who is still using it 60 days later. That is the measure that counts.

• Leaders stop asking for numbers, because they can already see them.

• By the end of your first year, a first version of the Company Brain answers real questions about Finance, and each

person who asks sees only what they are cleared to see.


WHO THIS IS FOR


• You have built products: 3 to 5 years at a software product company, on a product with real users at scale. That means 100k+ monthly active users, or heavy daily use by a large enterprise customer base. You have owned code in production, in front of real users, long after it shipped.

• You ship alone: You are comfortable as the only engineer in the room, and the only person on call for what you built.

• You work out new ground fast: A domain you do not know is interesting to you. You start without waiting for a spec or an expert.

• You are fluent with agents: One person cannot cover a whole company by hand. You use agent loops heavily and youare good at it.

• You write and speak clearly: Half this job is pulling a process out of a finance or delivery lead and giving it back to them correctly. You work remotely, so this matters a great deal.


HOW WE WILL ASSESS

• A design problem: Live. We give you a function of the B2B company, and you design the system for it. We watch how you break down a domain you do not know, how you size it, and what you leave out on purpose.

• A build exercise: Live and screen-shared, on your own setup, with your own agents. You build the way you normally build. We watch how you run the loop, when you step in, and what you decide to skip.

• Your work and your questions: We talk about what you have shipped before. You ask us whatever you want.


Communication is not a separate round. All three sessions are live, and how clearly you explain your thinking is part of how we judge you.


WHERE IT LEADS

You report to the CEO and CFO from your first day. Your charter covers the whole company. Nothing sits between you and production. Very few engineering jobs offer all three at once, and that is why this one exists.

In 18 months you will know how this company really runs: the data, the money, and the gaps between teams. The rolethen changes shape to fit whatever the biggest open problem is by then.

Read more
Remote, Bengaluru (Bangalore), Pune, Delhi, Gurugram, Noida
3 - 8 yrs
₹25L - ₹40L / yr
LangGraph
Web API
Information architecture
knowledge Graph
Semantic search

About Sentiaflow

Sentiaflow is an AI engineering and IT services company building production-grade agentic AI systems, sitting at the intersection of LLMs and real business operations — data, APIs, permissions, workflow state, human decisions, security, and measurable outcomes. Our initial domain focus is healthcare, particularly clinical trials, where reliability, traceability, and clear human-decision boundaries matter more than a slick demo.

Job Description

As a Level 1 engineer, you'll implement bounded parts of a production agentic workflow under an Agent Captain or senior engineer, connecting models to application services, tools, data sources, and human approval points.

You will:

  • Translate a scoped business workflow into typed inputs, outputs, states, actions, and escalation paths
  • Build backend services and tool integrations (Node.js/TypeScript or Python)
  • Use LLMs only where model judgment adds value; keep rules, validation, authorization, and workflow control in deterministic code
  • Design and validate structured model outputs before they touch downstream systems
  • Handle partial data, tool failures, duplicate events, retries, timeouts, rate limits
  • Add logging, traces, metrics, and decision records for diagnosability
  • Write tests and evaluation cases that check whether the full workflow behaves correctly — not just whether output sounds fluent
  • Protect sensitive data; participate in code, design, and release-readiness reviews
  • Explain implementation trade-offs clearly to engineers and stakeholders

Success in 6 months: own a bounded workflow module end-to-end, integrate models without letting probabilistic output bypass deterministic controls, produce release-ready evaluation evidence, and diagnose cross-boundary failures with less supervision.

Desired Skills

  • Approximately 3–6 years hands-on backend/application engineering experience, with demonstrable hands-on work building agentic systems — not just calling an LLM API from a backend service
  • LangGraph (or comparable agent orchestration framework) experience is required — building multi-step, stateful agent workflows with conditional branching, tool-calling loops, and recovery/retry logic, not a single-prompt wrapper
  • Deep RAG experience, including:
  • Chunking strategy design, embedding model selection, and retrieval evaluation (not just "connected a vector DB")
  • Hybrid search (dense + sparse/keyword), re-ranking, and query rewriting/decomposition
  • Handling retrieval failure modes: irrelevant context, stale data, contradictory sources, citation/grounding accuracy
  • Measuring RAG quality (precision/recall on retrieval, faithfulness/groundedness of generation) — not eyeballing outputs
  • Experience designing agent state machines / workflow graphs: tool selection, planning loops, human-in-the-loop interrupts, checkpointing, and state persistence across long-running workflows
  • Strong programming in JS/TypeScript (preferred), Python, Java, C#, or Go
  • Solid grasp of API design, databases, async processing, auth, testing, deployment
  • Comfort reasoning about state, retries, idempotency, concurrency, permissions, audit trails, failure recovery
  • Real production debugging experience, not just greenfield builds
  • Clear technical communication

We're looking for engineers who've actually built and tuned agentic/RAG systems in production — not those who've only wired together frameworks or prompted an LLM API.


Nice to have: experience with other orchestration frameworks (CrewAI, AutoGen, custom state machines), observability/eval tooling (LangSmith, Langfuse, custom trace pipelines), healthcare or regulated-industry background. Bachelors from IIT or NIT highly preferred.

Read more
company logo
Remote only
5 - 8 yrs
Best in industry
skill iconPython
skill iconReact.js
Artificial Intelligence (AI)

About Us


We believe the future of software development is AI-native — where engineers operate at a higher level of abstraction and quality remains non-negotiable. 

Incubyte is a software craft consultancy where the “how” of building software matters as much as the “what”.  


We partner with companies of all sizes, from helping enterprises build, scale, and modernize to early-stage founders bring their ideas to life. 

Our engineers operate in an AI-native development model, using AI as a collaborator across the SDLC to accelerate development while upholding the discipline of software craftsmanship. Guided by Software Craftsmanship and Extreme Programming practices, we build reliable, maintainable, and scalable systems with speed, without compromising quality. If this way of building software resonates with you, we’d like to talk. 


Our Guiding Principles 


These principles define how we work at Incubyte. They are non-negotiable. 


Relentless Pursuit of Quality with Pragmatism 


  We build high-quality systems without losing sight of delivery. 


Extreme Ownership 

  We take responsibility end-to-end for decisions, execution, and outcomes. 


Proactive Collaboration 


  We collaborate closely, challenge each other, and solve problems together. 


Active Pursuit of Mastery 


  We continuously improve our craft and raise our bar. 


Invite, Give, and Act on Feedback 


We seek, give, and act on feedback to get better every day. 


Ensuring Client Success 


We act as trusted partners and focus on real outcomes, not just output. 


Job Description


This is a remote position.


Experience Level


This role is ideal for engineers with total 5+ years of experience with a proven track record of shipping complex projects successfully.

An experienced individual contributor and leader who thrives in large, complex projects with widespread impact.


What You’ll Do as a Software Craftsperson 


  • Design and build high-quality, maintainable systems using disciplined engineering practices such as TDD, continuous refactoring, and pair programming 
  • Operate in an AI-native development model, using AI as a collaborator to explore architecture and design, accelerate development, and continuously improve systems while applying strong judgment to ensure that speed never compromises quality. 
  • Take end-to-end ownership of outcomes from problem understanding and system design to implementation, deployment, and operation in production 
  • Make thoughtful design decisions that balance simplicity, scalability, and long-term maintainability in real-world systems 
  • Maintain a high bar for engineering quality through rigorous testing, code reviews, and continuous feedback 
  • Investigate and resolve production issues, and implement systemic improvements to prevent recurrence 
  • Work directly with clients, navigate ambiguity, and translate business problems into well-designed technical solutions 
  • Contribute to improving team practices, tooling, and systems to raise the overall quality and effectiveness of engineering 



Requirements


What You’ll Bring 


  • 5+ years of experience building high-quality, production systems (flexible based on demonstrated capability) 
  • Strong fundamentals in software engineering, including object-oriented design, system design, and testing practices such as TDD 
  • Demonstrated ability to build simple, maintainable, and scalable systems with a focus on long-term reliability 
  • Proficiency in one or more modern technologies Python, React, AI, JavaScript, or TypeScript, with the ability to learn new technologies quickly 
  • Deep experience working with Git in collaborative environments, including managing shared codebases, conducting code reviews, and maintaining a high bar for quality 
  • Ability to operate effectively in an AI-native workflow using AI as a collaborator to explore solutions and accelerate development, while applying strong judgment to ensure correctness, quality, and maintainability 
  • Clear thinking and strong problem-solving ability, with the capacity to break down complex problems into simple, well-structured solutions 
  • A strong sense of ownership — you take responsibility for outcomes, care deeply about quality, and are not comfortable shipping work that does not meet your standards.



Benefits


Life at Incubyte ​


We are a remote-first company with structured flexibility. Teams commit to shared rhythms during core hours, ensuring smooth collaboration while maintaining autonomy. Twice a year, we come together in person for a co-working sprint and once a year for a retreat - with all travel expenses covered. 

 

Our environment is built for crafters: pairing, refactoring, experimenting with AI, and pushing the boundaries of software excellence. We are all lifelong learners, and our work is our passion. 


Benefits 



  • Dedicated learning & development budget. 
  • Sponsorship for conference talks. 
  • Comprehensive medical & term insurance. 
  • Employee-friendly leave policies. 
  • Home Office fund 
  • Medical Insurance 
Read more
company logo
Umama Sayed
Posted by Umama Sayed
Mumbai
5 - 8 yrs
Best in industry
skill iconPython
Large Language Models (LLM)
Artificial Intelligence (AI)
Prompt engineering
LangGraph
+6 more

Senior AI Engineer

Code Generation, Agent Architecture & LLM Systems

📍 Mumbai (On-site) | Full-time | 5+ years


About the Role:

Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.

We are hiring a Senior AI Engineer for a dedicated client engagement focused on building an AI-powered application builder platform - a product where users describe software in plain English and the system generates, previews, and iteratively refines working code.

The mandatory requirement for this role is hands-on production experience shipping LLM-powered systems with agent architectures, with experience in code generation or developer tooling contexts a strong advantage.


The role is product-focused and deeply hands-on. You will own everything between the user's prompt and correct code landing in the project: the agentic loop, code generation pipeline, context management, evaluation suite, and model cost strategy.

You will work alongside the Senior MLOps Engineer who operationalises the infrastructure around your system, and collaborate closely with backend, frontend, and DevOps engineers.


Responsibilities:


Agent Architecture

Design and own the agentic loop for the platform - request interpretation, planning, tool-calling sequence (read file, edit file, run build, search code, install package), and stop conditions.

Make and revisit architectural decisions on single-agent vs. multi-agent designs, including planner/executor splits and dedicated build-repair sub-agents.


Code Generation Pipeline

Own the end-to-end generation flow: task classification, context gathering, planning, targeted edits, verification, and commit.

Implement diff/search-replace-based file editing with fuzzy matching and fallback strategies.

Enforce scope discipline so the agent makes minimal diffs and does not modify code it was not asked to touch.


Self-Repair Loop

Build and tune the automated repair loop that pipes compiler, lint, build, and runtime errors back to the model with retry budgets and model escalation.

This loop is the primary quality lever - the difference between 60-70% and 90%+ build success rates.


Context Management

Build file-relevance retrieval so the agent sees the right files, not the whole codebase: dependency graphs, AST/tree-sitter-based chunking, embeddings, recency signals, and hybrid retrieval.

Implement conversation summarisation and memory for long sessions, and address long-project degradation through codebase summaries and periodic consistency passes.

Own token budgeting and prompt caching strategy.


Prompt Engineering as a Discipline

Own the system prompt and per-task prompt variants (new feature, bug fix, styling change).

Maintain few-shot examples and enforce coding conventions, stack rules, and prohibited behaviours such as no hardcoded secrets and no whole-file rewrites.

Version prompts like code with changelogs and rollback capability.


Evaluation and Quality Measurement

Design and own the evaluation suite: representative test prompts run on every prompt and model change, scored on build success rate, instruction adherence, and output quality including LLM-as-judge and visual/screenshot checks where relevant.

Define regression gates that block quality-degrading changes from shipping.

Treat evals the way engineers treat automated testing: versioned, automated, and tracked over time.

This responsibility is non-negotiable at this level.


Model Strategy and Cost

Design model routing - cheap and fast models for classification and small edits, frontier models for complex generation.

Drive cost optimisation through prompt caching, diff-based edits over full-file rewrites, and tighter context selection.

Track cost per agent run and tokens per task; evaluate new model releases against the eval suite and lead migrations when results justify it.


Safety and Reliability of Agent Behaviour

Defend against prompt injection from user content and fetched web content.

Ensure secrets never appear in generated client code.

Define what the agent's tools may and may not do in collaboration with the platform team.

Contribute to output moderation and abuse-pattern awareness.


Mentorship and Engineering Standards

Run code reviews, define engineering conventions for AI work, and raise the engineering bar across the AI team.

Work closely with the Senior MLOps Engineer on handoff of eval design, prompt configurations, and model routing logic.


Requirements:


Hands-on Production Ownership of LLM-Powered Systems with Agent Architectures (Mandatory)

Must have personally shipped and operated at least one complex production AI system - agentic, multi-step, or code generation - with end-to-end ownership of architecture, evaluation, and cost.

POCs, internal demos, and tutorial-grade work do not qualify.


5+ Years of Professional Software or AI Engineering Experience

With at least 3 years focused on LLM applications, AI engineering, or production AI systems.

Candidates with strong backend backgrounds and a clear, substantive pivot into LLM systems qualify.


Strong Python Proficiency and Service Development

Production-grade Python with FastAPI or equivalent: type hints, async patterns, streaming responses, testing, and packaging.

Not notebook-only.


Depth Across LLM APIs and Agent Systems

Production experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or open-weight models (vLLM, Ollama, Together).

Production experience with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.

Hands-on with tool calling, structured outputs, and multi-step reasoning.


Demonstrated, Systematic Evaluation Practice - Non-Negotiable

Must have built evaluation harnesses that gate production releases, not ad-hoc testing.

Hands-on with at least one of LangSmith, Langfuse, Promptfoo, Ragas, or DeepEval.

Candidates with no systematic answer to evaluation should not be considered at senior level regardless of other strengths.


Cost Discipline for Production AI

Track record of measurable cost optimisation on production AI features.

Able to speak in specifics: cost per request, savings achieved through caching or model routing, context reduction decisions.


AWS Working Knowledge

Hands-on with EC2, S3, IAM, and Docker.

Comfort with CI/CD workflows and deploying AI services.


Awareness of LLM Security Failure Modes

Familiar with prompt injection patterns, understands that system prompt rules alone are insufficient, and has experience with output validation and content safety in production.


Nice to Have

  • Experience with AST/tree-sitter tooling, diff-based editing systems, or compiler-adjacent work
  • MCP server authoring
  • Open-source AI contributions
  • Published technical writing on LLM systems
  • Multi-modal model experience
  • Fine-tuning exposure (LoRA, QLoRA, PEFT)
Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos