AI/ML Engineer - DevOps at Technoidentity · Hyderabad · 0 - 3 years · Profitable · Posted 16 Jul 2026

Supercharge Your Career as a AI DevOps Engineer at Technoidentity!
At Technoidentity, we're a Data & AI product engineering company with over 15 years of expertise in building durable digital products, intelligent enterprise solutions, and scalable Data & AI platforms. As we continue expanding globally, it's the perfect time to join our team of tech innovators and make a lasting impact.
What’s in it for You?
We are looking for an AI DevOps Engineer with 0–3 years of experience who is passionate about AI, Cloud, DevOps, and Automation. The role involves building, deploying, and managing AI-powered applications, LLM solutions, and cloud-native platforms while ensuring reliability, scalability, security, and observability.
What Will You Be Doing?
- Develop and deploy AI/ML and Generative AI solutions using Python.
- Build applications leveraging LLMs, RAG, and AI agents.
- Create and maintain CI/CD pipelines for AI applications.
- Deploy and manage workloads using Docker and Kubernetes.
- Support cloud platforms (AWS, Azure, or GCP).
- Implement Infrastructure as Code (Terraform) and automation workflows.
- Monitor applications using observability tools such as Prometheus, Grafana, and logging platforms.
- Collaborate with engineering teams to ensure system reliability, performance, and security.
- Contribute to MLOps practices, AI accelerators, and reusable frameworks.
Requirements
What Makes You the Perfect Fit?
- Python programming (mandatory)
- Understanding of Machine Learning, LLMs, Prompt Engineering, and RAG
- Experience with OpenAI, LangChain, LlamaIndex, or Hugging Face
- Docker, Kubernetes, Git, and CI/CD tools
- AWS, Azure, or GCP
- PostgreSQL; MongoDB and Vector Databases are a plus
- Basic knowledge of MLOps, Terraform, and workflow orchestration tools (Airflow/Temporal)
- Familiarity with observability and monitoring tools
Qualifications
- Bachelor's degree in Computer Science, AI, Data Science, IT, or related field
- 0–3 years of experience in AI/ML, Software Engineering, Cloud, DevOps, or related areas
Nice to Have
- Experience with Agentic AI frameworks
- Knowledge of MLOps and AI platform operations
- Exposure to enterprise-grade monitoring, reliability engineering, and security best practices

About Technoidentity
About
Founded a decade ago, a group of seasoned professionals came together to ask ‘How might we use technology to bring the world closer to social justice?’ A small team was put together to create incredible solutions for organisations working towards delivering large scale impact. Today, we’re a global organisation providing disruptive software solutions to our customers who are solving important problems in their industry and also the world we live in. At TechnoIdentity, we challenge our team of passionate technologists to create forward-looking solutions for our customers that improve operational efficiency and bolster their bottom line. We empower them with the tools, methodologies and decision-making to make real impact, fast. We invite creative minds and businesses to TechnoIdentity, let’s collaborate to transform lives through technology.
Life at TechnoIdentity is shaped by the belief that everyone deserves to bring their authentic self to work and as responsible adults, you will choose to spend your time wisely on things matters the most to you.
Tech stack
Similar jobs (10)
Job Description – AI Engineer (End-to-End Development & Deployment)
Role Summary
We are looking for an AI Engineer with hands-on experience in designing, developing, deploying, and maintaining Generative/Agentic AI solutions in production. The ideal candidate should have end-to-end ownership of AI applications, from development to deployment, monitoring, and optimization.
Key Responsibilities
● Design, build, and deploy Generative/Agentic AI solutions.
● Develop applications using LLMs, RAG, AI agents, and vector databases.
● Build scalable APIs and integrate AI solutions with enterprise applications.
● Implement CI/CD pipelines, containerization, and MLOps best practices.
● Monitor, optimize, and maintain production AI systems.
● Collaborate with cross-functional teams to deliver business-driven AI solutions.
Required Skills
● Strong programming skills in Python.
● Experience with vector databases (e.g., Pinecone, FAISS, ChromaDB) and graph memory systems
● Knowledge of atleast one agent development framework: Google ADK (preferred), LangChain/LangGraph/LlamaIndex, CrewAI
● Experience with LLMs, RAG, GenAI, AgenticAI Agents
● Hands-on experience with FastAPI, and REST APIs.
● Knowledge of Docker, Kubernetes, Git, CI/CD.
● Experience with AWS, Azure, or GCP.
● Experience with security compliance, monitoring and observability tools such as AWS CloudWatch, Azure Monitor, Google Cloud Monitoring.
Company Overview:
Planview is hiring a DevOps Engineer in Bengaluru, India to support Planview SaaS applications across the product line. You will work in a global, collaborative team — owning CI/CD pipelines, cloud infrastructure, and automation to keep deployments fast and systems reliable.
Responsibilities
- Build and maintain CI/CD pipelines in Jenkins for reliable, fast delivery.
- Manage containerized workloads on Docker and ECS — task definitions, services, and clusters.
- Provision and manage AWS infrastructure using Terraform (CloudFormation a plus).
- Automate configuration and deployment tasks using Python (Ansible a plus).
- Set up and maintain monitoring and alerting via New Relic (CloudWatch, Datadog, or Prometheus/Grafana a plus).
- Write Shell and Python scripts to automate operations and reduce manual work.
- Manage Git workflows — branching, merge strategies, and pull request reviews.
- Administer and support MSSQL databases underpinning the product line — backups, restores, and basic performance troubleshooting.
- Troubleshoot deployment, performance, and infrastructure issues with development teams.
- Participate in on-call rotations and drive incident response.
- Continuously improve infrastructure resilience and deployment speed.
- Apply AI-assisted engineering tools (e.g., GitHub Copilot, Claude Code) to speed up IaC authoring, pipeline debugging, and day-to-day scripting.
Qualifications
Must-Have Skills
- Experience: 4–6 years of experience in DevOps, SRE, or Infrastructure Engineering.
- OS: Linux & Windows administration (systemd, package management, log analysis).
- Cloud: AWS (Active Directory, ECS, EC2, CloudFront, S3, VPC, IAM, RDS, Lambda basics).
- Source Control: Git — branching, merge/rebase, PR reviews.
- CI/CD (Jenkins): Pipeline creation and basic Groovy scripting.
- Containerization (Docker + ECS): Task definitions, services, and clusters.
- IaC: Terraform.
- Monitoring: New Relic.
- Scripting: Bash and Python scripting for automation.
- Networking Basics: DNS, load balancers, security groups, VPNs.
- Logging: ELK stack / CloudWatch Logs.
- Database: MSSQL administration — backups, restores, basic performance troubleshooting.
- Infrastructure Automation: Hands-on experience automating infrastructure provisioning, configuration, and deployment end-to-end.
- AI-Assisted Engineering: Comfortable working with AI coding/DevOps assistants (e.g., GitHub Copilot, Claude Code) for IaC generation, scripting, and troubleshooting — verified via a mandatory AI proficiency assessment during interviews.
Nice-to-Have Skills
• Configuration Management: Ansible.
• Additional IaC: CloudFormation.
• Architecture: Knowledge of microservices architecture.
• Cloudflare: DNS, CDN, WAF.
• Artifact Repositories: Nexus, JFrog Artifactory, ECR.
• Other CI/CD Tools: GitHub Actions.
• AWS cost optimization / FinOps awareness.
• Datadog, CloudWatch, or Prometheus/Grafana.
• AIOps: Exposure to AI-driven anomaly detection, root-cause analysis, or incident triage (e.g., Dynatrace Davis AI, Datadog Bits AI, Harness AIDA).
• Database Basics: RDS backups, restores, performance tuning.
Amura’s Vision
We believe that the most under-appreciated route to releasing untapped human potential is to build a healthier body, and through which a better brain. This allows us to do more of everything that is important to each one of us.
Billions of healthier brains, sitting in healthier bodies, can take up more complex problems that defy solutions today, including many existential threats, and solve them in just a few decades.
Billions of healthier brains will make the world richer beyond what we can imagine today. The surplus wealth, combined with better human capabilities, will lead us to a new renaissance, giving us a richer and more beautiful culture.
These healthier brains will be equipped with deeper intellect, be less acrimonious, more magnanimous, and have a kinder outlook on the world, resulting in a world that is better than any previous time.
We find this vision of the future exhilarating. Our hopes and dreams are to create this future as quickly as possible and ensure that it is widely distributed and optimized to maximize all forms of human excellence.
Role Overview
We are looking for a highly skilled Senior DevOps Engineer (AI-Native Infrastructure & Platform Engineering) with deep expertise in AWS cloud infrastructure, automation, AI infrastructure operations, and modern DevOps/SRE practices.
This role goes beyond traditional DevOps and requires a seasoned specialist capable of building and operating AI-ready infrastructure platforms that support high-throughput APIs, LLM/AI workloads, GPU-based compute, data-intensive systems, real-time inference pipelines, and scalable ML platforms.
You will be responsible for architecting, automating, securing, and optimizing highly scalable and cost-efficient cloud environments that enable high-velocity engineering and AI teams. This is an ideal position for someone who combines technical ownership, an automation-first mindset, and a passion for developer productivity and platform reliability.
Key Responsibilities
Cloud Infrastructure & Platform Engineering (AWS)
- Architect, deploy, and manage highly scalable and secure infrastructure on AWS. Design cloud platforms supporting AI/ML workloads, data pipelines, real-time APIs, and high-concurrency backend systems.
- Hands-on expertise with key AWS services including EC2, ECS/EKS, Lambda, RDS, DynamoDB, S3, VPC, CloudFront, IAM, CloudWatch, and GPU-enabled instances.
- Build and maintain Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK.
- Design multi-AZ and multi-region architectures for high availability and disaster recovery (HA/DR).
- Build reusable platform templates and shared infrastructure modules.
AI/ML Infrastructure & MLOps
- Build and maintain infrastructure for LLM applications, AI inference workloads, model serving platforms, vector databases, and feature stores.
- Support GPU-based workloads and optimize compute/storage usage.
- Enable scalable deployment patterns for AI applications using Kubernetes/EKS. Collaborate with Data Science and ML Engineering teams on model deployment, training/tuning of models, CI/CD for ML systems, experiment environments, and reproducibility.
- Support orchestration and deployment of AI workflows and inference services while implementing observability and reliability for AI pipelines.
CI/CD, Automation & Developer Productivity
- Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline.
- Automate deployments, environment provisioning, and release workflows.
- Build self-service developer platforms, preview environments, and reusable deployment workflows to improve developer productivity.
- Implement automated patching, scaling, backups, cleanup workflows, and drift detection.
Containers, Kubernetes & Platform Reliability
- Manage Docker-based environments, containerized applications, and optimize workloads using Kubernetes (EKS) or ECS/Fargate.
- Manage autoscaling, cluster health, node pools, ingress, service mesh, and workload isolation.
- Optimize infrastructure for performance, resilience, and cost-efficiency.
- Implement progressive deployment strategies including blue/green, canary, and rolling deployments.
Observability, Incident Response & SRE Practices
- Implement observability stacks using CloudWatch, Prometheus, Grafana, ELK, Datadog, OpenTelemetry, or New Relic.
- Build actionable dashboards and intelligent alerting systems while defining and tracking SLIs, SLOs, and SLAs.
- Lead incident response, root cause analysis, and blameless postmortems to reduce operational toil and improve MTTR.
FinOps, Cost Governance & Security
- Continuously monitor and optimize cloud costs (compute utilization, storage lifecycle, GPU usage, and data transfer) using AWS Cost Explorer, Budgets, Trusted Advisor, CloudHealth, or Kubecost.
- Implement AWS security best practices for IAM, VPCs, security groups, NACLs, encryption, and manage secrets using KMS, SSM Parameter Store, or Vault.
- Build secure CI/CD pipelines with automated security checks, least-privilege access, audit logging, and ensure compliance readiness for ISO 27001, SOC2, and GDPR.
Collaboration, Leadership & Platform Culture
- Work closely with engineering, AI/ML, QA, product, and operations teams to drive a DevOps, SRE, GitOps, and automation-first culture.
- Mentor junior DevOps and Platform Engineers while creating and maintaining detailed runbooks, architecture diagrams, and platform documentation.
Skills & Qualifications
Must-Have:
- 7+ years of experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure Engineering.
- Strong expertise in AWS cloud architecture, services, and deep understanding of Kubernetes (EKS), containers, and cloud-native systems.
- Strong Infrastructure-as-Code expertise using Terraform, CloudFormation, or CDK. Strong Linux administration, networking, DNS, routing, and load balancing knowledge. Strong scripting/programming experience in Python, Bash, or Go (preferred). Experience with CI/CD automation, GitOps workflows, and observability platforms supporting scalable production systems.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Preferred / Nice-to-Have:
- Experience with AI/ML infrastructure, MLOps, model serving, vector databases, GPU orchestration, and inference optimization.
- Familiarity with Kafka, Redis, SQS, and event-driven systems.
- Exposure to platform engineering, internal developer platforms, and tools like ArgoCD, Flux, Helm, and OpenTelemetry.
- AWS Certifications: Solutions Architect, DevOps Engineer, or SysOps Administrator. Knowledge of distributed systems and large-scale platform operations.
Here are answers to some questions you may have
Where is your office?
Chennai (Velachery)
Work Model
Work from Office – because great stories are built in person!
Do you have an online presence?
https://amura.ai (we are @AmuraHealth on all social media)
Role: AI Developer
Experience: 3–4 Years
Employment Type: Full-Time
Location: Goregaon, Mumbai
About the Role
We are looking for an experienced AI Developer with 3–4 years of software development experience and strong hands-on exposure to Generative AI, AI Agents, Copilots, and AI-powered application development.
The candidate will be responsible for building production-ready AI solutions, developing agentic workflows, modernizing legacy applications, and integrating LLM capabilities into enterprise applications.
Key Responsibilities
- Design, develop, and deploy AI Agents and agentic workflows for enterprise use cases.
- Build AI Copilots and LLM-powered applications using modern AI frameworks and APIs.
- Develop RAG-based applications using embeddings, vector databases, and enterprise data.
- Work on legacy application migration and modernization, leveraging AI-assisted development and code transformation techniques.
- Analyze legacy codebases and design strategies for AI-driven migration, refactoring, and modernization.
- Integrate LLMs with enterprise applications, APIs, databases, and third-party systems.
- Implement tool calling, function calling, multi-agent workflows, and workflow automation.
- Perform prompt engineering, context optimization, model evaluation, and AI application testing.
- Take ownership of AI solutions from POC and prototyping through production deployment.
- Collaborate with product managers, architects, and engineering teams to convert business requirements into scalable AI solutions.
- Stay updated with emerging technologies in Generative AI, Agentic AI, LLMs, and AI-assisted software development.
Required Skills
- 3–4 years of professional software development experience.
- Strong proficiency in Python and/or JavaScript/TypeScript.
- Hands-on experience developing Generative AI / LLM-based applications.
- Strong understanding of AI Agents, RAG, Prompt Engineering, LLM APIs, and embeddings.
- Experience with frameworks such as LangChain, LangGraph, Semantic Kernel, AutoGen, or equivalent.
- Experience working with REST APIs, databases, Git, and cloud environments.
- Hands-on experience with vector databases such as Pinecone, Weaviate, Chroma, FAISS, or equivalent.
- Good understanding of software architecture, debugging, testing, and deployment practices.
Good to Have
- Experience with Microsoft Copilot / Copilot Studio.
- Experience working with Claude, OpenAI, Gemini, Azure OpenAI, or open-source LLMs.
- Experience in legacy application migration, modernization, or code conversion.
- Knowledge of Azure AI / AWS / Google Cloud AI services.
- Experience with MCP, multi-agent systems, tool calling, and AI orchestration.
- Experience building enterprise-grade AI solutions with focus on security, scalability, and performance.
Job Summary/ Job Opportunity:
This is an excellent opportunity for an ideal candidate with a high level of technical proficiency and meeting the below mentioned criteria -- • Strong experience in Machine Learning, Deep Learning, Generative AI, and Large Language Models (LLMs). • Hands-on experience building and deploying production-grade solutions using Azure OpenAI, OpenAI, LangChain, LangGraph, Semantic Kernel, LlamaIndex, and Agentic AI frameworks. • Strong expertise in Python, API development, microservices, and cloud-native architectures. • Experience designing and implementing RAG solutions, vector databases, embeddings, knowledge retrieval systems, and AI copilots. • Experience with Azure cloud services, MLOps, CI/CD pipelines, monitoring, and model lifecycle management. • Strong understanding of AI governance, responsible AI, security, compliance, and model evaluation frameworks. • Ability to lead technical discussions, provide architectural recommendations, mentor team members, and interact with business stakeholde
Key Objectives and Major Responsibilities:
• Design, develop, and implement scalable AI/ML and Generative AI solutions for enterprise applications. • Lead development of intelligent applications leveraging LLMs, RAG pipelines, AI agents, and document intelligence solutions. • Collaborate with business stakeholders, architects, and product teams to translate business requirements into technical solutions. • Design and optimize data pipelines, vector search solutions, embeddings, and retrieval mechanisms. • Build and maintain REST APIs, microservices, and cloud-native AI applications. • Ensure best practices in coding standards, performance optimization, security, scalability, and maintainability. • Drive AI solution deployment using MLOps practices, CI/CD pipelines, monitoring, and observability frameworks. • Perform code reviews, mentor junior developers, and contribute to capability building within the team
Key Capabilities and Competencies:
Knowledge, Skills, Qualification and Experience
• Degree in B.Tech/M.Tech (Computer Science/IT/Data Science) or related discipline preferred, with 3–4 years of relevant experience in AI/ML, GenAI and total 5-7 years of experience. • Proficiency in Python and hands-on experience with ML libraries (scikit-learn, TensorFlow, PyTorch) and GenAI frameworks/tools. • Strong understanding of machine learning, deep learning, LLMs, prompt engineering, and techniques like RAG and fine-tuning. • Experience with data processing, embeddings, vector databases, APIs, and building scalable AI driven applications. • Good communication skills, ability to work on multiple projects, and eagerness to learn and adapt to evolving AI technologies.
Experience - 4 to 6 year
Location – Ahmedabad/Pune/Indore
- Additional Job Description
Additional Job Description
Required Skills and Experience:
- Strong proficiency in Python and experience with ML/AI libraries (scikit-learn, TensorFlow, PyTorch, Hugging Face ecosystem).
- Hands-on experience with LLMs, RAG, vector databases, and retrieval pipelines.
- Practical experience deploying agentic workflows and building multi-step, tool-enabled agents.
- Experience using Garak (or similar LLM red-teaming/vulnerability scanners) to identify model weaknesses and harden deployments.
- Demonstrated experience implementing content filtering / moderation systems.
- Solid skills working with structured and unstructured data and advanced feature engineering.
- Familiarity with cloud GenAI platforms and services (Azure AI Services preferred; AWS/GCP acceptable).
- Experience building APIs/microservices; containerization (Docker), orchestration (Kubernetes).
- Strong understanding of model evaluation, performance profiling, inference cost optimization, and observability.
- Good knowledge of security, data governance, and privacy best practices for AI systems.
About Us
We believe the future of software development is AI-native — where engineers operate at a higher level of abstraction and quality remains non-negotiable.
Incubyte is a software craft consultancy where the “how” of building software matters as much as the “what”.
We partner with companies of all sizes, from helping enterprises build, scale, and modernize to early-stage founders bring their ideas to life.
Our engineers operate in an AI-native development model, using AI as a collaborator across the SDLC to accelerate development while upholding the discipline of software craftsmanship. Guided by Software Craftsmanship and Extreme Programming practices, we build reliable, maintainable, and scalable systems with speed, without compromising quality. If this way of building software resonates with you, we’d like to talk.
Our Guiding Principles
These principles define how we work at Incubyte. They are non-negotiable.
Relentless Pursuit of Quality with Pragmatism
We build high-quality systems without losing sight of delivery.
Extreme Ownership
We take responsibility end-to-end for decisions, execution, and outcomes.
Proactive Collaboration
We collaborate closely, challenge each other, and solve problems together.
Active Pursuit of Mastery
We continuously improve our craft and raise our bar.
Invite, Give, and Act on Feedback
We seek, give, and act on feedback to get better every day.
Ensuring Client Success
We act as trusted partners and focus on real outcomes, not just output.
Experience Level
This role is ideal for engineers with total 3+ years of experience with a proven track record of shipping complex projects successfully.
An experienced individual contributor and leader who thrives in large, complex projects with widespread impact.
What You’ll Do as a Software Craftsperson
- Design and build high-quality, maintainable systems using disciplined engineering practices such as TDD, continuous refactoring, and pair programming
- Operate in an AI-native development model, using AI as a collaborator to explore architecture and design, accelerate development, and continuously improve systems while applying strong judgment to ensure that speed never compromises quality
- Take end-to-end ownership of outcomes from problem understanding and system design to implementation, deployment, and operation in production
- Make thoughtful design decisions that balance simplicity, scalability, and long-term maintainability in real-world systems
- Maintain a high bar for engineering quality through rigorous testing, code reviews, and continuous feedback
- Investigate and resolve production issues, and implement systemic improvements to prevent recurrence
- Work directly with clients, navigate ambiguity, and translate business problems into well-designed technical solutions
- Contribute to improving team practices, tooling, and systems to raise the overall quality and effectiveness of engineering
Requirements
What You’ll Bring
- 3+ years of experience building high-quality, production systems (flexible based on demonstrated capability)
- Strong fundamentals in software engineering, including object-oriented design, system design, and testing practices such as TDD
- Demonstrated ability to build simple, maintainable, and scalable systems with a focus on long-term reliability
- Proficiency in one or more modern technologies, Python, PHP, JavaScript, or TypeScript, with the ability to learn new technologies quickly
- Deep experience working with Git in collaborative environments, including managing shared codebases, conducting code reviews, and maintaining a high bar for quality
- Ability to operate effectively in an AI-native workflow using AI as a collaborator to explore solutions and accelerate development, while applying strong judgment to ensure correctness, quality, and maintainability
- Clear thinking and strong problem-solving ability, with the capacity to break down complex problems into simple, well-structured solutions
- A strong sense of ownership — you take responsibility for outcomes, care deeply about quality, and are not comfortable shipping work that does not meet your standards.
Benefits
Life at Incubyte
We are a remote-first company with structured flexibility. Teams commit to shared rhythms during core hours, ensuring smooth collaboration while maintaining autonomy. Twice a year, we come together in person for a co-working sprint and once a year for a retreat - with all travel expenses covered.
Our environment is built for crafters: pairing, refactoring, experimenting with AI, and pushing the boundaries of software excellence. We are all lifelong learners, and our work is our passion.
Perks
- Dedicated learning & development budget.
- Sponsorship for conference talks.
- Comprehensive medical & term insurance.
- Employee-friendly leave policies.
- Home Office fund
- Medical Insurance
About the role
We are building AI systems that read, understand and act on real business documents, bank statements, financial reports, policy documents and forms and putting them into production where accuracy and cost both matters.
This is not a research role and it is not a prompt-writing role. You will own features end to end: pick and deploy open-source models, build the pipelines around them, measure whether they actually work on our documents, drive the cost per document down, and keep the whole thing running in production.
You will work closely with the engineering and product teams, and your work will be directly used by business users from day one.
What you will do
Deploy and evaluate open-source models
- Select, deploy and benchmark open-source LLMs and vision-language models for specific, narrow use cases not general chat.
- Build evaluation sets from real documents and define what "good" means numerically (field-level accuracy, extraction recall, hallucination rate) before shipping.
- Run structured comparisons between models and approaches, and write up the trade-offs so the team can make a decision.
- Apply quantization, batching and other optimizations to fit models into a sensible GPU budget.
Build and optimize AI orchestration
- Design multi-step pipelines that combine deterministic code, ML models and LLM calls and know when not to use an LLM.
- Optimize for latency, cost and reliability: caching, batching, request routing, fallback tiers, retries and graceful degradation.
- Instrument pipelines so failures are visible and traceable rather than silent.
Ship to production
- Package models and services with Docker, expose them behind clean APIs, and deploy them to our GPU and CPU infrastructure.
- Handle the unglamorous production concerns: cold starts, timeouts, concurrency limits, versioning, rollback and monitoring.
- Own on-call-style responsibility for the AI features you build, including cost tracking.
Must-have skills
Programming & engineering
- Strong Python: type hints, async/await, dataclasses/Pydantic, clean module design, testing.
- REST API development with FastAPI (or Flask/Django with a willingness to move to FastAPI).
- Git, code review discipline, and the ability to write code someone else can maintain.
- Comfortable in Linux and on the command line.
Machine learning fundamentals
- Working knowledge of PyTorch and the Hugging Face ecosystem (transformers, tokenizers, accelerate).
- Understanding of inference-time concepts: tokenization, context windows, batching, precision (FP16/BF16/INT8), memory footprint.
- Ability to read a model card and a paper well enough to judge whether a model fits a use case.
Document processing
- Hands-on experience with at least two of: pypdfium2, PyMuPDF, pdfplumber, pdfminer.six, Docling, Unstructured, Surya, DocTR, LayoutLM family.
- Practical OCR experience (Tesseract, PaddleOCR, or a cloud OCR) and an understanding of when OCR is the wrong tool.
- Experience extracting tables from PDFs and dealing with merged cells, multi-line rows, and inconsistent column layouts.
Strongly preferred
You will be a much stronger candidate with any of these. We do not expect all of them.
Model serving & optimization
- vLLM, TGI, Ollama, llama.cpp, or Triton Inference Server.
- Quantization formats and tooling: GGUF, AWQ, GPTQ, bitsandbytes, ONNX Runtime, INT8 export.
- Serverless GPU platforms: Modal, RunPod, Replicate, Baseten including cold-start and container-lifecycle management.
- LoRA / QLoRA fine-tuning with PEFT for narrow, task-specific improvements.
Vision-language models
- Practical use of open VLMs: Qwen2.5-VL, InternVL, Granite Vision, Molmo, Phi-Vision, or similar.
- Awareness of where VLMs hallucinate especially on numeric and financial content and patterns for constraining them (using the model for layout only, sourcing values from the text layer, constrained decoding).
Orchestration & pipelines
- Workflow orchestration: Dagster, Airflow, Prefect, or Temporal.
- Async job patterns: Celery, RQ, or platform-native spawn/poll patterns.
- LLM orchestration frameworks (LangGraph, LlamaIndex, Haystack) with the judgement to know when plain Python is a better answer.
- Structured output enforcement: Instructor, Outlines, XGrammar, JSON schema / tool-use modes.
Evaluation & observability
- Building golden datasets and regression suites for extraction tasks.
- Eval tooling: promptfoo, DeepEval, Ragas, or in-house harnesses.
- LLM tracing and monitoring: Langfuse, Arize Phoenix, LangSmith, OpenTelemetry.
Nice extras
- Rule engines and policy evaluation (Open Policy Agent / Rego, Drools, rule-engine).
- Experience in fintech, lending, insurance or accounting documents.
- Handling of PII and data-security practices in document pipelines.
- Contributions to open-source ML or document-processing projects.
Why join us
- Real production ownership from month one your work goes to actual users, not a demo.
- Genuinely hard technical problems in document AI, not wrappers over an API.
- Small team, short decision cycles, direct access to leadership.
- Budget and freedom to evaluate and adopt new open-source models as they land.
To apply: send your CV along with a short note on one AI system you have taken to production what it did, what the accuracy was, and what broke.
Job Description: AI Engineer – GenAI Platform Automation
Experience: 10+ Years
Location: Remote – Pan India
Employment Type: Haparz Payroll
Work Mode: Remote
Notice Period: Immediate / Short Notice Preferred
About the Role
We are looking for a senior AI Engineer – GenAI Platform Automation to lead automation initiatives across enterprise Generative AI, Data Science, Data Engineering, and Analytics platforms.
The role focuses on building scalable, secure, and self-service automation capabilities across infrastructure provisioning, CI/CD, cloud environments, AI workload deployment, governance, observability, and operational excellence. The ideal candidate will have strong hands-on experience in platform engineering, cloud automation, DevOps, Infrastructure-as-Code, Python, and enterprise GenAI ecosystems.
Key Responsibilities
- Lead end-to-end automation initiatives for enterprise GenAI, Data Science, Data Engineering, Metadata, Data Quality, Event Streaming, and Analytics platforms.
- Design self-service automation for infrastructure provisioning, environment onboarding, deployment, governance, monitoring, and operational workflows.
- Build automation capabilities supporting the AI lifecycle, including experimentation, model training, deployment, inference, observability, and lifecycle management.
- Develop scalable Infrastructure-as-Code solutions using Terraform and cloud-native automation frameworks.
- Design and maintain enterprise CI/CD pipelines, automated testing, deployment, and release processes using modern DevOps toolchains.
- Automate Kubernetes, containers, serverless, and distributed computing environments in collaboration with cloud and platform engineering teams.
- Develop automation solutions for GenAI and Agentic AI applications, including MCP-enabled services, API integrations, workflow automation, and event-driven architectures.
- Implement observability, monitoring, logging, tracing, alerting, automated remediation, and reliability engineering practices.
- Work with architecture, security, governance, engineering, and business teams to ensure enterprise standards and compliance requirements are met.
- Conduct technical design reviews, automation assessments, code reviews, and establish engineering best practices.
- Provide technical leadership and mentorship to engineering teams adopting automation-first and platform engineering practices.
What We’re Looking For
- 10+ years of hands-on experience in platform engineering, automation engineering, cloud engineering, DevOps, or distributed systems.
- Strong experience building enterprise self-service platforms supporting AI/ML, Data Science, Data Engineering, or Advanced Analytics workloads.
- Strong expertise in automation frameworks, CI/CD, DevOps, Infrastructure-as-Code, and software delivery lifecycle automation.
- Hands-on experience with Terraform and cloud-native infrastructure automation.
- Strong experience with Python for automation, orchestration, scripting, tooling, and platform engineering.
- Experience with Bitbucket, Bamboo, Jira, Confluence, or similar enterprise DevOps toolchains.
- Experience working with Kubernetes, containers, serverless platforms, YARN, and distributed processing environments.
- Knowledge of Generative AI and Agentic AI architectures, MCP frameworks, APIs, workflow automation, and enterprise AI platforms.
- Experience with event-driven architectures and technologies such as Kafka and streaming platforms.
- Strong understanding of cloud engineering, networking, security, scalability, resilience, and cost optimization.
- Experience implementing observability solutions covering monitoring, logging, tracing, alerting, and operational dashboards.
- Understanding of metadata management, data lineage, data governance, and semantic-layer concepts is highly valuable.
Good to Have
- Experience supporting enterprise GenAI platforms, AI governance, model management, and AI operationalization.
- Experience with GitOps, DevSecOps, Platform Engineering, and Reliability Engineering practices.
- Exposure to data governance, data quality, metadata management, and model lifecycle automation.
- Experience creating reusable internal developer platforms and self-service engineering tools at enterprise scale.
- Banking, AML, fraud detection, financial crime, or risk analytics domain experience is an advantage.
Strong AI/ML Engineer Profile
Mandatory (Experience) : Must have 3+ years of experience in software engineering with atleast 1+ years in GenAI application development and production deployment
Mandatory (GenAI Application Development): Must have proven experience building GenAI applications covering RAG pipelines, multi-agent systems, Text2SQL, and fine-tuning
Mandatory (Production GenAI Deployment): Must have expertise deploying production-grade GenAI applications including model evaluation, optimisation, and ownership of full production rollouts
Mandatory (ML & Data Science Tooling): Must have strong hands-on experience with core ML and data science tools including pandas, scikit-learn, and PyTorch
Mandatory (Cloud ML Infrastructure): Must have experience building and deploying production-grade ML workloads on at least one of AWS, Azure, or GCP
Mandatory (Communication): Must have strong English communication skills with the ability to work across time zones and collaborate cross-functionally with product, engineering, and business stakeholders
Mandatory (Note 1) : Role is Hybrid, WFH flexibility as well upto 6 days a month
Mandatory (Note 2) : CTC is inclusive of 10% variable
Mandatory (Note 3): Candidates should be available to join within May 31st or June first week max






