CUDA (Compute Unified Device Architecture) Developer at NeoGenCode Technologies Pvt Ltd · Remote only · 3 - 10 years · ₹2L - ₹15L / yr · Raised funding · Remote only · Posted 3 Feb 2025

CUDA (Compute Unified Device Architecture) Developer
Job Description :
Position Name : CUDA Developer
Experience : 5+ Years
Opportunity : Full-time, 8 hours/day (4-hour overlap with PST)
Notice Period : Immediate
Summary :
We are looking for an experienced CUDA (Compute Unified Device Architecture) Developer with strong expertise in parallel computing and performance optimization.
Responsibilities :
- Optimize and debug CUDA applications for high performance.
- Improve algorithm efficiency through effective parallelization.
- Stay updated with the latest CUDA advancements.
Requirements :
- Experience : 5+ Years in Software Development, including 3+ Years in CUDA and 5+ Years in C++.
- Technical Skills : Proficiency in C/C++, CUDA (12.0+ preferred), cuBLAS, cuDNN, and performance tuning.
- Other Skills : Strong problem-solving, teamwork, and communication skills.

About NeoGenCode Technologies Pvt Ltd
About
Welcome to Neogencode Technologies, an IT services and consulting firm that provides innovative solutions to help businesses achieve their goals. Our team of experienced professionals is committed to providing tailored services to meet the specific needs of each client. Our comprehensive range of services includes software development, web design and development, mobile app development, cloud computing, cybersecurity, digital marketing, and skilled resource acquisition. We specialize in helping our clients find the right skilled resources to meet their unique business needs. At Neogencode Technologies, we prioritize communication and collaboration with our clients, striving to understand their unique challenges and provide customized solutions that exceed their expectations. We value long-term partnerships with our clients and are committed to delivering exceptional service at every stage of the engagement. Whether you are a small business looking to improve your processes or a large enterprise seeking to stay ahead of the competition, Neogencode Technologies has the expertise and experience to help you succeed. Contact us today to learn more about how we can support your business growth and provide skilled resources to meet your business needs.
Candid answers by the company
IT & Engineering Talent Staffing
- Provides full-time and contract-based hiring, delivering handpicked, pre‑screened developers across tech stacks—ranging from web, mobile, AI/ML, Web3/blockchain.
- Maintains a bench o vetted candidates, offering fast delivery of interview-ready profiles—often within 24 hours.
- Offers payroll management, handling compliance, tax, attendance, and documentation for both contractors and full-time employees.
2. End-to-End Project Delivery
- Delivers full-stack development solutions: web, mobile, cloud, AI/ML, Blockchain/Web3.
- Manages entire project lifecycle—requirements gathering, design (UI/UX), development, deployment, and ongoing support .
3. Additional Offerings
- Expands into cybersecurity consulting, digital marketing, and cloud platform services (like AWS, GCP, Azure) .
- Provides strategic IT consulting to align technology solutions with business objectives
Similar jobs (7)

This is about a job opportunity with an established medtech company based out of Mysore only.
The technical competencies required are image & signal processing and algorithm development for imaging equipment like digital x-rays and others.

Qualifications & Requirements:
- 4+ years of experience in C++ application development.
- Hands-on experience with C++11 or above.
- Strong knowledge of object-oriented programming and software design.
- Deep understanding of STL, multi-threading, socket programming, and data structures.
- Solid grasp of Linux development and debugging techniques.
- Proficient in using GCC, GDB, and Makefile.
- Familiarity with Valgrind and similar analysis tools.
- Experience with version control tools like Git.
- Experience writing and maintaining automated tests.
- Experience in capital markets/trading domain is a plus.
Skills:
- Strong problem-solving and analytical thinking.
- Clear and effective communication.
- Self-driven with the ability to work independently.
- Passionate about high-quality software and strong engineering practices.
- Comfortable working in a fast-paced, collaborative environment.
Title : Senior AI Platform / MLOps Engineer
Experience : 6+ years
Work type : Chennai - Work from Office/other locations - Remote
Employment Type : Full Time
Notice Period : Immediate
Work Day :Mon to Fri
Key Responsibilities:
- Install, configure and operate OpenShift, NVIDIA GPU operator, OpenShift AI, and NIM microservices on 12× RTX PRO 6000 across two servers; single-node and HA control-plane topologies
- Serving configuration and tuning: quantized model deployment (FP8/FP4), replica balancing, batching, KV-cache and context management
- Azure GPU build environments: provisioning, cost control, parity with the on-prem stack via pinned container/model versions; cloud-to-factory migration with parity regression
- GitOps CI/CD, container registry, artifact/model versioning, environment promotion; observability and audit wiring (Splunk, Prometheus/Grafana)
- Benchmark automation: load harness, p50/p95/p99 latency, tokens/sec, GPU utilization; the capacity report data pipeline
- Platform upgrade procedure with evaluation-regression gates; deployment runbook as a first-class deliverable
Technical Skills:
- 6+ years infrastructure/platform engineering with 3+ years production Kubernetes; OpenShift experience strongly preferred
- Hands-on GPU inference serving in production: NIM, Triton, vLLM, or TensorRT-LLM — you have sized, deployed, and tuned LLM serving on real GPUs and can talk memory-bandwidth trade-offs
- GitOps fluency (ArgoCD/Flux), infrastructure-as-code, container internals; comfortable in air-gapped/proxy-restricted enterprise networks
- Observability depth: metrics, traces, log pipelines; has built performance test harnesses, not just run them
- Azure or AWS GPU compute operations experience
Strongly preferred
- NVIDIA GPU operator and AI Enterprise stack specifics; KServe; Milvus or pgvector operations; VAST/NFS/S3 storage integration; banking or other regulated-environment delivery
About Ampera:
Ampera Technologies, a purpose driven Digital IT Services with primary focus on supporting our client with their Data, AI / ML, Accessibility and other Digital IT needs. We also ensure that equal opportunities are provided to Persons with Disabilities Talent. Ampera Technologies has its Global Headquarters in Chicago, USA and its Global Delivery Center is based out of Chennai, India. We are actively expanding our Tech Delivery team in Chennai and across India. We offer exciting benefits for our teams, such as 1) Hybrid and Remote work options available, 2) Opportunity to work directly with our Global Enterprise Clients, 3) Opportunity to learn and implement evolving Technologies, 4) Comprehensive healthcare, and 5) Conducive environment for Persons with Disability Talent meeting Physical and Digital Accessibility standards
Position: Computer Vision Engineer
Experience: 2–3 Years
Location: Bengaluru, Karnataka
Employment Type: Full-time
About the Role
We are seeking a highly motivated Computer Vision Engineer to join our autonomy and avionics team. The role involves developing, implementing, and validating computer vision models and algorithms and pipelines for UAVs operating in both GNSS-available and GNSS-denied environments.
The ideal candidate should have a strong foundation in theory of deep learning and machine learning, strong understanding of electromagnetic spectrum, imaging fundamentals, camera principles, and mathematical concepts with hands-on experience in implementing these algorithms on embedded or real-time systems.
Key Responsibilities
- Design, develop, and optimise AI Models
- Make custom CNNs/ modify existing CNNs to suit specific problems at hand
- Handle end-to-end training flow
- Implement end to end inference pipelines on standard PCs as well as on embedded systems
- Understand performance benchmarks and assess the accuracy and inference times
- Implement traditional image processing algorithms
- Factor the code to leverage underlying hardware architecture
- Prune the networks for efficiency
- Integrate the system within the application framework using C++
- Work closely with perception, controls, embedded software, and systems engineering teams.
Required Qualifications
- B.E./B.Tech/M.E./M.Tech in Computer Science and Engineering, Electronics, ECE, Mechatronics, or a related discipline.
- 2–3 years of experience in relevant area
- Strong understanding of: Linear Algebra, Probability and Statistics, AI-ML-DL fundamentals, Image processing, Camera Functioning
- Strong programming skills in C++ and Python.
- Experience with MATLAB for algorithm development and validation.
- Familiarity with Linux development environments.
- Experience with Git version control.
Preferred Skills
- Experience with Camera, IMU Calibration and Synchronisation
- Experience with multi-sensor fusion.
- Experience working with NVIDIA devices
- Experience on FPGA will be an added advantage
- Full understanding of Git functionality
- Exposure to airborne software development processes and coding standards (e.g., MISRA C++).
Personal Attributes
- Strong analytical and problem-solving skills.
- Ability to work independently on challenging technical problems.
- Good communication and documentation skills.
- Passion for solving challenging problems
- Willingness to participate in field trials and flight testing.
- Team playwe
We are looking for an experienced Developer with experience in compiled languages and ready to work in Golang. You will be responsible for the Development of distributed services GUI and terminal application in Golang. Candidates should have strong understanding of Algorithms, Memory Management and Networking
Required Experience & Skills:
- C/C++/C#/Go/Rust.
- Networking, Database, Linux, Docker.
- Multiprocessing, Multithreading, Distributed computing, Thread synchronization techniques.
- Communication protocols, gRPC, MQTT, REST, Modbus, OPC...
- GUI Programming
- Testing, Debugging and performance optimization.
- REST APIs, MySQL, Redis
- Proficiency with Git for version control and collaborative development workflows.
Position Title: Real-Time Computer Vision & Edge AI Engineer (Founding Engineering Team / Core LLD)
Reporting Structure: High-Level AI Architect (Principal ML Scientist, Google)
Domain: Sub-16ms Edge AI, 3D Pose & Shape Estimation (SMPL-X), TensorRT C++ Inference, Zero-Copy Systems
Performance Benchmark: Hard locked 60 FPS (<16.6 ms total frame budget) on dedicated RTX hardware
1. Position Overview & Architecture
We are building a proprietary, ultra-low-latency spatial computing platform centered on high-fidelity 100% 3D Digital Twin architecture and real-time human digitization.
In this role, you will serve as the Low-Level Design (LLD) Core AI Engineer, working directly alongside a Lead AI Scientist from Google. Your primary mandate is to solve complex surface occlusion and volumetric estimation challenges by building an ultra-fast C++ inference pipeline. This system must accurately regress a subject's true underlying 3D body shape and skeletal pose directly from a live camera feed. You will deploy models that extract parametric data (SMPL-X shape/pose parameters) and bridge these joint rotations seamlessly into our Vulkan graphics engine via shared GPU memory.
System Architecture:
● Hardware Camera Ingestion: (V4L2 / GStreamer / CUDA)
↓ Raw RGB Frames (Zero CPU Copy)
● Edge AI Inference: (TensorRT / ONNX C++ API for 3D Pose Tracking, Kinematic Anchoring, SMPL-X Shape)
↓ 3D Skeletal Transforms & Shape Parameters
● Zero-Copy Shared Memory: (CUDA-Vulkan Bridge feeding directly into OpenRigLogic / MetaHuman Engine)
2. Key Responsibilities & Deliverables
A. Real-Time 3D Pose & Shape Estimation
● Deploy and optimize state-of-the-art 3D human body reconstruction models (e.g., Shapy, SMPLify-X, CLIFF) to accurately regress the user's underlying skeletal structure and body volume, effectively bypassing unpredictable surface topologies and complex environmental occlusions.
● Extract mathematically stable shape parameters (β) and pose parameters (θ) to drive the skeletal hierarchy of a high-fidelity digital avatar.
B. Edge Inference Pipeline (TensorRT)
● Translate Python-based research models into production-grade C++ inference engines using NVIDIA TensorRT and ONNX Runtime.
● Implement INT8/FP16 quantization, layer fusion, and custom CUDA plugins to ensure the entire AI inference pass executes within a strict <10 ms budget per frame.
C. Temporal Smoothing & Anti-Jitter Kinematics
● Implement highly optimized temporal filters (Kalman filters, One-Euro filters, optical flow tracking) in native C++ to eliminate all high-frequency jitter from the output joint rotations before they reach the graphics engine.
● Ensure kinematic constraints (e.g., fixed bone lengths) are strictly maintained to prevent the digital asset from stretching or warping dynamically.
D. Zero-Copy Ingestion & Engine Synchronization
● Build hardware-accelerated video capture pipelines using V4L2 or GStreamer to ingest raw camera frames directly into GPU memory.
● Bridge the output coordinate data and transformation matrices to the graphics team using POSIX shared memory and CUDA-Vulkan interop (VK_KHR_external_memory_fd), eliminating CPU staging overhead.
3. Technical Qualifications & Tech Stack
● Core Programming: Production-level Modern C++ (C++17/20), Python (strictly for model training/validation), and CUDA C/C++.
● AI & Acceleration Frameworks: NVIDIA TensorRT, ONNX Runtime (C++ API), PyTorch.
● Computer Vision Libraries: OpenCV (CUDA backend), MediaPipe C++ bindings.
● Mathematical Foundations: 3D Kinematics, Matrix Transformations, Quaternions/Euler angles, statistical body modeling (SMPL/SMPL-X architecture).
● Systems Architecture: Low-latency memory management, multi-threading (std::jthread, lock-free queues), SIMD vectorization.
4. Relevant Projects & Demonstrable Experience (Preferred)
Candidates will be preferred if they present functional codebases, GitHub repositories, or thesis work covering:
● Real-Time Body Fitting / Pose Estimation: Practical experience deploying 3D human pose or shape reconstruction models on live video feeds.
● TensorRT / C++ Deployment: Demonstrable experience stripping a PyTorch model out of Python and running it natively in C++ using TensorRT or ONNX, ideally with custom CUDA layers or INT8 calibration.
● High-Throughput Vision Pipelines: Built a C++ video processing pipeline that aggressively minimizes latency and avoids memory garbage collection pauses.
● Kinematics & Smoothing: Applied mathematical filters to raw sensor or AI data to produce smooth, mechanically accurate 3D rotations.
5. Compensation & Engagement Structure
● Compensation: ₹1,50,000 to ₹2,00,000/month
● Mentorship: Direct architectural guidance, algorithm review, and technical leadership from a Principal ML Scientist at Google.
● Hardware: Dedicated high-end workstation equipped with discrete NVIDIA RTX hardware.
Key Responsibilities:
· Architectural Leadership: Design and lead the development of robust, scalable AI architectures, ensuring high performance, reliability, and security.
· Applied Mathematics & Statistics: Apply statistical analysis, numerical computation, and mathematical modeling to derive insights from large-scale data and optimize model performance.
· Deep Learning Development: Design, train, and deploy advanced Deep Learning (DL) models.
· Technical Mentorship: Mentor engineering teams on best practices for AI/ML, coding standards, and architectural design.
· Model Optimization: Optimize models for speed, efficiency, and accuracy using techniques like pruning, quantization, or GPU acceleration.
· Strategy & Innovation: Evaluate and select appropriate AI frameworks, tools, and platforms, staying abreast of cutting-edge research and industry trends.
Qualifications:
Required:
· Education: Master's or PhD in Computer Science, Applied Mathematics, Statistics, Physics, or a related quantitative field.
· Experience: 10+ years of experience in software development, with at least 3-5 years in a Applied Mathematics and Deep learning.
· AI/ML Expertise: Proven experience designing and deploying deep learning models in production using frameworks.
· Mathematics/Statistics: Strong proficiency in linear algebra, calculus, probability, and statistical methods.
· Programming Skills: Expert-level coding skills in Python (NumPy, Pandas, Scikit-learn) and experience with languages like Java or C++.
Key Competencies:
- Strategic mindset with deep operational awareness.
- Excellent communication and stakeholder management skills.
- Ability to simplify complex technical concepts for executive reporting.
- Strong leadership, people development, and cross-functional influencing skills.
Bias for action and a relentless focus on continuous improvement.





