Data Engineer at NSEIT · Remote only · 7 - 12 years · ₹20L - ₹40L / yr (ESOP available) · Profitable · Remote only · Posted 21 Oct 2021
- Design AWS data ingestion frameworks and pipelines based on the specific needs driven by the Product Owners and user stories…
- Experience building Data Lake using AWS and Hands-on experience in S3, EKS, ECS, AWS Glue, AWS KMS, AWS Firehose, EMR
- Experience Apache Spark Programming with Databricks
- Experience working on NoSQL Databases such as Cassandra, HBase, and Elastic Search
- Hands on experience with leveraging CI/CD to rapidly build & test application code
- Expertise in Data governance and Data Quality
- Experience working with PCI Data and working with data scientists is a plus
- At least 4+ years of experience in the following Big Data frameworks: File Format (Parquet, AVRO, ORC), Resource Management, Distributed Processing and RDBMS
- 5+ years of experience on designing and developing Data Pipelines for Data Ingestion or Transformation using AWS technologies

About NSEIT
About
NSEIT is a global technology firm with a focus on the financial services industry. We are a vertical specialist organization with domain expertise and technology focus aligned to the needs of financial institutions. We offer Application Services, IT Enabled Services (Assessments), Testing Center of Excellence, Infrastructure Services, Integrated Security Response Center and Analytics as a Service primarily for the BFSI segment.
We are a 100% subsidiary of National Stock Exchange of India Limited (NSEIL). Being a part of the stock exchange our solutions inherently encapsulate industry strength, security, scalability, reliability and performance features.
Our focus on domain and key technologies enables us to use new trends in digital technologies like cloud computing, mobility and analytics while building solutions for our customers.
We are passionate about building innovative, futuristic and robust solutions for our customers. We have been assessed at Maturity Level 5 in Capability Maturity Model Integration for Development (CMMI® - DEV) v 1.3. We are also certified for ISO 9001:2015 for providing high quality products and services, and ISO 27001:2013 for our Information Security Management Systems.
Our offices are located in India and the US.
Connect with the team
Similar jobs (10)
About AuxoAI:
AuxoAI is a global platform-based services firm. We help companies—turn their strategies into practical digital and AI solutions. By understanding how our clients make decisions, we use digital and Artificial Intelligence (AI) technologies to drive growth, enhance their operations, improve customer experiences, and provide clear, actionable insights from their data. What We Do We work across various industries such as healthcare, high-tech, consumer packaged goods (CPG), finance etc., and in sales, marketing, and customer support functions.
We help our clients with accelerating their digital and AI journeys through:
• AI Application Development
• Data, Digital and Cloud acceleration using AI
• AI Native Product Engineering
We are seeking a skilled and experienced Data Engineer to join our dynamic team. The ideal candidate will have 6+ years of prior experience in data engineering, with a strong background in AWS (Amazon Web Services) technologies. This role offers an exciting opportunity to work on diverse projects, collaborating with cross-functional teams to design, build, and optimize data pipelines and infrastructure.
Responsibilities:
* Design, develop, and maintain scalable data pipelines and ETL processes leveraging AWS services such as S3, Glue, EMR, Lambda, and Redshift.
* Collaborate with data scientists and analysts to understand data requirements and implement solutions that support analytics and machine learning initiatives.
* Optimize data storage and retrieval mechanisms to ensure performance, reliability, and cost-effectiveness.
* Implement data governance and security best practices to ensure compliance and data integrity.
* Troubleshoot and debug data pipeline issues, providing timely resolution and proactive monitoring.
* Stay abreast of emerging technologies and industry trends, recommending innovative solutions to enhance data engineering capabilities.
Requirements :
* Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
* 6+ years of prior experience in data engineering, with a focus on designing and building data pipelines.
* Proficiency in AWS services, particularly S3, Glue, EMR, Lambda, and Redshift.
* Strong programming skills in languages such as Python, Java, or Scala.
* Experience with SQL and NoSQL databases, data warehousing concepts, and big data technologies.
* Familiarity with containerization technologies (e.g., Docker, Kubernetes) and orchestration tools (e.g., Apache Airflow) is a plus.
Data Engineer
Data Lakehouse & Platform Engineering
About the Role
We are hiring Data Engineer to own the lifecycle of our enterprise Data Lakehouse platform. We are looking for engineers who think in systems, make platform-level design decisions, and can build and operate a production-grade, multi-source lakehouse from the ground up, covering ingestion through consumption across a complex, multi-cloud source landscape.
You will be the technical authority for a platform that consolidates data from 18+ enterprise products (Costpoint, GovWin, Specpoint, Vantagepoint, and others) into a governed, medallion-architected data lake on AWS S3 with Apache Iceberg table format, orchestrated via AWS Step Functions, and queryable through AWS Athena and Trino. This role is end-to-end: you own ingestion, transformation, quality, orchestration, ML data supply, and BI consumption.
Key Responsibilities
• Architect and evolve the full medallion lakehouse — Bronze, Silver, and Gold layers — on AWS S3 with Apache Iceberg; own schema design, partitioning, compaction, and retention policies.
• Design and implement scalable Glue ETL (PySpark) pipelines for bronze_to_silver and silver_to_gold transformations, incorporating dbt for SQL-layer transformations where appropriate.
• Own and extend CDC ingestion via Fivetran; manage schema evolution, connector health, and sync reliability across 18+ source products.
• Build and maintain AWS Step Functions state machines and EventBridge schedules for end-to-end pipeline orchestration; implement Lambda-based quality and drift monitors.
• Govern the Glue Catalog and Lake Formation policies; enforce column-level security, row-level access controls, and audit logging to meet SOC2 and regulatory requirements.
• Architect the query layer — optimize Athena workgroups and partition pruning; plan and execute Trino-on-EKS deployment for sub-second analytics workloads.
• Partner with data science teams on SageMaker data supply: feature engineering pipelines, training dataset preparation, and model registry integration.
• Implement real-time and near-real-time streaming solutions using Kafka or Kinesis where sub-13-minute latency is required.
• Lead platform modernization initiatives: evaluate emerging formats (Iceberg vs. Delta Lake vs. Hudi), tooling, and cost optimization strategies.
• Establish and enforce data engineering best practices: code reviews, CI/CD for pipeline code, IaC (Terraform / CloudFormation), and incident response runbooks.
• Mentor and level up junior and mid-level data engineers; define team standards for pipeline design, testing, and documentation.
Required Qualifications
• Software or data engineering experience, with at least 4 years in an architect or technical lead capacity designing large-scale cloud data platforms.
• Deep, hands-on expertise with AWS data services: S3, Glue (PySpark ETL), Athena, Step Functions, Lambda, EventBridge, Lake Formation, SageMaker, and CloudWatch.
• Production experience with Apache Iceberg (or Delta Lake / Hudi) table formats — compaction, snapshot management, schema evolution, and time travel.
• Strong PySpark and Python skills; ability to write, review, and optimize distributed data processing jobs at scale.
• Hands-on experience with CDC-based ingestion platforms (Fivetran, Debezium, or equivalent) across heterogeneous source systems.
• Proven experience designing and implementing medallion (Bronze/Silver/Gold) or equivalent multi-hop lakehouse architectures.
• Experience with data pipeline orchestration: AWS Step Functions, Apache Airflow, or equivalent; event-driven pipeline design patterns.
• Strong SQL skills; experience with Athena, Trino, Presto, or equivalent query engines for large-scale analytical workloads.
• Familiarity with data governance tooling: catalog management (Glue Catalog, Apache Polaris/Iceberg REST), data lineage, access controls, and audit frameworks.
• Experience with Infrastructure as Code (Terraform or CloudFormation) for data platform provisioning and drift management.
• Solid understanding of dimensional modeling, schema design (star/snowflake), and data normalization for BI and analytics workloads.
• Bachelor's degree in Computer Science, Engineering, or a related field; or equivalent professional experience.
Preferred Qualifications
• Experience operating Trino or PrestoDB on Kubernetes (EKS); tuning for sub-second query latency and multi-tenant workloads.
• Familiarity with streaming platforms (Kafka, Kinesis, or Pub/Sub) and real-time lakehouse patterns.
• Experience with Apache Polaris or other Iceberg REST catalog implementations.
• Exposure to SageMaker MLOps pipelines, Model Registry, and feature store patterns for ML data supply.
• Experience with dbt (data build tool) for SQL-layer transformation and documentation in lakehouse environments.
• Government contracting or ERP domain knowledge (Costpoint, Deltek, Oracle, or similar enterprise platforms) is a strong plus.
• AWS certifications: Data Engineer Associate, Solutions Architect Professional, or equivalent.
What You Will Build
You will be a founding architect of a strategic, cross-product data platform that serves 18+ enterprise applications and their analytics, ML, and AI workloads. The decisions you make on schema, storage format, query layer, governance, and orchestration will shape the data foundation of the company for years. This is a high-impact, high-ownership role with direct visibility to senior leadership.
Job Title : Senior Data Engineer – Databricks
Experience : 14 to 20 Years
Location : HSR Layout, Bangalore
Work Mode : Hybrid – 3 Days WFO
Shift : 11:30 AM – 07:30 PM IST
Positions : 2
Notice Period : Immediate Joiners Only
Interview : 1 Technical Round + 2 Client Rounds
Role Overview :
We are looking for a Senior Data Engineer to build and lead enterprise-scale data platforms for a Switzerland-based commodity client.
The role requires a strong hands-on Data Engineering professional with expertise in Databricks, PySpark, Python, SQL, and AWS, along with technical leadership and stakeholder management experience.
Must-Have Skills :
- 14 to 20 years of Data Engineering experience
- Databricks & Apache Spark / PySpark
- Python & SQL
- AWS Cloud
- Lakehouse Architecture
- ETL / ELT & Distributed Data Processing
- Batch & Streaming Pipelines
- Data Pipeline Optimization & Data Modeling
- CDC & Incremental Processing
- Git, CI/CD & Testing
- Data Quality, Monitoring & Observability
- Technical Leadership & Stakeholder Management
Key Responsibilities :
- Design and build scalable data pipelines using Databricks, PySpark, Python, SQL, and AWS.
- Own data products from design through production.
- Develop batch / streaming pipelines and reusable ETL / ELT frameworks.
- Optimize pipelines for performance, scalability, reliability, and cost.
- Design scalable data architectures and data models.
- Implement data quality, monitoring, lineage, and CI/CD practices.
- Lead technical discussions and mentor engineering teams.
- Collaborate with business stakeholders, architects, product owners, and engineering teams.
- Remain hands-on while providing technical leadership.
Ideal Candidate :
A 14 to 20 years experienced, hands-on Data Engineering leader with strong Databricks + PySpark + AWS expertise, excellent communication, stakeholder management, and experience delivering enterprise-scale data platforms.
🔴 Super Urgent : Only Bangalore-based immediate joiners.
Description
We are looking for Senior Data Engineers to join our Data Platform team and build scalable, high-performance data platforms that power data processing, analytics, and downstream applications.
The ideal candidate will have strong experience in distributed data processing, ETL pipelines, and Big Data technologies, with hands-on expertise in Apache Spark and Python Scala.
You will be responsible for designing, developing, and optimizing large-scale data pipelines while collaborating closely with cross-functional engineering teams to build reliable, production-grade data solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL and data processing pipelines for large-scale datasets.
- Build and optimize distributed data applications using Apache Spark and Python Scala.
- Develop reliable, high-performance data pipelines for batch and streaming workloads.
- Design and manage data workflows using Apache Airflow.
- Build and operate data workloads on AWS, with strong usage of Amazon S3 for large-scale data storage.
- Work with large datasets to ensure data quality, consistency, reliability, and performance.
- Collaborate with engineering, product, analytics, and other platform teams to deliver robust data solutions.
- Optimize data workflows for scalability, reliability, performance, and cost efficiency.
- Troubleshoot production issues, identify bottlenecks, and continuously improve platform performance.
Requirements
Candidates who demonstrate:
- 5+ years of experience in Data Engineering, Big Data Engineering, or a similar role.
- Strong hands-on experience with Apache Spark and Scala.
- Experience designing, building, and maintaining large-scale ETL pipelines.
- Strong hands-on experience with AWS, particularly Amazon S3.
- Hands-on experience with Apache Airflow for workflow orchestration and scheduling.
- Strong SQL skills and a solid understanding of distributed data processing concepts.
- Experience working with batch and/or streaming data pipelines.
- Excellent debugging, problem-solving, and performance optimization skills.
- Strong communication and collaboration skills.
Good to Have
- Experience with Databricks and the broader Databricks data platform.
- Familiarity with streaming technologies such as Apache Kafka.
- Experience working on large-scale data platforms handling high-volume data workloads.
- Exposure to additional AWS data services and cloud-native data architectures.
Data Engineer
Location: Bengaluru, India (Hybrid)
Employment Type: Full-time
Experience: 3-5 years
Role Overview
What We’re Looking For:
- Bachelor’s degree in Computer Science/Engineering or equivalent experience required.
- Experience designing and shipping cloud services products.
- Experience driving and managing technical and architectural dependencies on AWS Cloud.
- A firm understanding of system architecture, cloud computing, PaaS/SaaS design principles, S3, DynamoDB, RDS mandatory.
- Experience in building or maintaining ETL processes and tools, i.e., AWS Glue or any open-source tool.
- Proven system-level design contribution to a current “Live” (in production / under daily high load) multi-region SaaS or PaaS offering.
- Proven experience with S3, DynamoDB, SQL, and AWS RDS services.
- Proficiency in programming languages such as Python.
- Strong analytical and problem-solving skills.
Required Skills & Experience
- Experience with Python, SQL, and data visualization/exploration tools.
- Familiarity with the AWS ecosystem, specifically S3, DynamoDB, and RDS.
- Communication skills, especially for explaining technical concepts to nontechnical business leaders.
- Ability to work on a dynamic, research-oriented team that has concurrent projects.
- Experience in AWS cost optimization (Savings Plans, Reserved Instances, Spot Instances) and governance frameworks.
- Experience developing solutions using infrastructure orchestration tools (SSM, automation account, Ansible, etc.).
- Excellent leadership, stakeholder management, and communication skills.
What We Offer
- Work with some of the brightest minds in the emerging EV industry.
- Make a tangible impact in reducing carbon emissions and enabling sustainable energy.
- Freedom to suggest, implement, and innovate on systems, processes, and technologies.
- Daily ownership in a high-growth, challenging environment.
- Flexible work environment with hybrid schedules and virtualization options.
- Competitive pay and benefits including health coverage, innovative PTO program, and performance bonuses.
Role Summary
We are hiring a Data Engineer / ML Data Pipeline Engineer to build and operate the data backbone of the Enterprise AI platform:
What You'll Own
- Ingestion & ETL/ELT pipelines for heterogeneous project folders (PDF drawings, SVG files, IFC models, BBS.json bar-bending-schedule data, Excel exports, and AI agent output JSON).
- AWS-based data architecture: S3 raw/staging/curated/outputs structuring, partitioning, versioning, and lifecycle management; querying via Athena/Glue and warehousing via Redshift or Snowflake as needed.
- Data validation frameworks: GUID cross-referencing between SVG and BBS data, schema enforcement, duplicate/orphan detection, reference integrity checks, and structured validation reporting.
- Agent run logging & observability: designing the database schema and pipelines that track every AI agent run (inputs, outputs, status, errors, cost, retries, reviewer feedback).
- AI Factory monitoring dashboards: operational dashboards (failure rates, retries, latency, data quality) and business dashboards (throughput, cost per run, rework rate) for Power BI/QuickSight or equivalent.
- ML data pipeline support: dataset preparation, labeling/annotation workflows, human-in-the-loop review tooling, and dataset versioning for models that classify or QC drawing issues.
- APIs: designing and building FastAPI/Flask endpoints to trigger validation runs and expose agent processing status to internal tools.
- Data quality & testing discipline: idempotent pipelines, quarantine/reject handling, regression and reconciliation testing, and root-cause debugging when pipelines or query performance degrade in production.
Key Skills — Non-Negotiable (Must-Have, Strong Level)
- Python — production-grade scripting: file/folder handling, JSON/schema processing, clean error handling, not just notebook-level scripting.
- SQL — strong hands-on ability, including GROUP BY/HAVING for duplicate detection, window functions, and daily aggregate/rate calculations (e.g., success-rate queries).
- AWS S3 data handling — practical experience structuring buckets for raw/staging/curated data, versioning, and avoiding overwrite issues at scale.
- Data validation — demonstrable experience building validation logic (set comparisons, duplicate/missing detection, structured pass/fail reporting), not just "I write assertions."
- ETL/ELT pipeline design — end-to-end ownership of at least one pipeline: source → transform → storage → validation → monitoring → business outcome, with clear articulation of what they personally built.
- Query/warehouse engine judgment — working knowledge of when to use Athena vs. Redshift vs. Snowflake (or equivalent), partitioning, clustering, sort/distribution keys, and storage format trade-offs (Parquet vs. JSON vs. CSV).
Key Skills — Good to Have
- Dashboarding — Power BI / QuickSight (or equivalent) fact/dimension table design, KPI cards, drill-downs; medium-to-strong level is a plus but trainable.
- FastAPI / Flask — building real endpoints with request/response schemas and basic error handling; especially valuable for validation-trigger and agent-status APIs.
- ML data pipeline experience — dataset labeling, annotation platform design, train/test/validation splitting, dataset versioning; strong on the pipeline/data side rather than model training itself.
- Human-in-the-loop / review tooling — experience building or contributing to browser-based labeling/review platforms (session persistence, label schema, export formats).
- Large-scale metadata querying — experience making file discovery fast across large volumes (1,000+ projects, thousands of files each) via metadata index tables, event-based ingestion, or catalog tools like AWS Glue.
Design, develop, and maintain ETL pipelines involving large-scale data.
Develop data processing and analytics applications primarily using PySpark and Python.
Build scalable and distributed data processing solutions using Apache Spark.
Develop and deploy data applications on AWS cloud.
Work with AWS services related to storage, compute, ETL, data warehousing, analytics, and streaming.
Implement distributed storage and processing solutions capable of handling high-volume datasets.
Design data processing applications with a focus on performance, scalability, reliability, and optimization.
Work with both SQL and NoSQL databases for data storage, processing, and analytics.
Write, optimize, and analyze SQL, HQL, and NoSQL queries.
Troubleshoot data pipeline and processing issues and ensure data quality and reliability.
Collaborate with data engineers, analysts, architects, and other technical teams to deliver data-driven solutions.
Company Name – Wissen Technology
Group of companies in India – Wissen Technology & Wissen Infotech
Work Location – Whitefield, Bangalore
Website and Company profile:
www.wissen.com
LinkedIn Page:
https://www.linkedin.com/company/wissen-technology/
While you may already know about Wissen and the company history, here is a quick rundown for you.
About Wissen Technology:
· The Wissen Group was founded in the year 2000. Wissen Technology, a part of Wissen Group, was established in the year 2015.
· Wissen Technology is a specialized technology company that delivers high-end consulting for organizations in the Banking & Finance, Telecom, and Healthcare domains. We help clients build world class products.
· Our workforce has highly skilled professionals, with leadership and senior management executives who have graduated from Ivy League Universities like Wharton, MIT, IITs, IIMs, and NITs and with rich work experience in some of the biggest companies in the world.
· Wissen Technology has grown its revenues by 400% in these five years without any external funding or investments.
· Globally present with offices US, India, UK, Australia, Mexico, and Canada.
· We offer an array of services including Application Development, Artificial Intelligence & Machine Learning, Big Data & Analytics, Visualization & Business Intelligence, Robotic Process Automation, Cloud, Mobility, Agile & DevOps, Quality Assurance & Test Automation.
· Wissen Technology has been certified as a Great Place to Work®.
· Wissen Technology has been voted as the Top 20 AI/ML vendor by CIO Insider in 2020.
· Over the years, Wissen Group has successfully delivered $650 million worth of projects for more than 20 of the Fortune 500 companies.
· We have served client across sectors like Banking, Telecom, Healthcare, Manufacturing, and Energy. They include likes of Morgan Stanley, Goldman Sachs, MSCI, StateStreet, Flipkart, Swiggy, Trafigura, GE to name a few.
About Role :
Key Responsibilities
- Build and maintain data transformation pipelines using java Spark
- Develop and optimize large-scale/CPU intensive data processing using Apache Spark
- Orchestrate workflows using Airflow
- Implement data quality checks, testing, and monitoring for pipeline. Good to have exposer into managing metadata, cataloguing, and lineage
- Support schema evolution, backfills, and incremental processing
- Ensure pipelines meet SLAs for freshness, reliability, and performance
- Expertise/working knowledge in Spark and HBase(semantic layer, virtual datasets, Reflections)
Required Skills & Qualifications
- Strong hands-on experience with
- HBase
- Apache Spark
- Experience with HBase or similar lakehouse query engines
- Airflow
- Understanding of data catalogs and lineage (e.g., OpenLineage, DataHub, Apache Polaris , openlineage)
- Proficiency in Java
- Experience with Git-based development and CI/CD
Nice-to-Have Skills
- OpenTable format/Iceberg ,Apache Arrow
- CDC-based analytics pipelines
- Cloud platforms (AWS)
- Kubernetes-based data platforms
Job description: Data Architect – Databricks / AWS
Job Summary
We are looking for an experienced Data Architect to define and drive the target data architecture for the Horizon MVP and its future evolution. The role will be responsible for designing a scalable, secure, governed cloud data platform covering ingestion, storage, processing, analytics, APIs, and downstream data consumption.
The architect will work closely with Data Engineering, Backend, DevOps, QA, and business stakeholders to establish architecture standards and ensure the platform is ready for advanced analytics, AI/ML, vector storage, and future LLM-based capabilities.
- Job Title: Data Architect
- Experience: 8+ Years
- Relevant Architecture Experience: 3+ Years in Data Architecture
- Location: Chennai / Pune
- Work Mode: Hybrid – 3 Days WFO
- Budget: Up to 24 LPA
- Payroll: Haparz
- Notice Period: Immediate Preferred
Key Responsibilities
- Define the target data architecture for the Horizon MVP and establish an architecture roadmap for future scalability.
- Design end-to-end architecture covering data ingestion, storage, processing, serving, reporting, APIs, and downstream applications.
- Establish canonical data models and schemas for travel signals, corridors, sources, evidence, scores, and generated insights.
- Define data normalization strategies for structured, semi-structured, and unstructured data from multiple external sources.
- Design and govern the Databricks platform architecture, including Unity Catalog, data schemas, access controls, and governance standards.
- Establish data-retention, lineage, data-quality, security, privacy, and compliance controls.
- Define secure integration patterns between Databricks, AWS PRODIGY, SharePoint, external APIs, and downstream applications.
- Design scalable data processing for 30-day signal windows, convergence/divergence scoring, corridor ranking, and spike detection.
- Define architecture patterns that support future vector storage, embeddings, LLM integration, and multi-year analytics.
- Design reliable batch and API-driven ingestion frameworks for structured and unstructured data.
- Review technical designs, identify architectural risks, and provide technical direction to engineering teams.
- Guide backend, data engineering, DevOps, and QA teams in implementing architecture standards.
- Ensure architecture decisions align with enterprise security, RBAC, PII handling, privacy, and operational requirements.
- Communicate architecture decisions, trade-offs, and technical recommendations effectively to technical and business stakeholders.
What We’re Looking For
- 8+ years of experience in data engineering, data platforms, or data architecture, with at least 3+ years in a Data Architect capacity.
- Strong hands-on experience designing cloud-based data platforms, lakehouses, or analytical platforms.
- Advanced knowledge of Databricks, Apache Spark/PySpark, Delta Lake, and Unity Catalog.
- Strong understanding of AWS data services, IAM, networking, and secure cloud integration patterns.
- Strong expertise in data modelling, metadata management, data lineage, data quality, retention, and governance.
- Experience architecting batch and API-based ingestion pipelines for structured, semi-structured, and unstructured data.
- Understanding of AI/ML workloads, feature pipelines, vector databases, embeddings, and LLM integration patterns.
- Experience designing APIs and downstream data-serving architectures.
- Strong knowledge of PII protection, RBAC, data privacy, and enterprise security controls.
- Excellent architectural communication and stakeholder-management skills.
Skills Referential (Required knowledge, skills and abilities)
Technical Skills:
Python
Pyspark
SQL
ETL Aws, Azure, gcp







