Data Engineer at dataeaze systems · Pune · 1 - 5 years · ₹3L - ₹10L / yr · Profitable · Posted 13 Oct 2023
- Core Java: advanced level competency, should have worked on projects with core Java development.
- Linux shell : advanced level competency, work experience with Linux shell scripting, knowledge and experience to use important shell commands
- Rdbms, SQL: advanced level competency, Should have expertise in SQL query language syntax, should be well versed with aggregations, joins of SQL query language.
- Data structures and problem solving: should have ability to use appropriate data structure.
- AWS cloud : Good to have experience with aws serverless toolset along with aws infra
- Data Engineering ecosystem : Good to have experience and knowledge of data engineering, ETL, data warehouse (any toolset)
- Hadoop, HDFS, YARN : Should have introduction to internal working of these toolsets
- HIVE, MapReduce, Spark: Good to have experience developing transformations using hive queries, MapReduce job implementation and Spark Job Implementation. Spark implementation in Scala will be plus point.
- Airflow, Oozie, Sqoop, Zookeeper, Kafka: Good to have knowledge about purpose and working of these technology toolsets. Working experience will be a plus point here.

Similar jobs (8)
Company Name – Wissen Technology
Group of companies in India – Wissen Technology & Wissen Infotech
Work Location – Whitefield, Bangalore
Website and Company profile:
www.wissen.com
LinkedIn Page:
https://www.linkedin.com/company/wissen-technology/
While you may already know about Wissen and the company history, here is a quick rundown for you.
About Wissen Technology:
· The Wissen Group was founded in the year 2000. Wissen Technology, a part of Wissen Group, was established in the year 2015.
· Wissen Technology is a specialized technology company that delivers high-end consulting for organizations in the Banking & Finance, Telecom, and Healthcare domains. We help clients build world class products.
· Our workforce has highly skilled professionals, with leadership and senior management executives who have graduated from Ivy League Universities like Wharton, MIT, IITs, IIMs, and NITs and with rich work experience in some of the biggest companies in the world.
· Wissen Technology has grown its revenues by 400% in these five years without any external funding or investments.
· Globally present with offices US, India, UK, Australia, Mexico, and Canada.
· We offer an array of services including Application Development, Artificial Intelligence & Machine Learning, Big Data & Analytics, Visualization & Business Intelligence, Robotic Process Automation, Cloud, Mobility, Agile & DevOps, Quality Assurance & Test Automation.
· Wissen Technology has been certified as a Great Place to Work®.
· Wissen Technology has been voted as the Top 20 AI/ML vendor by CIO Insider in 2020.
· Over the years, Wissen Group has successfully delivered $650 million worth of projects for more than 20 of the Fortune 500 companies.
· We have served client across sectors like Banking, Telecom, Healthcare, Manufacturing, and Energy. They include likes of Morgan Stanley, Goldman Sachs, MSCI, StateStreet, Flipkart, Swiggy, Trafigura, GE to name a few.
About Role :
Key Responsibilities
- Build and maintain data transformation pipelines using java Spark
- Develop and optimize large-scale/CPU intensive data processing using Apache Spark
- Orchestrate workflows using Airflow
- Implement data quality checks, testing, and monitoring for pipeline. Good to have exposer into managing metadata, cataloguing, and lineage
- Support schema evolution, backfills, and incremental processing
- Ensure pipelines meet SLAs for freshness, reliability, and performance
- Expertise/working knowledge in Spark and HBase(semantic layer, virtual datasets, Reflections)
Required Skills & Qualifications
- Strong hands-on experience with
- HBase
- Apache Spark
- Experience with HBase or similar lakehouse query engines
- Airflow
- Understanding of data catalogs and lineage (e.g., OpenLineage, DataHub, Apache Polaris , openlineage)
- Proficiency in Java
- Experience with Git-based development and CI/CD
Nice-to-Have Skills
- OpenTable format/Iceberg ,Apache Arrow
- CDC-based analytics pipelines
- Cloud platforms (AWS)
- Kubernetes-based data platforms
Must have experience in Java
Must have experience in Spark
Must have experience in ETL coding
Strong expertise in coding
Key Responsibilities
Build and maintain data transformation pipelines using java Spark
Develop and optimize large-scale/CPU intensive data processing using Apache Spark
Orchestrate workflows using Airflow
Implement data quality checks, testing, and monitoring for pipeline. Good to have exposer into managing metadata, cataloguing, and lineage
Support schema evolution, backfills, and incremental processing
Ensure pipelines meet SLAs for freshness, reliability, and performance
Expertise/working knowledge in Spark and HBase(semantic layer, virtual datasets, Reflections)
Required Skills & Qualifications
Strong hands-on experience with
HBase
Apache Spark
Experience with HBase or similar lakehouse query engines
Airflow
Understanding of data catalogs and lineage (e.g., OpenLineage, DataHub, Apache Polaris , openlineage)
Proficiency in Java
Experience with Git-based development and CI/CD
Nice-to-Have Skills
OpenTable format/Iceberg ,Apache Arrow
CDC-based analytics pipelines
Cloud platforms (AWS)
Kubernetes-based data platforms
4 - 10 years of experience in designing and buildingarchitecting highly resilient data platforms
∙Strong knowledge of data engineering, architecture and data modeling
∙Experience in platforms like Databricks and Snowflake
∙Experience on building applications on cloud (AWS or Azure or Google Cloud)
∙Strong analytical and problem-solving skills
∙Prior experience in developing data or computation intensive (e.g. grid based) backend applications is an
advantage
∙OOP design skills with an understanding or at least personal interest towards the concepts of Functional
Programming
∙Willingness to understand and enhance other people’s code, being able to work in an environment where
developers will oversee and work on wider components also dealing with older “legacy” code
∙Strong programming skills (Java/ Scala / Python) skills with the willingness to pick up the other language if not
already mastered at a sufficient level is important
∙Spring knowledge is an advantage, but in general willingness to learn, work with and even enhance in-house
developed frameworks is a must
∙Prior experience in working with Git, Bitbucket, Jenkins, working with PR-s, using JIRA, following the Scrum Agile
methodology is an advantage
∙Prior knowledge of financial products is an advantage
∙Bachelors or Masters in any relevant field of IT/Engineering area is an advantage
Description
We are looking for Senior Data Engineers to join our Data Platform team and build scalable, high-performance data platforms that power data processing, analytics, and downstream applications.
The ideal candidate will have strong experience in distributed data processing, ETL pipelines, and Big Data technologies, with hands-on expertise in Apache Spark and Python Scala.
You will be responsible for designing, developing, and optimizing large-scale data pipelines while collaborating closely with cross-functional engineering teams to build reliable, production-grade data solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL and data processing pipelines for large-scale datasets.
- Build and optimize distributed data applications using Apache Spark and Python Scala.
- Develop reliable, high-performance data pipelines for batch and streaming workloads.
- Design and manage data workflows using Apache Airflow.
- Build and operate data workloads on AWS, with strong usage of Amazon S3 for large-scale data storage.
- Work with large datasets to ensure data quality, consistency, reliability, and performance.
- Collaborate with engineering, product, analytics, and other platform teams to deliver robust data solutions.
- Optimize data workflows for scalability, reliability, performance, and cost efficiency.
- Troubleshoot production issues, identify bottlenecks, and continuously improve platform performance.
Requirements
Candidates who demonstrate:
- 5+ years of experience in Data Engineering, Big Data Engineering, or a similar role.
- Strong hands-on experience with Apache Spark and Scala.
- Experience designing, building, and maintaining large-scale ETL pipelines.
- Strong hands-on experience with AWS, particularly Amazon S3.
- Hands-on experience with Apache Airflow for workflow orchestration and scheduling.
- Strong SQL skills and a solid understanding of distributed data processing concepts.
- Experience working with batch and/or streaming data pipelines.
- Excellent debugging, problem-solving, and performance optimization skills.
- Strong communication and collaboration skills.
Good to Have
- Experience with Databricks and the broader Databricks data platform.
- Familiarity with streaming technologies such as Apache Kafka.
- Experience working on large-scale data platforms handling high-volume data workloads.
- Exposure to additional AWS data services and cloud-native data architectures.
Key Responsibilities
Build and maintain data transformation pipelines using java Spark
Develop and optimize large-scale/CPU intensive data processing using Apache Spark
Orchestrate workflows using Airflow
Implement data quality checks, testing, and monitoring for pipeline. Good to have exposer into managing metadata, cataloguing, and lineage
Support schema evolution, backfills, and incremental processing
Ensure pipelines meet SLAs for freshness, reliability, and performance
Expertise/working knowledge in Spark and HBase(semantic layer, virtual datasets, Reflections)
Required Skills & Qualifications
Strong hands-on experience with Apache Spark
Experience with HBase/SQL or similar lakehouse query engines
Airflow
Understanding of data catalogs and lineage (e.g., OpenLineage, DataHub, Apache Polaris , openlineage)
Proficiency in Java
Experience with Git-based development and CI/CD
Description
We are looking for Senior Data Engineers to join our AdTech team and build scalable, high-performance data platforms that power advertising insights and analytics. The ideal candidate will have strong experience in distributed data processing, ETL pipelines, and Big Data technologies, with hands-on expertise in Spark and Scala.
You will work on designing, developing, and optimizing large-scale data pipelines while collaborating closely with cross-functional engineering teams to build reliable, production-grade data solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL pipelines for large-scale data processing.
- Build and optimize distributed data applications using Spark and Scala.
- Develop reliable, high-performance data pipelines for batch and streaming workloads.
- Work with large datasets to ensure data quality, consistency, and performance.
- Collaborate with engineering, product, and analytics teams to deliver robust data solutions.
- Optimize data workflows for scalability, reliability, and cost efficiency.
- Deploy and manage data workloads in cloud and containerized environments.
- Troubleshoot production issues and continuously improve platform performance.
Requirements
Candidates who demonstrate:
- 5+ years of experience in Data Engineering or Big Data Engineering.
- Strong hands-on experience with Apache Spark and Scala.
- Experience building and maintaining ETL pipelines.
- Familiarity with Google Cloud Storage (GCS).
- Experience with Kubernetes (K8s).
- Strong SQL skills and understanding of distributed data processing.
- Excellent debugging, problem-solving, and performance optimization skills.
- Strong communication and collaboration skills.
Good to Have
- Experience with AWS and cloud-native data services.
- Familiarity with streaming technologies such as Kafka.
- Experience working on large-scale data platforms or AdTech systems.
- Exposure to orchestration tools such as Airflow.
Benefits
- Best-in-class salary: We hire strong talent and compensate accordingly.
- Proximity Talks: Meet and learn from designers, engineers, product leaders, and AI practitioners.
- Continuous learning: Work with a world-class team and stay close to the latest in AI, engineering, and product development.
- High-impact work: Build AI-first systems and products used at scale by global clients.
About Us
Proximity is the trusted technology, design, and consulting partner for some of the biggest Sports, Media, and Entertainment companies in the world. We’re headquartered in San Francisco and have offices in Palo Alto, Dubai, Mumbai, and Bangalore.
Since 2019, Proximity has built high-impact, scalable products used by millions of users every day. Today, we are a global team of engineers, designers, product managers, and experts solving complex problems and building cutting-edge technology at scale.
Job Summary:
- We are seeking an experienced Hadoop Engineer with strong hands-on expertise in MAPR and Hortonworks platform engineering and administration. The successful candidate will join our Hadoop Platform Engineering & Automation team to support our critical enterprise data lake platform. This role emphasizes platform support, patching, vulnerability remediation, L3 issue resolution, and automation initiatives utilizing Linux, Shell Scripting, and Ansible.
Responsibilities:
- Platform Engineering & Administration: Administer, maintain, and optimize MAPR and Hortonworks Hadoop clusters.
- Provide L3 support for complex technical issues related to cluster functioning, performance, and reliability.
- Monitor cluster health, capacity, performance, and security configurations.
- Patching & Vulnerability Remediation: Execute OS-level and platform-level patching across large-scale Hadoop clusters.
- Implement remediation for platform vulnerabilities in accordance with organizational InfoSec policies.
- Collaborate with other support teams to ensure compliance with standards and mitigation timelines.
- Automation Development: Develop, enhance, and maintain automation workflows using Linux, Shell Scripting, and Ansible.
- Automate recurring operational tasks including cluster patch deployment, configuration management, monitoring and ing integrations, and system health checks.
- Operational Support: Troubleshoot node failures, service crashes, cluster imbalance, and distributed computing issues.
- Perform root cause analysis for high-severity incidents.
- Ensure high availability and optimal performance of Hadoop platform services.
- Work closely with engineering teams to support consistency in deployment and configuration processes.
Mandatory Skills:
- Hands-on experience in Hadoop engineering and administration.
- Strong proficiency in MAPR and Hortonworks Administration.
- Deep understanding of Hadoop ecosystem components including HDFS, YARN, MapReduce, Hive, HBase, Spark, and Zookeeper.
- Experience with Linux system administration.
- Strong expertise in Shell Scripting and Ansible Automation.
- Experience with patching and security remediation for large-scale distributed systems.
- Understanding of configuration management, service orchestration, and cluster operations.
Preferred Skills:
- Exposure to DevOps tools such as Git.
- Experience with monitoring tools like Grafana, Prometheus, and Ambari.
- Understanding of ITIL processes including Incident, Change, and Problem management.
Qualifications:
- Strong analytical and troubleshooting skills for L3 support.
- Ability to work independently and in cross-functional teams.
- Excellent communication and documentation skills.
- Strong ownership mindset toward reliability and stability of platforms.






