Sr Hadoop Operations Engineer at Multinational Company providing energy & Automation digital · Hyderabad · 7 - 12 years · ₹12L - ₹24L / yr · Posted 26 Dec 2022

Sr Hadoop Operations Engineer
at Multinational Company providing energy & Automation digital
Skills

Similar jobs (4)
Job Summary:
- We are seeking an experienced Hadoop Engineer with strong hands-on expertise in MAPR and Hortonworks platform engineering and administration. The successful candidate will join our Hadoop Platform Engineering & Automation team to support our critical enterprise data lake platform. This role emphasizes platform support, patching, vulnerability remediation, L3 issue resolution, and automation initiatives utilizing Linux, Shell Scripting, and Ansible.
Responsibilities:
- Platform Engineering & Administration: Administer, maintain, and optimize MAPR and Hortonworks Hadoop clusters.
- Provide L3 support for complex technical issues related to cluster functioning, performance, and reliability.
- Monitor cluster health, capacity, performance, and security configurations.
- Patching & Vulnerability Remediation: Execute OS-level and platform-level patching across large-scale Hadoop clusters.
- Implement remediation for platform vulnerabilities in accordance with organizational InfoSec policies.
- Collaborate with other support teams to ensure compliance with standards and mitigation timelines.
- Automation Development: Develop, enhance, and maintain automation workflows using Linux, Shell Scripting, and Ansible.
- Automate recurring operational tasks including cluster patch deployment, configuration management, monitoring and ing integrations, and system health checks.
- Operational Support: Troubleshoot node failures, service crashes, cluster imbalance, and distributed computing issues.
- Perform root cause analysis for high-severity incidents.
- Ensure high availability and optimal performance of Hadoop platform services.
- Work closely with engineering teams to support consistency in deployment and configuration processes.
Mandatory Skills:
- Hands-on experience in Hadoop engineering and administration.
- Strong proficiency in MAPR and Hortonworks Administration.
- Deep understanding of Hadoop ecosystem components including HDFS, YARN, MapReduce, Hive, HBase, Spark, and Zookeeper.
- Experience with Linux system administration.
- Strong expertise in Shell Scripting and Ansible Automation.
- Experience with patching and security remediation for large-scale distributed systems.
- Understanding of configuration management, service orchestration, and cluster operations.
Preferred Skills:
- Exposure to DevOps tools such as Git.
- Experience with monitoring tools like Grafana, Prometheus, and Ambari.
- Understanding of ITIL processes including Incident, Change, and Problem management.
Qualifications:
- Strong analytical and troubleshooting skills for L3 support.
- Ability to work independently and in cross-functional teams.
- Excellent communication and documentation skills.
- Strong ownership mindset toward reliability and stability of platforms.
Description
We are looking for Senior Data Engineers to join our Data Platform team and build scalable, high-performance data platforms that power data processing, analytics, and downstream applications.
The ideal candidate will have strong experience in distributed data processing, ETL pipelines, and Big Data technologies, with hands-on expertise in Apache Spark and Python Scala.
You will be responsible for designing, developing, and optimizing large-scale data pipelines while collaborating closely with cross-functional engineering teams to build reliable, production-grade data solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL and data processing pipelines for large-scale datasets.
- Build and optimize distributed data applications using Apache Spark and Python Scala.
- Develop reliable, high-performance data pipelines for batch and streaming workloads.
- Design and manage data workflows using Apache Airflow.
- Build and operate data workloads on AWS, with strong usage of Amazon S3 for large-scale data storage.
- Work with large datasets to ensure data quality, consistency, reliability, and performance.
- Collaborate with engineering, product, analytics, and other platform teams to deliver robust data solutions.
- Optimize data workflows for scalability, reliability, performance, and cost efficiency.
- Troubleshoot production issues, identify bottlenecks, and continuously improve platform performance.
Requirements
Candidates who demonstrate:
- 5+ years of experience in Data Engineering, Big Data Engineering, or a similar role.
- Strong hands-on experience with Apache Spark and Scala.
- Experience designing, building, and maintaining large-scale ETL pipelines.
- Strong hands-on experience with AWS, particularly Amazon S3.
- Hands-on experience with Apache Airflow for workflow orchestration and scheduling.
- Strong SQL skills and a solid understanding of distributed data processing concepts.
- Experience working with batch and/or streaming data pipelines.
- Excellent debugging, problem-solving, and performance optimization skills.
- Strong communication and collaboration skills.
Good to Have
- Experience with Databricks and the broader Databricks data platform.
- Familiarity with streaming technologies such as Apache Kafka.
- Experience working on large-scale data platforms handling high-volume data workloads.
- Exposure to additional AWS data services and cloud-native data architectures.
The recruiter has not been active on this job recently. You may apply but please expect a delayed response.
Design, develop, and maintain ETL pipelines involving large-scale data.
Develop data processing and analytics applications primarily using PySpark and Python.
Build scalable and distributed data processing solutions using Apache Spark.
Develop and deploy data applications on AWS cloud.
Work with AWS services related to storage, compute, ETL, data warehousing, analytics, and streaming.
Implement distributed storage and processing solutions capable of handling high-volume datasets.
Design data processing applications with a focus on performance, scalability, reliability, and optimization.
Work with both SQL and NoSQL databases for data storage, processing, and analytics.
Write, optimize, and analyze SQL, HQL, and NoSQL queries.
Troubleshoot data pipeline and processing issues and ensure data quality and reliability.
Collaborate with data engineers, analysts, architects, and other technical teams to deliver data-driven solutions.
Company Name – Wissen Technology
Group of companies in India – Wissen Technology & Wissen Infotech
Work Location – Whitefield, Bangalore
Website and Company profile:
www.wissen.com
LinkedIn Page:
https://www.linkedin.com/company/wissen-technology/
While you may already know about Wissen and the company history, here is a quick rundown for you.
About Wissen Technology:
· The Wissen Group was founded in the year 2000. Wissen Technology, a part of Wissen Group, was established in the year 2015.
· Wissen Technology is a specialized technology company that delivers high-end consulting for organizations in the Banking & Finance, Telecom, and Healthcare domains. We help clients build world class products.
· Our workforce has highly skilled professionals, with leadership and senior management executives who have graduated from Ivy League Universities like Wharton, MIT, IITs, IIMs, and NITs and with rich work experience in some of the biggest companies in the world.
· Wissen Technology has grown its revenues by 400% in these five years without any external funding or investments.
· Globally present with offices US, India, UK, Australia, Mexico, and Canada.
· We offer an array of services including Application Development, Artificial Intelligence & Machine Learning, Big Data & Analytics, Visualization & Business Intelligence, Robotic Process Automation, Cloud, Mobility, Agile & DevOps, Quality Assurance & Test Automation.
· Wissen Technology has been certified as a Great Place to Work®.
· Wissen Technology has been voted as the Top 20 AI/ML vendor by CIO Insider in 2020.
· Over the years, Wissen Group has successfully delivered $650 million worth of projects for more than 20 of the Fortune 500 companies.
· We have served client across sectors like Banking, Telecom, Healthcare, Manufacturing, and Energy. They include likes of Morgan Stanley, Goldman Sachs, MSCI, StateStreet, Flipkart, Swiggy, Trafigura, GE to name a few.
About Role :
Key Responsibilities
- Build and maintain data transformation pipelines using java Spark
- Develop and optimize large-scale/CPU intensive data processing using Apache Spark
- Orchestrate workflows using Airflow
- Implement data quality checks, testing, and monitoring for pipeline. Good to have exposer into managing metadata, cataloguing, and lineage
- Support schema evolution, backfills, and incremental processing
- Ensure pipelines meet SLAs for freshness, reliability, and performance
- Expertise/working knowledge in Spark and HBase(semantic layer, virtual datasets, Reflections)
Required Skills & Qualifications
- Strong hands-on experience with
- HBase
- Apache Spark
- Experience with HBase or similar lakehouse query engines
- Airflow
- Understanding of data catalogs and lineage (e.g., OpenLineage, DataHub, Apache Polaris , openlineage)
- Proficiency in Java
- Experience with Git-based development and CI/CD
Nice-to-Have Skills
- OpenTable format/Iceberg ,Apache Arrow
- CDC-based analytics pipelines
- Cloud platforms (AWS)
- Kubernetes-based data platforms







