Windows Site Reliability Engineer at Zacks Research Pvt Ltd · Kolkata · 8 - 10 years · ₹8L - ₹10L / yr · Profitable · Posted 6 Jun 2023

Position: Windows SRE
Responsibilities:
- Windows Site Reliability Engineer with experience in managing large websites where Millions of customers hit
- Manage and monitor all installed systems and infrastructure
- Install, configure, test and maintain operating systems, application software and system management tools
- Among your responsibilities will be the installation and configuration of storage, servers, Microsoft servers (Cluster Services, File Services, Active Directory Services, Certificate Authority services), Virtual Infrastructure (Hyper-V), IIS, MS SQL and backup system
- Proactively ensure the highest levels of systems and infrastructure availability
- Monitor and test application performance for potential bottlenecks, identify possible solutions, and work with developers to implement those fixes
- Maintain security, backup, and redundancy strategies
- Write and maintain custom scripts to increase system efficiency and lower the human intervention time on any tasks
- Participate in the design of information and operational support systems
- Provide 2nd and 3rd level support
- Liaise and collaborate with vendors and Zacks personnel for problem resolution, decision making, knowledge sharing
Requirements:
- Minimum 5+ years of Windows support experience, 7, 8, 10, and Microsoft Server (all)
- Windows server expertise
- Familiar with WAN/LAN technologies
- Understanding of the OSI model
- Virtualization - MS Hyper-V, VMware, vSan
- Strong understanding of Internet protocols including HTTP(S), SSL, TCP, IP
- MS IIS administration and configuration
- MS Active Directory
- MS Storage Space
- DNS and DHCP
- SSL certificates and PKI
- Familiar with the ITIL framework
- Strong PowerShell experience
- Information Security experience a plus Other Qualifications
- Excellent attention to detail
Experience: 8-10 years

Similar jobs (9)
SRE – Network / DNS / Load Balancer
Experience: 7–11 Years
Location: Hyderabad
Work Mode: WFO
Availability: Immediate Joiner
Job Description:
- Strong experience in Site Reliability Engineering (SRE) with focus on infrastructure and application reliability.
- Hands-on experience with Network, DNS and Load Balancer troubleshooting and administration.
- Monitor system performance, availability, latency and infrastructure health.
- Troubleshoot network connectivity, DNS resolution, routing and load-balancing issues.
- Experience with load balancers such as F5, BIG-IP, HAProxy or similar technologies.
- Good understanding of TCP/IP, HTTP/HTTPS, LAN/WAN, SSL/TLS and networking concepts.
- Experience with DNS technologies such as BIND, Infoblox or equivalent.
- Work on incident management, root cause analysis and problem resolution.
- Collaborate with application, network, cloud and infrastructure teams to resolve production issues.
- Experience with monitoring and alerting tools such as Splunk, Grafana, Prometheus, AppDynamics or similar tools.
- Strong troubleshooting, production support and communication skills.
- Willingness to work from Hyderabad office (WFO) and join immediately.
Location : Mumbai
Big Picture (The Opportunity) :
Are you looking for an opportunity to advance your Career? & If you are able to maintain a positive attitude even when everything goes wrong, if you are detail oriented and self motivated with a passion to learn and improve your skills and knowledge, we have a perfect job for you !
What do we want from you ? (Our Expectations) :
- Zero to 2 years experience in Linux Operating System.
- Flexible working hours - able to support occasional nights, weekends, and call-ins, able to quickly adapt to a constantly changing faced paced environment. Open to travel to sites.
- An ideal person who is excited and motivated about running and supporting a production - grade critical infrastructure and looks for opportunities to improve processes with automation.
What are you required to do ? (Your Responsibilities) :
- Proactively maintain and develop all linux infrastructure technology to maintain a 24*7*365 uptime service.
- Engineering of systems administration-related various solutions for our various SAAS products and projects as well as operational needs.
- Proactively monitoring system performance and capacity planning.
- Providing technical support to customers for Applications, Operating systems, and networking.
- Will also be the first point of contact for our clients, where installations are placed on a permanent basis, for basic troubleshooting and problem solving.
- Fault finding, analysis and logging information for reporting of performance exceptions.
- Maintain best practices on managing systems and services across all environments.
Skills & Qualification Required (Add Value) :
- Graduate - Preferred to have Bachelor‘s Degree in Engineering, Computer Science or related field.
- He / She should be familiar with the installation and configuration of Linux operating systems and setup and operation of TCP/IP networking on Linux systems also familiar with Internet concepts including SMTP, IMAP, POP, HTTP, DNS, LDAP and related protocols.
- You should possess excellent communication skills .
- Knowledge of Email concepts, Helpdesk Concept , VoIP Concept , Cloud computing will be an added advantage.
Job Summary
We are looking for an experienced Site Reliability Engineer (SRE) with strong expertise in networking, DNS, load balancing, and hybrid cloud environments. The candidate will be responsible for maintaining the reliability, availability, performance, and scalability of production infrastructure and services.
The ideal candidate should have strong troubleshooting skills and experience working across network, cloud, infrastructure, and application environments.
Key Responsibilities
- Monitor and maintain the availability and reliability of production systems and services.
- Troubleshoot complex network, infrastructure, and application connectivity issues.
- Manage and troubleshoot DNS services, DNS resolution, records, and configuration issues.
- Configure, manage, and troubleshoot Load Balancers and traffic routing.
- Work with Layer 4 and Layer 7 networking and understand TCP/IP, HTTP/HTTPS, routing, and network connectivity.
- Support hybrid cloud environments involving on-premises infrastructure and public cloud platforms.
- Troubleshoot connectivity between on-premises data centers and cloud environments.
- Participate in production incidents, troubleshooting, root cause analysis (RCA), and problem management.
- Develop automation scripts and tools to reduce manual operational activities.
- Configure and maintain monitoring, alerting, and observability solutions.
- Work closely with Network, Cloud, DevOps, Security, and Application teams.
- Participate in on-call support and resolve production issues within defined SLAs.
- Document infrastructure, troubleshooting procedures, incident reports, and operational processes.
- Identify opportunities to improve system reliability, performance, and scalability.
Experience Required: 2–3 Years
Employment Type: Full-time
Location: Delhi
Department: IT / Infrastructure & Data Center Operations
About the Role
We are looking for a skilled Wintel Administrator (L2) with 2–3 years of hands-on experience to manage and support our data center infrastructure, which includes a virtualized VMware environment, physical servers, database servers, and network/security appliances across LAN and DMZ zones. The ideal candidate will handle day-to-day server administration, virtualization management, backups, monitoring, and L2-level troubleshooting, while working closely with the infrastructure team on ongoing projects.
Key Responsibilities
· Manage and maintain Active Directory, File Server, Web Server (IIS, Apache, Nginx), and Application Servers
· Administer and support VMware ESXi hosts and vCenter (multi-host, multi-cluster environment) — VM provisioning, migration, snapshots, and resource allocation
· Manage physical server hardware (Dell PowerEdge, HP ProLiant) including out-of-band management via iDRAC
· Install, configure, troubleshoot, and back up database servers — MS-SQL, MySQL, MongoDB
· Manage server backups and vReplication for disaster recovery of critical VMs and databases
· Support and maintain monitoring tools (e.g., Zabbix) to track server/host performance, uptime, and alerts
· Install and configure new Windows, Linux, and desktop systems
· Install and troubleshoot Windows/Linux applications and software (VB, C#, .NET, Node.js-based apps, etc.)
· Support DMZ and LAN zone server management, coordinating with network/security team on firewall, load balancer, and VPN (Site-to-Site) configurations
· Assist with basic public cloud tasks (AWS) where hybrid/DR connectivity is involved
· Manage and troubleshoot LAN/DMZ switches, WAN links, and connectivity issues at a basic level
· Assist with patch management, deployment, and server/firmware upgrades
· Monitor and maintain server performance, availability, and health across physical and virtual infrastructure
· Respond to and resolve end-user issues escalated from Desktop Support (L1)
· Maintain accurate documentation of server inventory, IP allocations, configurations, and network diagrams
· Work on ongoing IT infrastructure projects as assigned
Required Skills & Qualifications
· B.Tech or equivalent degree in Computer Science, IT, or related field
· Minimum 2–3 years of hands-on experience as a Windows/Wintel Administrator
· Strong working knowledge of Windows Server administration (2012/2016/2019/2022)
· Hands-on experience with VMware ESXi and vCenter administration (multi-vCenter environment preferred)
· Experience with physical server hardware management — Dell PowerEdge / HP ProLiant, iDRAC/ILO remote management
· Hands-on experience with IIS, Apache, or Nginx web server management
· Experience with MySQL, MS-SQL, and MongoDB — installation, configuration, troubleshooting, and backup
· Good understanding of Active Directory, DNS, DHCP, and Group Policy
· Working knowledge of firewall, load balancer, and basic network/VPN concepts (Site-to-Site VPN, WAN, DMZ)
· Experience with server/VM backup and replication solutions for disaster recovery
· Exposure to monitoring tools such as Zabbix or similar
· Basic exposure to cloud platforms (AWS) for hybrid/DR connectivity
· Familiarity with installing and troubleshooting applications like VB, C#, .NET
· Ability to troubleshoot server hardware, OS-level (Windows/Linux), and virtualization issues
· Experience with patch management and firmware/server upgrade cycles
Soft Skills
· Strong troubleshooting and analytical skills
· Good communication skills to coordinate with end users, network team, and cross-functional teams
· Ability to work independently as well as escalate/collaborate with L3 when required
· Willingness to work on-call/rotational shifts if required
· Detail-oriented with a proactive approach to server monitoring, backups, and uptime management
We are looking for a System Administrator to manage our servers and core IT infrastructure.
Responsibilities
- Administer Linux (RHEL) and Windows Server systems
- Manage VMware virtualisation and Active Directory
- Handle patching, user access and system hardening
- Monitor systems and resolve infrastructure issues
Requirements
- 1+ years of system administration experience
- Strong Linux and Windows Server skills
- RHCSA or MCSA certification is a plus
SRE / Success Engineering role focused on production operations, reliability, AWS infrastructure, monitoring, incident management, and platform support for the ZT platform.
Core responsibilities include:
- Production monitoring and debugging of live systems.
- Incident investigation, troubleshooting, and problem resolution.
- AWS cloud infrastructure support and maintenance.
- Deployment and operational support activities.
- Supporting a 24x7 production environment.
- Working with GitHub-based development workflows.
- Technical debt remediation and platform improvements.
- Customer issue investigation and support.
- Security and compliance-related work, including FedRAMP initiatives.
Preferred Skills:
AWS (especially S3 and EC2)
Strong debugging and troubleshooting skills
Site Reliability Engineering (SRE) experience
GitHub experience
Basic software development skills
TypeScript/JavaScript knowledge
C# preferred
AI experience is a plus.
Candidate should be a hands-on engineer with strong AWS, SRE, operational ownership, production support, and debugging capabilities, rather than a pure application or full-stack developer.
Job Description
We are looking for a Windows Administrator with strong hands-on experience in Windows Server administration, upgrades, patching, vulnerability remediation, and troubleshooting.
Key Responsibilities
- Perform Windows Server administration, upgrades, and patching.
- Plan and execute OS upgrades, security patches, and hotfixes.
- Identify and remediate Windows vulnerabilities and security issues.
- Troubleshoot server, OS, application, and performance-related issues.
- Monitor CPU, memory, disk, services, and overall server health.
- Perform incident management and root-cause analysis.
- Manage Active Directory, DNS, DHCP, and Group Policy.
- Support VMware / Hyper-V virtualized environments.
- Automate administrative and remediation activities using PowerShell.
- Coordinate with security, application, network, and cloud teams.
- Maintain system documentation and ensure compliance with operational standards.
Primary Skills
- Windows Server Administration
- Windows Upgrades & Patching
- Vulnerability Remediation
- Troubleshooting & RCA
- Active Directory
- DNS / DHCP / Group Policy
- PowerShell
- VMware / Hyper-V
Senior System Administrator
Location: Chennai
Experience: 10+ Years
Job Summary
We are looking for a skilled Senior System Administrator to manage and maintain enterprise infrastructure across Windows Server, VMware virtualization, and enterprise storage environments. The role will be responsible for ensuring high availability, performance, security, and reliability of IT infrastructure supporting business operations.
Key Responsibilities
Windows Server Administration
- Install, configure, maintain and troubleshoot Windows Server environments.
- Administer Active Directory, Group Policy, DNS, DHCP and File Services.
- Resolve server-related incidents and access issues.
VMware Administration
- Manage VMware vSphere, ESXi and vCenter environments.
- Perform VM migrations using vMotion.
- Manage snapshots and optimize VM resources.
Storage Administration
- Manage enterprise SAN/NAS storage environments.
- Work with storage platforms such as Dell EMC, NetApp, HPE or equivalent.
- Configure and manage LUNs, RAID, storage allocation and replication.
- Monitor storage performance and capacity.
- Support storage integration with backup environments.
Compliance & IT Operations
- Ensure infrastructure adherence to IT security policies and organizational standards.
- Support internal/external audits and compliance requirements, including ISO/SOX where applicable.
- Maintain infrastructure documentation, configurations and change records.
- Ensure patch compliance and vulnerability remediation.
- Support IT governance and implementation of standard operating procedures.
Required Skills
- Strong hands-on experience in Windows Server Administration.
- Good experience with Active Directory, Group Policy, DNS and DHCP.
- Hands-on experience with VMware vSphere, ESXi and vCenter.
- Practical knowledge of vMotion, VM snapshots and resource management.
- Experience with SAN/NAS storage administration.
- Knowledge of LUN, RAID, storage allocation and replication.
- Experience with infrastructure patching and vulnerability remediation.
- Good troubleshooting and incident-resolution skills.
- Good communication and documentation skills.
- Exposure to Linux administration.
Education & Experience
- Bachelor's degree in Computer Science, Information Technology, Electronics or related field; B.Tech/B.E/BCA/MCA or equivalent preferred.
- 10+ years of relevant infrastructure/system administration experience.
Good to Have
- Microsoft MCSA / MCSE / AZ-800 / AZ-801
- VMware VCP – Data Center Virtualization
- Dell EMC / NetApp / HPE Storage certifications
- ITIL Foundation
Key Technologies
Windows Server | Active Directory | GPO | DNS | DHCP | VMware | ESXi | vCenter | vSphere | vMotion | SAN | NAS | Dell EMC | NetApp | HPE Storage | LUN | RAID | Replication | Linux | Patch Management | Vulnerability Management
We're looking for a Site Reliability Engineer to keep our production systems fast, reliable, and scalable. Sitting at the intersection of software engineering and operations, you'll treat infrastructure as code, automate away toil, and build the observability that lets us catch problems before customers do. You'll own uptime and on-call for critical services, lead incident response and blameless postmortems, and continuously harden the platform against failure. This role suits an engineer who is as comfortable debugging a production incident at 2 a.m. as they are writing the automation that prevents the next one.
Key Responsibilities
- Own reliability, availability, and performance of production services, including on-call rotation
- Build and maintain monitoring, alerting, and observability (metrics, logs, traces)
- Automate deployments, scaling, and operational tasks to reduce manual toil
- Manage containerized workloads on Kubernetes and cloud infrastructure
- Design and maintain CI/CD pipelines for safe, frequent releases
- Lead incident response and drive blameless postmortems with clear follow-ups
- Perform capacity planning, performance tuning, and cost optimization
- Define and track SLIs/SLOs and error budgets with product teams
Requirements
- 3+ years in SRE, DevOps, or production-focused engineering
- Strong Linux administration and hands-on Kubernetes experience
- Solid experience with monitoring/observability tools (Prometheus, Grafana, ELK, or similar)
- Cloud experience with AWS, GCP, or Azure
- CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation)
- Proficient scripting in Python and/or Bash
Nice to have
- Experience with service meshes, Helm, or GitOps (ArgoCD/Flux)
- Background in high-traffic or distributed systems






