Cutshort logo
For Employers
Gmware Pvt Ltd  logo
Python Developer - web scraping
Python Developer - web scraping

Python Developer - web scraping at Gmware Pvt Ltd · Bengaluru (Bangalore) · 0 - 2 years · ₹3L - ₹4L / yr · Profitable · Posted 25 Sep 2025

Gmware Pvt Ltd 's logo

Python Developer - web scraping

Prerna Mittal's profile picture
Posted by Prerna Mittal
0 - 2 yrs
₹3L - ₹4L / yr
Bengaluru (Bangalore)
Skills
Web Scraping
Web crawling
scrapy
Beautiful Soup

We are seeking an experienced Web Scraping Engineer to data extraction efforts for our enterprise clients. In this role, you will be tasked with creating and maintaining robust, large-scale scraping systems for gathering structured data.


Responsibilities:


Develop and optimize custom web scraping tools and workflows.

Integrate scraping systems with data storage solutions like SQL and NoSQL databases.

Troubleshoot and resolve scraping challenges, including CAPTCHAs, rate limiting, and IP blocking.

Provide technical guidance on scraping best practices and standards.


Skills Required:


Expert in Python and scraping libraries such as Scrapy and BeautifulSoup.

Deep understanding of web scraping techniques and challenges (CAPTCHAs, anti-bot measures).

Experience with cloud platforms (AWS, Google Cloud).

Strong background in databases and data storage systems (SQL, MongoDB).

Read more
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos

About Gmware Pvt Ltd

Founded :
2017
Type :
Services
Size :
20-100
Stage :
Profitable

About

N/A

Company social profiles

linkedin

Similar jobs (10)

Commercify360
Viny Deshmukh
Posted by Viny Deshmukh
Gurugram
2 - 5 yrs
₹4.8L - ₹10L / yr
skill iconPython

About Us:

Datum builds market intelligence solutions for the retail industry. We transform public and proprietary data into actionable insights that help retail brands decide where to expand, compete, and grow. We are a small, fast-moving team that works closely with customers and

believes in shipping impactful products quickly.

About the Role

We're looking for a Full Stack Platform Engineer to build and own our end-to-end product ecosystem, including:

● Data acquisition through scalable web scraping and ETL pipelines

● Customer-facing analytics platform

● Geospatial analysis engine powering retail insights

You'll work across frontend, backend, data engineering, cloud infrastructure, and geospatial systems while collaborating directly with the founder and customers.

Key Responsibilities

● Build scalable web scrapers and ETL pipelines

● Develop customer-facing features using React/Next.js

● Design REST & WebSocket APIs

● Work with PostgreSQL/PostGIS and geospatial data

● Own deployment, CI/CD, monitoring, and cloud infrastructure

● Translate customer feedback into product features

Must-Have Skills

● 3–6 years of Full Stack development experience

● Python (FastAPI/Django)

● React or Next.js with JavaScript/TypeScript

● Web scraping using Scrapy, Playwright, Selenium, or BeautifulSoup

● PostgreSQL (PostGIS preferred)

● REST APIs, JWT/OAuth

● Docker, Git, AWS/GCP/Azure

● Understanding of proxy rotation and anti-bot techniques

Good to Have

● Mapbox, Leaflet, or Google Maps Platform

● Airflow, Redis, Kafka, or SQS

● Experience with large-scale scraping (Google Maps, Zomato, Justdial, etc.)

● Geospatial analytics or retail domain experience

● Startup experience with end-to-end ownership

Read more
VY SYSTEMS PRIVATE LIMITED
Santhanalakshmi A
Posted by Santhanalakshmi A
Bengaluru (Bangalore), Hyderabad
5 - 12 yrs
₹8L - ₹16L / yr
skill iconPython
PySpark
SQL

1st virtual , 2nd round F2F

Python pyspark, SQL, data engineer

5+yrs

Bang/hyderabad

immediate to 15days.

Read more
VY SYSTEMS PRIVATE LIMITED
Dharani S
Posted by Dharani S
Bengaluru (Bangalore)
5 - 9 yrs
₹3L - ₹20L / yr
skill iconPython
DevOps
PySpark

Job Description


We are looking for an experienced Data Engineer with strong expertise in Python, ETL, SQL, CI/CD, and DevOps to design, develop, and maintain scalable data pipelines and data processing solutions.


Key Responsibilities

  • Design, develop, and maintain scalable ETL/ELT data pipelines.
  • Develop data processing solutions using Python.
  • Write complex and optimized SQL queries, stored procedures, and data transformations.
  • Build and maintain data ingestion and integration workflows.
  • Implement data quality, validation, monitoring, and error-handling processes.
  • Develop and maintain CI/CD pipelines for data engineering applications.
  • Work with DevOps tools and practices for automated build, deployment, and infrastructure management.
  • Collaborate with data analysts, data scientists, software engineers, and business teams.
  • Optimize data pipelines for performance, reliability, and scalability.
  • Troubleshoot production data issues and ensure timely resolution.
  • Follow best practices for version control, code quality, testing, and deployment.


Mandatory Skills

  • Python
  • ETL
  • SQL
  • CI/CD
  • DevOps
  • Git / Version Control
  • Strong problem-solving and debugging skills


Read more
VY SYSTEMS PRIVATE LIMITED
Bengaluru (Bangalore), Hyderabad
5 - 10 yrs
₹18L - ₹28L / yr
skill iconPython
PySpark
SQL

1st virtual , 2nd round F2F

Python pyspark, SQL, data engineer

5+yrs

9+yrs

Bangalore/Hyderabad

immediate to 15days.

Read more
MS OUTSOURCING
Remote only
3 - 5 yrs
₹4L - ₹8L / yr
skill iconPython
skill iconDjango

We are looking for an experienced Software Engineer to join an AI engineering startup developing a document collection platform for accountants and professional services firms. 


Preference to candidates from Kerala, India.


The product eliminates the friction involved in gathering client files by automating document requests, centralising their collection, and organising incoming documents according to each organisation’s preferred folder structure.


The ideal candidate will be able to take ownership of work from start to finish, communicate clearly, and deliver high-quality solutions within tight timeframes.


What You’ll Work On


You’ll work with Python and Django daily, including models, views, templates, background jobs, and the wider product around them.


The frontend uses Django templates with HTMX and Alpine.js, built with Vite, TypeScript, and Tailwind CSS. The stack runs in Docker using PostgreSQL, Redis, RabbitMQ, and Celery.


You may also assist with ancillary projects, including custom integrations.


Project-based training will be provided.


Technology Stack


Backend: Python, Django, PostgreSQL, Celery, Redis and RabbitMQ

Frontend: Django Templates, HTMX, Alpine.js, Vite, TypeScript and Tailwind CSS

Infrastructure: Docker


Must Have

  • Strong Python and Django skills
  • Comfortable working with Docker
  • Fluent written and spoken English
  • Clear communication skills, including providing concise updates, asking honest questions, and writing information that others can act on
  • Evidence of exceptional ability—not simply a list of tools, but something challenging you have built or solved

Preference will be given to candidates with at least three years of relevant professional experience.


Nice to Have

  • Frontend experience with HTML, CSS and JavaScript
  • Experience with HTMX, Alpine.js, TypeScript or Tailwind CSS
  • Knowledge of PostgreSQL, Celery or pytest
  • Experience with integrations, including APIs, OAuth and cloud storage
  • Basic accounting knowledge

How to Apply

Please do not send a generic CV alone. Your application must include:

  1. Evidence of exceptional ability: Describe a project, open-source contribution, production system or challenging problem you solved. Include a link to the repository, write-up or demo where possible. Focus on your personal contribution by detailing the specific parts of the project where you played a critical role and explaining precisely what you built or solved.
  2. What you accomplished: Provide a short explanation of the outcome in your own words.
  3. The hardest part: Explain the hardest part of the problem and how you dealt with it.
  4. Your use of AI: Explain whether you use AI in your work and, if so, how you use it.
  5. Your professional experience and interests: Include a brief paragraph summarising your professional experience and general interests.

Applications that do not include the above Croissant details above will not be considered.


What We Offer

  • For the right candidate, salary will not be a constraint
  • Project-based training
  • A rewarding career with genuine opportunities for professional growth
  • The opportunity to work on an innovative AI-driven product
  • A remote, full-time position

Job Details and Application Submission


Location: Remote

Employment Type: Full-time

Contract: One-year contract, with the possibility of extension based on satisfactory performance

Probationary Period: Six months

Preferred Experience: Three or more years

Read more
VY SYSTEMS PRIVATE LIMITED
Hyderabad, Pune
5 - 9 yrs
₹18L - ₹20L / yr
PySpark
SQL
skill iconPython

Data Engineer Short Hiring Post


🚨 Hiring: Data Engineer

🔹 Experience: 5–9 Years

🔹 Location: Bangalore / Hyderabad

🔹 Skills: PySpark, Python, SQL, ETL, CI/CD, Data Modeling

🔹 Process: L1 Virtual → L2 F2F Karat Test

🔹 F2F: Bangalore / Hyderabad Location

🔹 Positions: Immediate requirement

⚠️ Note: Candidates must be available for F2F Karat immediately after L1.

#Hiring #DataEngineer #PySpark #Python #SQL #BangaloreJobs #HyderabadJobs #Mphasis #ImmediateJoiners

Read more
LH2 AI Labs
Bengaluru (Bangalore)
8 - 12 yrs
₹50L - ₹100L / yr
skill iconPython
Software engineering
Data engineering
OAuth
RESTful APIs
+1 more

About LH2 AI LabsLH2 AI Labs is an applied research lab solving data platform and curation challenges for foundation model development. We serve every frontier AI lab with the mission of delivering the best data to power the best models.


Our customers are the ones building the foundation models themselves and our work sits directly in the loop of how those systems improve. This is a rare opportunity to join a company at a defining moment in AI.


About the role:

This is a full-time, hands-on role where you will own the core infrastructure and systems that enable us to discover customers' data landscape, handle sensitive data like names, address securely and to clean and catalog them into training-ready datasets.

That path covers connectors and on-prem agentic discovery components, secrets and PII scrubbing, scale and cost efficient ingestion and clearance for use in model training. You lead a small team of 3-4 engineers and raise the quality bar in your pod.


What you'll own

  • Connectors and extraction - You'll build connectors for common SaaS and databases (Google Workspace, Slack, Postgres/MySQL, Jira, MongoDB) and for specific tools such as Razorpay, GreytHR, Keka, Zoho and LeadSquared. The rule is to build only where no good open-source connector or clean export exists.
  • On-prem scanning - You'll build a CLI or agent that runs at the data owner's end to sample and estimate the value of their data without shipping all of it to us.
  • Clearance pipeline - Secrets detection across full git history, fail-closed. PII redaction for code and conversational data, including Indian identifiers (PAN, Aadhaar, GSTIN, IFSC, UPI). Measurable recall and precision.
  • Lineage - Every output record must trace back to its raw source, and every lot must be revocable.
  • Delivery - Packaging, sampling for buyers, and supporting the supply and BD teams on data questions.


You have

  • 8+ years in backend or data engineering, with at least 2 years leading up to 4 or more engineers.
  • You've built ingestion or ETL pipelines that run in production.
  • Hands-on work with sensitive data such as PII, financial or health data, including redaction, masking or access controls.
  • The judgment to decide what to build and what to adopt.


Nice to have

  • Experience with Presidio, gitleaks/TruffleHog, dlt or Airbyte.
  • Knowledge of DPDP, GDPR, HIPAA or SOC 2.
  • You've shipped software that runs in customers' environments.


This is not a pure people-management role. You'll write code, review everything, and personally own the hardest problems in the pod.

Read more
QuaXigma IT solutions Private Limited
Tirupati, Chennai
5 - 10 yrs
Best in industry
SQL
skill iconPython
Stored Procedures
skill iconAmazon Web Services (AWS)
Microsoft Windows Azure
+10 more

About Us:

The QX Impact was launched with a mission to make A.I accessible and affordable and deliver AI Products/Solutions at scale for the enterprises by bringing the power of Data, AI, and Engineering to drive digital transformation. We believe without insights; businesses will continue to face challenges to better understand their customers and even lose them. Secondly, without insights businesses won't’ be able to deliver differentiated products/services; and finally, without insights, businesses can’t achieve a new level of “Operational Excellence” is crucial to remain competitive, meeting rising customer expectations, expanding markets, and digitalization.


Job Summary:

We are looking for a Senior Data Engineer who is creative, collaborative, and adaptable to join our agile team of data scientists, engineers, and UX developers. The role focuses on building and maintaining robust data pipelines to support advanced analytics, data science, and BI solutions.

As a Senior Data Engineer, you will work with internal and external data, collaborate with data scientists, and contribute to the design, development, and deployment of innovative solutions.


Key Responsibilities:

  • Design, develop, test, and maintain optimal data pipeline and ETL architectures.
  • Map out data systems and define/design required integrations, ETL, BI, and AI systems/processes.
  • Prepare and optimize data for predictive and prescriptive modeling.
  • Collaborate with teams to integrate ERP data into the enterprise data lake, ensuring seamless flow and quality.
  • Enhance cloud data infrastructure on AWS or Azure for scalability and performance.
  • Utilize big data tools and frameworks to optimize data acquisition and preparation.
  • Build architectures to move data to/from data lakes and data warehouses for advanced analytics.
  • Develop and curate data models for analytics, dashboards, and reports.
  • Conduct code reviews, maintain production-level code, and implement testing approaches.
  • Monitor, troubleshoot, and resolve data ingestion workflows to maintain reliability and uptime.
  • Drive innovation and implement efficient new approaches to data engineering tasks.


Must-Have Skills:

  • Bachelor’s degree in Computer Science, Mathematics, Engineering, or a related field.
  • 5+ years of experience working with enterprise data platforms, including building and managing data lakes.
  • 3–5 years of experience designing and implementing data warehouse solutions.
  • Expertise in SQL, including developing stored procedures (SP) and applying advanced data design concepts.
  • Proficiency in Spark (Python/Scala) and Spark Streaming for real-time data pipelines.
  • Experience with AWS or Azure services (e.g., AWS Glue, Azure Data Factory, Redshift, Snowflake).
  • Familiarity with big data tools such as Apache Kafka, Apache Spark, or Flink.
  • Hands-on experience with orchestration tools (e.g., Apache Airflow, Prefect).
  • Knowledge of CI/CD processes, version control (e.g., Git, Jenkins), and deployment automation.
  • Strong problem-solving, communication, and collaboration skills.


Good-to-Have Skills:

  • Experience in integrating ERP data into data lakes.
  • Experience with traditional ETL tools (e.g., Talend, Pentaho).


Competencies:

  • Tech Savvy - Anticipating and adopting innovations in business-building digital and technology applications.
  • Self-Development - Actively seeking new ways to grow and be challenged using both formal and informal development channels.
  • Action Oriented - Taking on new opportunities and tough challenges with a sense of urgency, high energy, and enthusiasm.
  • Customer Focus - Building strong customer relationships and delivering customer-centric solutions.
  • Optimize Work Processes - Knowing the most effective and efficient processes to get things done, with a focus on continuous improvement.


Why Join Us?

  • Be part of a collaborative and agile team driving cutting-edge AI and data engineering solutions.
  • Work on impactful projects that make a difference across industries.
  • Opportunities for professional growth and continuous learning.
  • Competitive salary and benefits package.


Application Details

Ready to make an impact? Apply today and become part of the QX Impact team!


Read more
Wissen Technology
at Wissen Technology
4 recruiters
Shivangi Bhattacharyya
Posted by Shivangi Bhattacharyya
Bengaluru (Bangalore)
5 - 15 yrs
Best in industry
skill iconPython
PySpark
Snowflake
Data Structures
Generative AI
+1 more

Job Description: Python + AI

Company: Wissen Technology

Location: Bangalore, India

Experience: 5+Years

Employment Type: Full-Time

Role: Python + AI / Data Engineer


About the Role

Wissen Technology is looking for experienced Python + AI / Data Engineering professionals to join our technology team in Bangalore. The ideal candidate will have strong hands-on experience in Python, Artificial Intelligence, Generative AI, PySpark, Snowflake, and data pipeline development.

The candidate should be capable of designing and developing scalable data and AI solutions, building robust ETL/ELT pipelines, working with large datasets, and integrating AI/ML capabilities into enterprise applications.


Key Responsibilities

  • Design, develop, and maintain scalable data pipelines using Python and PySpark.
  • Develop robust ETL/ELT pipelines for processing large volumes of structured and unstructured data.
  • Build and optimize data processing solutions using Apache Spark / PySpark.
  • Develop data ingestion and transformation pipelines into Snowflake.
  • Design and implement scalable Snowflake data models, tables, views, and SQL transformations.
  • Work with batch and, where applicable, real-time data processing pipelines.
  • Build and integrate AI and Generative AI solutions using Python.
  • Develop LLM-based applications, RAG solutions, AI agents, and AI-powered services.
  • Integrate AI models with enterprise data platforms and data pipelines.
  • Develop REST APIs and microservices using FastAPI, Flask, or Django.
  • Perform data cleansing, transformation, validation, and quality checks.
  • Optimize PySpark jobs, SQL queries, Snowflake workloads, and data pipelines for performance and scalability.
  • Implement data pipeline monitoring, logging, error handling, and alerting.
  • Work with cloud platforms such as AWS, Azure, or GCP.
  • Collaborate with Data Scientists, Data Engineers, Software Engineers, Architects, and business stakeholders.
  • Participate in technical design, architecture, code reviews, and production support.
  • Mentor junior engineers and contribute to engineering best practices.


Preferred Qualifications

  • Bachelor's or master's degree in computer science, Engineering, Data Science, Artificial Intelligence, or a related field.
  • Experience working on enterprise-scale AI and data engineering projects.
  • Experience combining Python + PySpark + Snowflake + AI/GenAI in production environments.
  • Experience with Databricks is an advantage.
  • Experience with AI Agents / Agentic AI and tool/function calling.
  • Knowledge of distributed systems and cloud-native architecture.
  • Experience leading technical initiatives or mentoring engineering teams.


Read more
Proximity Works
at Proximity Works
1 video
5 recruiters
Tushar Vaghela
Posted by Tushar Vaghela
Bengaluru (Bangalore)
5 - 10 yrs
Best in industry
skill iconPython
skill iconScala
Apache Spark
Apache Kafka
databricks
+1 more

Description


We are looking for Senior Data Engineers to join our Data Platform team and build scalable, high-performance data platforms that power data processing, analytics, and downstream applications.

The ideal candidate will have strong experience in distributed data processing, ETL pipelines, and Big Data technologies, with hands-on expertise in Apache Spark and Python Scala.

You will be responsible for designing, developing, and optimizing large-scale data pipelines while collaborating closely with cross-functional engineering teams to build reliable, production-grade data solutions.



Key Responsibilities

  • Design, develop, and maintain scalable ETL and data processing pipelines for large-scale datasets.
  • Build and optimize distributed data applications using Apache Spark and Python Scala.
  • Develop reliable, high-performance data pipelines for batch and streaming workloads.
  • Design and manage data workflows using Apache Airflow.
  • Build and operate data workloads on AWS, with strong usage of Amazon S3 for large-scale data storage.
  • Work with large datasets to ensure data quality, consistency, reliability, and performance.
  • Collaborate with engineering, product, analytics, and other platform teams to deliver robust data solutions.
  • Optimize data workflows for scalability, reliability, performance, and cost efficiency.
  • Troubleshoot production issues, identify bottlenecks, and continuously improve platform performance.



Requirements

Candidates who demonstrate:

  • 5+ years of experience in Data Engineering, Big Data Engineering, or a similar role.
  • Strong hands-on experience with Apache Spark and Scala.
  • Experience designing, building, and maintaining large-scale ETL pipelines.
  • Strong hands-on experience with AWS, particularly Amazon S3.
  • Hands-on experience with Apache Airflow for workflow orchestration and scheduling.
  • Strong SQL skills and a solid understanding of distributed data processing concepts.
  • Experience working with batch and/or streaming data pipelines.
  • Excellent debugging, problem-solving, and performance optimization skills.
  • Strong communication and collaboration skills.


Good to Have

  • Experience with Databricks and the broader Databricks data platform.
  • Familiarity with streaming technologies such as Apache Kafka.
  • Experience working on large-scale data platforms handling high-volume data workloads.
  • Exposure to additional AWS data services and cloud-native data architectures.
Read more
Why apply to jobs via Cutshort
people_solving_puzzle
Personalized job matches
Stop wasting time. Get matched with jobs that meet your skills, aspirations and preferences.
people_verifying_people
Verified hiring teams
See actual hiring teams, find common social connections or connect with them directly.
ai_chip
Move faster with AI
We use AI to get you faster responses, recommendations and unmatched user experience.
Did not find a job you were looking for?
icon
Search for relevant jobs from 10000+ companies such as Google, Amazon & Uber actively hiring on Cutshort.
companies logo
companies logo
companies logo
companies logo
companies logo
Get to hear about interesting companies hiring right now
Company logo
Company logo
Company logo
Company logo
Company logo
Linkedin iconFollow Cutshort
Users love Cutshort
Read about what our users have to say about finding their next opportunity on Cutshort.
Shubham Vishwakarma's profile image

Shubham Vishwakarma

Full Stack Developer - Averlon
I had an amazing experience. It was a delight getting interviewed via Cutshort. The entire end to end process was amazing. I would like to mention Reshika, she was just amazing wrt guiding me through the process. Thank you team.
Companies hiring on Cutshort
companies logos