Python Developer - web scraping at Gmware Pvt Ltd · Bengaluru (Bangalore) · 0 - 2 years · ₹3L - ₹4L / yr · Profitable · Posted 25 Sep 2025

We are seeking an experienced Web Scraping Engineer to data extraction efforts for our enterprise clients. In this role, you will be tasked with creating and maintaining robust, large-scale scraping systems for gathering structured data.
Responsibilities:
Develop and optimize custom web scraping tools and workflows.
Integrate scraping systems with data storage solutions like SQL and NoSQL databases.
Troubleshoot and resolve scraping challenges, including CAPTCHAs, rate limiting, and IP blocking.
Provide technical guidance on scraping best practices and standards.
Skills Required:
Expert in Python and scraping libraries such as Scrapy and BeautifulSoup.
Deep understanding of web scraping techniques and challenges (CAPTCHAs, anti-bot measures).
Experience with cloud platforms (AWS, Google Cloud).
Strong background in databases and data storage systems (SQL, MongoDB).

Similar jobs (10)
About Us:
Datum builds market intelligence solutions for the retail industry. We transform public and proprietary data into actionable insights that help retail brands decide where to expand, compete, and grow. We are a small, fast-moving team that works closely with customers and
believes in shipping impactful products quickly.
About the Role
We're looking for a Full Stack Platform Engineer to build and own our end-to-end product ecosystem, including:
● Data acquisition through scalable web scraping and ETL pipelines
● Customer-facing analytics platform
● Geospatial analysis engine powering retail insights
You'll work across frontend, backend, data engineering, cloud infrastructure, and geospatial systems while collaborating directly with the founder and customers.
Key Responsibilities
● Build scalable web scrapers and ETL pipelines
● Develop customer-facing features using React/Next.js
● Design REST & WebSocket APIs
● Work with PostgreSQL/PostGIS and geospatial data
● Own deployment, CI/CD, monitoring, and cloud infrastructure
● Translate customer feedback into product features
Must-Have Skills
● 3–6 years of Full Stack development experience
● Python (FastAPI/Django)
● React or Next.js with JavaScript/TypeScript
● Web scraping using Scrapy, Playwright, Selenium, or BeautifulSoup
● PostgreSQL (PostGIS preferred)
● REST APIs, JWT/OAuth
● Docker, Git, AWS/GCP/Azure
● Understanding of proxy rotation and anti-bot techniques
Good to Have
● Mapbox, Leaflet, or Google Maps Platform
● Airflow, Redis, Kafka, or SQS
● Experience with large-scale scraping (Google Maps, Zomato, Justdial, etc.)
● Geospatial analytics or retail domain experience
● Startup experience with end-to-end ownership
1st virtual , 2nd round F2F
Python pyspark, SQL, data engineer
5+yrs
Bang/hyderabad
immediate to 15days.
Job Description
We are looking for an experienced Data Engineer with strong expertise in Python, ETL, SQL, CI/CD, and DevOps to design, develop, and maintain scalable data pipelines and data processing solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines.
- Develop data processing solutions using Python.
- Write complex and optimized SQL queries, stored procedures, and data transformations.
- Build and maintain data ingestion and integration workflows.
- Implement data quality, validation, monitoring, and error-handling processes.
- Develop and maintain CI/CD pipelines for data engineering applications.
- Work with DevOps tools and practices for automated build, deployment, and infrastructure management.
- Collaborate with data analysts, data scientists, software engineers, and business teams.
- Optimize data pipelines for performance, reliability, and scalability.
- Troubleshoot production data issues and ensure timely resolution.
- Follow best practices for version control, code quality, testing, and deployment.
Mandatory Skills
- Python
- ETL
- SQL
- CI/CD
- DevOps
- Git / Version Control
- Strong problem-solving and debugging skills
1st virtual , 2nd round F2F
Python pyspark, SQL, data engineer
5+yrs
9+yrs
Bangalore/Hyderabad
immediate to 15days.
We are looking for an experienced Software Engineer to join an AI engineering startup developing a document collection platform for accountants and professional services firms.
Preference to candidates from Kerala, India.
The product eliminates the friction involved in gathering client files by automating document requests, centralising their collection, and organising incoming documents according to each organisation’s preferred folder structure.
The ideal candidate will be able to take ownership of work from start to finish, communicate clearly, and deliver high-quality solutions within tight timeframes.
What You’ll Work On
You’ll work with Python and Django daily, including models, views, templates, background jobs, and the wider product around them.
The frontend uses Django templates with HTMX and Alpine.js, built with Vite, TypeScript, and Tailwind CSS. The stack runs in Docker using PostgreSQL, Redis, RabbitMQ, and Celery.
You may also assist with ancillary projects, including custom integrations.
Project-based training will be provided.
Technology Stack
Backend: Python, Django, PostgreSQL, Celery, Redis and RabbitMQ
Frontend: Django Templates, HTMX, Alpine.js, Vite, TypeScript and Tailwind CSS
Infrastructure: Docker
Must Have
- Strong Python and Django skills
- Comfortable working with Docker
- Fluent written and spoken English
- Clear communication skills, including providing concise updates, asking honest questions, and writing information that others can act on
- Evidence of exceptional ability—not simply a list of tools, but something challenging you have built or solved
Preference will be given to candidates with at least three years of relevant professional experience.
Nice to Have
- Frontend experience with HTML, CSS and JavaScript
- Experience with HTMX, Alpine.js, TypeScript or Tailwind CSS
- Knowledge of PostgreSQL, Celery or pytest
- Experience with integrations, including APIs, OAuth and cloud storage
- Basic accounting knowledge
How to Apply
Please do not send a generic CV alone. Your application must include:
- Evidence of exceptional ability: Describe a project, open-source contribution, production system or challenging problem you solved. Include a link to the repository, write-up or demo where possible. Focus on your personal contribution by detailing the specific parts of the project where you played a critical role and explaining precisely what you built or solved.
- What you accomplished: Provide a short explanation of the outcome in your own words.
- The hardest part: Explain the hardest part of the problem and how you dealt with it.
- Your use of AI: Explain whether you use AI in your work and, if so, how you use it.
- Your professional experience and interests: Include a brief paragraph summarising your professional experience and general interests.
Applications that do not include the above Croissant details above will not be considered.
What We Offer
- For the right candidate, salary will not be a constraint
- Project-based training
- A rewarding career with genuine opportunities for professional growth
- The opportunity to work on an innovative AI-driven product
- A remote, full-time position
Job Details and Application Submission
Location: Remote
Employment Type: Full-time
Contract: One-year contract, with the possibility of extension based on satisfactory performance
Probationary Period: Six months
Preferred Experience: Three or more years
Data Engineer Short Hiring Post
🚨 Hiring: Data Engineer
🔹 Experience: 5–9 Years
🔹 Location: Bangalore / Hyderabad
🔹 Skills: PySpark, Python, SQL, ETL, CI/CD, Data Modeling
🔹 Process: L1 Virtual → L2 F2F Karat Test
🔹 F2F: Bangalore / Hyderabad Location
🔹 Positions: Immediate requirement
⚠️ Note: Candidates must be available for F2F Karat immediately after L1.
#Hiring #DataEngineer #PySpark #Python #SQL #BangaloreJobs #HyderabadJobs #Mphasis #ImmediateJoiners
About LH2 AI LabsLH2 AI Labs is an applied research lab solving data platform and curation challenges for foundation model development. We serve every frontier AI lab with the mission of delivering the best data to power the best models.
Our customers are the ones building the foundation models themselves and our work sits directly in the loop of how those systems improve. This is a rare opportunity to join a company at a defining moment in AI.
About the role:
This is a full-time, hands-on role where you will own the core infrastructure and systems that enable us to discover customers' data landscape, handle sensitive data like names, address securely and to clean and catalog them into training-ready datasets.
That path covers connectors and on-prem agentic discovery components, secrets and PII scrubbing, scale and cost efficient ingestion and clearance for use in model training. You lead a small team of 3-4 engineers and raise the quality bar in your pod.
What you'll own
- Connectors and extraction - You'll build connectors for common SaaS and databases (Google Workspace, Slack, Postgres/MySQL, Jira, MongoDB) and for specific tools such as Razorpay, GreytHR, Keka, Zoho and LeadSquared. The rule is to build only where no good open-source connector or clean export exists.
- On-prem scanning - You'll build a CLI or agent that runs at the data owner's end to sample and estimate the value of their data without shipping all of it to us.
- Clearance pipeline - Secrets detection across full git history, fail-closed. PII redaction for code and conversational data, including Indian identifiers (PAN, Aadhaar, GSTIN, IFSC, UPI). Measurable recall and precision.
- Lineage - Every output record must trace back to its raw source, and every lot must be revocable.
- Delivery - Packaging, sampling for buyers, and supporting the supply and BD teams on data questions.
You have
- 8+ years in backend or data engineering, with at least 2 years leading up to 4 or more engineers.
- You've built ingestion or ETL pipelines that run in production.
- Hands-on work with sensitive data such as PII, financial or health data, including redaction, masking or access controls.
- The judgment to decide what to build and what to adopt.
Nice to have
- Experience with Presidio, gitleaks/TruffleHog, dlt or Airbyte.
- Knowledge of DPDP, GDPR, HIPAA or SOC 2.
- You've shipped software that runs in customers' environments.
This is not a pure people-management role. You'll write code, review everything, and personally own the hardest problems in the pod.
About Us:
The QX Impact was launched with a mission to make A.I accessible and affordable and deliver AI Products/Solutions at scale for the enterprises by bringing the power of Data, AI, and Engineering to drive digital transformation. We believe without insights; businesses will continue to face challenges to better understand their customers and even lose them. Secondly, without insights businesses won't’ be able to deliver differentiated products/services; and finally, without insights, businesses can’t achieve a new level of “Operational Excellence” is crucial to remain competitive, meeting rising customer expectations, expanding markets, and digitalization.
Job Summary:
We are looking for a Senior Data Engineer who is creative, collaborative, and adaptable to join our agile team of data scientists, engineers, and UX developers. The role focuses on building and maintaining robust data pipelines to support advanced analytics, data science, and BI solutions.
As a Senior Data Engineer, you will work with internal and external data, collaborate with data scientists, and contribute to the design, development, and deployment of innovative solutions.
Key Responsibilities:
- Design, develop, test, and maintain optimal data pipeline and ETL architectures.
- Map out data systems and define/design required integrations, ETL, BI, and AI systems/processes.
- Prepare and optimize data for predictive and prescriptive modeling.
- Collaborate with teams to integrate ERP data into the enterprise data lake, ensuring seamless flow and quality.
- Enhance cloud data infrastructure on AWS or Azure for scalability and performance.
- Utilize big data tools and frameworks to optimize data acquisition and preparation.
- Build architectures to move data to/from data lakes and data warehouses for advanced analytics.
- Develop and curate data models for analytics, dashboards, and reports.
- Conduct code reviews, maintain production-level code, and implement testing approaches.
- Monitor, troubleshoot, and resolve data ingestion workflows to maintain reliability and uptime.
- Drive innovation and implement efficient new approaches to data engineering tasks.
Must-Have Skills:
- Bachelor’s degree in Computer Science, Mathematics, Engineering, or a related field.
- 5+ years of experience working with enterprise data platforms, including building and managing data lakes.
- 3–5 years of experience designing and implementing data warehouse solutions.
- Expertise in SQL, including developing stored procedures (SP) and applying advanced data design concepts.
- Proficiency in Spark (Python/Scala) and Spark Streaming for real-time data pipelines.
- Experience with AWS or Azure services (e.g., AWS Glue, Azure Data Factory, Redshift, Snowflake).
- Familiarity with big data tools such as Apache Kafka, Apache Spark, or Flink.
- Hands-on experience with orchestration tools (e.g., Apache Airflow, Prefect).
- Knowledge of CI/CD processes, version control (e.g., Git, Jenkins), and deployment automation.
- Strong problem-solving, communication, and collaboration skills.
Good-to-Have Skills:
- Experience in integrating ERP data into data lakes.
- Experience with traditional ETL tools (e.g., Talend, Pentaho).
Competencies:
- Tech Savvy - Anticipating and adopting innovations in business-building digital and technology applications.
- Self-Development - Actively seeking new ways to grow and be challenged using both formal and informal development channels.
- Action Oriented - Taking on new opportunities and tough challenges with a sense of urgency, high energy, and enthusiasm.
- Customer Focus - Building strong customer relationships and delivering customer-centric solutions.
- Optimize Work Processes - Knowing the most effective and efficient processes to get things done, with a focus on continuous improvement.
Why Join Us?
- Be part of a collaborative and agile team driving cutting-edge AI and data engineering solutions.
- Work on impactful projects that make a difference across industries.
- Opportunities for professional growth and continuous learning.
- Competitive salary and benefits package.
Application Details
Ready to make an impact? Apply today and become part of the QX Impact team!
Job Description: Python + AI
Company: Wissen Technology
Location: Bangalore, India
Experience: 5+Years
Employment Type: Full-Time
Role: Python + AI / Data Engineer
About the Role
Wissen Technology is looking for experienced Python + AI / Data Engineering professionals to join our technology team in Bangalore. The ideal candidate will have strong hands-on experience in Python, Artificial Intelligence, Generative AI, PySpark, Snowflake, and data pipeline development.
The candidate should be capable of designing and developing scalable data and AI solutions, building robust ETL/ELT pipelines, working with large datasets, and integrating AI/ML capabilities into enterprise applications.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using Python and PySpark.
- Develop robust ETL/ELT pipelines for processing large volumes of structured and unstructured data.
- Build and optimize data processing solutions using Apache Spark / PySpark.
- Develop data ingestion and transformation pipelines into Snowflake.
- Design and implement scalable Snowflake data models, tables, views, and SQL transformations.
- Work with batch and, where applicable, real-time data processing pipelines.
- Build and integrate AI and Generative AI solutions using Python.
- Develop LLM-based applications, RAG solutions, AI agents, and AI-powered services.
- Integrate AI models with enterprise data platforms and data pipelines.
- Develop REST APIs and microservices using FastAPI, Flask, or Django.
- Perform data cleansing, transformation, validation, and quality checks.
- Optimize PySpark jobs, SQL queries, Snowflake workloads, and data pipelines for performance and scalability.
- Implement data pipeline monitoring, logging, error handling, and alerting.
- Work with cloud platforms such as AWS, Azure, or GCP.
- Collaborate with Data Scientists, Data Engineers, Software Engineers, Architects, and business stakeholders.
- Participate in technical design, architecture, code reviews, and production support.
- Mentor junior engineers and contribute to engineering best practices.
Preferred Qualifications
- Bachelor's or master's degree in computer science, Engineering, Data Science, Artificial Intelligence, or a related field.
- Experience working on enterprise-scale AI and data engineering projects.
- Experience combining Python + PySpark + Snowflake + AI/GenAI in production environments.
- Experience with Databricks is an advantage.
- Experience with AI Agents / Agentic AI and tool/function calling.
- Knowledge of distributed systems and cloud-native architecture.
- Experience leading technical initiatives or mentoring engineering teams.
Description
We are looking for Senior Data Engineers to join our Data Platform team and build scalable, high-performance data platforms that power data processing, analytics, and downstream applications.
The ideal candidate will have strong experience in distributed data processing, ETL pipelines, and Big Data technologies, with hands-on expertise in Apache Spark and Python Scala.
You will be responsible for designing, developing, and optimizing large-scale data pipelines while collaborating closely with cross-functional engineering teams to build reliable, production-grade data solutions.
Key Responsibilities
- Design, develop, and maintain scalable ETL and data processing pipelines for large-scale datasets.
- Build and optimize distributed data applications using Apache Spark and Python Scala.
- Develop reliable, high-performance data pipelines for batch and streaming workloads.
- Design and manage data workflows using Apache Airflow.
- Build and operate data workloads on AWS, with strong usage of Amazon S3 for large-scale data storage.
- Work with large datasets to ensure data quality, consistency, reliability, and performance.
- Collaborate with engineering, product, analytics, and other platform teams to deliver robust data solutions.
- Optimize data workflows for scalability, reliability, performance, and cost efficiency.
- Troubleshoot production issues, identify bottlenecks, and continuously improve platform performance.
Requirements
Candidates who demonstrate:
- 5+ years of experience in Data Engineering, Big Data Engineering, or a similar role.
- Strong hands-on experience with Apache Spark and Scala.
- Experience designing, building, and maintaining large-scale ETL pipelines.
- Strong hands-on experience with AWS, particularly Amazon S3.
- Hands-on experience with Apache Airflow for workflow orchestration and scheduling.
- Strong SQL skills and a solid understanding of distributed data processing concepts.
- Experience working with batch and/or streaming data pipelines.
- Excellent debugging, problem-solving, and performance optimization skills.
- Strong communication and collaboration skills.
Good to Have
- Experience with Databricks and the broader Databricks data platform.
- Familiarity with streaming technologies such as Apache Kafka.
- Experience working on large-scale data platforms handling high-volume data workloads.
- Exposure to additional AWS data services and cloud-native data architectures.






