Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

 
 

Latest commit

 

History

16 Commits

Folders and files

Repository files navigation

Web Scraper

Overview

This project scrapes startup data from the Y Combinator website using Python. It extracts both dynamic fields (company name, batch, short description) and static fields (founder names, LinkedIn URLs).

Approach

  • Used Selenium and BeautifulSoup in Python to scrape dynamic company data and static founder details.
  • Leveraged ThreadPoolExecutor for concurrent scraping to improve speed and efficiency.
  • Implemented fallback parsing logic to ensure robustness against HTML variations.

Tech Stack

  • Python
  • Selenium
  • BeautifulSoup
  • pandas
  • concurrent.futures (ThreadPoolExecutor)

References

Output

  • yc_dynamic_data.csv: Contains scraped dynamic fields
  • yc_full_data.csv: Contains dynamic + static fields

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages