This project scrapes startup data from the Y Combinator website using Python. It extracts both dynamic fields (company name, batch, short description) and static fields (founder names, LinkedIn URLs).
- Used Selenium and BeautifulSoup in Python to scrape dynamic company data and static founder details.
- Leveraged ThreadPoolExecutor for concurrent scraping to improve speed and efficiency.
- Implemented fallback parsing logic to ensure robustness against HTML variations.
- Python
- Selenium
- BeautifulSoup
- pandas
- concurrent.futures (ThreadPoolExecutor)
- BeautifulSoup Web Scraping Full Course (freeCodeCamp)
- Selenium Course for Beginners (freeCodeCamp)
- Multithreading with ThreadPoolExecutor (Corey Schafer)
- Python Selenium for Beginners (The PyCoach)
yc_dynamic_data.csv: Contains scraped dynamic fieldsyc_full_data.csv: Contains dynamic + static fields