This repository contains a solution for the Gitcoin Grants Round 23 Predictive Funding Challenge. The goal is to predict how much funding each project will receive in Gitcoin Grants Round 23, including both community contributions and matching fund payouts.
The challenge requires predicting the total funding each project will receive in Gitcoin Grants Round 23. The predictions are evaluated using the root mean squared error (RMSE) between the predicted and actual funding amounts.
There are four rounds in GG23:
- WEB3 INFRA - Web3 Infrastructure
- DEV TOOLING - Developer Tooling and Libraries
- DAPPS & APPS - dApps and Apps
- MATURE BUILDERS - A new round with no community contributions
For the first three rounds, we need to predict both the allocation of the $200,000 matching pool and the community contributions. For the MATURE BUILDERS round, we only need to predict how the $600,000 will be distributed among the 30 competing projects.
Our solution uses a multi-model approach with the following components:
- Data Processing: Clean and preprocess historical data from past Gitcoin rounds.
- Feature Engineering: Extract relevant features from historical data and merge them with current projects.
- Model Training: Train separate models for each round type using various regression algorithms.
- Ensemble Predictions: Combine predictions from multiple models to improve accuracy.
- Special Handling for MATURE BUILDERS: Use a different approach for the new round with no historical data.
GG23-Predictive-Funding-Challenge/
├── data/ # Data directory
│ ├── example/ # Example data files
│ └── README.md # Data documentation
├── docs/ # Documentation
│ └── writeup_template.md # Challenge writeup template
├── models/ # Trained models directory
│ └── README.md # Models documentation
├── notebooks/ # Jupyter notebooks
│ └── exploratory_data_analysis.ipynb # Data exploration
├── scripts/ # Shell scripts
│ ├── setup.sh # Environment setup
│ ├── copy_data.sh # Data copying
│ ├── run_pipeline.sh # Pipeline execution
│ └── init_git.sh # Git initialization
├── src/ # Source code
│ ├── check_environment.py # Environment check
│ ├── check_submission.py # Submission validation
│ ├── data_processor.py # Data processing
│ ├── evaluate_submission.py # Submission evaluation
│ ├── fix_submission.py # Submission fixing
│ ├── generate_sample_submission.py # Sample generation
│ ├── main.py # Main script
│ ├── model_trainer.py # Model training
│ └── predictor.py # Prediction generation
├── .gitignore # Git ignore file
├── CONTRIBUTING.md # Contribution guidelines
├── LICENSE # License file
├── README.md # This file
└── requirements.txt # Python dependencies
- Python 3.8 or higher
- pip package manager
- Clone this repository:
git clone https://github.com/yourusername/GG23-Predictive-Funding-Challenge.git
cd GG23-Predictive-Funding-Challenge- Set up the environment:
chmod +x scripts/setup.sh
./scripts/setup.sh- Activate the virtual environment:
source venv/bin/activate-
Download the required data files:
GG Allocation Since GG18.csv- Historical dataprojects_Mar_31.csv- Current projects
-
Place the data files in the
data/directory -
Run the data copy script:
chmod +x scripts/copy_data.sh
./scripts/copy_data.shRun the complete pipeline with:
chmod +x scripts/run_pipeline.sh
./scripts/run_pipeline.shThis will:
- Process the data
- Train models for each round type
- Generate predictions
- Create a submission.csv file
You can run individual steps of the pipeline using the main.py script:
# Train models only
python src/main.py --train
# Generate predictions only
python src/main.py --predict
# Customize data and output directories
python src/main.py --train --predict --data-dir /path/to/data --models-dir /path/to/models --output-dir /path/to/outputOur solution uses the following regression models:
- Ridge Regression
- Lasso Regression
- Random Forest
- Gradient Boosting
- XGBoost
- LightGBM
For each round type, we train separate models and combine their predictions using a weighted ensemble approach. The weights are determined based on the performance of each model on the validation set.
For the MATURE BUILDERS round, which has no historical data, we use a different approach:
- Look for these projects in other rounds
- Use the average matching amount as a score
- Normalize scores and allocate the $600,000 pool based on these scores
Please see CONTRIBUTING.md for guidelines on how to contribute to this project.
This project is licensed under the MIT License - see the LICENSE file for details.
- Gitcoin for organizing the challenge
- The open-source community for providing the tools and libraries used in this project