This repository contains activities and exercises completed as part of the Foundations of Data Science: K-Means Clustering in Python by University of London. The course focuses on understanding and applying the K-Means clustering algorithm on real-world datasets using Python.
This project uses the Banknote Authentication Dataset to explore how K-Means clustering can group data points based on variance, skewness, and other features. While the dataset contains labels, they were only used to evaluate clustering performance—not to train the model.
Key Concepts:
- Feature scaling using MinMaxScaler
- Applying KMeans to unlabeled data
- Evaluating clustering accuracy with known labels
- Visualizing clusters and decision boundaries
In this project, the World Happiness Report data is used to find patterns among countries based on happiness scores, GDP per capita, and other indicators.
Key Concepts:
- Handling real-world datasets with pandas
- Normalizing features
- Selecting optimal k using the elbow method
- Analyzing clustering results geographically and statistically
- Python (Jupyter Notebook)
- pandas, numpy
- scikit-learn (KMeans, MinMaxScaler)
- matplotlib, seaborn for data visualization
- Clone the repository
- Open the notebooks in Jupyter or VS Code with Python support or open this through Google Colab
- Run the cells to explore each clustering analysis step-by-step