A constrained ranking system that predicts exact 1-through-68 seed assignments for all NCAA March Madness tournament teams using pairwise comparison, ensemble blending, combinatorial optimization, and domain-specific post-processing.
Predicting the NCAA Selection Committee's seeding decisions is a challenging constrained ranking problem: each of 68 teams must receive a unique seed, rankings are relative rather than absolute, and the committee's decision process incorporates substantial subjective judgment. We present a four-stage pipeline that transforms this into a tractable learning-to-rank task. Stage 1 converts the problem from 68 point predictions into ~4,500 pairwise comparisons per season, increasing effective training data by 66×. Stage 2 blends three classifiers (Logistic Regression and XGBoost) with cross-validation-optimized weights. Stage 3 applies the Hungarian algorithm to enforce the one-team-per-seed constraint via globally optimal assignment. Stage 4 applies domain-specific zone corrections targeting systematic committee biases. On leave-one-season-out cross-validation across 5 seasons (340 teams), with all label-derived features built inside the evaluation fold, the full pipeline achieves 78.0% exact-match accuracy (71/91) with RMSE 2.689; the learning core alone (no hand-tuned corrections) achieves 59.3%. On a genuinely prospective test — the 2025–26 committee field seeded by the model frozen before that season existed — the pairwise ensemble places 50 of 68 teams within ±2 seeds (RMSE 2.48), while both the learned committee-Ridge step and the hand-tuned corrections reduce accuracy out of sample. An earlier version of this pipeline reported 91.2%; that figure was inflated by two features that aggregated ground-truth seeds across all cohorts, and the correction is documented in full in the accompanying paper and in revision_experiments/leakage_check.py.
Keywords: learning-to-rank, Hungarian algorithm, constrained optimization, ensemble methods, sports analytics, NCAA basketball
|
| Metric | Pairwise base | + Ridge step | Full pipeline | |:-------|------:|------:| | Exact Matches | 14 / 68 | 5 / 68 | 6 / 68 | | Within ±2 Seeds | 50 / 68 (73.5%) | 48 / 68 | 47 / 68 | | MAE | 1.82 | — | 2.38 | | RMSE | 2.48 | 2.98 | 3.08 | | Squared Error | 418 | 602 | 646 | |
| Season | Teams | Exact | Accuracy | RMSE |
|---|---|---|---|---|
| 2020–21 | 18 | 16 | 88.9% | 0.333 |
| 2021–22 | 17 | 11 | 64.7% | 6.049 |
| 2022–23 | 21 | 17 | 81.0% | 0.976 |
| 2023–24 | 21 | 15 | 71.4% | 0.756 |
| 2024–25 | 14 | 12 | 85.7% | 0.378 |
| Total | 91 | 71 | 78.0% | 2.689 |
┌──────────────────────────────────┐
│ 20 Raw Statistical Features │
│ (NET, SOS, W-L, Quads, Conf.) │
└──────────────┬───────────────────┘
│
┌──────────────▼───────────────────┐
│ Feature Engineering (→ 68 dims) │
│ Ratios · Composites · Context │
└──────────────┬───────────────────┘
│
┌───────────────────┼───────────────────┐
▼ ▼ ▼
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
│ LR (C=5.0) │ │ LR (C=0.5) │ │ XGBoost │
│ 68 feats │ │ Top-25 feats │ │ 68 feats │
│ Adj-pairs ≤30 │ │ All pairs │ │ All pairs │
│ Weight: 64% │ │ Weight: 28% │ │ Weight: 8% │
└────────┬────────┘ └────────┬────────┘ └────────┬────────┘
└───────────────────┼───────────────────┘
▼
┌──────────────────────────────────┐
│ Dual-Hungarian Ensemble │
│ 75% Pairwise + 25% Ridge(α=10) │
│ ↓ Hungarian Assignment (p=0.15) │
│ Globally optimal 1-to-68 mapping │
└──────────────┬───────────────────┘
│
┌──────────────▼───────────────────┐
│ 7 Zone Corrections + AQ↔AL Swap │
│ Domain-specific post-processing │
└──────────────┬───────────────────┘
│
┌──────────────▼───────────────────┐
│ Final Seed Assignments 1–68 │
└──────────────────────────────────┘
Instead of directly regressing seed values, we frame the problem as pairwise learning-to-rank: for each pair of teams
Three pairwise classifiers are blended with learned weights:
- Component 1 (64%): Logistic Regression (C=5.0), full 68 features, adjacent pairs only (gap ≤ 30)
- Component 2 (28%): Logistic Regression (C=0.5), top-25 features, all pairs
- Component 3 (8%): XGBoost (depth=4, 300 trees, lr=0.05), full features, all pairs
Pairwise scores yield a continuous ranking, but valid seeds require a discrete bijection from teams to
Seven seed-range-specific correction rules address systematic committee biases (e.g., mid-major auto-qualifiers under-seeded, power-conference at-large teams over-seeded). Corrections only re-order teams within assigned seeds — they cannot introduce or remove assignments.
Two audits govern how these numbers should be read. First, a leakage audit: two "historical prior" features averaged the true seed over all cohorts, so held-out teams contributed to their own feature; rebuilding them inside each fold moves the headline from 91.2% to 78.0% and accounts for the entire difference (revision_experiments/leakage_check.py). Second, a transfer audit: under walk-forward validation and on the prospective 2025–26 season, the pairwise learning core generalizes while the hand-tuned zone corrections do not. Key findings:
| Finding | Detail |
|---|---|
| The pairwise base transfers | On the never-seen 2025–26 field the pairwise ensemble alone reaches RMSE 2.48 (50/68 within ±2), better than any retrospective walk-forward step |
| The layers on top of it do not | The learned committee-Ridge step turns 14 exact matches into 5 and raises SE from 418 to 602; the zone corrections and swap then give 6/68 and SE 646. A deployment should use the pairwise base alone |
| Pairwise reduction is what matters | Every pairwise method (RankSVM, RankNet, this ensemble) beats every pointwise/listwise baseline by 2.5–4× in SE — including TabPFN v2 (2025), the strongest pointwise model tested (SE 12,564 vs RankSVM 5,062); a RankNet baseline outperforms this ensemble on both criteria, and the paper claims no algorithmic superiority |
| Metric choice matters | Sorting by NET alone ties the learning core on exact matches (14/68) but carries 80% more squared error (752 vs 418) |
| Projected fields are not real fields | The 2025–26 data file shipped with a projected field; 10 conference-tournament upsets changed it. data/NCAA_2026_Data_truefield.csv carries the committee's actual field and bid types |
├── README.md # This file
├── LICENSE # MIT License
├── CITATION.cff # Citation metadata
├── requirements.txt # Python dependencies
├── ncaa_2026_model.py # Core model: features, training, inference
├── generate_kaggle_submission.py # Competition submission (pooled priors; see note below)
├── predict_2026.py # End-to-end 2025–26 prediction runner
├── run_baselines.py # Classical LTR baselines (fold-internal)
│
├── revision_experiments/ # Everything reported in the IEEE Access paper
│ ├── pipeline.py # fold-internal feature construction + staged pipeline
│ ├── leakage_check.py # quantifies the pooled-vs-fold-internal effect
│ ├── decomposition_analysis.py # learning-vs-rules split, when corrections help/fail
│ ├── walk_forward.py # expanding-window chronological validation
│ ├── prospective_2026.py # frozen-model test on the real 2025–26 field
│ ├── baselines_extended.py # 10 baselines incl. MLP, RankNet, TabNet, TabPFN v2
│ ├── nested_loso.py / blend_hpo.py # robustness and sensitivity analyses
│ ├── feature_importance.py # leak-free combined importance ranking
│ └── results/VERIFIED_NUMBERS.md # the single source of truth for every number
│
├── data/
│ ├── NCAA_Seed_Training_Set2.0.csv # 249 labeled teams (2020–2024)
│ ├── NCAA_Seed_Test_Set2.0.csv # 91 labeled teams (held-out seasons)
│ ├── NCAA Statistics.xlsx # Full D-I stats, 2025–26 season
│ ├── NCAA_2026_Data.csv # 2025–26 input with the PROJECTED field (pre-Selection Sunday)
│ ├── NCAA_2026_Data_truefield.csv # 2025–26 input with the committee's actual field and bid types
│ └── TRUE_SEEDS_2025-26.csv # Committee's overall 1–68 seed list (ground truth)
│
└── output/ # Model predictions & submissions
├── submission_kaggle.csv
└── 2026/ # 2025–26 bracket predictions
git clone https://github.com/Om-singhaI/NCAA.git
cd NCAA
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txtpython revision_experiments/leakage_check.py # pooled vs fold-internal priors (Table 3)
python revision_experiments/paper_tables.py # headline, per-cohort, ablation, errors
python revision_experiments/decomposition_analysis.py # learning vs rules; when corrections help/fail
python revision_experiments/walk_forward.py # chronological validation
python revision_experiments/prospective_2026.py # frozen model on the real 2025–26 field
python revision_experiments/baselines_extended.py # ten baselines under one protocol
python revision_experiments/baselines_extended.py --only "TabPFN v2" # (needs tabpfn==2.2.1)Every number in the paper is listed with its provenance in revision_experiments/results/VERIFIED_NUMBERS.md.
Note on
generate_kaggle_submission.py. This is the original competition script. It builds the conference-bid prior features once over all labeled teams, which is appropriate for producing a submission but leaks held-out labels into the evaluation and reproduces the superseded 91.2% figure. It is kept for the competition record; userevision_experiments/for any evaluation.
# 1. Place NCAA Statistics Excel in data/
# 2. Generate predictions:
python predict_2026.py
# → output/2026/seed_selections_2026.txt
# → output/2026/submission_2026.csvEach team record contains 20 statistical features:
| Feature | Description |
|---|---|
NET Rank |
NCAA Evaluation Tool ranking (primary Selection Committee metric) |
NETSOS |
NET Strength of Schedule |
AvgOppNETRank |
Average opponent NET ranking |
PrevNET |
Previous season's NET ranking |
WL, Conf.Record, RoadWL |
Win-loss records (overall, conference, road) |
Quadrant1–Quadrant4 |
Record against each quality quadrant |
Conference, Bid Type |
Conference affiliation, auto-qualifier (AQ) vs at-large (AL) |
These 20 raw features are engineered into 68 model features across six categories: raw rankings, parsed win-loss metrics, quadrant quality scores, composite ratings, bid-type interactions, and historical context features.
This model evolved through 50 iterations, each targeting specific failure modes:
| Version | Architecture Change | Exact Match | RMSE | SE |
|---|---|---|---|---|
| v27 | Pairwise LR baseline | 67/91 | 2.31 | 487 |
| v45c | + Feature engineering (68 feats) | 66/91 | 1.60 | 233 |
| v46 | + Zone corrections (5 zones) | 67/91 | 1.20 | 132 |
| v47 | + Dual-Hungarian ensemble | 73/91 | 1.02 | 94 |
| v48 | + Zones 6–7 refinement | 76/91 | 0.94 | 80 |
| v49 | + AQ↔AL swap rule | 81/91 | 0.42 | 16 |
| v50 | + Zone parameter tuning | 83/91 | 0.39 | 14 |
| paper | Fold-internal label-derived features (leak removed) | 71/91 | 2.69 | 658 |
The v27–v50 rows are the development log as it was kept at the time and were computed with the pooled prior features; they are retained for the record, not as results.
If you use this work in your research, please cite:
@software{singhal2026ncaa,
author = {Singhal, Om},
title = {Pairwise Learning-to-Rank with Hungarian Assignment for {NCAA} Tournament Seed Prediction},
year = {2026},
url = {https://github.com/Om-singhaI/NCAA},
note = {NCAA Final Four Analytics Challenge}
}This project is licensed under the MIT License — see LICENSE for details.
- ESPN bracketology projections for 2025–26 field composition
- The Hungarian algorithm implementation via SciPy (
linear_sum_assignment)