| Paper | Authors | Year | Reference | Links | Suggested by | Key takeaways | |
|---|---|---|---|---|---|---|---|
| #1 | I-JEPA Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture |
Mahmoud Assran et al. | 2023 | arXiv:2301.08243 | Code Slides | Orian Sharoni & Luke Dzwonczyk | View key takeawaysCore JEPA idea: learn representations by predicting the latent representation of a target region from a context region, rather than reconstructing pixels. This provides a foundation for non-generative self-supervised learning and subsequent JEPA models. |
| #1 | Music-JEPA Learning a World Model of Sound from Action |
Ziyu Wang et al. | 2026 | arXiv:2607.22000 | Demo Slides | Orian Sharoni & Luke Dzwonczyk | View key takeawaysJEPA applied to music/audio: models music as an action-conditioned dynamical system. Audio is treated as the state and piano-roll information as the action, with applications including beat tracking, composer identification, key estimation, and piano transcription. |
| #2 | --- | --- | --- | --- | --- | --- | --- |
| #3 | ImprovID Machine learning of artistic fingerprints in jazz |
Huw Cheston, Reuben Bance & Peter M. C. Harrison | 2026 | Nature Machine Intelligence 8, 1261–1274 | Code Demo | Pierre Sant-Germier | View key takeawaysArtistic fingerprint / style identification: uses machine learning to identify individual jazz pianists from their performances, capturing stylistic characteristics across melodic, harmonic, rhythmic, and dynamic dimensions. |