An interactive, glassmorphic web dashboard designed to visualize and inspect transformer self-attention mechanisms. The application interfaces with a FastAPI backend to extract attention matrices and token probability branches from a pre-trained Qwen-0.5B model, as well as a custom, randomly-initialized PyTorch Multihead Attention layer.
-
Query-Key Attention Heatmap: Interactive matrix grid displaying attention weight signatures. Clicking any cell slides open a mathematical breakdown panel showing the step-by-step vector projections (
$q_i, k_j$ ), dot-product similarity, scaling factor ($\sqrt{d_k}$ ), and softmax normalization. - Bézier Connection Map: Token-to-token interactive connection paths. Hovering over a token highlights its specific attention flow and dims irrelevant tokens.
- Multihead Grid View: A side-by-side overview of all 14 attention heads. Users can toggle individual head visibility to compare focus patterns across heads.
- Next-Token Probability Tree (D3.js): A collapsible tree rendering token generation probability branches under sampling constraints (Temperature, Top-K, Top-P). Clicking a leaf node appends that token to the prompt and triggers a new forward pass.
- Frontend: HTML5, Vanilla CSS3 (Obsidian Dark Theme + Glassmorphism), Modern ES6 JavaScript, D3.js (v7)
- Backend: Python, FastAPI, PyTorch, Hugging Face Transformers
Ensure you have Python 3.9+ and pip installed. Run:
pip install -r requirements.txtFrom the root directory, navigate to:
cd transformer-attention-visualizer
python server.pyOn startup, the server will cache and initialize the Qwen/Qwen2.5-0.5B-Instruct model and configure the custom MHA layer on GPU (CUDA) if available, defaulting to CPU otherwise.
Open your web browser and navigate to:
http://127.0.0.1:8000
Use the preset buttons or type custom prompts to analyze attention signatures!