Hi, I ran the command for benchmarking tokens/sec of the 80M parameter M2-Bert model on a 40GB A100 machine. The command I used:
python benchmark_fwd.py yamls/pretrain/monarch-mixer-pretrain-786dim-80m-parameters.yaml max_seq_len=512 device_train_microbatch_size=32 model.model_config.use_flash_mm=True
The output:
Using Monarch Mixer for Sequence Mixing: True
hyena_filter_dropout: 0.2
Using Flash MM Sequence Mixing (no bwd pass!)
-- Bidirectional: True
-- Using Long Conv Residual: True
-- Hyena w: 10
-- Hyena w mod: 1
Batch size: 32
max seq len: 512
Running forward pass...
- Forward pass
fn_amp(*inputs, **kwinputs)
73.16 ms
1 measurement, 30 runs, 48 threads
Time: 0.07315629470006874
Tokens/ms: 223.95885504005173
The reported number for this configuration in the paper is 386.3 tokens/ms. Could this difference be because of the torch or CUDA version mismatch or something else?
I have my environment in a Docker image with the tag harshauwm163/m2:0.02 on dockerhub, if you would like to reproduce my results.
I'm more interested in the improvement due to the monarch matrices. Within the same environment, I tried comparing the throughput of FFN with block diagonal Monarch matrices vs a dense MLP with the same hidden and intermediate dimensions.
The MLP with the Monarch Matrices is about 1.36x faster than its Dense counterpart with 4x lower number of parameters. Does this makes sense or should I expect more improvements in terms of tokens/ms?
dense parameters: 7080192, mmmlp parameters: 1771776
ratio = 0.25024406117800197
dense model benchmark:
- Forward pass
fn_amp(*inputs, **kwinputs)
2.19 ms
1 measurement, 3000 runs , 48 threads
Time: 0.002187790456334672
Tokens/ms: 7488.834203733128
monarch matrix mlp benchmark:
- Forward pass
fn_amp(*inputs, **kwinputs)
1.60 ms
1 measurement, 3000 runs , 48 threads
Time: 0.0016031975863346208
Tokens/ms: 10219.576264120147
Hi, I ran the command for benchmarking tokens/sec of the 80M parameter M2-Bert model on a 40GB A100 machine. The command I used:
python benchmark_fwd.py yamls/pretrain/monarch-mixer-pretrain-786dim-80m-parameters.yaml max_seq_len=512 device_train_microbatch_size=32 model.model_config.use_flash_mm=TrueThe output:
The reported number for this configuration in the paper is 386.3 tokens/ms. Could this difference be because of the torch or CUDA version mismatch or something else?
I have my environment in a Docker image with the tag
harshauwm163/m2:0.02on dockerhub, if you would like to reproduce my results.I'm more interested in the improvement due to the monarch matrices. Within the same environment, I tried comparing the throughput of FFN with block diagonal Monarch matrices vs a dense MLP with the same hidden and intermediate dimensions.
The MLP with the Monarch Matrices is about 1.36x faster than its Dense counterpart with 4x lower number of parameters. Does this makes sense or should I expect more improvements in terms of tokens/ms?
dense parameters: 7080192, mmmlp parameters: 1771776 ratio = 0.25024406117800197 dense model benchmark: - Forward pass fn_amp(*inputs, **kwinputs) 2.19 ms 1 measurement, 3000 runs , 48 threads Time: 0.002187790456334672 Tokens/ms: 7488.834203733128 monarch matrix mlp benchmark: - Forward pass fn_amp(*inputs, **kwinputs) 1.60 ms 1 measurement, 3000 runs , 48 threads Time: 0.0016031975863346208 Tokens/ms: 10219.576264120147