Skip to content

Unable to reproduce benchmarking results from the paper #44

Description

@harshaUwm163

Hi, I ran the command for benchmarking tokens/sec of the 80M parameter M2-Bert model on a 40GB A100 machine. The command I used:

python benchmark_fwd.py yamls/pretrain/monarch-mixer-pretrain-786dim-80m-parameters.yaml max_seq_len=512 device_train_microbatch_size=32 model.model_config.use_flash_mm=True

The output:

Using Monarch Mixer for Sequence Mixing: True
hyena_filter_dropout: 0.2
Using Flash MM Sequence Mixing (no bwd pass!)
-- Bidirectional: True
-- Using Long Conv Residual: True
-- Hyena w: 10
-- Hyena w mod: 1
Batch size:  32
max seq len:  512
Running forward pass...
 - Forward pass

fn_amp(*inputs, **kwinputs)
  73.16 ms
  1 measurement, 30 runs, 48 threads
Time:  0.07315629470006874
Tokens/ms:  223.95885504005173

The reported number for this configuration in the paper is 386.3 tokens/ms. Could this difference be because of the torch or CUDA version mismatch or something else?

I have my environment in a Docker image with the tag harshauwm163/m2:0.02 on dockerhub, if you would like to reproduce my results.


I'm more interested in the improvement due to the monarch matrices. Within the same environment, I tried comparing the throughput of FFN with block diagonal Monarch matrices vs a dense MLP with the same hidden and intermediate dimensions.

The MLP with the Monarch Matrices is about 1.36x faster than its Dense counterpart with 4x lower number of parameters. Does this makes sense or should I expect more improvements in terms of tokens/ms?

dense parameters: 7080192, mmmlp parameters: 1771776                                                                                                                                                                                                                                                                   
ratio = 0.25024406117800197     
                                                                                                                                                                                                                                                                                        
dense model benchmark:                                                                                                                                                                                                                                                                                                 
 - Forward pass                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      
fn_amp(*inputs, **kwinputs)                                                                                                                                                                                                                                                                                            
  2.19 ms                                                                                                                                                                                                                                                                                                              
  1 measurement, 3000 runs , 48 threads                                                                                                                                                                                                                                                                                
Time:  0.002187790456334672                                                                                                                                                                                                                                                                                            
Tokens/ms:  7488.834203733128   

                                                                                                                                                                                                                                                                                       
monarch matrix mlp benchmark:                                                                                                                                                                                                                                                                                          
 - Forward pass                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                     
fn_amp(*inputs, **kwinputs)                                                                                                                                                                                                                                                                                            
  1.60 ms                                                                                                                                                                                                                                                                                                              
  1 measurement, 3000 runs , 48 threads                                                                                                                                                                                                                                                                                
Time:  0.0016031975863346208                                                                                                                                                                                                                                                                                           
Tokens/ms:  10219.576264120147        

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions