Skip to content

Repository files navigation

Automatic benchmark generation suite

Suite to generate benchmarks across different WaterLily versions, Julia versions (using juliaup), backends, cases, and cases sizes using the benchmark.sh script. Profiling tools are also provided, see README_profiling.md.

TL;DR

Default benchmark for quick validations on the TGV and Jelly test cases.

sh benchmark.sh
julia --project compare.jl

On a laptop, or any machine whose clocks change with temperature or power, add -r 3 to repeat the sweep in separate processes (see Measurement methodology).

NOTE: For Ubuntu users: The default shell (/bin/sh) in Ubuntu points to dash, and not bash. Executing the sh benchmark.sh calls dash and results in a [Bad Substitution Error](benchmark.sh: 2: Bad substitution). The simple workaround is to call the shell script directly ./benchmark.sh followed by the desired arguments. This applies to all the examples below.

NOTE: For Windows users: Use git-bash to run the benchmarks. Run git-bach or open Git Bash from the Windows search bar and then you can follow the other instructions in this readme.

Detailed usage example

sh benchmark.sh -v "release 1.11" -t "1 4" -b "Array CuArray" -c "tgv jelly" -p "6,7 5,6" -s "25 25" -ft "Float32 Float64"
julia --project compare.jl --data_dir="data" --plot_dir="plots" --patterns="tgv jelly" --sort=1

runs both the TGV and jelly benchmarks (-c) using the current WaterLily version available in $WATERLILY_DIR, using 2 different Julia versions (latest release version and latest 1.11, noting that these need to be available in juliaup), 3 different backends (CPUx01, CPUx04, CUDA). Note that $WATERLILY_DIR needs to be available in the environmental variables, or otherwise the argument -wd "my/waterlily/dir" can be specified. The cases size -p, number of time steps -s, and float type -ft are bash (ordered) arrays which specify each benchmark case (respectively). They must either be equally sized to -c (one value per case) or be a single value that is propagated to every case; omitted ones use their defaults. In this example, the TGV case is run with Float32 precision and the jelly case in Float64. The default benchmarks launch (sh benchmark.sh) is equivalent to:

sh benchmark.sh -wd "$WATERLILY_DIR" -w "" -v "release" -t "4" -b "Array CuArray" -c "tgv jelly" -p "6,7 5,6" -s "25 25" -ft "Float32 Float32"

Note that -w or --waterlily can be used to pass different WaterLily versions by using commit hashes, tags, or branch names, eg. -w "master v1.2.0". An empty -w argument will benchmark the current state of $WATERLILY_DIR.

Benchmarks are then post-processed using the compare.jl script. Plots can be generated by passing the --plot_dir argument. Note that --patterns can be passed to post-process only certain benchmarks. Alternatively, benchmark files can also be passed directly as arguments with

julia --project compare.jl --plot_dir="plots" $(find data/ \( -name "tgv*json" -o -name "jelly*json" \) -printf "%T@ %Tc %p\n" | sort -n | awk '{print $7}')

Usage information

The accepted command line arguments are (parenthesis for short version):

  • Backend arguments: --waterlily(-w), --waterlily_dir(-wd), --data_dir(-dd), --versions(-v), --backends(-b), --threads(-t), --update(-u). Respectively: List of WaterLily git hashes/tags/branches to test, WaterLily.jl directory, path to store the benchmark data, Julia version, backend types, number of threads (when --backends contains Array), and whether to update the Julia environment before each run (Pkg.develop $WATERLILY_DIR + Pkg.update). --update accepts true/false/1/0 and defaults to false; set -u true (or -u 1) when the WaterLily/Julia version changes between runs. The version/backend/threads arguments accept a list of different parameters, for example:
    -w "fae590d e22ad41" -v "1.8.5 1.9.4" -b "Array CuArray" -t "1 6"
    would generate benchmark for all these combinations of parameters.
  • Case arguments: --cases(-c), --log2p(-p), --max_steps(-s), --ftype(-ft). The --cases argument specifies which cases to benchmark, and it can be again a list of different cases. The name of the cases needs to be defined in cases.jl, for example tgv or jelly. The current available cases are "tgv sphere cylinder donut jelly". Hence, to add a new case first define the function that returns a Simulation in cases.jl, and then it can be called using the --cases(-c) list argument. Case size, number of time steps, and float data type are then defined for each case (-p, -s, -ft, respectively). Each of -p, -s, -ft may either provide one value per case (an array matching the length of -c), or a single value that is broadcast to every case. Omitted arguments use the defaults at the top of benchmark.sh: the per-case size map DEF_LOG2P (currently tgv→6,7, jelly→5,6, sphere→3,4, cylinder→4,5) and the uniform DEF_MAXSTEPS (25) and DEF_FTYPE (Float32). So -c "tgv sphere" -ft "Float64" runs tgv at 6,7 and sphere at 3,4, both in Float64. A case with no entry in DEF_LOG2P (e.g. donut) errors on an omitted -p, with a hint to add it there or pass -p explicitly; -s/-ft always default. Omitting -c runs only the default sweep (tgv jelly).
  • BiotSavartBCs pairing (optional): --biotsavart(-bs), --biotsavart_dir(-bsd). Some cases (e.g. jelly) solve the pressure Poisson equation through BiotSavartBCs.jl's own mom_project! rather than WaterLily's solver!, so benchmarking a change that spans both packages (such as the Poisson stopping criterion) requires switching BiotSavartBCs in tandem with WaterLily. Pass -bs a list of BiotSavartBCs branches/tags/hashes paired 1:1 with -w (e.g. -w "master poisson-rms-tol" -bs "main combined-tol") and point -bsd (or $BIOTSAVART_DIR) at a local BiotSavartBCs clone. Each run checks out the WaterLily version first, then its paired BiotSavartBCs version (order matters when the BiotSavartBCs branch needs new WaterLily symbols). Omit -bs to use the BiotSavartBCs pinned in Project.toml's [sources].
  • Repetitions: --repeats(-r). Repeats the whole sweep N times (default 1), each benchmark in a new Julia process, reversing the Julia/WaterLily version order on even repetitions. The files get a _r<n> suffix and compare.jl merges the repetitions of a benchmark into a single row (the Reps column shows how many were merged). Use it when the machine state can change between processes, which is typical on laptops (CPU frequency, thermal and power limits): the runs of a single process share that state, so with Reps = 1 the Noise column cannot show it. See Measurement methodology.
  • Developed-flow checkpoints: --developed(-dev). By default the suite times sim_step! starting from a developed flow loaded from checkpoints/<case>_<log2p>_<ftype>.jld2 (git-LFS) instead of from the startup transient, so the measurement reflects steady-state cost. Generate the checkpoints with julia --project=. develop.jl (which advances each case to its develop_time in util.jl) and render a vorticity image per case with julia --project=. visualize.jl. A requested case/size with no stored checkpoint is a hard error; pass -dev "" to time the transient instead, or -dev <dir> to read checkpoints from a different directory.

The following command

sh benchmark.sh -v "release" -t "1 3 6" -b "Array CuArray" -c "tgv sphere" -p "6,7,8 5,6" -s "10 100" -ft "Float64 Float32"

would allow running benchmarks with 4 backends: CPUx01 (serial), CPUx03, CPUx06, GPU-NVIDIA. Additionally, two benchmarks would be tested, tgv and sphere, with different sizes, number of time steps, and float type, each. This would result into 1 Julia version x (3 Array + 1 CuArray) backends x (3 TGV sizes + 2 jelly sizes) = 20 benchmarks.

Benchmarks are saved in JSON format with the following nomenclature: casename_sizes_maxsteps_ftype_backend_waterlilyHEADhash_juliaversion.json. Benchmarks can be finally compared using compare.jl as follows

julia --project compare.jl benchmark_1.json benchmark_2.json benchmark_3.json ...

or by using pattern syntax

julia --project compare.jl --data_dir="data/benchmark" --patterns="tgv*CPU"

for which only TGV benchmarks on a CPU backend found in the "data/benchmark" directory would be processed. The following syntax would produce equivalent results:

julia --project compare.jl $(find data/benchmark -name "tgv*CPU.json" -printf "%T@ %Tc %p\n" | sort -n | awk '{print $7}') --sort=1

by taking the tgv JSON files, sort them by creation time, and pass them as arguments to the compare.jl program. Multiple ppaterns can also be specified with --patterns="tgv jelly" for example.

The --speedup_base="<backend>,<waterlily hash/ref>,<julia version>" argument (or a subset of it, ie. "<backend>,<julia version>") can be passed to reference speed-ups: speedup_x = time(benchmark_<backend>) / time(benchmark_x). Tokens are matched case-sensitively against the tags stored in the JSON files, so use CPUx04 (not cpux04); if no match is found the script prints the available (Backend, WaterLily ref, Julia, FP) quadruples and the unmatched tokens. The --sort=<column> argument can also be used when running the comparison. It will sort the benchmark table rows by the values of that column, given by (the start of) its name, case-insensitive: --sort=Backend sorts by backend alphabetically; --sort=Mean sorts by mean step; --sort=Median sorts by median step; --sort=Speedup sorts by speedup; --sort=Δ sorts by Δ; --sort=Signif sorts by significance. A column index (--sort=<1 to 14>, counting the hidden GC column) also works, but it changes whenever the columns do. The speedup baseline row is highlighted in blue, and the fastest run per backend is highlighted in green. The --gc (or --gc=1) flag shows the GC fraction column, which is hidden by default. Last, a --backend_color=<colorscheme> can be passed to use a certain color scheme if plotting results (ie. passing the --plot_dir=<plot_dir> argument).

Measurement methodology

What is timed. One sample is a single sim_step! followed by a backend synchronisation. A run is max_steps consecutive steps (default 25), and each benchmark does 5 runs. With developed flows (the default), every run is first reset to the same checkpoint, so all runs time the same steps of the flow; with -dev "" the runs march on continuously from the startup transient. Before the runs, the case is warmed up from the same state for at least 50 steps and 2 seconds (JIT, device clocks), and GC.gc() is called.

Reference time. For each run, take the mean of its per-step times, i.e. the run time divided by max_steps. The reference, shown as Mean, is the minimum of the run means, which discards runs slowed down from outside, as BenchmarkTools' minimum does: a step slowed down by an interrupt or a GC pause raises the mean of its run, and the minimum drops that run. Every run times the same steps, so the mean counts all the work of those steps, also when steps do different amounts of work. In jelly, for example, a pressure solve needs a varying number of Biot-Savart boundary updates, so its step times come in levels.

The minimum of the run medians is still shown, as Median. It describes the typical step and ignores occasional slow steps, but it is not a measure of cost when steps do different amounts of work: when about half the steps sit on each level, one step changing level moves the median by the whole gap between levels, while the total work barely changes. In the example below, Median of jelly on CPUx01 jumps by 17% from v1.7.0 to v1.8.0 while Mean changes by 1%. When all steps do the same work, as in tgv or sphere, Mean and Median are close.

Table columns. An example table from a comparison of WaterLily releases (jelly at log2p = 5, master as speedup baseline, 3 repetitions):

▶ log2p = 5
┌────────────┬───────────┬────────┬─────────┬───────┬────────┬────────┬─────────────┬─────────┬───────┬──────────────┬─────────┬──────┐
│  Backend   │ WaterLily │ Julia  │   FP    │ Alloc │  Mean  │ Median │    Cost     │ Speedup │ Noise │    Δ ± σ     │ Signif  │ Reps │
│            │           │        │         │  [k]  │  [ms]  │  [ms]  │ [ns/DOF/dt] │         │  [%]  │     [%]      │ [|Δ|/σ] │      │
├────────────┼───────────┼────────┼─────────┼───────┼────────┼────────┼─────────────┼─────────┼───────┼──────────────┼─────────┼──────┤
│     CPUx01 │    master │ 1.11.5 │ Float32 │   0.7 │ 157.99 │ 165.25 │     1205.41 │    1.00 │   0.4 │            - │       - │    3 │
│     CPUx01 │    v1.6.1 │ 1.11.5 │ Float32 │   0.8 │ 159.65 │ 143.64 │     1218.04 │    0.99 │   0.6 │  +1.0 ±  0.7 │    1.5  │    3 │
│     CPUx01 │    v1.7.0 │ 1.11.5 │ Float32 │   0.8 │ 160.19 │ 144.18 │     1222.13 │    0.99 │   0.5 │  +1.4 ±  0.6 │    2.2  │    3 │
│     CPUx01 │    v1.8.0 │ 1.11.5 │ Float32 │   0.8 │ 161.27 │ 168.63 │     1230.37 │    0.98 │   0.5 │  +2.1 ±  0.7 │    3.1  │    3 │
│     CPUx04 │    master │ 1.11.5 │ Float32 │  33.2 │  56.61 │  57.74 │      431.90 │    2.79 │   1.6 │            - │       - │    3 │
│     CPUx04 │    v1.6.1 │ 1.11.5 │ Float32 │  51.6 │  63.00 │  57.98 │      480.61 │    2.51 │   1.2 │ +11.3 ±  2.0 │    5.7  │    3 │
│     CPUx04 │    v1.7.0 │ 1.11.5 │ Float32 │  53.6 │  64.39 │  59.71 │      491.28 │    2.45 │   1.1 │ +13.7 ±  1.9 │    7.2  │    3 │
│     CPUx04 │    v1.8.0 │ 1.11.5 │ Float32 │  62.3 │  65.31 │  66.21 │      498.29 │    2.42 │   1.5 │ +15.4 ±  2.2 │    7.0  │    3 │
│ GPU-NVIDIA │    master │ 1.11.5 │ Float32 │  36.0 │   7.67 │   7.80 │       58.55 │   20.59 │   5.6 │            - │       - │    3 │
│ GPU-NVIDIA │    v1.6.1 │ 1.11.5 │ Float32 │  31.9 │   7.81 │   7.25 │       59.61 │   20.22 │   1.4 │  +1.8 ±  5.8 │    0.3  │    3 │
│ GPU-NVIDIA │    v1.7.0 │ 1.11.5 │ Float32 │  34.0 │   7.44 │   6.77 │       56.76 │   21.24 │   4.6 │  -3.1 ±  7.3 │    0.4  │    3 │
│ GPU-NVIDIA │    v1.8.0 │ 1.11.5 │ Float32 │  42.6 │   8.29 │   8.37 │       63.26 │   19.05 │   4.5 │  +8.1 ±  7.2 │    1.1  │    3 │
└────────────┴───────────┴────────┴─────────┴───────┴────────┴────────┴─────────────┴─────────┴───────┴──────────────┴─────────┴──────┘
  • Alloc [k]: allocations per step, averaged over a block of max_steps steps, divided by 1000. Only meaningful on the SIMD backend (CPUx01); the KernelAbstractions backends report kernel-launch bookkeeping.
  • GC [%] (hidden by default; show with --gc): GC fraction of the fastest step.
  • Mean [ms]: the reference time per step, the minimum over runs of the mean step of each run. It is the point estimate used for cost, speedup and Δ.
  • Median [ms]: the minimum over runs of the median step of each run, i.e. the typical step. It differs from Mean when steps do different amounts of work, see Reference time.
  • Cost [ns/DOF/dt]: Mean divided by the number of cells, i.e. the cost per cell and per time step.
  • Speedup: Mean(speedup_base) / Mean. The speedup baseline is a single row for the whole table, the first one by default (set it with --speedup_base), so it also compares across backends.
  • Noise [%]: scatter of the row's measurement relative to Mean (one standard deviation), i.e. how much Mean changes from run to run and from process to process, see below.
  • Δ ± σ [%]: cost difference against the reference row, which is the row of the same backend matching the speedup baseline's (WaterLily ref, Julia, FP), e.g. the master row of that backend when comparing master against a PR. The reference row itself prints -. Positive means slower. σ is the scatter of Δ, σ = sqrt(Noise_row² + Noise_ref²), because Δ is the difference of two noisy rows.
  • Signif [|Δ|/σ]: significance of Δ (it is not Δ divided by the row's own Noise). Below about 1 the Δ is indistinguishable from scatter; believe it from about 2 to 3. A * marks a value where the row or its reference row has Reps = 1: σ then only covers the scatter within a process.
  • Reps: number of repetitions (separate processes) merged into the row, see --repeats.

Noise and repetitions. All the runs of a benchmark happen inside one Julia process, so they share any machine state that outlives the process (CPU frequency, thermal or power state). On a power- or thermal-limited machine such as a laptop, two processes measuring the same code a few minutes apart have been seen to differ by 14% while each reported a noise below 1%, and no run of the slow process reached the level of the fast one. A single process cannot measure this scatter, so with Reps = 1 a small Noise and a large Signif do not make a Δ trustworthy on such machines. With --repeats N each benchmark is measured in N separate processes, and:

  • Mean is the minimum over the processes of their reference times.
  • Noise with Reps = 1 is the std of the 5 run means, divided by Mean.
  • Noise with Reps > 1 is the larger of (a) the std of the per-process reference times and (b) the largest std of the run means of a process, divided by Mean. (a) measures how much the reported number changes from one process to the next; (b) is a floor, since (a) is poorly estimated from 2 or 3 processes. Both are standard deviations, so the column means the same with and without repetitions.

Example with 3 repetitions, run means in ms:

          Process 1   Process 2   Process 3
  Run 1     36.4        41.3        36.5
  Run 2     36.5        41.4        36.6
  Run 3     36.6        41.5        36.7
  Run 4     36.4        41.3        36.5
  Run 5     36.5        41.4        36.6

  per-process reference times (min of each column): 36.4, 41.3, 36.5  ->  Mean = 36.4
  (a) between processes: std(36.4, 41.3, 36.5) = 2.80 ms
  (b) within a process:  largest std of a column = 0.08 ms
  Noise = max(2.80, 0.08) / 36.4 = 7.7%

Three repetitions are a better default than two on a laptop: with two, (a) rests on a single difference, and both processes can land in the same machine state by chance.

About

Automated benchmark suite

Resources

Stars

5 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages