Repository navigation
vello_gpu: Optimize full-sized texture sampling - #1964
laurenz-canva wants to merge 1 commit into
Conversation
048b162 to
42ba050
Compare
42ba050 to
bfec171
Compare
5b66f7a to
9eaef9e
Compare
9eaef9e to
c31e48a
Compare
fa12eb3 to
597aa52
Compare
| /// **It is important that the provided width and height match the actual dimensions of the | ||
| /// texture, otherwise, rendering might be corrupted!** | ||
| Full { | ||
| /// Width of the bound texture view. | ||
| width: u16, | ||
| /// Height of the bound texture view. | ||
| height: u16, | ||
| }, |
There was a problem hiding this comment.
In theory, we could derive that information ourselves using textureDimensions in the shader. However, I faintly remember this being a performance footgun. Hence why I think it's better to require that information to be passed along.
| ); | ||
| out.sample_xy = get_native_image_translate(image_texel1) | ||
| + get_native_image_transform(image_texel0) * pos; | ||
| out.payload = get_native_image_opacity(image_texel1); |
There was a problem hiding this comment.
Bit annoying to abuse the payload like this. 😅 But it avoids having to do another texture sample in the fragment shader, as all the metadata we need is passed directly.
| (paint_and_rect_flag >> EXTERNAL_TEXTURE_SLOT_SHIFT) & 0x3u; | ||
|
|
||
| if external_texture_slot == 0u { | ||
| final_color = paint_alpha * textureSampleLevel(external_texture_0, external_sampler_0, sample_xy, 0.0); |
There was a problem hiding this comment.
I think it's better to leave this as is, and revisit once we properly support mip-chains.
597aa52 to
340a2ca
Compare
340a2ca to
f9e3ed5
Compare
f9e3ed5 to
3df9f0b
Compare
3df9f0b to
5fde82a
Compare
5fde82a to
c54ee4d
Compare
c54ee4d to
bf7f46d
Compare
bf7f46d to
b835478
Compare
b835478 to
fa4fcb7
Compare
fa4fcb7 to
3e7e39c
Compare
| if image.tint.is_some_and(|tint| { | ||
| tint.mode != TintMode::Multiply || tint.color.components[..3] != [1.0; 3] |
There was a problem hiding this comment.
In the future, we cna likely support Multiply tinting as well, by storing the full rgba instead of just the opacity.
…ender#1963) This PR makes the following changes: - It adds a shim in the test suite so that we can run most external texture tests (the once that sample the full texture) with Vello CPU as well. By doing so, Vello CPU can serve as the reference, and we increase the coverage of tests that compare against Vello CPU and GPU. - It fixes the issue where, instead of extending each tap of bilinear/bicubic samples, we only extended the original sample location. This is wrong, and is the reason why a couple of image tests were inconsistent between CPU and GPU. In addition to that, it also ports some changes to the extend logic that we recently applied to Vello CPU, to make the results more numerically stable and the code more comparable to Vello CPU. **This does unfortunately regress performance: Rendering a full-sized image on my Android phone goes from around 73FPS to 66FPS.** See the below videos. Before: https://github.com/user-attachments/assets/0660b553-9076-4a1d-88ec-88328b034501 After: https://github.com/user-attachments/assets/66fd5fcd-f266-4eb1-b739-29bb78ee8783 But the previous approach was just fundamentally wrong. However, the good news is that with linebender#1964, I will introduce a fast path that uses GPU-native bilinear sampling for images that are sampled from the whole external texture, **which will not only undo this slowdown, but in fact make image rendering even faster compared to current main, when using external textures and sampling the whole texture! See that PR for more information.** Since rendering whole images is the most common operation (except for glyph caching, but this is experimental right now, anyway. And nearest-neighbor sampling is less affected than bilinear sampling), in my opinion this is a trade-off worth taking. Using the image atlas will unfortunately stay slower for now, even with linebender#1964, but I think that's something we have to accept for now. Once we revisit glyph caching, we can figure out how to best fix this. We could also improve this in the future by restricting the possible image sampling modes for atlas images (for example, only allowing extend mode `Pad`). But the problem is simply that when we sample from an arbitrary subregion of an atlas, we have to emulate correct extension ourselves, which is much slower than letting the hardware do it, so this should be avoided. --------- Co-authored-by: Laurenz Stampfl <laurenz.stampfl+github@gmail.com>
This PR adds a fast path for images where we want to sample the whole texture. When this is the case, we can make use of WebGL sampler objects to define whether to use NN/bilinear sampling as well as which extend mode to use, without us having to reimplement these operations in software! This is much faster, see this video and compare to the FPS numbers from #1963 instead. We now fully saturate the 90FPS of this device 🎉!
IMG_1338.MOV
The only two downsides are that
image-bicubicfeature to instead allow compiling out all non-native image operations.ImageQuality::Low.Once this (and the preceding PR) have been approved, I will make sure to rerun tests on other devices to ensure we are not creating any regressions by making the shader bigger.