-DFILAMENT_SKIP_SAMPLES=ON with CMake
-Pfilament_skip_samples with gradle
This change also renames CMake options specific to Filament
to avoid clashes with subprojects.
- minor uniform optimizations (saves ~0.1ms)
- fix some missing highp precision qualifers
- make it easier to tweak the blur filter
- make sure samples's radius goes from 0 to 1
- update comments
After the native async functionality landed, there was no way for web
clients to be notified that the decoding has finished. This PR changes
the existing `onDone` callback so that it gets called after all textures
have been decoded. (Previously it was called after downloading rather
than decoding.)
This has the side effect of simplifying the API because clients no
longer need to call a finalize function.
During the SSAO pass, we pack the decoded depth to the GB channels of
the AO texture, which reduces our bandwidth requirements during the
blurring pass, as well as some ALU usage.
It puts SSAO at about 2ms on pixel4/1080p.
Currently it only selects how many samples are used.
Low,medium,high and ultra respectively map to 7,11,16 and 32 samples.
The default is "low", which is sufficient for most mobile applications.
This is achieved by computing a small 2x2 box blur in the AO pass
taking advantage of quad shading, which allows to halve the size
of the bilateral blur kernel.
Reducing the size of the blur kernel has a side effect to kick Adreno
gpu into direct mode, which apparently is much faster here.
Overall we go from 3.4ms to 2.4ms on Pixel4.
The quality is impacted, but not severely. This probably assumes
GPU that have working derivatives.
- got rid of all the precomputed samples for now, it makes things
much more simple, and actually doesn't slow things down. Use the
analytic offsets (spiral) instead.
- use textureLod instead of texelFetch in places where we need to do
how own clamping, turns out it's faster (as measured on pixel4).
- use a interleaved gradient noise instead of a hardcoded blue noise,
the difference isn't big.
Overall we're actually 0.2ms faster out of 3.5ms. It's not immediately
clear why, might just be noise.
* Add support for ccache
To benefit from ccache, first install ccache (for instance with
brew install ccache on macOS). On my machine a clean build takes
~5 minutes the first time. A second clean build with ccache enabled
takes 34 seconds.
* Enable more aggressive caching
This is 466 KB insted of the usual 850 KB. It assumes that users do not
need spec-gloss, non-lit, or a special transparency mode.
(uncompressed SO size on arm64)
We can add a new maven project for this next week.
This doesn't affect quality significantly, but saves 30% of gpu time
on the upsampling, overall it's about 0.1 ms. This brings The bloom
effect below 2ms on Pixel4.
We untonemap/tonemap the first level of blur to reduce the very high
frequencies due to HDR highlights in the image, this produce a softer
image at higher roughness. This can programmatically be turned off,
but that setting is not exposed.
The very first downsample stage costs us a lot because it reads from
a full-res texture. We mitigate this by always doing a blit to
1/4 res. Blits are quite a bit faster, this saves about 1ms on
Pixel4.
math/mathwfd.h forward declares all {mat|vec}{2|3|4}<> classes,
which allows us to remove their respective #include in a lot of
our public headers.
Our math headers are full of templates, so this should help build times
a bit.
Also we want to keep the public headers as minimalist as possible.
To trigger an exception users could either shrink the window or enlarge
the sidebar. Constraining both of these is somewhat tricky so let's
just clamp the viewport width at a low level.
There was actually no need to give special treatment to leading drive
designators since they effectively form the first path segment anyway.
To help prevent regressions, I added a few unit tests in a previous CL.
The motivation for the CL is to remove a dependency on `locale.cpp`,
which can result in shorter build times and reduced binary sizes.