Fix#1165
In order to use transparent views, post-processing had to be disabled
because we were not able to blit post-processed buffers with blending.
This is now fixed by simply reverting to a quad.
A side effect of this, is improved performance when a scaling blit is
needed when MSAA is active.
This also introduced new "bugs", or at least weird behaviors: the
system needs to know when blending is needed, and this has to be
based on heuristics (unless we add a new api). Currently the heuristic
is that the COLOR buffer is not discarded and the view is cleared with
an alpha value (or not cleared).
1) When scaling is enabled, we're using a blit at the end of the frame,
however, if MSAA is also enabled, this blit is not allowed, so we
need an explicit MSAA resolve.
2) When rendering directly into the default render target with MSAA
enabled, we need an explicit resolve ONLY IF there is a format
conversion, because those are also not allowed.
Before, we were doing this resolve regardless.
3) The test for "rendering directly into the default target" had
become wrong because, the Render target can now be a user texture
(which is assumed to have the proper MSAA-ness and a depth buffer),
so no intermediate buffer should be needed for those. We now
check that render target is actually the default render target
We used to have this functionality, but it we removed it because we
didn't have a use for it and it was costly.
This new implementation is only internal, and would (will?) be useful
if we wanted to generate a special buffer for SSAO for instance (instead
of just depth). This implementation, overrides the material at the time
the commands are executed, which is much cheaper than the previous
implementation.
For the street light glb from the Khronos suite, this reduces texture
loading time from 290 ms to 110 ms.
Stay tuned for another feature: notification callbacks.
Related to #1876.
- we add an 'intensity' parameter that allows to control the strength
of the AO effect. This is useful for aesthetic reasons.
- the default intensity is now 2x that of before this change, which
makes the intensity parameter match AO papers this implementation
is inspired of.
- bias is now z-dependent, which reduces some self shadowing wrt the
depth.
This is needed because filament doesn't make strong assumptions about
the units used (e.g. meters), so we can't hardcode a distance in the
shader. But also, this is a cheap change.
ssaogen now produces better formatted tables, directly usable in
shader code.
SAO shader now squares the sample radius outside of `tapLocation`, also
added some comments.
Previously, Streams had two modes (native and texid), this adds a third
mode called "acquired", which allows for copy-free synchronized external
textures in OpenGL and paves the way for Vulkan.
The native mode is now deprecated but texid mode needs to stay around
until all clients can be upgraded.
In an early prototype, this functionality was added directly into
Texture but required quite a bit of additional state tracking, so the
Stream API seemed like a better fit.
This API is probably not necessary for Metal due to Metal's shared
ownership semantics.
This has been tested with a new Android sample that will be added in a
subsequent PR.
RenderPass's API now doesn't need to know about FScene, instead we
pass the geometry info as parameters.
This is one step towards keeping the "scene" concept more away from
rendering.
This fixes an oversight in #1641 for scenes that have only 1 primitive
when the depth prepass is enabled. In such cases, the same prim is
rendered twice in a row: once for depth and once for color.
We had enhanced the main render loop to override the primitive's
backface culling state when the material instance changed, but that was
not sufficient since the same material instance can be used for depth
and color passes.
Fixes#1872.
New internal API to add custom commands (lambda) before or after each
pass. These commands can be added at any point and will be sorted
properly.
e.g.:
appendCustomCommand(Pass::COLOR, CustomCommand::PROLOG, 0, [](){
// custom code
});
Also split appendSortedCommands() in two functions to make it easier
to insert custom commands.
On Debug builds, HeapAllocatorArena needs LockingPolicy::Mutex because
it uses a TrackingPolicy, which needs to be synchronized.
On Release builds, HeapAllocatorArena doesn't need a LockingPolicy
because HeapAllocator is intrinsically synchronized as it relies on
heap allocations (i.e.: malloc/free)
- this is to allow the use of C++17 std::invoke
- don't std::move() each argument of the tuple<> instead
std::move() the tuple itself and perfectly-forward each argument.
- handle return values
- handle synchronous commands
The "clampNoV" shader function is ignoring the passed dot product argument and recomputing the dot product itself. This will result in the wrong NoV value being used for clear coat IBL computations.
TrackingPolicy::Debug didn't store the base pointer of the Area, and
instead relied on the first allocation to discover it, however, because
of alignment, the first allocation may not match the base pointer.
Because of that there could be an overflow in onRewind(), i.e. we
could rewind to a pointer before the (wrongly computed) base. This
overflow caused the debug memset to go awry and stomped on memory.
This is fixed by passing the base pointer to the constructor of the
TrackingPolicy. This base pointer could be nullptr with certain
allocators, but in that case, onReset/onRewind should never be called;
and this is enforced at compile time.
Also fixed a (luckily) harmless buffer overflow when preparing the
dynamic lights, if the number of lights wasn't a multiple of 4. This
was harmless because we use a linear allocator, so overflows are not
really overflows.
Clients who do not yet have this fix can usually work around this issue
by calling the non-array overload of `MaterialInstance::setParameter`
when the array size is 1.
This was caught by ASAN.
Our algorithm header has many one-liners that compute the "next power of
two length / 2" but they all have the caveat that if the input is
already POT, then the "/ 2" part does not occur.
Usually we deal with this by testing the difference against zero.
However in `partition_point` we were skipping the test, thus causing a
potential out-of-bounds access.
I fixed `partition_point` and added a few more tests for non-POT cases.
This doesn't change the user facing API because it already used a
callback.
With this change we implement readPixels() with a PBO, which won't
stall the GPU. Then we periodically check a fence to see when the
command is completed, at which point we map the PBO and copy the data
to the user buffer.
This is useful for unit tests. The content of the swapchain can be
read to main memory with Renderer::readPixels().
This PR doesn't implement this new feature in the following backends and
platforms (TODO):
VulkanDriver.cpp
MetalDriver.mm
PlatformWGL.cpp
PlatformWebGL.cpp
PlatformCocoaTouchGL.mm