Compare commits

...

18 Commits

Author SHA1 Message Date
Run Yu
d105dc01b1 This is a test for pushing branch to github 2025-10-17 16:16:15 -04:00
Filament Bot
40d533ada1 [automated] Updating /docs due to commit b6df2b9
Full commit hash is b6df2b9b35

DOCS_ALLOW_DIRECT_EDITS
2025-10-17 06:20:51 +00:00
Powei Feng
b6df2b9b35 renderdiff: add transmission and ssao tests (#9232)
- Add transmission and ssao tests
 - Small mods to README.md
 - Add transmission models to gltf list
 - Change focal length to zoom in for more details

RDIFF_BRANCH=pf/renderdiff-add-transmission
2025-10-16 23:17:07 -07:00
Powei Feng
c379b31267 webgpu: fix picking (#9318)
Context:
 - Picking has switched to use integer textures and this texture
   is meant to have only RG components (#8917).
 - WebGPU only supports 4 component textures, and so we need to
   to a blit-based transform to turn 2 components into 4.
 - WebGPUBlitter supports only float textures

In this commit, we add support for integer textures to
WebGPUBlitter. In doing so, we refactor parts of the shader
construction code in the blitter class and introduce a couple
helper functions for texture formats.

This fixes picking (verifiable with gltf_viewer)

Co-authored-by: Ben Doherty <bendoherty@google.com>
2025-10-17 04:44:12 +00:00
Powei Feng
b03909c07a github: add hash id to commit message action (#9334)
We add the commit hash as output of get-commit-msg
2025-10-17 04:05:52 +00:00
Doris Wu
712c03b17b buffer update opt: Add uniform batching option for Material Instances (#9324) 2025-10-17 00:22:00 +00:00
Mathias Agopian
9dae0748d5 update perfetto sdk to 51.2 and add update script (#9335) 2025-10-16 16:00:25 -07:00
Powei Feng
fa02a7fa3b backend: add two command line arguments to backend test (#9329)
`--headless_only`
Headless-only mode indicates that we will only use headless
swapchains.  This is particularly useful if we want to test in
a CI (continuous integration) environment.

`--ci`
Having CONTINUOUS_INTEGRATION as an OS will allow us to have
CI-only temporary exceptions. Though it's expected that some
exceptions will be unfixable and will remain as workarounds.
2025-10-16 18:54:38 +00:00
Powei Feng
a068d3df79 github: fix get-commit-msg (#9330)
If we assign the message (which might contain quotes) to an
environment variable and then echo it, this should prevent the
problem of having a double quote (").

Also fix a problem in docs script for checking TAGs in commits.
2025-10-16 11:33:12 -07:00
Powei Feng
f64c087ffe backend: fix mapped buffer test for vk (#9328)
The descriptor set needs to be rebound when rendering two
separate render passes.
2025-10-16 06:26:51 +00:00
Doris Wu
9561137d53 buffer update opt: Change to use feature flag instead of engine flag (#9327)
* Revert "buffer update opt: Add a flag to guard the feature (#9322)"

This reverts commit 49c4a5d62c.

* Update feature flags

* feedback

* Update naming
2025-10-16 07:13:30 +08:00
Powei Feng
d0efda21ad webgpu: fix deadlock for readPixels (#9319)
To enable the readPixels callback to be called, we need to make
move to 'AllowSpontaneous' mode in order to not have to call
Instance::ProcessEvents() to force process the callback.
2025-10-15 11:11:05 -07:00
Mathias Agopian
a5b047f93d fix matc's -l option (#9321)
MaterialParser did the feature level check (-l) too early, before the
material was parsed, so -l0 would always fail.
2025-10-15 09:42:19 -07:00
Doris Wu
06c4ed4e6b fix: material sandbox scene crashes during shutdown (#9325) 2025-10-15 15:34:50 +00:00
Doris Wu
49c4a5d62c buffer update opt: Add a flag to guard the feature (#9322) 2025-10-15 23:14:53 +08:00
Doris Wu
c757cc3629 buffer update opt: Introduce BufferAllocator and its tests (#9304)
* Introduce AllocationStrategy and its tests

* feedback

* feedback

* Add stress test

* Fix the build

* Minor fix

* Rename the class

* Rename

* Make const
2025-10-15 06:25:42 +00:00
rafadevai
dbdf8f672b GL: Check if GL_EXT_texture_filter_anisotropic in supported in GLES (#9317)
In GLES the anisotropic filtering featuring is always
disabled even in the cases where is supported by the hardware
because the check for the extension is missing.
2025-10-14 23:03:25 -07:00
Filament Bot
54bd888374 [automated] Updating /docs due to commit b7dea28
Full commit hash is b7dea28cc5

DOCS_ALLOW_DIRECT_EDITS
2025-10-14 21:28:53 +00:00
46 changed files with 14643 additions and 2322 deletions

View File

@@ -1,43 +1,53 @@
# This action retrieves the latest commit message from a push or pull_request event
# and makes it available as an output variable named 'msg'.
name: 'Get commit message'
outputs:
msg:
value: ${{ steps.action_output.outputs.msg }}
hash:
value: ${{ steps.action_output.outputs.hash }}
runs:
using: "composite"
steps:
- name: Find commit message (on push)
if: github.event_name == 'push'
shell: bash
env:
COMMIT_MESSAGE: ${{ github.event.head_commit.message }}
run: |
AUTHOR_NAME="${{ github.event.head_commit.author.name }}"
AUTHOR_EMAIL="${{ github.event.head_commit.author.email }}"
TSTAMP="${{ github.event.head_commit.timestamp }}"
echo "commit ${{ github.event.head_commit.id }}" >> /tmp/commit_msg.txt
HASH="${{ github.event.head_commit.id }}"
echo "commit $HASH" >> /tmp/commit_msg.txt
echo "Author: ${AUTHOR_NAME}<${AUTHOR_EMAIL}>" >> /tmp/commit_msg.txt
echo "Date: ${TSTAMP}" >> /tmp/commit_msg.txt
echo "" >> /tmp/commit_msg.txt
echo "${{ github.event.head_commit.message }}" >> /tmp/commit_msg.txt
echo "$COMMIT_MESSAGE" >> /tmp/commit_msg.txt
echo "$HASH" > /tmp/commit_hash.txt
- name: Find commit message (PR)
shell: bash
id: checkout_code
if: github.event_name == 'pull_request'
run: |
echo "+++++ head commit message +++++"
echo "$(git log -1 --no-merges)"
echo "+++++++++++++++++++++++++++++++"
echo "hash=$(git rev-parse HEAD)" >> "$GITHUB_OUTPUT"
git checkout ${{ github.event.pull_request.head.sha }}
echo "$(git log -1 --no-merges)" >> /tmp/commit_msg.txt
BEFORE_HASH=$(git rev-parse HEAD)
echo "hash=$BEFORE_HASH" >> "$GITHUB_OUTPUT"
# Next we will checkout the actual head (not the merge commits) of the PR
AFTER_HASH="${{ github.event.pull_request.head.sha }}"
git checkout $AFTER_HASH
COMMIT_MESSAGE=$(git log -1 --no-merges)
echo "$COMMIT_MESSAGE" > /tmp/commit_msg.txt
echo "$AFTER_HASH" > /tmp/commit_hash.txt
- shell: bash
id: action_output
run: |
# Get the commit message
DELIMITER="EOF_FILE_CONTENT_$(date +%s)" # Using timestamp to make it more unique
echo "msg<<$DELIMITER" >> "$GITHUB_OUTPUT"
cat /tmp/commit_msg.txt >> "$GITHUB_OUTPUT"
echo "$DELIMITER" >> "$GITHUB_OUTPUT"
echo "----- got commit message ---"
cat /tmp/commit_msg.txt
echo "----------------------------"
# Get the commit hash
echo "hash=$(cat /tmp/commit_hash.txt)" >> "$GITHUB_OUTPUT"
- name: Cleanup Find commit message (PR)
shell: bash
if: github.event_name == 'pull_request'

View File

@@ -23,7 +23,7 @@ jobs:
GH_TOKEN: ${{ secrets.FILAMENTBOT_TOKEN }}
run: |
GOLDEN_BRANCH=$(echo "${{ steps.get_commit_msg.outputs.msg }}" | python3 test/renderdiff/src/commit_msg.py)
COMMIT_HASH=$(echo "${{ steps.get_commit_msg.outputs.msg }}" | head -n 1 | sed "s/commit //g")
COMMIT_HASH="${{ steps.get_commit_msg.outputs.hash }}"
if [[ "${GOLDEN_BRANCH}" != "main" ]]; then
git config --global user.email "filament.bot@gmail.com"
git config --global user.name "Filament Bot"
@@ -51,7 +51,7 @@ jobs:
GH_TOKEN: ${{ secrets.FILAMENTBOT_TOKEN }}
run: |
bash docs_src/build/install_mdbook.sh && source ~/.bashrc
COMMIT_HASH=$(echo "${{ steps.get_commit_msg.outputs.msg }}" | head -n 1 | sed "s/commit //g")
COMMIT_HASH="${{ steps.get_commit_msg.outputs.hash }}"
git config --global user.email "filament.bot@gmail.com"
git config --global user.name "Filament Bot"
git config --global credential.helper cache

View File

@@ -108,8 +108,7 @@ jobs:
uses: ./.github/actions/get-commit-msg
- name: Check for manual edits to /docs
run: |
COMMIT_ID=$(echo "${{ steps.get_commit_msg.outputs.msg }}" | head -n 1 | sed "s/commit //g")
bash docs_src/build/presubmit_check.sh ${COMMIT_ID}
bash docs_src/build/presubmit_check.sh ${{ steps.get_commit_msg.outputs.hash }}
test-renderdiff:
name: test-renderdiff
@@ -121,7 +120,7 @@ jobs:
- uses: ./.github/actions/mac-prereq
- uses: ./.github/actions/get-gltf-assets
- uses: ./.github/actions/get-mesa
- uses: ./.github/actions/get-vulkan-sdk
- uses: ./.github/actions/get-vulkan-sdk
- id: get_commit_msg
uses: ./.github/actions/get-commit-msg
- name: Prerequisites
@@ -132,12 +131,14 @@ jobs:
shell: bash
- name: Render and compare
id: render_compare
env:
COMMIT_MESSAGE: ${{ steps.get_commit_msg.outputs.msg }}
run: |
ls ./gltf/Models
TEST_DIR=test/renderdiff
source ${TEST_DIR}/src/preamble.sh
start_
GOLDEN_BRANCH=$(echo "${{ steps.get_commit_msg.outputs.msg }}" | python3 ${TEST_DIR}/src/commit_msg.py)
GOLDEN_BRANCH=$(echo "${COMMIT_MESSAGE}" | python3 ${TEST_DIR}/src/commit_msg.py)
bash ${TEST_DIR}/generate.sh && \
python3 ${TEST_DIR}/src/golden_manager.py \
--branch=${GOLDEN_BRANCH} \

View File

@@ -181,7 +181,7 @@ important for <code>matc</code> (material compiler).</p>
}
dependencies {
implementation 'com.google.android.filament:filament-android:1.65.4'
implementation 'com.google.android.filament:filament-android:1.66.0'
}
</code></pre>
<p>Here are all the libraries available in the group <code>com.google.android.filament</code>:</p>
@@ -196,7 +196,7 @@ dependencies {
</div>
<h3 id="ios"><a class="header" href="#ios">iOS</a></h3>
<p>iOS projects can use CocoaPods to install the latest release:</p>
<pre><code class="language-shell">pod 'Filament', '~&gt; 1.65.4'
<pre><code class="language-shell">pod 'Filament', '~&gt; 1.66.0'
</code></pre>
<h2 id="documentation"><a class="header" href="#documentation">Documentation</a></h2>
<ul>

View File

@@ -159,16 +159,16 @@
<div id="content" class="content">
<main>
<h1 id="rendering-difference-test"><a class="header" href="#rendering-difference-test">Rendering Difference Test</a></h1>
<p>We created a few scripts to run <code>gltf_viewer</code> and produce headless renderings.</p>
<p>This tool is a collections of scripts to run <code>gltf_viewer</code> and produce headless renderings.</p>
<p>This is mainly useful for continuous integration where GPUs are generally not available on cloud
machines. To perform software rasterization, these scripts are centered around <a href="https://docs.mesa3d.org">Mesa</a>'s
software rasterizers, but nothing bars us from using another rasterizer like <a href="https://github.com/google/swiftshader">SwiftShader</a>.
Additionally, we should be able to use GPUs where available (though this is more of a future
work).</p>
<p>The script <code>render.py</code> contains the core logic for taking input parameters (such as the test
description file) and then running gltf_viewer to produce the renderings.</p>
description file) and then running <code>gltf_viewer</code> to produce the renderings.</p>
<p>In the <code>test</code> directory is a list of test descriptions that are specified in json. Please see
<code>sample.json</code> to parse the structure.</p>
<code>sample.json</code> to glean the structure.</p>
<h2 id="setting-up-python"><a class="header" href="#setting-up-python">Setting up python</a></h2>
<p>The <code>renderdiff</code> project uses <code>python</code> extensively. To install the dependencies for producing
renderings, do the following step</p>

File diff suppressed because one or more lines are too long

File diff suppressed because one or more lines are too long

View File

@@ -77,11 +77,11 @@ def check_has_source_edits(commit_hash, printing=True):
# Returns true in a given TAG is found in the commit msg
def commit_msg_has_tag(commit_hash, tag, printing=True):
res, ret = execute(f'git log --pretty=%B {commit_hash}', cwd=ROOT_DIR)
res, ret = execute(f'git log -n1 --pretty=%B {commit_hash}', cwd=ROOT_DIR)
for l in ret.split('\n'):
if tag == l.strip():
if printing:
print(f'Found tag={tag} in commit message')
print(f'Found tag={tag} in commit={commit_hash}')
return True
return False

View File

@@ -117,6 +117,7 @@ set(SRCS
src/components/LightManager.cpp
src/components/RenderableManager.cpp
src/components/TransformManager.cpp
src/details/BufferAllocator.cpp
src/details/BufferObject.cpp
src/details/Camera.cpp
src/details/ColorGrading.cpp
@@ -196,6 +197,7 @@ set(PRIVATE_HDRS
src/components/LightManager.h
src/components/RenderableManager.h
src/components/TransformManager.h
src/details/BufferAllocator.h
src/details/BufferObject.h
src/details/Camera.h
src/details/ColorGrading.h

View File

@@ -760,6 +760,7 @@ void OpenGLContext::initExtensionsGLES(Extensions* ext, GLint major, GLint minor
ext->EXT_texture_compression_rgtc = exts.has("GL_EXT_texture_compression_rgtc"sv);
ext->EXT_texture_compression_bptc = exts.has("GL_EXT_texture_compression_bptc"sv);
ext->EXT_texture_cube_map_array = exts.has("GL_EXT_texture_cube_map_array"sv) || exts.has("GL_OES_texture_cube_map_array"sv);
ext->EXT_texture_filter_anisotropic = exts.has("GL_EXT_texture_filter_anisotropic"sv);
ext->GOOGLE_cpp_style_line_directive = exts.has("GL_GOOGLE_cpp_style_line_directive"sv);
ext->KHR_debug = exts.has("GL_KHR_debug"sv);
ext->KHR_parallel_shader_compile = exts.has("GL_KHR_parallel_shader_compile"sv);

View File

@@ -84,128 +84,12 @@ constexpr uint32_t UNIFORM_BINDING_INDEX{ 1 };
constexpr uint32_t SAMPLER_BINDING_INDEX{ 2 };
constexpr std::string_view VERTEX_SHADER_ENTRY_POINT{ "vertexShaderMain" };
constexpr std::string_view FRAGMENT_SHADER_ENTRY_POINT{ "fragmentShaderMain" };
// note that the placeholders below must start and end with this prefix and suffix and should not
// otherwise be present in the template:
constexpr std::string_view PLACEHOLDER_PREFIX{ "{{" };
constexpr std::string_view PLACEHOLDER_SUFFIX{ "}}" };
// texture_2d<f32> or texture_multisampled_2d<f32> or texture_3d<f32> or
// texture_depth_2d or texture_depth_multisampled_2d
constexpr std::string_view TEXTURE_TYPE_PLACEHOLDER{ "TEXTURE_TYPE" };
// texture_2d<f32> or texture_depth_2d
[[maybe_unused]] constexpr std::string_view TEXTURE_2D_TYPE_PLACEHOLDER { "TEXTURE_2D_TYPE" };
// "" for no sampler, otherwise "@group(0) @binding(2) var sourceSampler: sampler;"
constexpr std::string_view SAMPLER_DECLARATION_PLACEHOLDER{ "SAMPLER_DECLARATION" };
constexpr std::string_view FRAGMENT_SHADER_SNIPPET_PLACEHOLDER{ "FRAGMENT_SHADER_SNIPPET" };
// @location(0) vec4<f32> or @builtin(frag_depth) f32
constexpr std::string_view FRAGMENT_RETURN_ATTRIBUTE_AND_TYPE_PLACEHOLDER{
"FRAGMENT_RETURN_ATTRIBUTE_AND_TYPE"
};
constexpr std::string_view SHADER_SOURCE_TEMPLATE{ R"(
struct BlitFragmentShaderArgs {
depthPlane: u32,
scale: vec2<f32>,
sourceOffset: vec2<u32>,
destinationOffset: vec2<u32>,
};
@group(0) @binding(0) var sourceTexture: {{TEXTURE_TYPE}};
@group(0) @binding(1) var<uniform> fragmentShaderArgs: BlitFragmentShaderArgs;
{{SAMPLER_DECLARATION}}
fn getUnnormalizedSourceTextureCoordinates(position: vec2<f32>) -> vec2<f32> {
// These coordinates match the Vulkan vkCmdBlitImage spec:
// https://www.khronos.org/registry/vulkan/specs/1.2-extensions/man/html/vkCmdBlitImage.html
let uvOffset: vec2<f32> = position - vec2<f32>(f32(fragmentShaderArgs.destinationOffset.x),f32(fragmentShaderArgs.destinationOffset.y));
let uvScaled: vec2<f32> = uvOffset * fragmentShaderArgs.scale;
return uvScaled + vec2<f32>(f32(fragmentShaderArgs.sourceOffset.x), f32(fragmentShaderArgs.sourceOffset.y));
}
fn normalize2dSourceTextureCoordinates(
unnormalizedSourceTextureCoordinates: vec2<f32>,
sourceDimensions: vec2<u32>) -> vec2<f32> {
return unnormalizedSourceTextureCoordinates /
vec2<f32>(f32(sourceDimensions.x), f32(sourceDimensions.y));
}
fn getNormalized2dSourceTextureCoordinates(
position: vec2<f32>,
sourceTexture: {{TEXTURE_2D_TYPE}}) -> vec2<f32> {
let sourceDimensions: vec2<u32> = textureDimensions(sourceTexture);
return normalize2dSourceTextureCoordinates(
getUnnormalizedSourceTextureCoordinates(position),
sourceDimensions
);
}
fn normalize3dSourceTextureCoordinates(
unnormalizedSourceTextureCoordinates: vec2<f32>,
sourceDimensions: vec3<u32>) -> vec3<f32> {
let uvNormalized: vec2<f32> = normalize2dSourceTextureCoordinates(
unnormalizedSourceTextureCoordinates,
sourceDimensions.xy
);
return vec3<f32>(
uvNormalized,
(f32(fragmentShaderArgs.depthPlane) + 0.5) / f32(sourceDimensions.z)
);
}
fn getNormalized3dSourceTextureCoordinates(
position: vec2<f32>,
sourceTexture: texture_3d<f32>) -> vec3<f32> {
let sourceDimensions: vec3<u32> = textureDimensions(sourceTexture);
return normalize3dSourceTextureCoordinates(
getUnnormalizedSourceTextureCoordinates(position),
sourceDimensions
);
}
@vertex
fn vertexShaderMain(@builtin(vertex_index) vertexIndex: u32) -> @builtin(position) vec4<f32> {
let fullScreenTriangleVertices = array<vec2<f32>, 3>(
vec2<f32>(-1.0, -1.0),
vec2<f32>( 3.0, -1.0),
vec2<f32>(-1.0, 3.0)
);
return vec4<f32>(fullScreenTriangleVertices[vertexIndex].xy, 0.0, 1.0);
}
{{FRAGMENT_SHADER_SNIPPET}}
)" };
constexpr std::string_view FRAGMENT_SHADER_SNIPPET_MSAA_INPUT_TEMPLATE{ R"(
@fragment
fn fragmentShaderMain(@builtin(position) position: vec4<f32>) -> {{FRAGMENT_RETURN_ATTRIBUTE_AND_TYPE}} {
var color: vec4<f32> = vec4<f32>(0.0, 0.0, 0.0, 0.0);
let numberOfSamples: u32 = textureNumSamples(sourceTexture);
let coordinatesF: vec2<f32> = getUnnormalizedSourceTextureCoordinates(position.xy);
let coordinates: vec2<u32> = vec2<u32>(u32(coordinatesF.x), u32(coordinatesF.y));
for (var sampleIndex: u32 = 0; sampleIndex < numberOfSamples; sampleIndex++) {
color += textureLoad(sourceTexture, coordinates, sampleIndex);
}
color /= f32(numberOfSamples);
return color;
}
)" };
constexpr std::string_view FRAGMENT_SHADER_SNIPPET_3D_INPUT_TEMPLATE{ R"(
@fragment
fn fragmentShaderMain(@builtin(position) position: vec4<f32>) -> {{FRAGMENT_RETURN_ATTRIBUTE_AND_TYPE}} {
let coordinates: vec3<f32> = getNormalized3dSourceTextureCoordinates(position.xy, sourceTexture);
return textureSample(sourceTexture, sourceSampler, coordinates);
}
)" };
constexpr std::string_view FRAGMENT_SHADER_SNIPPET_2D_INPUT_TEMPLATE{ R"(
@fragment
fn fragmentShaderMain(@builtin(position) position: vec4<f32>) -> {{FRAGMENT_RETURN_ATTRIBUTE_AND_TYPE}} {
let coordinates: vec2<f32> = getNormalized2dSourceTextureCoordinates(position.xy, sourceTexture);
return textureSample(sourceTexture, sourceSampler, coordinates);
}
)" };
} // namespace
// Include the wgsl template sources
#include "WebGPUBlitter_wgsl.inc"
// It can perform a direct memory copy if the formats are compatible and no scaling or format
// conversion is needed. Otherwise, it uses a render pass with a custom shader to perform
// the blit. This allows for scaling, format conversion, and resolving multisampled textures.
@@ -395,11 +279,25 @@ void WebGPUBlitter::blit(wgpu::Queue const& queue, wgpu::CommandEncoder const& c
const bool multisampledSource{ args.source.texture.GetSampleCount() > 1 };
const bool depthSource{ hasDepth(args.source.texture.GetFormat()) };
const bool depthDestination{ hasDepth(args.destination.texture.GetFormat()) };
wgpu::TextureSampleType srcSampleType = {};
if (depthSource) {
srcSampleType = wgpu::TextureSampleType::Depth;
} else if (multisampledSource) {
srcSampleType = wgpu::TextureSampleType::UnfilterableFloat;
} else if (isIntFormat(args.source.texture.GetFormat())) {
srcSampleType = wgpu::TextureSampleType::Sint;
} else if (isUIntFormat(args.source.texture.GetFormat())) {
srcSampleType = wgpu::TextureSampleType::Uint;
} else {
srcSampleType = wgpu::TextureSampleType::Float;
}
const PipelineLayoutKey pipelineLayoutKey{
.sourceDimension = sourceDimension,
.sourceSampleType = srcSampleType,
.filterType = args.filter,
.multisampledSource = multisampledSource,
.depthSource = depthSource,
};
const wgpu::BindGroupDescriptor textureBindGroupDescriptor{
.label = "blit_texture_bind_group",
@@ -454,10 +352,11 @@ void WebGPUBlitter::blit(wgpu::Queue const& queue, wgpu::CommandEncoder const& c
<< "Failed to create render pass encoder for blit.";
const RenderPipelineKey renderPipelineKey{
.sourceDimension = sourceDimension,
.sourceTextureFormat = args.source.texture.GetFormat(),
.sourceSampleType = srcSampleType,
.destinationTextureFormat = args.destination.texture.GetFormat(),
.sourceSampleCount = static_cast<uint8_t>(args.source.texture.GetSampleCount()),
.filterType = args.filter,
.depthSource = depthSource,
};
renderPassEncoder.SetPipeline(getOrCreateRenderPipeline(renderPipelineKey));
renderPassEncoder.SetBindGroup(TEXTURE_BIND_GROUP_INDEX, textureBindGroup);
@@ -557,9 +456,9 @@ wgpu::RenderPipeline WebGPUBlitter::createRenderPipeline(RenderPipelineKey const
};
const ShaderModuleKey shaderModuleKey{
.sourceDimension = key.sourceDimension,
.sourceTextureFormat = key.sourceTextureFormat,
.destinationTextureFormat = key.destinationTextureFormat,
.multisampledSource = key.sourceSampleCount > 1,
.depthSource = key.depthSource,
.depthDestination = hasDepth(key.destinationTextureFormat),
};
wgpu::ShaderModule const& shaderModule{ getOrCreateShaderModule(shaderModuleKey) };
const wgpu::FragmentState fragmentState{
@@ -572,9 +471,9 @@ wgpu::RenderPipeline WebGPUBlitter::createRenderPipeline(RenderPipelineKey const
};
const PipelineLayoutKey pipelineLayoutKey{
.sourceDimension = key.sourceDimension,
.sourceSampleType = key.sourceSampleType,
.filterType = key.filterType,
.multisampledSource = key.sourceSampleCount > 1,
.depthSource = key.depthSource,
};
const wgpu::RenderPipelineDescriptor pipelineDescriptor{
.label = "render_pass_blit_pipeline",
@@ -663,13 +562,7 @@ wgpu::BindGroupLayout WebGPUBlitter::createTextureBindGroupLayout(PipelineLayout
.binding = TEXTURE_BINDING_INDEX,
.visibility = wgpu::ShaderStage::Fragment,
.texture = {
.sampleType = key.depthSource
? wgpu::TextureSampleType::Depth
: (key.multisampledSource
? wgpu::TextureSampleType::UnfilterableFloat
: wgpu::TextureSampleType::Float), // only F32 scalar sample
// type supported for now
// (aside from depth)
.sampleType = key.sourceSampleType,
.viewDimension = key.sourceDimension,
.multisampled = key.multisampledSource,
},
@@ -695,7 +588,7 @@ wgpu::BindGroupLayout WebGPUBlitter::createTextureBindGroupLayout(PipelineLayout
};
const wgpu::BindGroupLayoutDescriptor textureBindGroupLayoutDescriptor{
.label = "render_pass_blit_texture_bind_group_layout",
// TODO, doesnt make any sense but gets rid of the error. Are the entries 0 based?
// No need for the last entry (sampler) if source is multisampled.
.entryCount = key.multisampledSource ? (MAX_TEXTURE_BIND_GROUP_ENTRY_SIZE - 1)
: (MAX_TEXTURE_BIND_GROUP_ENTRY_SIZE),
.entries = bindGroupLayoutEntries,
@@ -720,63 +613,109 @@ wgpu::ShaderModule const& WebGPUBlitter::getOrCreateShaderModule(ShaderModuleKey
// The shader source is generated from a template, with placeholders filled in based on the
// blit configuration (e.g., texture type, sample count).
wgpu::ShaderModule WebGPUBlitter::createShaderModule(ShaderModuleKey const& key) {
std::string_view textureType;
if (key.depthSource) {
if (key.multisampledSource) {
textureType = "texture_depth_multisampled_2d";
} else {
textureType = "texture_depth_2d";
}
bool const isDepthSrc = hasDepth(key.sourceTextureFormat);
bool const isDepthDst = hasDepth(key.destinationTextureFormat);
std::string_view vecDim;
if (key.sourceDimension == wgpu::TextureViewDimension::e3D) {
vecDim = "vec3";
} else {
if (key.multisampledSource) {
textureType = "texture_multisampled_2d<f32>";
} else {
if (key.sourceDimension == wgpu::TextureViewDimension::e3D) {
textureType = "texture_3d<f32>";
} else {
textureType = "texture_2d<f32>";
}
}
vecDim = "vec2";
}
const std::string_view texture2dType{ key.depthSource ? "texture_depth_2d"
: "texture_2d<f32>" };
// we don't declare a sampler in the shader or pipeline for the multisampled case
const std::string_view samplerDeclaration{
key.multisampledSource ? "" : "@group(0) @binding(2) var sourceSampler: sampler;"
};
const std::string_view fragmentReturnAttributeAndType{
key.depthDestination ? "@builtin(frag_depth) f32" : "@location(0) vec4<f32>"
};
const std::unordered_map<std::string_view, std::string_view>
valueByPlaceholderNameForFragmentSnippet{
{ FRAGMENT_RETURN_ATTRIBUTE_AND_TYPE_PLACEHOLDER, fragmentReturnAttributeAndType },
};
std::string fragmentShaderSnippet;
if (key.multisampledSource) {
fragmentShaderSnippet = webgpuutils::processPlaceholderTemplate(
FRAGMENT_SHADER_SNIPPET_MSAA_INPUT_TEMPLATE, PLACEHOLDER_PREFIX, PLACEHOLDER_SUFFIX,
valueByPlaceholderNameForFragmentSnippet);
std::string_view srcPrimType;
std::string_view sampleImpl;
if (isIntFormat(key.sourceTextureFormat)) {
srcPrimType = "<i32>";
sampleImpl = INT_TEXTURE_SAMPLE_IMPL_TEMPLATE;
} else if (isUIntFormat(key.sourceTextureFormat)) {
srcPrimType = "<u32>";
sampleImpl = INT_TEXTURE_SAMPLE_IMPL_TEMPLATE;
} else {
if (key.sourceDimension == wgpu::TextureViewDimension::e3D) {
fragmentShaderSnippet = webgpuutils::processPlaceholderTemplate(
FRAGMENT_SHADER_SNIPPET_3D_INPUT_TEMPLATE, PLACEHOLDER_PREFIX,
PLACEHOLDER_SUFFIX, valueByPlaceholderNameForFragmentSnippet);
} else {
fragmentShaderSnippet = webgpuutils::processPlaceholderTemplate(
FRAGMENT_SHADER_SNIPPET_2D_INPUT_TEMPLATE, PLACEHOLDER_PREFIX,
PLACEHOLDER_SUFFIX, valueByPlaceholderNameForFragmentSnippet);
}
srcPrimType = "<f32>";
sampleImpl = FLOAT_TEXTURE_SAMPLE_IMPL_TEMPLATE;
}
const std::unordered_map<std::string_view, std::string_view> valueByPlaceholderNameForModule{
{ TEXTURE_TYPE_PLACEHOLDER, textureType },
{ TEXTURE_2D_TYPE_PLACEHOLDER, texture2dType },
{ SAMPLER_DECLARATION_PLACEHOLDER, samplerDeclaration },
{ FRAGMENT_SHADER_SNIPPET_PLACEHOLDER, fragmentShaderSnippet },
std::string_view fragmentTemplate;
std::string_view srcTextureType;
if (isDepthSrc && key.multisampledSource) {
srcTextureType = "texture_depth_multisampled_2d";
fragmentTemplate = FRAGMENT_SHADER_SNIPPET_2D_INPUT_TEMPLATE;
} else if (isDepthSrc && !key.multisampledSource) {
srcTextureType = "texture_depth_2d";
fragmentTemplate = FRAGMENT_SHADER_SNIPPET_2D_INPUT_TEMPLATE;
} else if (key.multisampledSource) {
srcTextureType = "texture_multisampled_2d{{SRC_PRIM_TYPE}}";
fragmentTemplate = FRAGMENT_SHADER_SNIPPET_MSAA_INPUT_TEMPLATE;
} else if (key.sourceDimension == wgpu::TextureViewDimension::e3D) {
srcTextureType = "texture_3d{{SRC_PRIM_TYPE}}";
fragmentTemplate = FRAGMENT_SHADER_SNIPPET_3D_INPUT_TEMPLATE;
} else {
srcTextureType = "texture_2d{{SRC_PRIM_TYPE}}";
fragmentTemplate = FRAGMENT_SHADER_SNIPPET_2D_INPUT_TEMPLATE;
}
std::string_view dstPrimType;
if (isIntFormat(key.destinationTextureFormat)) {
dstPrimType = "<i32>";
} else if (isUIntFormat(key.destinationTextureFormat)) {
dstPrimType = "<u32>";
} else {
dstPrimType = "<f32>";
}
std::string_view retAttribute;
std::string_view retType;
if (isDepthDst) {
retAttribute = "@builtin(frag_depth)";
retType = "f32";
} else {
retAttribute = "@location(0)";
retType = "vec4{{DST_PRIM_TYPE}}";
}
using ReplacementMap = std::unordered_map<std::string_view, std::string_view>;
auto const replace = [&](std::string_view src, ReplacementMap const& replacements) {
return webgpuutils::processPlaceholderTemplate(
src, PLACEHOLDER_PREFIX, PLACEHOLDER_SUFFIX, replacements);
};
const std::string shaderSource{ webgpuutils::processPlaceholderTemplate(SHADER_SOURCE_TEMPLATE,
PLACEHOLDER_PREFIX, PLACEHOLDER_SUFFIX, valueByPlaceholderNameForModule) };
// The dependency chain is shader module -> fragment source -> sampling function
// Hence we need to fill in the templates in reverse order.
// Fill in sampling implementation in the sampling func
std::string source = replace(TEXTURE_SAMPLE_IMPL_MAIN_TEMPLATE, {
{ TEXTURE_SAMPLE_IMPL, sampleImpl },
});
// Fill in sampling func in the fragment
source = replace(fragmentTemplate, {
{ TEXTURE_SAMPLE_IMPL_MAIN, source },
});
// Fill in fragment in the shader module, first pass
source = replace(SHADER_SOURCE_TEMPLATE, {
{ FRAGMENT_SHADER_SNIPPET, source },
});
// Fill in fragment in the shader module, second pass
source = replace(source, {
{ TEXTURE_TYPE, srcTextureType },
{ RET_ATTRIBUTE, retAttribute },
{ RET_TYPE, retType },
});
// Fill in fragment in the shader module, third pass
// TEXTURE_TYPE, RET_ATTRIBUTE, RET_TYPE have a dependency on SRC_PRIM_TYPE, DST_PRIM_TYPE,
// VECTOR_DIM so we need another pass to fill them in.
source = replace(source, {
{ VECTOR_DIM, vecDim },
{ SRC_PRIM_TYPE, srcPrimType },
{ DST_PRIM_TYPE, dstPrimType },
});
wgpu::ShaderModuleWGSLDescriptor wgslDescriptor{};
wgslDescriptor.code = shaderSource.data();
wgslDescriptor.code = source.data();
const wgpu::ShaderModuleDescriptor shaderModuleDescriptor{
.nextInChain = &wgslDescriptor,
.label = "render_pass_blit_shaders",

View File

@@ -97,11 +97,12 @@ private:
struct RenderPipelineKey {
wgpu::TextureViewDimension sourceDimension; // 4 bytes
wgpu::TextureFormat sourceTextureFormat; // 4
wgpu::TextureSampleType sourceSampleType; // 4
wgpu::TextureFormat destinationTextureFormat; // 4
uint8_t sourceSampleCount; // 1
SamplerMagFilter filterType; // 1
bool depthSource; // 1
uint8_t padding = 0; // 1
uint8_t padding[2] = {}; // 2
bool operator==(const RenderPipelineKey& other) const;
using Hasher = utils::hash::MurmurHashFn<RenderPipelineKey>;
@@ -109,10 +110,10 @@ private:
struct PipelineLayoutKey {
wgpu::TextureViewDimension sourceDimension; // 4 bytes
wgpu::TextureSampleType sourceSampleType; // 4
SamplerMagFilter filterType; // 1
bool multisampledSource; // 1
bool depthSource; // 1
uint8_t padding = 0; // 1
uint8_t padding[2] = {}; // 2
bool operator==(const PipelineLayoutKey& other) const;
using Hasher = utils::hash::MurmurHashFn<PipelineLayoutKey>;
@@ -120,20 +121,20 @@ private:
struct ShaderModuleKey {
wgpu::TextureViewDimension sourceDimension; // 4 bytes
wgpu::TextureFormat sourceTextureFormat; // 4
wgpu::TextureFormat destinationTextureFormat; // 4
bool multisampledSource; // 1
bool depthSource; // 1
bool depthDestination; // 1
uint8_t padding = 0; // 1
uint8_t padding0[3] = {}; // 3
bool operator==(const ShaderModuleKey& other) const;
using Hasher = utils::hash::MurmurHashFn<ShaderModuleKey>;
};
static_assert(sizeof(RenderPipelineKey) == 12,
static_assert(sizeof(RenderPipelineKey) == 20,
"RenderPipelineKey must not have implicit padding.");
static_assert(sizeof(PipelineLayoutKey) == 8,
static_assert(sizeof(PipelineLayoutKey) == 12,
"PipelineLayoutKey must not have implicit padding.");
static_assert(sizeof(ShaderModuleKey) == 8, "ShaderModuleKey must not have implicit padding.");
static_assert(sizeof(ShaderModuleKey) == 16, "ShaderModuleKey must not have implicit padding.");
wgpu::Device mDevice;
wgpu::Sampler mNearestSampler{ nullptr };

View File

@@ -0,0 +1,174 @@
/*
* Copyright (C) 2025 The Android Open Source Project
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
/*
* NOTE: This file is meant to be included with #include in WebGPUBlitter.cpp
* It is separated out for clarity.
*/
namespace {
// note that the placeholders below must start and end with this prefix and suffix and should not
// otherwise be present in the template:
constexpr std::string_view PLACEHOLDER_PREFIX{ "{{" };
constexpr std::string_view PLACEHOLDER_SUFFIX{ "}}" };
#define PLACEHOLDER_DEF(name) constexpr std::string_view name = #name
PLACEHOLDER_DEF(FRAGMENT_SHADER_SNIPPET);
PLACEHOLDER_DEF(TEXTURE_SAMPLE_IMPL);
PLACEHOLDER_DEF(TEXTURE_SAMPLE_IMPL_MAIN);
// texture_multisampled_2d<f32>
// texture_2d<f32/i32/u32>
// texture_3d<f32/i32/u32>
// texture_depth_2d
// texture_depth_multisampled_2d
PLACEHOLDER_DEF(TEXTURE_TYPE);
// vec2, vec3
PLACEHOLDER_DEF(VECTOR_DIM);
// vec_<f32>, vec_<u32>, vec_<i32>, f32
PLACEHOLDER_DEF(RET_TYPE);
// @builtin(frag_depth), @location(0)
PLACEHOLDER_DEF(RET_ATTRIBUTE);
// <f32>, <u32>, <i32>
PLACEHOLDER_DEF(DST_PRIM_TYPE);
PLACEHOLDER_DEF(SRC_PRIM_TYPE);
#undef PLACEHOLDER_DEF
constexpr std::string_view SHADER_SOURCE_TEMPLATE{ R"(
struct BlitFragmentShaderArgs {
depthPlane: u32,
scale: vec2<f32>,
sourceOffset: vec2<u32>,
destinationOffset: vec2<u32>,
};
@group(0) @binding(0) var sourceTexture: {{TEXTURE_TYPE}};
@group(0) @binding(1) var<uniform> fragmentShaderArgs: BlitFragmentShaderArgs;
fn getUnnormalizedSourceTextureCoordinates(position: vec2<f32>) -> vec2<f32> {
// These coordinates match the Vulkan vkCmdBlitImage spec:
// https://www.khronos.org/registry/vulkan/specs/1.2-extensions/man/html/vkCmdBlitImage.html
let uvOffset: vec2<f32> = position - vec2<f32>(f32(fragmentShaderArgs.destinationOffset.x),f32(fragmentShaderArgs.destinationOffset.y));
let uvScaled: vec2<f32> = uvOffset * fragmentShaderArgs.scale;
return uvScaled + vec2<f32>(f32(fragmentShaderArgs.sourceOffset.x), f32(fragmentShaderArgs.sourceOffset.y));
}
fn normalize2dSourceTextureCoordinates(
unnormalizedSourceTextureCoordinates: vec2<f32>,
sourceDimensions: vec2<u32>) -> vec2<f32> {
return unnormalizedSourceTextureCoordinates /
vec2<f32>(f32(sourceDimensions.x), f32(sourceDimensions.y));
}
@vertex
fn vertexShaderMain(@builtin(vertex_index) vertexIndex: u32) -> @builtin(position) vec4<f32> {
let fullScreenTriangleVertices = array<vec2<f32>, 3>(
vec2<f32>(-1.0, -1.0),
vec2<f32>( 3.0, -1.0),
vec2<f32>(-1.0, 3.0)
);
return vec4<f32>(fullScreenTriangleVertices[vertexIndex].xy, 0.0, 1.0);
}
{{FRAGMENT_SHADER_SNIPPET}}
)"
};
constexpr std::string_view FLOAT_TEXTURE_SAMPLE_IMPL_TEMPLATE{ R"(
return textureSample(sourceTexture, sourceSampler, coordinates);
)" };
constexpr std::string_view INT_TEXTURE_SAMPLE_IMPL_TEMPLATE { R"(
let texelCoords = {{VECTOR_DIM}}<u32>(coordinates * {{VECTOR_DIM}}<f32>(sourceDimensions));
return textureLoad(sourceTexture, texelCoords, 0);
)"};
constexpr std::string_view TEXTURE_SAMPLE_IMPL_MAIN_TEMPLATE { R"(
// Assumes that "@group(0) @binding(2) var sourceSampler: sampler;" has been declared in the main program;
fn sampleTextureImpl(sourceTexture: {{TEXTURE_TYPE}}, sourceDimensions: {{VECTOR_DIM}}<u32>,
coordinates: {{VECTOR_DIM}}<f32>) -> vec4{{DST_PRIM_TYPE}} {
{{TEXTURE_SAMPLE_IMPL}}
}
)"};
constexpr std::string_view FRAGMENT_SHADER_SNIPPET_MSAA_INPUT_TEMPLATE{ R"(
@fragment
fn fragmentShaderMain(@builtin(position) position: vec4<f32>) -> {{RET_ATTRIBUTE}} {{RET_TYPE}} {
var color: vec4<f32> = vec4<f32>(0.0, 0.0, 0.0, 0.0);
let numberOfSamples: u32 = textureNumSamples(sourceTexture);
let coordinatesF: vec2<f32> = getUnnormalizedSourceTextureCoordinates(position.xy);
let coordinates: vec2<u32> = vec2<u32>(u32(coordinatesF.x), u32(coordinatesF.y));
for (var sampleIndex: u32 = 0; sampleIndex < numberOfSamples; sampleIndex++) {
color += textureLoad(sourceTexture, coordinates, sampleIndex);
}
color /= f32(numberOfSamples);
return color;
}
)" };
constexpr std::string_view FRAGMENT_SHADER_SNIPPET_3D_INPUT_TEMPLATE{ R"(
@group(0) @binding(2) var sourceSampler: sampler;
{{TEXTURE_SAMPLE_IMPL_MAIN}}
fn normalize3dSourceTextureCoordinates(
unnormalizedSourceTextureCoordinates: vec2<f32>,
sourceDimensions: vec3<u32>) -> vec3<f32> {
let uvNormalized: vec2<f32> = normalize2dSourceTextureCoordinates(
unnormalizedSourceTextureCoordinates,
sourceDimensions.xy
);
return vec3<f32>(
uvNormalized,
(f32(fragmentShaderArgs.depthPlane) + 0.5) / f32(sourceDimensions.z)
);
}
@fragment
fn fragmentShaderMain(@builtin(position) position: vec4<f32>) -> {{RET_ATTRIBUTE}} {{RET_TYPE}} {
let sourceDimensions: vec3<u32> = textureDimensions(sourceTexture);
let coordinates: vec3<f32> = normalize3dSourceTextureCoordinates(
getUnnormalizedSourceTextureCoordinates(position),
sourceDimensions
);
return sampleTextureImpl(sourceTexture, sourceDimensions, coordinates);
}
)" };
constexpr std::string_view FRAGMENT_SHADER_SNIPPET_2D_INPUT_TEMPLATE { R"(
@group(0) @binding(2) var sourceSampler: sampler;
{{TEXTURE_SAMPLE_IMPL_MAIN}}
@fragment
fn fragmentShaderMain(@builtin(position) position: vec4<f32>) -> {{RET_ATTRIBUTE}} {{RET_TYPE}} {
let sourceDimensions: vec2<u32> = textureDimensions(sourceTexture);
let coordinates: vec2<f32> = normalize2dSourceTextureCoordinates(
getUnnormalizedSourceTextureCoordinates(position.xy),
sourceDimensions
);
return sampleTextureImpl(sourceTexture, sourceDimensions, coordinates);
}
)" };
} // namespace

View File

@@ -446,6 +446,7 @@ void WebGPUDriver::createSwapChainHeadlessR(Handle<HwSwapChain> sch, uint32_t wi
void WebGPUDriver::createVertexBufferInfoR(Handle<HwVertexBufferInfo> vertexBufferInfoHandle,
const uint8_t bufferCount, const uint8_t attributeCount, const AttributeArray attributes,
utils::ImmutableCString&& tag) {
// Hello world! This is a test for pushing branch to Github
FWGPU_SYSTRACE_SCOPE();
constructHandle<WebGPUVertexBufferInfo>(vertexBufferInfoHandle, bufferCount, attributeCount,
attributes, mDeviceLimits);
@@ -1497,7 +1498,7 @@ void WebGPUDriver::readPixels(Handle<HwRenderTarget> sourceRenderTargetHandle, c
mReadPixelMapsCounter.startTask();
userData->buffer.MapAsync(
wgpu::MapMode::Read, 0, bufferSize, wgpu::CallbackMode::AllowProcessEvents,
wgpu::MapMode::Read, 0, bufferSize, wgpu::CallbackMode::AllowSpontaneous,
[](wgpu::MapAsyncStatus status, const char* message, UserData* userdata) {
std::unique_ptr<UserData> data(static_cast<UserData*>(userdata));
if (UTILS_LIKELY(status == wgpu::MapAsyncStatus::Success)) {

View File

@@ -41,6 +41,45 @@ namespace filament::backend {
textureFormat == wgpu::TextureFormat::Depth32FloatStencil8;
}
[[nodiscard]] constexpr bool isUIntFormat(const wgpu::TextureFormat format) {
// see https://www.w3.org/TR/webgpu/#texture-formats
// and https://www.w3.org/TR/webgpu/#texture-format-caps
switch (format) {
case wgpu::TextureFormat::R8Uint:
case wgpu::TextureFormat::R16Uint:
case wgpu::TextureFormat::RG8Uint:
case wgpu::TextureFormat::R32Uint:
case wgpu::TextureFormat::RG16Uint:
case wgpu::TextureFormat::RGBA8Uint:
case wgpu::TextureFormat::RGB10A2Uint:
case wgpu::TextureFormat::RG32Uint:
case wgpu::TextureFormat::RGBA16Uint:
case wgpu::TextureFormat::RGBA32Uint:
return true;
default:
return false;
}
}
[[nodiscard]] constexpr bool isIntFormat(const wgpu::TextureFormat format) {
// see https://www.w3.org/TR/webgpu/#texture-formats
// and https://www.w3.org/TR/webgpu/#texture-format-caps
switch (format) {
case wgpu::TextureFormat::R8Sint:
case wgpu::TextureFormat::R16Sint:
case wgpu::TextureFormat::RG8Sint:
case wgpu::TextureFormat::R32Sint:
case wgpu::TextureFormat::RG16Sint:
case wgpu::TextureFormat::RGBA8Sint:
case wgpu::TextureFormat::RG32Sint:
case wgpu::TextureFormat::RGBA16Sint:
case wgpu::TextureFormat::RGBA32Sint:
return true;
default:
return false;
}
}
[[nodiscard]] constexpr std::string_view toString(const PixelDataFormat format) {
switch (format) {
case PixelDataFormat::R: return "R";

View File

@@ -52,15 +52,14 @@ std::string processPlaceholderTemplate(std::string_view const& stringTemplate,
"Malformed source with missing suffix to placeholder");
const std::string_view placeholderName{ std::string_view(sourceData + positionOfPlaceholder,
positionAfterPlaceholder - positionOfPlaceholder) };
if (const auto iter{ valueByPlaceholderName.find(placeholderName) };
iter == valueByPlaceholderName.end()) {
// wrapping placeholderName in a std::string to null terminate the C string....
PANIC_POSTCONDITION("Found placeholder '%s' in template, but this is not present in "
"valueByPlaceholderName",
std::string{ placeholderName }.data());
} else {
iter != valueByPlaceholderName.end()) {
const std::string_view value{ iter->second };
out << value;
} else {
// It's ok to not replace a template item. We assume replacement can take multiple passes.
out << placeholderPrefix << placeholderName << placeholderSuffix;
}
// update the cursor for after the placeholder...
positionCursorInTemplateString = positionAfterPlaceholder + placeholderSuffix.size();

View File

@@ -23,13 +23,17 @@
namespace test {
Backend parseArgumentsForBackend(int argc, char* argv[]) {
Backend backend = Backend::OPENGL;
TestArguments parseArguments(int argc, char* argv[]) {
TestArguments arguments = {};
arguments.backend = Backend::OPENGL;
// The first colon in OPTSTR turns on silent error reporting. This is important, as the
// arguments may also contain gtest parameters we don't know about.
static constexpr const char* OPTSTR = ":a:";
static constexpr const char* OPTSTR = ":a:kc";
static const struct option OPTIONS[] = {
{ "api", required_argument, nullptr, 'a' },
{ "headless_only", no_argument, nullptr, 'k' },
{ "ci", no_argument, nullptr, 'c' },
{ nullptr, 0, nullptr, 0 } // termination of the option list
};
@@ -41,13 +45,13 @@ Backend parseArgumentsForBackend(int argc, char* argv[]) {
switch (opt) {
case 'a':
if (arg == "opengl") {
backend = Backend::OPENGL;
arguments.backend = Backend::OPENGL;
} else if (arg == "vulkan") {
backend = Backend::VULKAN;
arguments.backend = Backend::VULKAN;
} else if (arg == "metal") {
backend = Backend::METAL;
arguments.backend = Backend::METAL;
} else if (arg == "webgpu") {
backend = Backend::WEBGPU;
arguments.backend = Backend::WEBGPU;
} else {
std::cerr << "Unrecognized target API. Must be 'opengl'|'vulkan'|'metal'|'webgpu'."
<< std::endl
@@ -55,10 +59,16 @@ Backend parseArgumentsForBackend(int argc, char* argv[]) {
<< std::endl;
}
break;
case 'k':
arguments.headlessOnly = true;
break;
case 'c':
arguments.isContinuousIntegration = true;
break;
}
}
return backend;
return arguments;
}
} // namespace test

View File

@@ -41,6 +41,7 @@ enum class OperatingSystem: uint8_t {
LINUX = 2,
// Also represents iOS phones.
APPLE = 3,
CONTINUOUS_INTEGRATION = 4,
// TODO: When tests support windows add it here.
};
@@ -73,11 +74,17 @@ void initTests(Backend backend, OperatingSystem operatingSystem, bool isMobile,
*/
int runTests();
struct TestArguments {
Backend backend;
bool headlessOnly = false;
bool isContinuousIntegration = false;
};
/**
* A utility method that can be invoked by test runners to parse arguments.
* Looks through the provided command-line arguments and finds any -a <backend> arguments.
*/
Backend parseArgumentsForBackend(int argc, char* argv[]);
TestArguments parseArguments(int argc, char* argv[]);
} // namespace test

View File

@@ -43,7 +43,10 @@ std::array<test::Backend, 2> const VALID_BACKENDS{
}// namespace
int main(int argc, char* argv[]) {
auto backend = test::parseArgumentsForBackend(argc, argv);
const auto arguments = test::parseArguments(argc, argv);
const auto backend = arguments.backend;
// Note that Linux is headless-only.
if (!std::any_of(VALID_BACKENDS.begin(), VALID_BACKENDS.end(),
[backend](test::Backend validBackend) { return backend == validBackend; })) {
@@ -51,6 +54,8 @@ int main(int argc, char* argv[]) {
return 1;
}
test::initTests(backend, test::OperatingSystem::LINUX, false, argc, argv);
const auto operatingSystem = arguments.isContinuousIntegration ?
test::OperatingSystem::CONTINUOUS_INTEGRATION : test::OperatingSystem::LINUX;
test::initTests(backend, operatingSystem, false, argc, argv);
return test::runTests();
}

View File

@@ -32,29 +32,34 @@ test::NativeView getNativeView() {
@interface AppDelegate : NSObject <NSApplicationDelegate>
@property test::Backend backend;
@property bool headlessOnly;
@end
@implementation AppDelegate
- (void)applicationDidFinishLaunching:(NSNotification *)aNotification {
NSView* view = [self createView];
if (self.backend == test::Backend::OPENGL) {
nativeView.ptr = (void*) view;
if (self.headlessOnly) {
nativeView.ptr = nullptr;
nativeView.width = test::WINDOW_WIDTH;
nativeView.height = test::WINDOW_HEIGHT;
} else {
NSView* view = [self createView];
switch (self.backend) {
case test::Backend::OPENGL:
nativeView.ptr = (void*) view;
break;
case test::Backend::METAL:
case test::Backend::VULKAN:
case test::Backend::WEBGPU:
case test::Backend::NOOP:
nativeView.ptr = (void*) view.layer;
break;
}
CGSize drawableSize = ((CAMetalLayer*) view.layer).drawableSize;
nativeView.width = static_cast<size_t>(drawableSize.width);
nativeView.height = static_cast<size_t>(drawableSize.height);
}
if (self.backend == test::Backend::METAL) {
nativeView.ptr = (void*) view.layer;
}
if (self.backend == test::Backend::VULKAN) {
nativeView.ptr = (void*) view.layer;
}
if (self.backend == test::Backend::WEBGPU) {
nativeView.ptr = (void*) view.layer;
}
CGSize drawableSize = ((CAMetalLayer*) view.layer).drawableSize;
nativeView.width = static_cast<size_t>(drawableSize.width);
nativeView.height = static_cast<size_t>(drawableSize.height);
exit(test::runTests());
}
@@ -100,11 +105,25 @@ test::NativeView getNativeView() {
@end
int main(int argc, char* argv[]) {
auto backend = test::parseArgumentsForBackend(argc, argv);
test::initTests(backend, test::OperatingSystem::APPLE, false, argc, argv);
AppDelegate* delegate = [AppDelegate new];
delegate.backend = backend;
const auto arguments = test::parseArguments(argc, argv);
const auto operatingSystem = arguments.isContinuousIntegration ?
test::OperatingSystem::CONTINUOUS_INTEGRATION : test::OperatingSystem::APPLE;
test::initTests(arguments.backend, operatingSystem, false, argc, argv);
NSApplication* app = [NSApplication sharedApplication];
AppDelegate* delegate = [AppDelegate new];
delegate.backend = arguments.backend;
delegate.headlessOnly = arguments.headlessOnly;
[app setDelegate:delegate];
if (arguments.headlessOnly) {
// In headless mode, we don't want to start the NSApplication event loop.
// Instead, we can manually "finish" launching the app, which will trigger the tests to run.
[app finishLaunching];
[delegate applicationDidFinishLaunching:nil];
// The line above calls exit(), so we should not reach here.
return 0;
}
[app run];
}

View File

@@ -158,7 +158,7 @@ TEST_F(BlitTest, ColorMagnify) {
constexpr int kNumLevels = 3;
// Create a SwapChain and make it current. We don't really use it so the res doesn't matter.
auto swapChain = mCleanup.add(api.createSwapChainHeadless(256, 256, 0));
auto swapChain = mCleanup.add(createSwapChain());
api.makeCurrent(swapChain, swapChain);
// Create a source texture.
@@ -222,7 +222,7 @@ TEST_F(BlitTest, ColorMinify) {
constexpr int kNumLevels = 3;
// Create a SwapChain and make it current. We don't really use it so the res doesn't matter.
auto swapChain = mCleanup.add(api.createSwapChainHeadless(256, 256, 0));
auto swapChain = mCleanup.add(createSwapChain());
api.makeCurrent(swapChain, swapChain);
// Create a source texture.
@@ -371,7 +371,7 @@ TEST_F(BlitTest, Blit2DTextureArray) {
constexpr int kDstTexLayer = 0;
// Create a SwapChain and make it current. We don't really use it so the res doesn't matter.
auto swapChain = mCleanup.add(api.createSwapChainHeadless(256, 256, 0));
auto swapChain = mCleanup.add(createSwapChain());
api.makeCurrent(swapChain, swapChain);
// Create a source texture.
@@ -437,7 +437,7 @@ TEST_F(BlitTest, BlitRegion) {
constexpr int kDstLevel = 0;
// Create a SwapChain and make it current. We don't really use it so the res doesn't matter.
auto swapChain = mCleanup.add(api.createSwapChainHeadless(256, 256, 0));
auto swapChain = mCleanup.add(createSwapChain());
api.makeCurrent(swapChain, swapChain);
// Create a source texture.

View File

@@ -103,7 +103,7 @@ TEST_F(BackendTest, FrameCompletedCallback) {
Cleanup cleanup(api);
// Create a SwapChain.
auto swapChain = cleanup.add(api.createSwapChainHeadless(256, 256, 0));
auto swapChain = cleanup.add(createSwapChain());
int callbackCountA = 0;
api.setFrameCompletedCallback(swapChain, nullptr,

View File

@@ -22,6 +22,7 @@
#include "SharedShaders.h"
#include "SharedShadersConstants.h"
#include "Skip.h"
#include "Workarounds.h"
#include <backend/BufferDescriptor.h>
#include <backend/DriverEnums.h>
@@ -55,7 +56,7 @@ protected:
auto colorTexture = cleanup.add(
api.createTexture(SamplerType::SAMPLER_2D, 1, TextureFormat::RGBA8, 1,
screenWidth(), screenHeight(), 1,
TextureUsage::COLOR_ATTACHMENT | TextureUsage::SAMPLEABLE));
TextureUsage::COLOR_ATTACHMENT | TextureUsage::SAMPLEABLE TEXTURE_USAGE_READ_PIXELS));
auto renderTarget = cleanup.add(api.createRenderTarget(TargetBufferFlags::COLOR0,
screenWidth(), screenHeight(), 1, 1, { { colorTexture } }, {}, {}));
return renderTarget;
@@ -90,19 +91,27 @@ protected:
return { vbih, vbh, ibh };
}
Shader getSimpleShader(Cleanup& cleanup, math::float4 const& color) {
std::pair<Shader, DescriptorSetHandle> getSimpleShader(Cleanup& cleanup,
math::float4 const& color) {
auto& api = getDriverApi();
Shader const shader = SharedShaders::makeShader(getDriverApi(), cleanup, ShaderRequest{
.mVertexType = VertexShaderType::Simple,
.mFragmentType = FragmentShaderType::SolidColored,
.mUniformType = ShaderUniformType::Simple
});
Shader const shader = SharedShaders::makeShader(getDriverApi(), cleanup,
ShaderRequest{
.mVertexType = VertexShaderType::Simple,
.mFragmentType = FragmentShaderType::SolidColored,
.mUniformType = ShaderUniformType::Simple,
});
auto descSet = shader.createDescriptorSet(api);
UniformBindingConfig uboBindingConfig = {
.descriptorSet = descSet,
};
auto const ubuffer = cleanup.add(api.createBufferObject(sizeof(SimpleMaterialParams),
BufferObjectBinding::UNIFORM, BufferUsage::DYNAMIC_BIT));
shader.bindUniform<SimpleMaterialParams>(api, ubuffer);
// This will bind the UBO to descSet and also bind descSet. But we also need to manually
// bind descset every begin/end frame.
shader.bindUniform<SimpleMaterialParams>(api, ubuffer, uboBindingConfig);
shader.uploadUniform(api, ubuffer, SimpleMaterialParams{ .color = color });
return shader;
return { shader, descSet };
}
void copyData(const void* data, size_t const size, size_t const offset,
@@ -122,7 +131,9 @@ protected:
int64_t const frame,
SwapChainHandle swapChain,
RenderTargetHandle renderTarget,
RenderPrimitiveHandle renderPrimitive, PipelineState const& state) {
RenderPrimitiveHandle renderPrimitive,
DescriptorSetHandle descSet,
PipelineState const& state) {
auto& api = getDriverApi();
RenderPassParams params = getClearColorRenderPass({0,0,0,1});
params.viewport = getFullViewport();
@@ -130,6 +141,8 @@ protected:
api.beginFrame(frame, 0, 0);
api.beginRenderPass(renderTarget, params);
api.bindPipeline(state);
// The binding index is always 0 for the simple shaders we're using.
api.bindDescriptorSet(descSet, 0, {});
api.bindRenderPrimitive(renderPrimitive);
api.draw2(0, count, 1);
api.endRenderPass();
@@ -185,7 +198,7 @@ protected:
flushAndWait();
EXPECT_EQ(callbackExecuted, 1);
Shader const shader = getSimpleShader(cleanup, color);
auto [shader, descset] = getSimpleShader(cleanup, color);
auto [vbih, vbh, ibh] = setupGeometryBuffer(cleanup, 3, stride, baseVertex, bufferObject);
@@ -197,7 +210,7 @@ protected:
state.primitiveType = PrimitiveType::TRIANGLES;
state.vertexBufferInfo = vbih;
render(3, 0, swapChain, renderTarget, renderPrimitive, state);
render(3, 0, swapChain, renderTarget, renderPrimitive, descset, state);
EXPECT_IMAGE(renderTarget, ScreenshotParams(screenWidth(), screenHeight(), screenshotName, 0));
}
@@ -288,7 +301,7 @@ TEST_F(MemoryMappedTest, MultipleCopies) {
flushAndWait();
EXPECT_EQ(callbacksExecuted, 3);
Shader const shader = getSimpleShader(cleanup, { 1, 1, 0, 1 });
auto [shader, descset] = getSimpleShader(cleanup, { 1, 1, 0, 1 });
auto [vbih, vbh, ibh] = setupGeometryBuffer(cleanup, 9, sizeof(math::float2), 0, bufferObject);
@@ -300,7 +313,7 @@ TEST_F(MemoryMappedTest, MultipleCopies) {
state.primitiveType = PrimitiveType::TRIANGLES;
state.vertexBufferInfo = vbih;
render(9, 0, swapChain, renderTarget, renderPrimitive, state);
render(9, 0, swapChain, renderTarget, renderPrimitive, descset, state);
EXPECT_IMAGE(renderTarget, ScreenshotParams(screenWidth(), screenHeight(), "MultipleCopies", 0));
}
@@ -345,7 +358,7 @@ TEST_F(MemoryMappedTest, UpdatePartial) {
}
Shader const shader = getSimpleShader(cleanup, { 1, 1, 0, 1 });
auto [shader, descset] = getSimpleShader(cleanup, { 1, 1, 0, 1 });
auto [vbih, vbh, ibh] = setupGeometryBuffer(cleanup, 9, sizeof(math::float2), 0, bufferObject);
@@ -357,7 +370,7 @@ TEST_F(MemoryMappedTest, UpdatePartial) {
state.primitiveType = PrimitiveType::TRIANGLES;
state.vertexBufferInfo = vbih;
render(9, 0, swapChain, renderTarget, renderPrimitive, state);
render(9, 0, swapChain, renderTarget, renderPrimitive, descset, state);
EXPECT_IMAGE(renderTarget, ScreenshotParams(screenWidth(), screenHeight(), "UpdatePartial_before", 0));
@@ -380,7 +393,7 @@ TEST_F(MemoryMappedTest, UpdatePartial) {
api.makeCurrent(swapChain, swapChain);
// Second render, after update
render(9, 1, swapChain, renderTarget, renderPrimitive, state);
render(9, 1, swapChain, renderTarget, renderPrimitive, descset, state);
EXPECT_IMAGE(renderTarget, ScreenshotParams(screenWidth(), screenHeight(), "UpdatePartial_after", 0));
}

View File

@@ -385,8 +385,7 @@ TEST_F(ReadPixelsTest, ReadPixelsPerformance) {
Cleanup cleanup(api);
// Create a platform-specific SwapChain and make it current.
auto swapChain = cleanup.add(
api.createSwapChainHeadless(renderTargetSize, renderTargetSize, 0));
auto swapChain = cleanup.add(createSwapChain());
api.makeCurrent(swapChain, swapChain);
Shader shader = SharedShaders::makeShader(api, cleanup, ShaderRequest{

View File

@@ -70,7 +70,7 @@ TEST_F(BackendTest, ScissorViewportRegion) {
// executeCommands().
{
// Create a SwapChain and make it current. We don't really use it so the res doesn't matter.
auto swapChain = cleanup.add(api.createSwapChainHeadless(256, 256, 0));
auto swapChain = cleanup.add(createSwapChain());
api.makeCurrent(swapChain, swapChain);
Shader shader = SharedShaders::makeShader(api, cleanup,
@@ -149,7 +149,7 @@ TEST_F(BackendTest, ScissorViewportEdgeCases) {
// executeCommands().
{
// Create a SwapChain and make it current. We don't really use it so the res doesn't matter.
auto swapChain = cleanup.add(api.createSwapChainHeadless(256, 256, 0));
auto swapChain = cleanup.add(createSwapChain());
api.makeCurrent(swapChain, swapChain);
Shader shader = SharedShaders::makeShader(api, cleanup, ShaderRequest{

View File

@@ -33,8 +33,7 @@ TEST_F(BackendTest, TestTemplate) {
auto& api = getDriverApi();
Cleanup cleanup(api);
auto swapChain =
cleanup.add(api.createSwapChainHeadless(kRenderTargetSize, kRenderTargetSize, 0));
auto swapChain = cleanup.add(createSwapChain());
api.makeCurrent(swapChain, swapChain);
RenderTargetHandle renderTarget = cleanup.add(api.createDefaultRenderTarget());

View File

@@ -0,0 +1,234 @@
/*
* Copyright (C) 2025 The Android Open Source Project
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#include "details/BufferAllocator.h"
#include <utils/Panic.h>
#include <utils/debug.h>
namespace filament {
namespace {
static bool isValidId(BufferAllocator::AllocationId id) {
return id != BufferAllocator::UNALLOCATED && id != BufferAllocator::REALLOCATION_REQUIRED;
}
#ifndef NDEBUG
constexpr static bool isPowerOfTwo(uint32_t n) {
return (n > 0) && ((n & (n - 1)) == 0);
}
#endif
} // anonymous namespace
BufferAllocator::BufferAllocator(allocation_size_t totalSize, allocation_size_t slotSize)
: mTotalSize(totalSize),
mSlotSize(slotSize) {
assert_invariant(mSlotSize > 0);
assert_invariant(isPowerOfTwo(mSlotSize));
reset(mTotalSize);
}
void BufferAllocator::reset(allocation_size_t newTotalSize) {
assert_invariant(newTotalSize % mSlotSize == 0);
mTotalSize = newTotalSize;
mSlotPool.clear();
mFreeList.clear();
mOffsetMap.clear();
// Initialize the pool with a single large free slot.
mSlotPool.emplace_back(InternalSlotNode{
.mSlot = {
.offset = 0,
.slotSize = mTotalSize,
.isAllocated = false,
.gpuUseCount = 0
}
});
InternalSlotNode* firstNode = &mSlotPool.front();
// Add this initial free slot to the free list and offset map.
auto freeListIter = mFreeList.emplace(newTotalSize, firstNode);
auto offsetMapIter = mOffsetMap.emplace(0, firstNode);
firstNode->mSlotPoolIterator = mSlotPool.begin();
firstNode->mFreeListIterator = freeListIter;
firstNode->mOffsetMapIterator = offsetMapIter.first;
}
std::pair<BufferAllocator::AllocationId, BufferAllocator::allocation_size_t>
BufferAllocator::allocate(allocation_size_t size) noexcept {
if (size == 0) {
return { UNALLOCATED, 0 };
}
const allocation_size_t alignedSize = alignUp(size);
auto bestFitIter = mFreeList.lower_bound(alignedSize);
if (bestFitIter == mFreeList.end()) {
return { REALLOCATION_REQUIRED, 0 };
}
InternalSlotNode* targetNode = bestFitIter->second;
const allocation_size_t originalSlotSize = targetNode->mSlot.slotSize;
mFreeList.erase(bestFitIter);
targetNode->mFreeListIterator = mFreeList.end();
targetNode->mSlot.isAllocated = true;
// Split the slot if it is larger than what we need.
if (originalSlotSize > alignedSize) {
targetNode->mSlot.slotSize = alignedSize;
allocation_size_t remainingSize = originalSlotSize - alignedSize;
allocation_size_t newSlotOffset = targetNode->mSlot.offset + alignedSize;
assert_invariant(remainingSize % mSlotSize == 0);
assert_invariant(newSlotOffset % mSlotSize == 0);
// Create a new node for the remaining free space.
auto insertPos = std::next(targetNode->mSlotPoolIterator);
auto newNodeIter = mSlotPool.emplace(insertPos, InternalSlotNode{
.mSlot = {
.offset = newSlotOffset,
.slotSize = remainingSize,
.isAllocated = false,
.gpuUseCount = 0
}
});
InternalSlotNode* newNode = &(*newNodeIter);
// Add the new free slot to our tracking maps.
auto freeListIter = mFreeList.emplace(remainingSize, newNode);
auto offsetMapIter = mOffsetMap.emplace(newSlotOffset, newNode);
newNode->mSlotPoolIterator = newNodeIter;
newNode->mFreeListIterator = freeListIter;
newNode->mOffsetMapIterator = offsetMapIter.first;
}
auto allocationId = calculateIdByOffset(targetNode->mSlot.offset);
return { allocationId, targetNode->mSlot.offset };
}
BufferAllocator::InternalSlotNode* BufferAllocator::getNodeById(
AllocationId id) const noexcept {
if (!isValidId(id)) {
return nullptr;
}
auto offset = getAllocationOffset(id);
auto iter = mOffsetMap.find(offset);
// We cannot find the corresponding node in the map.
if (iter == mOffsetMap.end()) {
return nullptr;
}
return iter->second;
}
void BufferAllocator::retire(AllocationId id) {
auto targetNode = getNodeById(id);
assert_invariant(targetNode != nullptr);
targetNode->mSlot.isAllocated = false;
}
void BufferAllocator::acquireGpu(AllocationId id) {
auto targetNode = getNodeById(id);
assert_invariant(targetNode != nullptr);
targetNode->mSlot.gpuUseCount++;
}
void BufferAllocator::releaseGpu(AllocationId id) {
auto targetNode = getNodeById(id);
assert_invariant(targetNode != nullptr);
assert_invariant(targetNode->mSlot.gpuUseCount > 0);
targetNode->mSlot.gpuUseCount--;
}
void BufferAllocator::releaseFreeSlots() {
auto curr = mSlotPool.begin();
while (curr != mSlotPool.end()) {
if (!curr->mSlot.isFree()) {
++curr;
continue;
}
auto next = std::next(curr);
bool merged = false;
while (next != mSlotPool.end() && next->mSlot.isFree()) {
merged = true;
// Combine the size of free slots
curr->mSlot.slotSize += next->mSlot.slotSize;
assert_invariant(curr->mSlot.slotSize % mSlotSize == 0);
// Erase the merged slot from all maps
if (next->mFreeListIterator != mFreeList.end()) {
mFreeList.erase(next->mFreeListIterator);
}
mOffsetMap.erase(next->mOffsetMapIterator);
next = mSlotPool.erase(next);
}
// If we performed any merge, the current block's size has changed.
// We need to update its position in the mFreeList.
if (curr->mFreeListIterator != mFreeList.end()) {
// If it's already in the free list and we merged, we need to update it.
if (merged) {
mFreeList.erase(curr->mFreeListIterator);
curr->mFreeListIterator = mFreeList.emplace(curr->mSlot.slotSize, &(*curr));
}
} else {
// If it's not in the free list, it must be a newly freed block. Add it.
curr->mFreeListIterator = mFreeList.emplace(curr->mSlot.slotSize, &(*curr));
}
curr = next;
}
}
BufferAllocator::allocation_size_t BufferAllocator::getTotalSize() const noexcept {
return mTotalSize;
}
BufferAllocator::allocation_size_t
BufferAllocator::getAllocationOffset(AllocationId id) const {
assert_invariant(isValidId(id));
return (id - 1) * mSlotSize;
}
BufferAllocator::AllocationId BufferAllocator::calculateIdByOffset(
allocation_size_t offset) const {
assert_invariant(offset % mSlotSize == 0);
// The ID is 1-based since we use 0 for UNALLOCATED.
return (offset / mSlotSize) + 1;
}
BufferAllocator::allocation_size_t BufferAllocator::alignUp(
allocation_size_t size) const noexcept {
if (size == 0) return 0;
return (size + mSlotSize - 1) & ~(mSlotSize - 1);
}
} // namespace filament

View File

@@ -0,0 +1,118 @@
/*
* Copyright (C) 2025 The Android Open Source Project
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#ifndef TNT_FILAMENT_DETAILS_BUFFERALLOCATOR_H
#define TNT_FILAMENT_DETAILS_BUFFERALLOCATOR_H
#include <cstdint>
#include <list>
#include <map>
#include <unordered_map>
namespace filament {
// This class is NOT thread-safe.
//
// It internally manages shared state (e.g., mSlotPool, mFreeList, mOffsetMap) without any
// synchronization primitives. Concurrent access from multiple threads to the same
// BufferAllocator instance will result in data races and undefined behavior.
//
// If an instance of this class is to be shared between threads, all calls to its member
// functions MUST be protected by external synchronization (e.g., a std::mutex).
class BufferAllocator {
public:
using allocation_size_t = uint32_t;
using AllocationId = uint32_t;
static constexpr AllocationId UNALLOCATED = 0;
static constexpr AllocationId REALLOCATION_REQUIRED = ~0u;
struct Slot {
const allocation_size_t offset; // 4 bytes
allocation_size_t slotSize; // 4 bytes
bool isAllocated; // 1 byte
char padding[3]; // 3 bytes
uint32_t gpuUseCount; // 4 bytes
[[nodiscard]] bool isFree() const noexcept {
return !isAllocated && gpuUseCount == 0;
}
};
// `slotSize` is derived from the GPU's uniform buffer offset alignment requirement,
// which can be up to 256 bytes.
explicit BufferAllocator(allocation_size_t totalSize,
allocation_size_t slotSize);
BufferAllocator(BufferAllocator const&) = delete;
BufferAllocator(BufferAllocator&&) = delete;
// Allocate a new slot and return its id and slot offset in the UBO.
// If the returned id is not valid, that means there's no large enough slot for allocation.
[[nodiscard]] std::pair<AllocationId, allocation_size_t> allocate(
allocation_size_t size) noexcept;
// Call it when MaterialInstance gives up the ownership of the allocation.
// We don't release the slot immediately in this function even if it is not being used,
// the release is centralized in releaseFreeSlots().
void retire(AllocationId id);
// Increments the GPU read-lock.
void acquireGpu(AllocationId id);
// Decrements the GPU read-lock.
// We don't release the slot immediately in this function even if it is not being used,
// the release is centralized in releaseFreeSlots().
void releaseGpu(AllocationId id);
// Traverse all slots and free all slots that are not being used by both CPU and GPU.
// Perform the merge at the same time.
void releaseFreeSlots();
// Resets the allocator to its initial state with a new total size.
// All existing allocations are cleared.
void reset(allocation_size_t newTotalSize);
// Size of the UBO in bytes.
[[nodiscard]] allocation_size_t getTotalSize() const noexcept;
// Query the allocation offset by AllocationId.
allocation_size_t getAllocationOffset(AllocationId id) const;
private:
[[nodiscard]] AllocationId calculateIdByOffset(allocation_size_t offset) const;
[[nodiscard]] allocation_size_t alignUp(allocation_size_t size) const noexcept;
// Having an internal node type holding the base slot node and additional information.
struct InternalSlotNode {
Slot mSlot;
std::list<InternalSlotNode>::iterator mSlotPoolIterator;
std::multimap<allocation_size_t, InternalSlotNode*>::iterator mFreeListIterator;
std::unordered_map<allocation_size_t, InternalSlotNode*>::iterator mOffsetMapIterator;
};
[[nodiscard]] InternalSlotNode* getNodeById(AllocationId id) const noexcept;
allocation_size_t mTotalSize;
const allocation_size_t mSlotSize; // Size of a single slot in bytes.
std::list<InternalSlotNode> mSlotPool; // All slots, including both allocated and freed
std::multimap</*slot size*/allocation_size_t, InternalSlotNode*> mFreeList;
std::unordered_map</*slot offset*/allocation_size_t, InternalSlotNode*> mOffsetMap;
};
} // namespace filament
#endif // TNT_FILAMENT_DETAILS_BUFFERALLOCATOR_H

View File

@@ -776,6 +776,7 @@ public:
} backend;
struct {
bool check_crc32_after_loading = false;
bool enable_material_instance_uniform_batching = false;
} material;
} features;
@@ -825,6 +826,9 @@ public:
{ "material.check_crc32_after_loading",
"Verify the checksum of package data when a material is loaded.",
&features.material.check_crc32_after_loading, false },
{ "material.enable_material_instance_uniform_batching",
"Make all MaterialInstances share a common large uniform buffer and use sub-allocations within it.",
&features.material.enable_material_instance_uniform_batching, false },
}};
utils::Slice<const FeatureFlag> getFeatureFlags() const noexcept {

View File

@@ -60,8 +60,13 @@ namespace filament {
using namespace backend;
FMaterialInstance::FMaterialInstance(FEngine& engine, FMaterial const* material,
const char* name) noexcept
: mMaterial(material),
const char* name) noexcept
: FMaterialInstance(engine, material, name,
engine.features.material.enable_material_instance_uniform_batching) {
}
FMaterialInstance::FMaterialInstance(FEngine& engine, FMaterial const* material,
const char* name, bool useUboBatching) noexcept : mMaterial(material),
mDescriptorSet("MaterialInstance", material->getDescriptorSetLayout()),
mCulling(CullingMode::BACK),
mShadowCulling(CullingMode::BACK),
@@ -71,6 +76,7 @@ FMaterialInstance::FMaterialInstance(FEngine& engine, FMaterial const* material,
mHasScissor(false),
mIsDoubleSided(false),
mIsDefaultInstance(false),
mUseUboBatching(useUboBatching),
mTransparencyMode(TransparencyMode::DEFAULT),
mName(name ? CString(name) : material->getName()) {
@@ -80,12 +86,16 @@ FMaterialInstance::FMaterialInstance(FEngine& engine, FMaterial const* material,
// expected by the per-material descriptor-set layout
size_t const uboSize = std::max(size_t(16), material->getUniformInterfaceBlock().getSize());
mUniforms = UniformBuffer(uboSize);
mUbHandle = driver.createBufferObject(mUniforms.getSize(), BufferObjectBinding::UNIFORM,
BufferUsage::STATIC, utils::ImmutableCString{ material->getName().c_str_safe() });
// set the UBO, always descriptor 0
mDescriptorSet.setBuffer(material->getDescriptorSetLayout(),
0, mUbHandle, 0, mUniforms.getSize());
if (mUseUboBatching) {
mUboData = BufferAllocator::UNALLOCATED;
} else {
mUboData = driver.createBufferObject(mUniforms.getSize(), BufferObjectBinding::UNIFORM,
BufferUsage::STATIC, ImmutableCString{ material->getName().c_str_safe() });
// set the UBO, always descriptor 0
mDescriptorSet.setBuffer(material->getDescriptorSetLayout(),
0, std::get<Handle<HwBufferObject>>(mUboData), 0, mUniforms.getSize());
}
const RasterState& rasterState = material->getRasterState();
// At the moment, only MaterialInstances have a stencil state, but in the future it should be
@@ -140,6 +150,7 @@ FMaterialInstance::FMaterialInstance(FEngine& engine,
mHasScissor(false),
mIsDoubleSided(other->mIsDoubleSided),
mIsDefaultInstance(false),
mUseUboBatching(other->mUseUboBatching),
mScissorRect(other->mScissorRect),
mName(name ? CString(name) : other->mName) {
@@ -147,12 +158,16 @@ FMaterialInstance::FMaterialInstance(FEngine& engine,
FMaterial const* const material = other->getMaterial();
mUniforms.setUniforms(other->getUniformBuffer());
mUbHandle = driver.createBufferObject(mUniforms.getSize(), BufferObjectBinding::UNIFORM,
BufferUsage::DYNAMIC, ImmutableCString{ material->getName().c_str_safe() });
// set the UBO, always descriptor 0
mDescriptorSet.setBuffer(mMaterial->getDescriptorSetLayout(),
0, mUbHandle, 0, mUniforms.getSize());
if (mUseUboBatching) {
mUboData = BufferAllocator::UNALLOCATED;
} else {
mUboData = driver.createBufferObject(mUniforms.getSize(), BufferObjectBinding::UNIFORM,
BufferUsage::DYNAMIC, ImmutableCString{ material->getName().c_str_safe() });
// set the UBO, always descriptor 0
mDescriptorSet.setBuffer(material->getDescriptorSetLayout(),
0, std::get<Handle<HwBufferObject>>(mUboData), 0, mUniforms.getSize());
}
if (material->hasDoubleSidedCapability()) {
setDoubleSided(mIsDoubleSided);
@@ -173,7 +188,7 @@ FMaterialInstance::FMaterialInstance(FEngine& engine,
material->getId(), material->generateMaterialInstanceId());
// If the original descriptor set has been commited, the copy needs to commit as well.
if (other->mDescriptorSet.getHandle()) {
if (!mUseUboBatching && other->mDescriptorSet.getHandle()) {
mDescriptorSet.commitSlow(mMaterial->getDescriptorSetLayout(), driver);
}
}
@@ -190,13 +205,16 @@ FMaterialInstance::~FMaterialInstance() noexcept = default;
void FMaterialInstance::terminate(FEngine& engine) {
FEngine::DriverApi& driver = engine.getDriverApi();
mDescriptorSet.terminate(driver);
driver.destroyBufferObject(mUbHandle);
auto* ubHandle = std::get_if<Handle<HwBufferObject>>(&mUboData);
if (ubHandle){
driver.destroyBufferObject(*ubHandle);
}
}
void FMaterialInstance::commitStreamUniformAssociations(FEngine::DriverApi& driver) {
mHasStreamUniformAssociations = false;
if (!mTextureParameters.empty()) {
backend::BufferObjectStreamDescriptor descriptor;
BufferObjectStreamDescriptor descriptor;
for (auto const& [binding, p]: mTextureParameters) {
ssize_t offset = mMaterial->getUniformInterfaceBlock().getTransformFieldOffset(binding);
if (offset >= 0) {
@@ -206,7 +224,12 @@ void FMaterialInstance::commitStreamUniformAssociations(FEngine::DriverApi& driv
}
}
if (descriptor.mStreams.size() > 0) {
driver.registerBufferObjectStreams(mUbHandle, std::move(descriptor));
// UBO batching is incompatible with stream uniform associations because streams require a
// dedicated UBO handle, not a sub-allocation. Assert that any instance here uses its own UBO.
assert_invariant(!mUseUboBatching);
auto const* ubHandle = std::get_if<Handle<HwBufferObject>>(&mUboData);
assert_invariant(ubHandle);
driver.registerBufferObjectStreams(*ubHandle, std::move(descriptor));
}
}
}
@@ -217,11 +240,17 @@ void FMaterialInstance::commit(FEngine& engine) const {
}
}
void FMaterialInstance::commit(DriverApi& driver) const {
// update uniforms if needed
void FMaterialInstance::commit(FEngine::DriverApi& driver) const {
if (mUniforms.isDirty() || mHasStreamUniformAssociations) {
mUniforms.clean();
driver.updateBufferObject(mUbHandle, mUniforms.toBufferDescriptor(driver), 0);
if (mUseUboBatching) {
// TODO: update the content by `copyToMemoryMappedBuffer`
}
else {
auto* ubHandle = std::get_if<Handle<HwBufferObject>>(&mUboData);
assert_invariant(ubHandle != nullptr);
driver.updateBufferObject(*ubHandle, mUniforms.toBufferDescriptor(driver), 0);
}
}
if (!mTextureParameters.empty()) {
for (auto const& [binding, p]: mTextureParameters) {
@@ -411,6 +440,21 @@ void FMaterialInstance::use(FEngine::DriverApi& driver, Variant variant) const {
mDescriptorSet.bind(driver, DescriptorSetBindingPoints::PER_MATERIAL);
}
void FMaterialInstance::assignUboAllocation(
const Handle<HwBufferObject>& ubHandle,
BufferAllocator::AllocationId id,
BufferAllocator::allocation_size_t offset) {
assert_invariant(mUseUboBatching);
mUboData = id;
mDescriptorSet.setBuffer(mMaterial->getDescriptorSetLayout(), 0, ubHandle, offset,
mUniforms.getSize());
}
BufferAllocator::AllocationId FMaterialInstance::getAllocationId() const noexcept {
auto const* allocationId = std::get_if<BufferAllocator::AllocationId>(&mUboData);
return allocationId ? *allocationId : BufferAllocator::UNALLOCATED;
}
void FMaterialInstance::fixMissingSamplers() const {
// Here we check that all declared sampler parameters are set, this is required by
// Vulkan and Metal; GL is more permissive. If a sampler parameter is not set, we will

View File

@@ -23,6 +23,7 @@
#include "ds/DescriptorSet.h"
#include "details/BufferAllocator.h"
#include "details/Engine.h"
#include "private/backend/DriverApi.h"
@@ -57,6 +58,9 @@ class FMaterialInstance : public MaterialInstance {
public:
FMaterialInstance(FEngine& engine, FMaterial const* material,
const char* name) noexcept;
// Use this constructor when you need to override the ubo batching flag for an individual MI.
FMaterialInstance(FEngine& engine, FMaterial const* material,
const char* name, bool useUboBatching) noexcept;
FMaterialInstance(FEngine& engine, FMaterialInstance const* other, const char* name);
FMaterialInstance(const FMaterialInstance& rhs) = delete;
FMaterialInstance& operator=(const FMaterialInstance& rhs) = delete;
@@ -75,6 +79,12 @@ public:
void use(FEngine::DriverApi& driver, Variant variant = {}) const;
void assignUboAllocation(const backend::Handle<backend::HwBufferObject>& ubHandle,
BufferAllocator::AllocationId id,
BufferAllocator::allocation_size_t offset);
BufferAllocator::AllocationId getAllocationId() const noexcept;
FMaterial const* getMaterial() const noexcept { return mMaterial; }
uint64_t getSortingKey() const noexcept { return mMaterialSortingKey; }
@@ -234,6 +244,8 @@ public:
return mIsDefaultInstance;
}
bool isUsingUboBatching() const noexcept { return mUseUboBatching; }
// Called by the engine to ensure that unset samplers are initialized with placedholders.
void fixMissingSamplers() const;
@@ -274,7 +286,7 @@ private:
backend::SamplerParams params;
};
backend::Handle<backend::HwBufferObject> mUbHandle;
std::variant<BufferAllocator::AllocationId, backend::Handle<backend::HwBufferObject>> mUboData;
tsl::robin_map<backend::descriptor_binding_t, TextureParameter> mTextureParameters;
mutable DescriptorSet mDescriptorSet;
UniformBuffer mUniforms;
@@ -296,6 +308,7 @@ private:
bool mHasScissor : 1;
bool mIsDoubleSided : 1;
bool mIsDefaultInstance : 1;
bool mUseUboBatching : 1;
TransparencyMode mTransparencyMode : 2;
uint64_t mMaterialSortingKey = 0;

View File

@@ -43,6 +43,8 @@ list(APPEND RESGEN_SOURCE ${DUMMY_SRC})
if (TNT_DEV)
add_executable(test_${TARGET}
filament_AtlasAllocator_test.cpp
test_BufferAllocator.cpp
test_BufferAllocatorStress.cpp
test_CircularQueue.cpp
filament_test_exposure.cpp
filament_rendering_test.cpp

View File

@@ -0,0 +1,415 @@
/*
* Copyright (C) 2025 The Android Open Source Project
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#include <gtest/gtest.h>
#include "../src/details/BufferAllocator.h"
#include "utils/Panic.h"
#include <utility>
#include <vector>
using namespace filament;
namespace {
class BufferAllocatorTest : public ::testing::Test {
protected:
// We use a total size of 1024 and a slot size (alignment) of 64.
// This gives us 1024 / 64 = 16 total possible slots if aligned.
static constexpr BufferAllocator::allocation_size_t TOTAL_SIZE = 1024;
static constexpr BufferAllocator::allocation_size_t SLOT_SIZE = 64;
BufferAllocatorTest() : mAllocator(TOTAL_SIZE, SLOT_SIZE) {}
BufferAllocator mAllocator;
};
TEST_F(BufferAllocatorTest, ConstructorFailure) {
// The constructor requires slotSize to be a power of two.
constexpr BufferAllocator::allocation_size_t NON_POT_SLOT_SIZE = 60;
EXPECT_DEATH(BufferAllocator(TOTAL_SIZE, NON_POT_SLOT_SIZE), "failed assertion");
}
TEST_F(BufferAllocatorTest, InitialState) {
EXPECT_EQ(mAllocator.getTotalSize(), TOTAL_SIZE);
// Initially, there should be one large free block of the total size.
// Let's try to allocate the whole thing.
auto [id, offset] = mAllocator.allocate(TOTAL_SIZE);
EXPECT_EQ(id, 1); // The first ID should be 1.
EXPECT_EQ(offset, 0);
}
TEST_F(BufferAllocatorTest, SimpleAllocation) {
// Allocate 100 bytes, which should be aligned up to 128 (2 * 64).
auto [id, offset] = mAllocator.allocate(100);
EXPECT_EQ(id, 1);
EXPECT_EQ(offset, 0);
EXPECT_EQ(mAllocator.getAllocationOffset(id), offset);
// Try to allocate again. The next allocation should start after the first one.
auto [id2, offset2] = mAllocator.allocate(50); // Aligns to 64
EXPECT_EQ(id2, 3); // ID is (128 / 64) + 1 = 3
EXPECT_EQ(offset2, 128);
EXPECT_EQ(mAllocator.getAllocationOffset(id2), offset2);
}
TEST_F(BufferAllocatorTest, AllocateZeroSize) {
// Allocating zero bytes should return an unallocated ID and not affect the state.
auto [id, offset] = mAllocator.allocate(0);
EXPECT_EQ(id, BufferAllocator::UNALLOCATED);
EXPECT_EQ(offset, 0);
// The allocator should still be in its initial state, able to allocate the full size.
auto [fullId, fullOffset] = mAllocator.allocate(TOTAL_SIZE);
EXPECT_EQ(fullId, 1);
EXPECT_EQ(fullOffset, 0);
}
TEST_F(BufferAllocatorTest, AllocateAll) {
auto [id1, offset1] = mAllocator.allocate(512);
EXPECT_EQ(id1, 1);
EXPECT_EQ(offset1, 0);
auto [id2, offset2] = mAllocator.allocate(512);
EXPECT_EQ(id2, 9); // ID is (512 / 64) + 1 = 9
EXPECT_EQ(offset2, 512);
// The buffer is now full. The next allocation should fail.
auto [id3, offset3] = mAllocator.allocate(1);
EXPECT_EQ(id3, BufferAllocator::REALLOCATION_REQUIRED);
}
TEST_F(BufferAllocatorTest, AllocationFailure) {
auto [id1, offset1] = mAllocator.allocate(TOTAL_SIZE - 1); // Allocate most of the buffer.
EXPECT_EQ(id1, 1);
EXPECT_EQ(offset1, 0);
// Try to allocate more than the remaining space.
auto [id2, offset2] = mAllocator.allocate(100);
EXPECT_EQ(id2, BufferAllocator::REALLOCATION_REQUIRED);
EXPECT_EQ(offset2, 0);
}
TEST_F(BufferAllocatorTest, AllocationLifecycle) {
// 1. Allocate
auto [id, offset] = mAllocator.allocate(128);
EXPECT_EQ(id, 1);
// 2. Retire from CPU side
mAllocator.retire(id);
// 3. The slot is not free yet because the GPU might still be using it.
// Let's simulate the GPU acquiring and releasing it.
mAllocator.acquireGpu(id);
mAllocator.releaseGpu(id);
// Now the slot should be considered free, but it's not merged yet.
// Let's try to allocate something else to see where it goes.
auto [id2, offset2] = mAllocator.allocate(200); // Aligns to 256
EXPECT_EQ(id2, 3); // It should use the space after the first retired slot.
EXPECT_EQ(offset2, 128);
// Now, let's release the free slots.
mAllocator.releaseFreeSlots();
// The first slot (128 bytes) should now be available again.
// Let's try to allocate something that fits in it.
auto [id3, offset3] = mAllocator.allocate(100); // Aligns to 128
EXPECT_EQ(id3, 1); // It should reuse the first slot.
EXPECT_EQ(offset3, 0);
}
TEST_F(BufferAllocatorTest, RetireThenReleaseGpu) {
// 1. Allocate a block and a dummy block next to it.
auto [id1, offset1] = mAllocator.allocate(128);
auto [id2, offset2] = mAllocator.allocate(128);
EXPECT_EQ(id1, 1);
EXPECT_EQ(id2, 3);
EXPECT_EQ(offset1, 0);
EXPECT_EQ(offset2, 128);
// 2. Retire the first block (CPU is done), then acquire it for the GPU.
mAllocator.retire(id1);
mAllocator.acquireGpu(id1); // gpuUseCount = 1
// 3. The slot is now free from CPU but locked by GPU. releaseFreeSlots should do nothing to it.
mAllocator.releaseFreeSlots();
// 4. Try to allocate the same space. It should fail, and the allocation should go
// to the next available free space.
auto [id3, offset3] = mAllocator.allocate(64);
EXPECT_EQ(id3, 5);
EXPECT_EQ(offset3, 256); // It should be allocated after id2.
// 5. Now, release the GPU lock.
mAllocator.releaseGpu(id1); // gpuUseCount = 0
// 6. Call releaseFreeSlots again. This time it should be freed.
mAllocator.releaseFreeSlots();
// 7. The original slot should now be available for allocation.
auto [id4, offset4] = mAllocator.allocate(64);
EXPECT_EQ(id4, 1);
EXPECT_EQ(offset4, 0); // Success! It reuses the first slot.
}
TEST_F(BufferAllocatorTest, MultipleGpuAcquires) {
// 1. Allocate a block.
auto [id, offset] = mAllocator.allocate(128);
EXPECT_EQ(id, 1);
EXPECT_EQ(offset, 0);
// 2. Retire from CPU, then acquire multiple times for GPU (e.g., used in 3 command buffers).
mAllocator.retire(id);
mAllocator.acquireGpu(id); // gpuUseCount = 1
mAllocator.acquireGpu(id); // gpuUseCount = 2
mAllocator.acquireGpu(id); // gpuUseCount = 3
// 3. Release GPU lock once. The slot should still be locked.
mAllocator.releaseGpu(id); // gpuUseCount = 2
mAllocator.releaseFreeSlots();
auto [failId1, failOffset1] = mAllocator.allocate(64);
EXPECT_NE(failOffset1, 0); // Should not be able to allocate at offset 0.
EXPECT_NE(failId1, 1);
// 4. Release GPU lock again. The slot should still be locked.
mAllocator.releaseGpu(id); // gpuUseCount = 1
mAllocator.releaseFreeSlots();
auto [failId2, failOffset2] = mAllocator.allocate(64);
EXPECT_NE(failOffset2, 0); // Still cannot allocate at offset 0.
EXPECT_NE(failId2, 1);
// 5. Final release. The lock count is now 0.
mAllocator.releaseGpu(id); // gpuUseCount = 0
// 6. Now it should be freed and available.
mAllocator.releaseFreeSlots();
auto [successId, successOffset] = mAllocator.allocate(64);
EXPECT_EQ(successOffset, 0);
EXPECT_EQ(successId, 1);
}
TEST_F(BufferAllocatorTest, GpuPanicOnUnderflow) {
// 1. Allocate a block and acquire it.
auto [id, _] = mAllocator.allocate(128);
mAllocator.acquireGpu(id);
// 2. Release it once, which is fine.
mAllocator.releaseGpu(id);
// 3. Releasing it again when the count is 0 should trigger a failed assertion.
EXPECT_DEATH(mAllocator.releaseGpu(id), "failed assertion");
}
TEST_F(BufferAllocatorTest, MergeFreeSlots) {
// Allocate three blocks
auto [id1, offset1] = mAllocator.allocate(128); // Slot 0-127
auto [id2, offset2] = mAllocator.allocate(128); // Slot 128-255
auto [id3, offset3] = mAllocator.allocate(128); // Slot 256-383
EXPECT_EQ(id1, 1);
EXPECT_EQ(id2, 3);
EXPECT_EQ(id3, 5);
EXPECT_EQ(offset1, 0);
EXPECT_EQ(offset2, 128);
EXPECT_EQ(offset3, 256);
// Retire the first and third blocks
mAllocator.retire(id1);
mAllocator.retire(id3);
// At this point, we have: [Free, Allocated, Free, Free (remaining)]
// releaseFreeSlots should not merge slot 1 and 3 because they are not adjacent.
// It should merge slot 3 and slot 4 instead.
mAllocator.releaseFreeSlots();
// At this point, we have: [Free, Allocated, Free (remaining)]
// Let's verify by trying to allocate 200. It should go into the merged slot.
auto [id4, offset4] = mAllocator.allocate(200); // Aligns to 256
EXPECT_EQ(id4, 5);
EXPECT_EQ(offset4, 256);
// Now, retire the middle block
mAllocator.retire(id2);
// Now we have: [Free, Free, Allocated, Free (remaining)]
// The first two blocks are now adjacent and free.
mAllocator.releaseFreeSlots();
// After merging, the first 256 bytes should be one large free block.
// Let's try to allocate something that requires this merged space.
auto [id5, offset5] = mAllocator.allocate(200); // Aligns to 256
EXPECT_EQ(id5, 1);
EXPECT_EQ(offset5, 0);
}
TEST_F(BufferAllocatorTest, MergeAllSlots) {
// Allocate the entire buffer in small chunks.
constexpr BufferAllocator::allocation_size_t CHUNK_SIZE = 128;
constexpr uint32_t NUM_CHUNKS = TOTAL_SIZE / CHUNK_SIZE;
std::vector<BufferAllocator::AllocationId> ids;
for (uint32_t i = 0; i < NUM_CHUNKS; ++i) {
auto [id, offset] = mAllocator.allocate(CHUNK_SIZE);
ASSERT_NE(id, BufferAllocator::REALLOCATION_REQUIRED);
ids.push_back(id);
}
// The buffer should be full.
auto [failId, failOffset] = mAllocator.allocate(1);
EXPECT_EQ(failId, BufferAllocator::REALLOCATION_REQUIRED);
EXPECT_EQ(failOffset, 0);
// Retire all chunks.
for (auto id : ids) {
mAllocator.retire(id);
}
// Release and merge.
mAllocator.releaseFreeSlots();
// Now, the allocator should be back to its initial state with one large free block.
// We should be able to allocate the entire buffer again.
auto [fullId, fullOffset] = mAllocator.allocate(TOTAL_SIZE);
EXPECT_EQ(fullId, 1);
EXPECT_EQ(fullOffset, 0);
}
TEST_F(BufferAllocatorTest, NoMergePossible) {
// Allocate blocks in an alternating pattern.
auto [id1, offset1] = mAllocator.allocate(128);
auto [id2, offset2] = mAllocator.allocate(128);
auto [id3, offset3] = mAllocator.allocate(128);
auto [id4, offset4] = mAllocator.allocate(128);
EXPECT_EQ(id1, 1);
EXPECT_EQ(id2, 3);
EXPECT_EQ(id3, 5);
EXPECT_EQ(id4, 7);
EXPECT_EQ(offset1, 0);
EXPECT_EQ(offset2, 128);
EXPECT_EQ(offset3, 256);
EXPECT_EQ(offset4, 384);
// Retire the first and third blocks, leaving allocated blocks in between.
mAllocator.retire(id1);
mAllocator.retire(id3);
// At this point, we have: [Free, Allocated, Free, Allocated, Free (remaining)]
mAllocator.releaseFreeSlots();
// Now, let's try to re-allocate the two 128-byte slots.
auto [id5, offset5] = mAllocator.allocate(128);
auto [id6, offset6] = mAllocator.allocate(128);
EXPECT_EQ(offset5, 0);
EXPECT_EQ(offset6, 256);
}
TEST_F(BufferAllocatorTest, Reset) {
auto [id1, offset1] = mAllocator.allocate(100);
auto [id2, offset2] = mAllocator.allocate(200);
EXPECT_EQ(id1, 1);
EXPECT_EQ(id2, 3);
EXPECT_EQ(offset1, 0);
EXPECT_EQ(offset2, 128);
// Reset the allocator to a new size.
mAllocator.reset(2048);
EXPECT_EQ(mAllocator.getTotalSize(), 2048);
// After reset, all previous allocations should be gone,
// and we should be able to allocate the entire new size.
auto [id3, offset3] = mAllocator.allocate(2048);
EXPECT_EQ(id3, 1);
EXPECT_EQ(offset3, 0);
}
TEST_F(BufferAllocatorTest, ResetWithInvalidSize) {
// Reset to a size which is not a power of two.
EXPECT_DEATH(mAllocator.reset(123), "failed assertion");
}
TEST_F(BufferAllocatorTest, ResetWithGpuLock) {
// 1. Allocate a block and acquire a GPU lock on it.
auto [id1, offset1] = mAllocator.allocate(128);
EXPECT_EQ(id1, 1);
mAllocator.acquireGpu(id1); // gpuUseCount = 1
// 2. Call reset. This should disregard the GPU lock and clear everything.
constexpr BufferAllocator::allocation_size_t NEW_TOTAL_SIZE = 4096;
mAllocator.reset(NEW_TOTAL_SIZE);
// 3. Verify the allocator is in a pristine state with the new size.
EXPECT_EQ(mAllocator.getTotalSize(), NEW_TOTAL_SIZE);
// 4. The strongest verification is to allocate the entire new size, which should
// succeed, proving that the old GPU-locked block is gone.
auto [id2, offset2] = mAllocator.allocate(NEW_TOTAL_SIZE);
EXPECT_EQ(id2, 1);
EXPECT_EQ(offset2, 0);
}
TEST_F(BufferAllocatorTest, InvalidOperations) {
// These operations on invalid IDs should not crash and should be handled gracefully.
EXPECT_DEATH(mAllocator.retire(BufferAllocator::UNALLOCATED), "failed assertion");
EXPECT_DEATH(mAllocator.retire(999), "failed assertion"); // Non-existent ID
EXPECT_DEATH(mAllocator.acquireGpu(BufferAllocator::UNALLOCATED), "failed assertion");
EXPECT_DEATH(mAllocator.acquireGpu(999), "failed assertion");
EXPECT_DEATH(mAllocator.releaseGpu(BufferAllocator::UNALLOCATED), "failed assertion");
EXPECT_DEATH(mAllocator.releaseGpu(999), "failed assertion");
// Check that an invalid offset query panics in debug/testing builds.
EXPECT_DEATH(mAllocator.getAllocationOffset(BufferAllocator::UNALLOCATED), "failed assertion");
EXPECT_DEATH(mAllocator.getAllocationOffset(BufferAllocator::REALLOCATION_REQUIRED),
"failed assertion");
}
TEST_F(BufferAllocatorTest, ComplexScenario) {
std::vector<BufferAllocator::AllocationId> ids;
// 1. Allocate 4 blocks of 256 bytes
for (int i = 0; i < 4; ++i) {
auto [id, offset] = mAllocator.allocate(256);
EXPECT_EQ(offset, i * 256);
ids.push_back(id);
}
// Buffer should be full
auto [failId1, failOffset1] = mAllocator.allocate(1);
EXPECT_EQ(failId1, BufferAllocator::REALLOCATION_REQUIRED);
// 2. Retire the 2nd and 3rd blocks (ids[1] and ids[2])
mAllocator.retire(ids[1]);
mAllocator.retire(ids[2]);
// 3. Release free slots. This should merge the two retired blocks.
mAllocator.releaseFreeSlots();
// We now have a free block of 512 bytes at offset 256.
// Let's allocate 512 bytes. It should fit perfectly.
auto [id, offset] = mAllocator.allocate(512);
EXPECT_EQ(id, 5); // (256 / 64) + 1 = 5
EXPECT_EQ(offset, 256);
// 4. The buffer should be full again.
auto [failId2,failOffset2] = mAllocator.allocate(1);
EXPECT_EQ(failId2, BufferAllocator::REALLOCATION_REQUIRED);
EXPECT_EQ(failOffset2,0);
}
} // anonymous namespace

View File

@@ -0,0 +1,122 @@
/*
* Copyright (C) 2025 The Android Open Source Project
*
* Licensed under the Apache License, Version 2.0 (the "License");
* you may not use this file except in compliance with the License.
* You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing, software
* distributed under the License is distributed on an "AS IS" BASIS,
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
* See the License for the specific language governing permissions and
* limitations under the License.
*/
#include <gtest/gtest.h>
#include "../src/details/BufferAllocator.h"
#include "utils/Panic.h"
#include <algorithm>
#include <random>
#include <vector>
using namespace filament;
namespace {
class BufferAllocatorStressTest : public ::testing::Test {
protected:
// We use a total size of 1024 * 64 and a slot size (alignment) of 64.
// This gives us 1024 total possible slots if aligned.
static constexpr BufferAllocator::allocation_size_t SLOT_SIZE = 64;
static constexpr BufferAllocator::allocation_size_t SLOT_COUNT = 4096;
static constexpr BufferAllocator::allocation_size_t TOTAL_SIZE = SLOT_COUNT * SLOT_SIZE;
BufferAllocatorStressTest() : mAllocator(TOTAL_SIZE, SLOT_SIZE) {
}
BufferAllocator mAllocator;
};
TEST_F(BufferAllocatorStressTest, StressTest) {
// Many operations.
constexpr int operationCount = 5000;
// 1. Prepare the random number distributions.
std::random_device rd;
std::mt19937 gen(rd());
std::uniform_int_distribution<> slotCountDistrib(1, SLOT_COUNT);
// Random the slot offset we allocate within a slot.
std::uniform_int_distribution<> slotOffsetDistrib(0, SLOT_SIZE - 1);
// Random the operation we perform: 0,1,2 -> allocate, 3 -> retire, 4 -> release slots
// We bias the operations to make more allocations than releases.
std::uniform_int_distribution operationDistrib(0, 4);
// 2. Randomly do some actions and expect to have no crash
std::vector<BufferAllocator::AllocationId> ids;
for (int i = 0; i < operationCount; i++) {
switch (operationDistrib(gen)) {
case 0: // Allocate
case 1: // Allocate
case 2: // Allocate
{
// Random the slot count + its offset.
const BufferAllocator::allocation_size_t size =
slotCountDistrib(gen) * SLOT_SIZE - slotOffsetDistrib(gen);
if (auto [id, _] = mAllocator.allocate(size);
id != BufferAllocator::REALLOCATION_REQUIRED && id !=
BufferAllocator::UNALLOCATED) {
ids.push_back(id);
}
break;
}
case 3: // Allocate
{
if (ids.empty()) {
continue;
}
// Retire a random slot.
std::uniform_int_distribution<uint32_t> idDistrib(0, ids.size() - 1);
uint32_t indexToRetire = idDistrib(gen);
BufferAllocator::AllocationId idToRetire = ids[indexToRetire];
mAllocator.retire(idToRetire);
// Remove the retired id from the list.
std::swap(ids[indexToRetire], ids.back());
ids.pop_back();
break;
}
case 4: // Allocate
{
mAllocator.releaseFreeSlots();
break;
}
default:
break;
}
}
// 3. Retire all remaining allocations.
for (const auto& id: ids) {
mAllocator.retire(id);
}
ids.clear();
// 4. Release and merge everything.
mAllocator.releaseFreeSlots();
// 5. The allocator should now be in a pristine state.
// A final allocation of the total size should succeed.
auto [finalId, finalOffset] = mAllocator.allocate(TOTAL_SIZE);
EXPECT_NE(finalId, BufferAllocator::REALLOCATION_REQUIRED);
EXPECT_NE(finalId, BufferAllocator::UNALLOCATED);
EXPECT_EQ(finalOffset, 0);
}
} // anonymous namespace

View File

@@ -459,13 +459,6 @@ utils::Status MaterialParser::processTemplateSubstitutions(
utils::Status MaterialParser::parse(filamat::MaterialBuilder& builder,
const Config& config, ssize_t& size, std::unique_ptr<const char[]>& buffer) {
if (builder.getFeatureLevel() > config.getFeatureLevel()) {
utils::io::sstream errorMessage;
errorMessage << "Material feature level (" << +builder.getFeatureLevel()
<< ") is higher than maximum allowed (" << +config.getFeatureLevel() << ")";
return utils::Status::invalidArgument(errorMessage.c_str());
}
// Before attempting an expensive lex, let's find out if we were sent pure JSON.
utils::Status parsedStatus;
if (utils::Status validJson = isValidJsonStart(buffer.get(), size_t(size));
@@ -479,6 +472,13 @@ utils::Status MaterialParser::parse(filamat::MaterialBuilder& builder,
return parsedStatus;
}
if (builder.getFeatureLevel() > config.getFeatureLevel()) {
utils::io::sstream errorMessage;
errorMessage << "Material feature level (" << +builder.getFeatureLevel()
<< ") is higher than maximum allowed (" << +config.getFeatureLevel() << ")";
return utils::Status::invalidArgument(errorMessage.c_str());
}
switch (config.getReflectionTarget()) {
case Config::Metadata::NONE:
break;

View File

@@ -191,6 +191,8 @@ static int handleCommandLineArgments(int argc, char* argv[], Config* config) {
}
static void cleanup(Engine* engine, View*, Scene*) {
g_meshSet.reset(nullptr);
for (const auto& material : g_meshMaterialInstances) {
engine->destroy(material.second);
}
@@ -203,8 +205,6 @@ static void cleanup(Engine* engine, View*, Scene*) {
engine->destroy(i);
}
g_meshSet.reset(nullptr);
engine->destroy(g_params.light);
engine->destroy(g_params.spotLight);
engine->destroy(g_colorGrading);

View File

@@ -1,6 +1,6 @@
# Rendering Difference Test
We created a few scripts to run `gltf_viewer` and produce headless renderings.
This tool is a collections of scripts to run `gltf_viewer` and produce headless renderings.
This is mainly useful for continuous integration where GPUs are generally not available on cloud
machines. To perform software rasterization, these scripts are centered around [Mesa]'s
@@ -9,10 +9,10 @@ Additionally, we should be able to use GPUs where available (though this is more
work).
The script `render.py` contains the core logic for taking input parameters (such as the test
description file) and then running gltf_viewer to produce the renderings.
description file) and then running `gltf_viewer` to produce the renderings.
In the `test` directory is a list of test descriptions that are specified in json. Please see
`sample.json` to parse the structure.
`sample.json` to glean the structure.
## Setting up python

View File

@@ -229,7 +229,7 @@ if __name__ == '__main__':
important_print(f'Successfully compared {success_count} / {len(results)} images')
if tolerance_used_details:
pstr = 'Tolerance-based passes:'
pstr = 'Passed:'
for detail in tolerance_used_details:
pstr += '\n' + detail
important_print(pstr)
@@ -237,7 +237,7 @@ if __name__ == '__main__':
if failed_details:
pstr = 'Failed:'
for detail in failed_details:
pstr = '\n' + detail
pstr += '\n' + detail
important_print(pstr)
if len(failed) > 0:
exit(1)

View File

@@ -74,7 +74,7 @@ def _get_deletes_updates(update_dir, golden_dir):
for fpath in base.intersection(new):
base_fpath = os.path.join(golden_dir, fpath)
new_fpath = os.path.join(update_dir, fpath)
if (ext == 'tif' and not same_image(new_fpath, base_fpath)) or \
if (ext == 'tif' and not same_image(new_fpath, base_fpath)[0]) or \
(ext == 'json' and _file_as_str(new_fpath) != _file_as_str(base_fpath)):
update.append(fpath)

View File

@@ -1,12 +1,13 @@
AlphaBlendModeTest
AttenuationTest
Box
BoomBoxWithAxes
Box
BoxInterleaved
BoxTextured
BoxTexturedNonPowerOfTwo
Duck
IridescenceSuzanne
IridescentDishWithOlives
Lantern
MetalRoughSpheres
NegativeScaleTest
@@ -17,4 +18,5 @@ Sponza
Suzanne
TextureCoordinateTest
TextureSettingsTest
TransmissionRoughnessTest
TwoSidedPlane

View File

@@ -5,9 +5,9 @@
"presets": [
{
"name": "base",
"models": ["Box", "BoxTextured", "Duck", "lucy", "FlightHelmet"],
"models": ["Box", "BoxTextured", "FlightHelmet", "lucy"],
"rendering": {
"viewer.cameraFocusDistance": 0,
"viewer.cameraFocalLength": 35.0,
"view.postProcessingEnabled": true,
"view.dithering": "NONE"
},
@@ -17,27 +17,54 @@
}
},
{
"name": "helmet_only",
"name": "bloom_models",
"models": ["DamagedHelmet"],
"rendering": {}
},
{
"name": "ssao_models",
"models": ["Duck", "lucy"],
"rendering": {}
},
{
"name": "transmission_models",
"models": ["TransmissionRoughnessTest", "IridescentDishWithOlives"],
"rendering": {}
}
],
"tests": [
{
"name": "Bloom",
"description": "Testing bloom",
"apply_presets": ["base", "helmet_only"],
"description": "Bloom",
"apply_presets": ["base", "bloom_models"],
"rendering": {
"view.bloom.enabled": true
}
},
{
"name": "MSAA",
"description": "Testing multisampling anti-aliasing",
"description": "Multisampling anti-aliasing",
"apply_presets": ["base"],
"rendering": {
"view.msaa.enabled": true
}
},
{
"name": "Transimssion",
"description": "transmission",
"apply_presets": ["base", "transmission_models"],
"rendering": {
"viewer.cameraFocalLength": 52.0,
"view.screenSpaceReflections.enabled": true
}
},
{
"name": "SSAO",
"description": "Screen space ambient occulsion",
"apply_presets": ["base", "ssao_models"],
"rendering": {
"view.ssao.enabled": true
}
}
]
}

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

12
third_party/perfetto/tnt/README.md vendored Normal file
View File

@@ -0,0 +1,12 @@
To update the Perfetto library, run the following command from the tnt directory:
```shell
./update_perfetto.sh <version_tag>
```
For example:
```shell
./update_perfetto.sh v51.2
```
**Useful links:**
https://perfetto.dev/
https://perfetto.dev/docs/instrumentation/tracing-sdk

View File

@@ -1,8 +0,0 @@
Useful links:
https://perfetto.dev/
https://perfetto.dev/docs/instrumentation/tracing-sdk
Perfetto was fetched as:
git clone https://github.com/google/perfetto.git -b v50.1
Only the sdk/ directory is needed.

86
third_party/perfetto/tnt/update_perfetto.sh vendored Executable file
View File

@@ -0,0 +1,86 @@
#!/bin/bash
set -euo pipefail
function usage() {
echo "Usage: $0 [-h] <version_tag>"
echo " -h: Display this help message"
echo " <version_tag>: The Perfetto version to download (e.g., v50.1)"
exit 1
}
if [[ $# -ne 1 || "$1" == "-h" ]]; then
usage
fi
DOWNLOAD_REF=$1
DOWNLOAD_URL="https://github.com/google/perfetto/archive/refs/tags/${DOWNLOAD_REF}.zip"
ZIP_FILE_NAME="${DOWNLOAD_REF}.zip"
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
THIRD_PARTY_DIR="$(dirname "$(dirname "$SCRIPT_DIR")")"
if ! command -v curl &> /dev/null; then
echo "Error: curl is not installed." >&2
exit 1
fi
if ! command -v unzip &> /dev/null; then
echo "Error: unzip is not installed." >&2
exit 1
fi
if ! command -v rsync &> /dev/null; then
echo "Error: rsync is not installed." >&2
exit 1
fi
cd "${THIRD_PARTY_DIR}"
echo "Downloading Perfetto ${DOWNLOAD_REF}..."
if ! curl -f -L -o "${ZIP_FILE_NAME}" "${DOWNLOAD_URL}"; then
echo "Error: Failed to download Perfetto ${DOWNLOAD_REF}." >&2
echo "Please check the version number and your network connection." >&2
rm -f "${ZIP_FILE_NAME}"
exit 1
fi
echo "Unzipping..."
TEMP_DIR="perfetto_temp_unzip"
# Clean up temp dir from previous runs, if any.
rm -rf "${TEMP_DIR}"
mkdir "${TEMP_DIR}"
if ! unzip -q "${ZIP_FILE_NAME}" -d "${TEMP_DIR}"; then
echo "Error: Failed to unzip the downloaded file." >&2
rm -f "${ZIP_FILE_NAME}"
rm -rf "${TEMP_DIR}"
exit 1
fi
# The archive contains a single directory. Find it.
EXTRACTED_DIR=$(find "${TEMP_DIR}" -mindepth 1 -maxdepth 1 -type d)
if [ -z "${EXTRACTED_DIR}" ] || [ ! -d "${EXTRACTED_DIR}" ]; then
echo "Error: Could not find the extracted directory inside the archive." >&2
rm -f "${ZIP_FILE_NAME}"
rm -rf "${TEMP_DIR}"
exit 1
fi
echo "Replacing existing perfetto contents with new version..."
# If perfetto_new exists, remove it.
rm -rf perfetto_new
mv "${EXTRACTED_DIR}" perfetto_new
# The perfetto directory should contain the content of the sdk directory from the archive, and our tnt directory.
rsync -a --delete perfetto_new/sdk/ perfetto/perfetto/
echo "Cleaning up..."
rm -rf "${ZIP_FILE_NAME}" perfetto_new "${TEMP_DIR}"
echo "Staging changes..."
git add perfetto
echo "Done. Please commit the changes."