Publications and Presentations
PBRT 2: Creating and Working with Large Scenes
Today, Blender is undeniably one of the best tools out there for a large swathe of graphics related tasks. However, some work can be done to make it even more useful when working with the complex scenes that are often used in graphics research.
During BCON24, I demonstrated how to create basic importers and exporters with the Python API using the PBRT format as a basis. In this talk, I will present various improvements to the importer that allows PBRT scenes to be rendered with Cycles at a comparable image quality, and by extension, leverage the existing Blender ecosystem to export these scenes to other systems via other formats, such as glTF.
Additionally, as some of these scenes contain numerous, or complex objects, techniques for working with them in a way that maintains user-interface responsiveness are highlighted.
@online{WaldemarsonBCON26, title = {PBRT 2: Creating and Working with Large Scenes}, year = {2026}, organization = {The Blender Foundation, Youtube}, author = {Gustaf Waldemarson}, url = {https://youtu.be/iao8H8jPrqA}, }
Rendering Small Things: Hardware Micromaps and Particles
Faculty opponent: Associate Professor Jeppe Frisvad.
In computer graphics, there are numerous aspects that must be considered when rendering images of virtual scenes: What physical light-generating phenomena do we care about? How should object and material surfaces be described? And how should these be stored to ensure as fast and efficient image rendering as possible?
As a part of this thesis, a method for rendering images of scenes lit by virtual atomic particles traveling at superluminal speeds is presented that handles both the particle interactions and light generation in a unified ray-tracing framework. However, this form of ray-tracing can be very time-consuming. Thus, this work includes an investigation into accelerating one aspect of this process: Parallelizing the construction of the spatial split bounding volume hierarchy in a simple and straightforward way with the OpenMP framework. A similar ray-tracing process is then optimized for real-time rendering of partially-transparent triangle meshes by efficiently leveraging and compressing a structure known as micromaps that enables ray-tracing to work more efficiently for various alpha-masked geometries such as grass and foliage. This structure is subsequently generalized and extended to arbitrary surface attributes, with a thorough analysis of its performance, quality trade-offs, and potential future use-cases in the hardware accelerated real-time ray-tracing pipeline.
@phdthesis{WaldemarsonPhD2026, title = "Rendering Small Things: Hardware Micromaps and Particles", abstract = "In computer graphics, there are numerous aspects that must be considered when rendering images of virtual scenes: What physical light-generating phenomena do we care about? How should object and material surfaces be described? And how should these be stored to ensure as fast and efficient image rendering as possible? As a part of this thesis, a method for rendering images of scenes lit by virtual atomic particles traveling at superluminal speeds is presented that handles both the particle interactions and light generation in a unified ray-tracing framework. However, this form of ray-tracing can be very time-consuming. Thus, this work includes an investigation into accelerating one aspect of this process: Parallelizing the construction of the spatial split bounding volume hierarchy in a simple and straightforward way with the OpenMP framework. A similar ray-tracing process is then optimized for real-time rendering of partially-transparent triangle meshes by efficiently leveraging and compressing a structure known as micromaps that enables ray-tracing to work more efficiently for various alpha-masked geometries such as grass and foliage. This structure is subsequently generalized and extended to arbitrary surface attributes, with a thorough analysis of its performance, quality trade-offs, and potential future use-cases in the hardware accelerated real-time ray-tracing pipeline.", keywords = "Rendering, Ray-Tracing, Acceleration Structures, Micromaps, Particles", author = "Gustaf Waldemarson", year = "2026", month = apr, day = "24", language = "English", isbn = "978-91-8104-872-8", series = "Dissertation", publisher = "Department of Computer Science, Lund University", type = "Doctoral Thesis (compilation)", school = "Department of Computer Science", url = "https://portal.research.lu.se/en/publications/rendering-small-things-hardware-micromaps-and-particles/", }
Real-Time Path-Tracing of Ray-Tracing Benchmarks of Yore
Ray-tracing has long been used as tool for generating realistic images, and has accumulated a number of different benchmarks over the years, primarily focusing on generating high-quality images. However, we are now moving into the era of real-time animated ray-tracing, thus, some old benchmarks that, at the time, did not garner much attention, can be a good starting point for trialing techniques that target animated content. Thus, in this work we take a look back at the Benchmark for Animated Ray Tracing (BART) and adapts it to a world of real-time path-tracing with the help of the Spatiotemporal Variance-Guided Filtering (SVGF).
@online{WaldemarsonHPG25Poster, title = {Real-Time Path-Tracing of Ray-Tracing benchmarks of Yore}, author = {Gustaf Waldemarson}, note = {Poster presented during the High-Performance Graphics conference 2025 student competetition: \url{https://highperformancegraphics.org/2025/student-competition}}, url = {https://gustafwaldemarson.com/misc/hpg25/benchmarks.pdf}, year = {2025}, }
-
- Poster
Parallel Axis Split Tasks for Bounding Volume Construction with OpenMP
Proceedings of the 20th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications - (Volume 1)
Many algorithms in computer graphics make use of acceleration structures such as Bounding Volume Hierarchies (BVHs) to speed up performance critical tasks, such as collision detection or ray-tracing. However, while the typical algorithms for constructing BVHs are relatively simple, actually implementing them for performance critical systems is still challenging. Further, to construct them as quickly as possible, it is also desirable to parallelize the process. To that end, parallelization APIs such as OpenMP® can be leveraged to greatly simplify this matter. However, BVH construction is not a trivially parallelizable problem. Thus, in this paper we propose a method of using OpenMP® tasking to further parallelize the spatial splitting algorithm and thus improve construction performance. We evaluate the proposed way and compare it with other ways of using OpenMP®, finding that some of these work well to improve the construction time by between 3 and 5 times on an 8-core machine with a minimal amount of work and negligible quality reduction of the final BVH.
@conference{grapp25, author = {Gustaf Waldemarson and Michael Doggett}, title = {Parallel Axis Split Tasks for Bounding Volume Construction with OpenMP}, booktitle = {Proceedings of the 20th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications - GRAPP}, year = {2025}, pages = {347-354}, publisher = {SciTePress}, organization = {INSTICC}, doi = {10.5220/0013317100003912}, isbn = {978-989-758-728-3}, issn = {2184-4321}, }
PBRT: Create your own Importers and Exporters
Blender is arguably one of the best tools for performing a wide variety graphics work, and while Eevee and Cycles are great renderers, sometimes it is desirable to create a more specialized rendering engine for some particular task. However, getting data into these engines can sometimes be challenging. Thus, in this talk I will present some of the work I have done to connect Blender with the well known research renderer: PBRT, allowing Blender to export scenes in the native PBRT format or import existing PBRT scenes into Blender for further editing. Furthermore, this work should be easily adaptable such that you can use it you create your own importer, exporter or even renderer if you want to!
@online{waldemarsonBCON24, title = {PBRT: Create your own Importers and Exporters}, year = {2024}, organization = {The Blender Foundation, Youtube}, author = {Gustaf Waldemarson}, url = {https://youtu.be/BEbscsBRIx0}, }
Succinct Opacity Micromaps
Proceedings of the ACM on Computer Graphics and Interactive Techniques, Volume 7, Issue 3
Alpha masked geometry such as foliage has long been one of the trickier things to render efficiently, both for rasterization based approaches and for hardware accelerated ray-tracing. Recently, a new type of primitive was introduced to the Vulkan® and DirectX® ray-tracing APIs that promises to alleviate this issue: Opacity Micromaps, a structure that uses a bit of extra memory as hints to the pipeline when it should actually call the AnyHit-shader. In this paper, we extend this primitive with a novel compression method that uses the concept of succinct 4-way trees to reduce the memory footprint by up to 110 times, including an algorithm for looking up micromap values directly from this compressed form. Further, we perform a comprehensive analysis of the generated micromaps to demonstrate their performance in terms of both memory footprint and frame render time compared to a number of similar structures. Finally, we highlight some aspects of the extension that developers and artists should be aware of to make the most out of it.
@article{WaldemarsonHPG24, author = {Waldemarson, Gustaf and Doggett, Michael}, title = {Succinct Opacity Micromaps}, year = {2024}, issue_date = {August 2024}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, volume = {7}, number = {3}, url = {https://doi.org/10.1145/3675385}, doi = {10.1145/3675385}, abstract = {Alpha masked geometry such as foliage has long been one of the trickier things to render efficiently, both for rasterization based approaches and for hardware accelerated ray-tracing. Recently, a new type of primitive was introduced to the Vulkan® and DirectX® ray-tracing APIs that promises to alleviate this issue: Opacity Micromaps, a structure that uses a bit of extra memory as hints to the pipeline when it should actually call the AnyHit-shader. In this paper, we extend this primitive with a novel compression method that uses the concept of succinct 4-way trees to reduce the memory footprint by up to 110 times, including an algorithm for looking up micromap values directly from this compressed form. Further, we perform a comprehensive analysis of the generated micromaps to demonstrate their performance in terms of both memory footprint and frame render time compared to a number of similar structures. Finally, we highlight some aspects of the extension that developers and artists should be aware of to make the most out of it.}, journal = {Proc. ACM Comput. Graph. Interact. Tech.}, month = {aug}, articleno = {45}, numpages = {18}, keywords = {Compression, Opacity Micromaps, Ray Tracing} }
Handling Custom Data in glTF Files with Exporter/Importer Plugins
Blender is arguably one of the best tools for developing a wide variety of assets, and glTF is often a good export format for real-time graphics applications given how much of it supports and how extensible it is. However, handling glTF files with application specific content is still somewhat tricky in Blender, often requiring pretty deep knowledge of both the glTF format and the Blender scripting API.
During this presentation, the basic developer view of the glTF Blender IO add-on is presented with a focus on using it to create custom importer and exporters plugins to extract or embed almost arbitrary data from or into glTF files.
@online{waldemarsonBCON23, title = {Handling Custom Data in glTF Files with Exporter/Importer Plugins}, year = {2023}, organization = {The Blender Foundation, Youtube}, author = {Gustaf Waldemarson}, url = {https://youtu.be/4fBGM8qc21M?t=1783}, }
Photon Mapping Superluminal Particles
One type of light source that remains largely unexplored in the field of light transport rendering is the light generated by superluminal particles, a phenomenon more commonly known as Cherenkov radiation.
By re-purposing the Frank-Tamm equation for rendering, the energy output of these particles can be estimated and consequently mapped to photons, making it possible to visualize the brilliant blue light characteristic of the effect.
In this paper we extend a stochastic progressive photon mapper and simulate the emission of superluminal particles from a source object close to a medium with a high index of refraction. In practice, the source is treated as a new kind of light source, allowing us to efficiently reuse existing photon mapping methods.
@inproceedings {s.20201004, booktitle = {Eurographics 2020 - Short Papers}, editor = {Wilkie, Alexander and Banterle, Francesco}, title = {{Photon Mapping Superluminal Particles}}, author = {Waldemarson, Gustaf and Doggett, Michael}, year = {2020}, publisher = {The Eurographics Association}, ISSN = {1017-4656}, ISBN = {978-3-03868-101-4}, DOI = {10.2312/egs.20201004} }
Packet Ray Tracing with the ARM NEON Architecture
The ray tracing algorithm has long been used to create near photorealistic images and even simple ray tracers can simulate effects such as shadows, reflections and refractions where other methods struggle. Today, the state-of-the-art CPU ray tracers are typically those that are able to efficiently leverage datalevel parallelism at the instruction level by using SIMD extensions to allow the processor to trace multiple rays simultaneously, rather than one at a time. Within the ARM processor, the SIMD architecture is known as NEON and it is used to speed up data-parallel applications such as multimedia decoding and computer graphics. Thus, in this thesis we have investigated the implementation and performance of a ray tracer that utilizes the NEON architecture to trace packets of rays similar to modern ray tracing frameworks. To determine the efficiency of it we also developed a single-ray tracer as a reference and compared them in regard to both runtime and power consumption. In the end, we found that the optimized ray tracer scales better and performs around 150 – 300% better than the reference single-ray tracer.
@misc{WaldemarsonMasterThesis2014, abstract = {{The ray tracing algorithm has long been used to create near photorealistic images and even simple ray tracers can simulate effects such as shadows, reflections and refractions where other methods struggle. Today, the state-of-the-art CPU ray tracers are typically those that are able to efficiently leverage datalevel parallelism at the instruction level by using SIMD extensions to allow the processor to trace multiple rays simultaneously, rather than one at a time. Within the ARM processor, the SIMD architecture is known as NEON and it is used to speed up data-parallel applications such as multimedia decoding and computer graphics. Thus, in this thesis we have investigated the implementation and performance of a ray tracer that utilizes the NEON architecture to trace packets of rays similar to modern ray tracing frameworks. To determine the efficiency of it we also developed a single-ray tracer as a reference and compared them in regard to both runtime and power consumption. In the end, we found that the optimized ray tracer scales better and performs around 150 – 300\% better than the reference single-ray tracer.}}, author = {{Waldemarson, Gustaf}}, issn = {{1650-2884}}, language = {{eng}}, note = {{Student Paper}}, title = {{Packet Ray Tracing with the ARM NEON Architecture}}, url = {https://lup.lub.lu.se/student-papers/search/publication/4730664}, year = {{2014}}, }
Reflective Variance Shadow Maps
Lighting is a very important phenomenon in computer graphics, while basic direct lighting and shadowing algorithms provide an acceptable approximation to the light in the real world, they are unable to model more subtle lighting phenomena – such as color bleeding and indirect illumination. Recently, a new use of the classic Shadow Map has lead to a relatively simple algorithm, capable of introducing single bounce indirect illumination called Reflective Shadow Maps. This paper focuses on the implementation of this algorithm, as well as that of another extension of the normal shadow map called Variance Shadow Maps which can significantly smoothen the shadows in a wide variety of scenes.
@misc{WaldemarsonOguz2012, abstract = {{Lighting is a very important phenomenon in computer graphics, while basic direct lighting and shadowing algorithms provide an acceptable approximation to the light in the real world, they are unable to model more subtle lighting phenomena – such as color bleeding and indirect illumination. Recently, a new use of the classic Shadow Map has lead to a relatively simple algorithm, capable of introducing single bounce indirect illumination called Reflective Shadow Maps. This paper focuses on the implementation of this algorithm, as well as that of another extension of the normal shadow map called Variance Shadow Maps which can significantly smoothen the shadows in a wide variety of scenes.}}, author = {{Waldemarson, Gustaf, Taskin, Oguz}}, language = {{eng}}, note = {{Student Paper}}, title = {{Reflective Variance Shadow Maps}}, year = {{2012}}, }
Patents
A graphics processing system that is operable to perform ray tracing using micromaps is disclosed. A tree representation of a micromap is generated, and when it is desired to determine whether and/or how a ray interacts with a sub-region of a primitive, the tree representation of the micromap is traversed to determine a property value for the sub-region of the primitive.
@patent{waldemarsonPatent2026, title = {Ray tracing graphics processing using micromaps}, author = {Waldemarson, Gustaf}, year = {2026}, month = feb # "~3", url = {https://patents.google.com/patent/US12541921B2/en}, number = {US12541921B2}, assignee = {ARM Ltd}, keywords = {patents}, note = {US Patent 12,541,921} }
A graphics processing system that is operable to perform ray tracing using micromaps is disclosed. A tree representation of a micromap is generated, and when it is desired to determine whether and/or how a ray interacts with a sub-region of a primitive, the tree representation of the micromap is traversed to determine a property value for the sub-region of the primitive.
@patent{waldemarsonPatentApp2026-1, title = {Graphics processing with micromap memory footprint reduction}, author = {Waldemarson, Gustaf}, year = {2026}, month = apr # "~23", url = {https://patents.google.com/patent/US20260112070A1/en}, number = {US20260112070A1}, assignee = {ARM Ltd}, keywords = {patents}, note = {US Patent App. 18/918,807} }
A graphics processing system that provides a micromap defining property values for sub-regions of a primitive of a scene to be rendered; and generates a first tree representation of the micromap; and applies a first barycentric rotational transform to the micromap and for the first barycentric rotational transform generates a second tree representation of the micromap; and selects one of the first or second tree representations for processing as the selected tree representation. The selected tree representation of a micromap is used when it is desired to determine whether and/or how a ray interacts with a sub-region of a primitive, the tree representation of the micromap is traversed to determine a property value for the sub-region of the primitive.
@patent{waldemarsonPatentApp2026-2, title = {Graphics processing with barycentric rotations}, author = {Waldemarson, Gustaf}, year = {2026}, month = apr # "~23", url = {https://patents.google.com/patent/US20260112112A1/en}, number = {US20260112112A1}, assignee = {ARM Ltd}, keywords = {patents}, note = {US Patent App. 18/918,815} }











