feat(ugc): ray_backend=hiprt traces on the GPU with HIPRT and Orochi (optional build)

HIPRT (MIT) behind the CMake option DLU_HIPRT (off): its headers come from its
SDK (HIPRT_ROOT, else ROCm's /opt/rocm) and are copied next to the servers; its
library is loaded when first used (hiprtew), as HIP or CUDA are by Orochi
(MIT, fetched pinned by hash; CUDA when its toolkit is found). The trace
kernels (nearest hit skipping the triangle a ray leaves, any hit) are compiled
the first time and kept in cache/hiprt. One GPU context for the process
(hiprt_device picks the GPU); the workers take turns on it. When HIPRT, the
GPU or a scene's upload fails, Embree is used instead, and the UGC server logs
why at start.

For a GPU the rays go in batches (UgcRays::Scene gets batch queries; the CPU
backends answer them a ray at a time):
- hidden faces: with a batch backend the paths are traced side by side, a
  bounce at a time (the path code split into Start, Scatter and Bounce, the one
  by one tracing unchanged); the same paths with the same random numbers, so
  the same triangles are decided (tested with builtin side by side)
- the occlusion bake and the denoised icons' traced occlusion always ask in
  batches (the same rays, the same results)

Its symbols are hidden: the servers export theirs (-rdynamic), and HIPRT's
library, which has an Orochi of its own, would otherwise call ours.

Check: configure with -DDLU_HIPRT=ON on a machine with ROCm (or HIPRT's SDK)
and an AMD RDNA or NVIDIA GPU; UgcServer --make-model x.lxfml out hiprt; the
UGC tests (hits, hidden faces and occlusion against builtin); a build without
it leaves everything as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Aaron Kimbrell
2026-09-29 11:57:37 -05:00
parent 2e2e8153e2
commit 07dbf30dc8
13 changed files with 719 additions and 92 deletions

View File

@@ -12,6 +12,10 @@
#include <embree4/rtcore.h>
#ifdef DLU_HIPRT
#include "UgcRaysHiprt.h"
#endif
namespace {
using UgcRays::Hit;
using UgcRays::INF;
@@ -528,6 +532,14 @@ namespace {
}
namespace UgcRays {
void Scene::Closest(const Ray* rays, Hit* hits, size_t count) const {
for (size_t i = 0; i < count; i++) hits[i] = Closest(rays[i].origin, rays[i].direction, rays[i].skip, rays[i].maxT);
}
void Scene::Occluded(const Ray* rays, uint8_t* occluded, size_t count) const {
for (size_t i = 0; i < count; i++) occluded[i] = Occluded(rays[i].origin, rays[i].direction, rays[i].minT, rays[i].maxT) ? 1 : 0;
}
std::string_view Name(eBackend backend) {
switch (backend) {
case eBackend::EMBREE: return "embree";
@@ -544,6 +556,9 @@ namespace UgcRays {
}
bool Available(eBackend backend) {
#ifdef DLU_HIPRT
if (backend == eBackend::HIPRT) return UgcRaysHiprt::Available();
#endif
return backend == eBackend::BUILTIN || backend == eBackend::EMBREE;
}
@@ -551,10 +566,32 @@ namespace UgcRays {
return Available(wanted) ? wanted : eBackend::EMBREE;
}
std::string Problem(eBackend backend) {
if (Available(backend)) return {};
#ifdef DLU_HIPRT
if (backend == eBackend::HIPRT) return UgcRaysHiprt::Problem();
#endif
return "the server was built without it (DLU_HIPRT)";
}
std::unique_ptr<Scene> Make(eBackend backend, const UgcModel::Mesh& mesh) {
switch (Resolve(backend)) {
#ifdef DLU_HIPRT
case eBackend::HIPRT:
// A GPU that fails now (out of memory, ...) leaves the job to Embree
if (auto scene = UgcRaysHiprt::Make(mesh)) return scene;
return std::make_unique<EmbreeScene>(mesh);
#endif
case eBackend::EMBREE: return std::make_unique<EmbreeScene>(mesh);
default: return std::make_unique<BuiltinScene>(mesh);
}
}
void SetGpuDevice(int index) {
#ifdef DLU_HIPRT
UgcRaysHiprt::SetDevice(index);
#else
(void)index;
#endif
}
}