feat(ugc): ray_backend=hiprt traces on the GPU with HIPRT and Orochi (optional build)

HIPRT (MIT) behind the CMake option DLU_HIPRT (off): its headers come from its
SDK (HIPRT_ROOT, else ROCm's /opt/rocm) and are copied next to the servers; its
library is loaded when first used (hiprtew), as HIP or CUDA are by Orochi
(MIT, fetched pinned by hash; CUDA when its toolkit is found). The trace
kernels (nearest hit skipping the triangle a ray leaves, any hit) are compiled
the first time and kept in cache/hiprt. One GPU context for the process
(hiprt_device picks the GPU); the workers take turns on it. When HIPRT, the
GPU or a scene's upload fails, Embree is used instead, and the UGC server logs
why at start.

For a GPU the rays go in batches (UgcRays::Scene gets batch queries; the CPU
backends answer them a ray at a time):
- hidden faces: with a batch backend the paths are traced side by side, a
  bounce at a time (the path code split into Start, Scatter and Bounce, the one
  by one tracing unchanged); the same paths with the same random numbers, so
  the same triangles are decided (tested with builtin side by side)
- the occlusion bake and the denoised icons' traced occlusion always ask in
  batches (the same rays, the same results)

Its symbols are hidden: the servers export theirs (-rdynamic), and HIPRT's
library, which has an Orochi of its own, would otherwise call ours.

Check: configure with -DDLU_HIPRT=ON on a machine with ROCm (or HIPRT's SDK)
and an AMD RDNA or NVIDIA GPU; UgcServer --make-model x.lxfml out hiprt; the
UGC tests (hits, hidden faces and occlusion against builtin); a build without
it leaves everything as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Aaron Kimbrell
2026-09-29 11:57:37 -05:00
parent 2e2e8153e2
commit 07dbf30dc8
13 changed files with 719 additions and 92 deletions

View File

@@ -164,6 +164,8 @@ namespace {
// What traces the rays of the hidden faces' paths and of the occlusion (the icon's too)
settings.hsr.rays = UgcRays::Parse(Game::config->GetValue("ray_backend")).value_or(UgcRays::eBackend::BUILTIN);
settings.ao.rays = settings.hsr.rays;
// The GPU hiprt uses (read before it is first used; changing it takes a restart)
UgcRays::SetGpuDevice(std::max(Setting<int32_t>("hiprt_device", 0), 0));
settings.ao.enabled = Setting<int32_t>("bake_ao", 1) != 0;
settings.ao.distance = Setting<float>("ao_distance", 5.0f);
settings.ao.samples = std::clamp(Setting<int32_t>("ao_samples", 64), 1, 1024);
@@ -186,6 +188,17 @@ namespace {
return settings;
}
// The processing options the settings pick, and what is used instead when this build or machine can't
// (main thread: the GPU is set up here the first time)
void LogProcessingOptions(const UgcJobs::Settings& settings) {
const auto made = UgcJobs::MadeWith(settings);
LOG("Processing options: %s", UgcProcessOptions::ToString(made).c_str());
if (UgcRays::Resolve(settings.hsr.rays) != settings.hsr.rays) {
LOG("ray_backend=%s can't be used (%s): embree instead", std::string(UgcRays::Name(settings.hsr.rays)).c_str(), UgcRays::Problem(settings.hsr.rays).c_str());
}
if (!UgcRender::Available(settings.icon.denoise)) LOG("denoise=%s can't be used (the server was built without DLU_OIDN): off instead", std::string(UgcRender::Name(settings.icon.denoise)).c_str());
}
UgcProcessor::Limits ReadLimits() {
UgcProcessor::Limits limits;
const auto cores = std::max<unsigned>(std::thread::hardware_concurrency(), 1);
@@ -691,6 +704,7 @@ int main(int argc, char** argv) {
processorConfig.threads = threads > 0 ? threads : std::max<size_t>(std::thread::hardware_concurrency() / 2, 1);
UgcProcessor processor(processorConfig, storage, library, ReadSettings());
processor.Configure(ReadSettings(), ReadLimits());
LogProcessingOptions(ReadSettings());
g_Processor = &processor;
// Sent with the traffic reports to the dashboard (Diagnostics)
TrafficStats::Local().SetGauge("workers_busy", [&processor] { return static_cast<double>(processor.Busy()); });