Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

1 Commit
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

mruby-gpu-kernel

A Ruby block, compiled into a Vulkan compute shader. An add-on for mruby-gpu-narray.

a = GPU::SFloat.new(1024).seq
a.map { |x| x * 2 + sin(x) }     # this block becomes one compute shader

Why

In mruby-gpu-narray, every operator is its own dispatch. a * 2 + 1 - 3 + 4 builds four command buffers, submits four times, and waits on a fence four times — and the waiting dominates: a 1,048,576-element operation takes about as long as an 8-element one, because almost none of the time is arithmetic.

Written as a block, the same expression becomes one shader, one submit, one wait.

1024 elements, Apple M5 time
a * 2 + 1 - 3 + 4 (4 dispatches) 1.02–1.12 ms
one traced kernel (1 dispatch) 0.20–0.23 ms
~5× faster

Adding more operators to the block does not make it slower — the cost was the round trips, not the arithmetic.

How — tracing, not parsing

mruby has no RubyVM::AbstractSyntaxTree, and driving the mruby parser from C to read a block is heavy. So the block is never parsed. It is called once, with a proxy object in place of the data. Every operator applied to that proxy records itself and returns a new node, so running the block builds an expression tree instead of computing anything.

GPU::Expr.wrap(
  GPU::KernelContext.new.instance_exec(GPU::Expr.new(:var, ["x"])) { |x| x * 2 + 1 }
)
#=> (+ (* x 2.0) 1.0)

That tree is printed as GLSL, handed to glslangValidator, and the resulting SPIR-V becomes a compute pipeline. Nothing is hidden — you can read exactly what was generated:

puts GPU.kernel_source { |x| x * 2 + sin(x) }
#version 450

layout(local_size_x = 256) in;

layout(set = 0, binding = 0) buffer BufA { float a[]; };
layout(set = 0, binding = 1) buffer BufB { float b[]; };

layout(push_constant) uniform PushConstants { uint n; } pc;

// generated by GPU.kernel
void main() {
    uint idx = gl_GlobalInvocationID.x;
    if (idx >= pc.n) { return; }
    float x = a[idx];
    b[idx] = ((x * 2.0) + sin(x));
}

API

a.map { |x| x * 2 + 1 }        # trace, compile, run -> a new GPU::SFloat

k = GPU.kernel { |x| x * 2 }   # a reusable kernel; identical source compiles once
k.call(a)
k.glsl                         # the exact GLSL that was compiled

GPU.kernel_source { |x| x }    # trace only, no compilation (tests, slides)

Inside a block you have GLSL's element-wise built-ins under their GLSL names: sin, cos, tan, asin, acos, atan, sinh, cosh, tanh, exp, exp2, log, log2, sqrt, inversesqrt, abs, sign, floor, ceil, fract, radians, degrees, min, max, pow, mod, step, clamp, mix, smoothstep.

Two rules

  • The block runs via instance_exec — that is what lets you write sin(x) instead of x.sin. Methods of the enclosing object are therefore not visible inside the block. Local variables still are.
  • Keep the expression on the left. 2 * x raises TypeError, because mruby has no numeric coercion — the same rule the arrays themselves follow.

Not yet

One input and one output, element-wise only. No control flow, no reductions, no multi-array kernels. A shader that fails to compile raises GPU::CompileError carrying glslang's own message.

Requirements

  • mruby-gpu-narray — pulled in automatically as a dependency via mgem-list.
  • A Vulkan 1.1+ loader and a compute-capable device.
  • glslangValidator on PATH at run time (override with GLSLANG=). The base gem needs it only at build time; this one shells out to it whenever a new block is traced.

The only link dependency is the Vulkan loader (-lvulkan).

Build

Add just this gem — mruby-gpu-narray comes along as a dependency, resolved through mgem-list:

MRuby::Build.new do |conf|
  toolchain :clang          # :gcc on the Pi
  conf.gembox 'default'
  conf.gem '/path/to/mruby-gpu-kernel'
  # macOS only:
  conf.cc.include_paths     << '/opt/homebrew/include'
  conf.linker.library_paths << '/opt/homebrew/lib'
  conf.enable_test
end
cd /path/to/mruby && MRUBY_CONFIG=/path/to/mruby-gpu-kernel/build_config.reference.rb rake
./build/host/bin/mruby /path/to/mruby-gpu-kernel/test/kernel_test.rb    # ALL TESTS PASSED
./build/host/bin/mruby /path/to/mruby-gpu-kernel/examples/kernel_dsl.rb

To develop against a local checkout of the base gem instead of the published one, list it before this gem in build_config.rb.

How it fits together

mrblib/gpu_kernel.rb   Expr / KernelContext / GPU.kernel / NArray#map
   │                   block -> expression tree -> GLSL
   ▼
src/gpu_kernel.c       GPU::Kernel: glslangValidator -> SPIR-V -> VkPipeline -> dispatch
   │
   ├─ uses mruby-gpu-narray's g_ctx (device, queue, layouts, pools)
   └─ uses its create_buffer / wrap_buffer / gpu_buffer_type

This gem compiles against the base gem's src/gpu_internal.h and links against symbols it already exports, so the base gem needs no changes. The dispatch routine here is a copy of the base gem's dispatch_compute(): upstream's version is keyed to its internal PipeId enum and cannot bind a pipeline built elsewhere. When upstream grows a pipeline-taking variant, this copy should go away.

Verified on

  • macOS — Apple M5 GPU via MoltenVK, Vulkan 1.1: 20/20 tests pass.
  • Raspberry Pi 5 — not yet run.

License

MIT © 2026 Yuji Teshima

About

Turn a Ruby block into a Vulkan compute shader — an mrbgem add-on for mruby-gpu-narray. na.map { |x| x * 2 + sin(x) } runs on the GPU.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages