package rake
Install
dune-project
Dependency
Authors
Maintainers
Sources
sha256=d3fd9d46d5c2352fe220831de1ccee832ef8a80255a95ba5e6d4aed9f183a601
sha512=afc912788fd47ed1c8a605f3094357ae361f314d4bde5824b18d86020d9b4bfd8ad7fb358952fe7ae29ca1593d18bfeccd20c65da3a95397e19251aac6f13be4
Description
Rake makes predicated SIMD execution a first-class language construct. Data rakes through tine declarations (mask patterns), enabling divergent parallel computations on physical vector targets and the WebAssembly SIMD virtual machine.
Added to opam-repository:
README
Rake
Rake is a programming language for SIMD kernels. Its values are racks, one vector register each on a physical CPU target and one v128 value in the WebAssembly virtual machine. Outside slow code, the compiler uses vector instructions wherever the selected profile supports the source operation, or it refuses the program. It never replaces rack work with hidden scalar loops or helper calls. The language and its documentation are at rake-lang.org.
What Rust does for safety with unsafe {}, Rake does for speed with slow {}. A run explicitly enters scalar code at slow { and resumes vector work at }. Slow blocks are available in 0.7.0 and in the playground.
Release: 0.7.0.
scratch advance(positions: f32s, velocities: f32s) -> f32s:
| scaled <| velocities * <0.5>
| moved <| positions + scaled
moved
tine #valid(values: f32s) means values >= <0.0>
rake safe_root(values: f32s) -> f32s:
through #valid(values) into rooted:
sqrt(values)
sweep:
| #valid(values) => rooted
| #valid(values) gaps => <0.0><0.5> is a uniform, one scalar shared by every lane. #valid is a reusable tine predicate, and through computes roots only in its selected lanes. gaps selects the complementary lanes, including NaNs. The sweep picks each lane's result. | name <| expression binds a stage of a fused computation: on AVX2 and NEON, advance compiles to one fused multiply-add.
Targets
Profile | Rack | What compiles |
|---|---|---|
| one 128-bit XMM register, 4 32-bit lanes | float racks and a 32-bit integer subset, as assembly |
| one 256-bit YMM register, 8 32-bit lanes | float racks and a 32-bit integer subset, as assembly |
| one 512-bit ZMM register, 16 32-bit lanes | float racks and a 32-bit integer subset, as assembly |
| one 128-bit vector register, 4 32-bit lanes | float racks and a 32-bit integer subset, as assembly |
| one | scratches and rakes over float and integer racks, runs over memory, and whole programs with scalar |
| as | adds the relaxed SIMD operations, by opt-in |
The scalar fallback remains WIP (work in progress), and the compiler rejects code for it. General native runs are also WIP. Rake 0.7.0 compiles slow orchestration with register kernels as native C and objects. It embeds Rake's selected assembly, checks the kernels in the final object, and supports uniform f32, i32, u32 and bool parameters and f32, bool, i32 or u32 results at the boundary from slow code. Other native scalar kernel boundaries remain WIP. Platform C imports and exports, header-backed unions, typed C callbacks and process arguments are supported too. SSE2, AVX2, AVX-512 and NEON also support f32s, i32s and u32s traversals that yield a matching stream or update one mutable stack column, including a separate destination with its own record layout, with checked partial racks. Explicit widen reads signed or unsigned byte and 16-bit columns into 32-bit working racks. Stored columns stay compact, and the tail reads only existing records. Native bitcast between i32s and u32s preserves every lane's bits. to_f32 converts either integer type numerically, and to_i32 or to_u32 rounds and saturates floats to the corresponding integer range. These conversions also compile in native streams. The physical profiles also compile direct uniform f32 comparisons as vector selections, including within through masks and stream tails. Native i32s and u32s support wrapping add/subtract and bitwise AND/OR/XOR. Signed i32s comparisons produce masks for selection and mask reductions. sum, product, minimum and maximum reduce an i32s or u32s rack to one result of the matching scalar type. The corresponding scan_ operations keep every inclusive prefix in a rack. The operation reference lists the implemented subset and remaining work. These CPU and C ABI additions are included in 0.7.0. On wasm-simd128, a whole program becomes one C file with a C entry point, and every selected instruction is written as one wasm_simd128.h intrinsic.
A whole program has vector runs and scalar slow code. Slow code never holds a rack, and a scalar becomes a rack only where it is marked:
pack Samples {
f32: value;
u8: quality;
}
run weigh(input: stack Samples, <count: i64>, <scale: f32>) -> f32:
for chunk in input using f32s up to <count>:
let quality = to_f32(widen(chunk.quality))
yield chunk.value * <scale> + quality
run running_sum(x: []f32, out: mut []f32, <n: i32>):
total := <0.0>
for <i: i32> from <0> up to <n> by <4>:
total <- total + x[<i>]
out[<i>] <- total
state calls: i32 := 0
slow main() -> i32:
values: [8]f32 := [1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0]
qualities: [8]u8 := [0, 1, 0, 1, 0, 1, 0, 1]
weighed: [8]f32 := [0.0; 8]
weigh(stack Samples { value: values, quality: qualities }, <8>, <0.5>, weighed)
sums: [8]f32 := [0.0; 8]
running_sum(weighed, sums, <8>)
calls <- calls + 1
return i32(sums[7] * 4.0) + callsUsing the compiler
The opam submission is awaiting review. Its package tests run portable compiler checks, without requiring toolchains for every output target. Producing and verifying objects needs the target tools described in the compiler reference.
The compiler is rakec, written in OCaml. The Nix development shell has everything it needs:
nix develop --command dune build
nix develop --command dune exec rakec -- --interpret program.rk
nix develop --command dune exec rakec -- --emit-c --target x86-avx2 -o program.c program.rk
nix develop --command dune exec rakec -- --emit-asm --target x86-avx2 -o kernels.s kernels.rk
nix develop --command dune exec rakec -- --verify-native --target wasm-simd128 -o program.o program.rk--interpret runs main in Rake's executable semantics. --emit-c writes a C unit for a whole program or register kernels. On a physical target, that unit embeds Rake-selected assembly, while the platform C compiler lowers slow code. --emit-asm writes physical-target register-kernel assembly. --verify-native builds an object, disassembles it and checks its vector functions against the profile's rules. A scratch or rake contains only register work from the profile's instruction list, with no calls and no stack. On x86 and AArch64 each rack is one whole physical register, and the object has exactly the fused multiply-adds the compiler selected. On WebAssembly, Rake keeps to the virtual machine's v128 values and SIMD instructions. It doesn't try to replace the runtime's physical register allocation. A wasm run also contains the uniform address, loop and bounds work written in the source, but it never scalarises rack work. rakec --help lists every mode, --print-targets the profiles and --print-capabilities each language feature's status.
Documentation
The syntax reference lists every form, and the pages it links define what each means. Goals states the language's promises, the backend how the compiler keeps them, and the roadmap what comes next. The playground is an interactive tutorial in twelve lessons, the glossary defines the terms, and Rake and other languages and GPU profiles set Rake beside ISPC and Bend 2. The GPU contract is a design, not an implemented backend. Its first proposed profile maps racks across NVIDIA warp lanes, emits PTX and verifies the ahead-of-time cubin. It separates checked execution properties from occupancy, memory costs and elapsed time, which still need measurement. The rakec command lists its modes and options, test/README.md describes the tests, and CHANGELOG.md the releases.
Licence
MIT, in LICENSE.
Dependencies (8)
-
cmdliner
>= "1.2.0" -
js_of_ocaml-ppx
>= "6.0.0" -
js_of_ocaml
>= "6.0.0" -
ppx_deriving
>= "6.0.0" -
menhir
>= "20231231" - base-unix
-
ocaml
>= "5.1.0" -
dune
>= "3.20"
Dev Dependencies (1)
-
odoc
with-doc
Used by
None
Conflicts
None