package rake
sectionYPositions = computeSectionYPositions($el), 10)"
x-init="setTimeout(() => sectionYPositions = computeSectionYPositions($el), 10)"
>
On This Page
Vector-first language for predictable SIMD
Install
dune-project
Dependency
Authors
Maintainers
Sources
rake-0.7.0.tar.gz
sha256=d3fd9d46d5c2352fe220831de1ccee832ef8a80255a95ba5e6d4aed9f183a601
sha512=afc912788fd47ed1c8a605f3094357ae361f314d4bde5824b18d86020d9b4bfd8ad7fb358952fe7ae29ca1593d18bfeccd20c65da3a95397e19251aac6f13be4
doc/CHANGELOG.html
Changelog
A Rake version identifies one compiler, Tree-sitter grammar, documentation set and website. Releases before 1.0 may still change the source language and the binary boundaries between versions. A design in the documentation gets a version only when the compiler implements it and the tests cover it.
0.7.0
- Package tests run without target assemblers or disassemblers. Separate native and WebAssembly object suites retain mandatory verification checks, and the release gate runs both. Native GNU ELF tools accept explicit executable overrides and common distribution cross-toolchain prefixes. Package CI checks Linux, macOS and Windows separately from object checks.
- Native register kernels bitcast between
f32s,i32sandu32swithout changing lane bits. Reinterpreting an input uses its register or a vector register copy, preserving zero signs and NaN payloads without floating-point exceptions. - Register kernels accept unary minus on
f32andi32uniforms, including reduced and extracted values. Packed negation preserves float zero signs and NaN payloads, and signed results wrap to 32 bits. Negation under a partial lane mask remains WIP. - Register kernels accept arithmetic on
f32,i32andu32uniforms, including results of reductions, extraction andbitmask. Physical profiles broadcast those values and use packed instructions, then retain the completed scalar result. Float arithmetic rounds to binary32 at each step, and integer addition, subtraction and multiplication wrap to 32 bits. Arithmetic under a partial lane mask remains WIP. scan_sum,scan_product,scan_minimumandscan_maximumaccepti32sandu32son every CPU profile. Packed stages combine inclusive prefixes across the full rack, preserving wrapping arithmetic and signed or unsigned extrema. A sequential C oracle checks every lane and reuse of the input rack. Scans and reductions at other integer widths remain WIP.sum,product,minimumandmaximumaccepti32sandu32son every CPU profile. Packed shuffle-and-combine trees preserve wrapping arithmetic and signedness, with one completed result crossing the scalar C ABI. Independent C checks overflow boundaries and reuse of the input rack. Scans and reductions at other integer widths remain WIP.- Native stream bodies support local rack assignments, arrays of rack locations and unrolled fixed-count
repeat, including nested repeats. Ordered SSA bindings preserve earlier snapshots without duplicating their expressions. Each traversal chunk starts with fresh locations, and each unrolled copy has its own locals. Runtime-counted inner loops remain WIP. --emit-cproduces a C translation unit for WebAssembly, native whole programs or native register kernels.--emit-asmproduces physical-target register-kernel assembly and directs whole-program and WebAssembly users to--emit-c. Native C output embeds Rake-selected assembly and delegates slow lowering to the platform C compiler. C ABI interoperation remains independent of the output format.- Uniform conditions compose with
and,orandnotin physical register kernels and native streams. Chained immutable Boolean bindings retain that structure. Short-circuit masks sanitise skipped floating-point comparisons, including inside through masks and partial racks. Independent C checks exact choices and signalling-NaN exception flags, with guarded stream inputs and column updates. Mixed-program and WebAssembly checks cover the composed condition too. WebAssembly C emission also keeps parameter identifiers distinct from C keywords, generated temporaries and typedefs. - Native
i32sandu32sshifts accept runtime uniformu32counts on SSE2, AVX2, AVX-512F and NEON, including native streams. Vector operations normalise the count modulo 32 in one allocated temporary, preserving live inputs and the original count. Final-object checks require the complete selected sequence with full-width data operands. Independent C checks all counts, high count bits, extracted values, masked calls and retained inputs. Guarded streams check full racks, tails, column updates and separate destinations. Mixed-program checks cover the scalar C boundary and WebAssembly. - Literal-index
extractandinsertsupporti32sandu32son SSE2, AVX2, AVX-512F and NEON. They reuse the float lane transfers without numerical conversion. Scalar integer values cross the platform's integer C ABI and remain in vector registers inside kernels. Independent C checks every lane, signed boundaries, unsigned high bits, rebroadcasts and preservation of still-live racks and replacement uniforms. Mixed-program checks cover slow callers and WebAssembly. Native stream participation remains WIP. to_u32convertsf32sto unsigned 32-bit racks on all five CPU profiles, including native streams. It rounds to nearest with ties to even, clamps to the unsigned range and converts NaNs to zero. SSE2 and AVX2 use packed signed conversion on two ranges, restoring bit 31 after exact subtraction. AVX-512F and NEON use packed unsigned conversion after sanitising the range. WebAssembly rounds then uses unsigned saturating conversion. Four allocated native temporaries preserve masks and safe values, with spills rejected. Independent bit oracles check results and masked exceptions, including retained inputs, guarded tails, destinations and exact in-place updates.to_f32acceptsu32son all five CPU profiles, in register kernels and streams. SSE2 and AVX2 convert two exactly represented 16-bit halves before one rounded addition. AVX-512F, NEON and WebAssembly select unsigned vector conversion instructions. The integer-bit oracle checks rounding across bit 31, retained inputs, masked exceptions, compact widening and guarded stream tails. Generated WebAssembly C explicitly discards unused ABI parameters, including unneeded tail masks, so strict unused-parameter warnings remain enabled.- Signed
to_f32andto_i32compile on SSE2, AVX2, AVX-512F and NEON, in register kernels and streams. Packed operations preserve nearest-even rounding, saturating float-to-integer results and NaN-to-zero conversion. Masked operands become zero before conversion. Independent integer-bit oracles check boundary and raw values, retained inputs, exception flags and guarded tails. Final-object fixtures reject scalar and narrowed instructions. The interpreter now clamps infinities and out-of-range values before scalar integer conversion. - Native streams explicitly widen
i8andi16columns intoi32s, oru8andu16columns intou32s, on SSE2, AVX2, AVX-512F and NEON. Packed extension preserves signed values, with guarded compact transfers in partial racks. Independent C checks all four stored types, wrapping sums, unsigned products, single-column updates and separate destinations for every count from zero through 65, with arrays ending at guard pages. Outputs remain 32-bit, and general native runs remain WIP. - Native
bitcastbetweeni32sandu32spreserves lane bits in register kernels and streams. Allocation reuses a dying input or copies it through a vector register, without numerical conversion. The interpreter now keeps the selected profile's width through bitcasts and integer broadcasts. - WebAssembly object verification accepts narrower packed comparisons and extending multiplies only when source-derived operand bounds justify them. Compact column types and literals establish those bounds. Unproved arithmetic drops them. Independently compiled packed objects reject without the bounds and pass with them. LLVM's six extending-load spellings are normalised to their WebAssembly instruction identifiers. Compact streams pass numerical and object checks in both addressing modes.
- Native streams traverse
f32,i32andu32columns on SSE2, AVX2, AVX-512F and NEON. One to four same-width columns can contribute to a stream or single-column update, retaining each column's type and the existing guarded-tail contract. Independent C checks signed wrapping arithmetic, unsigned boundaries, literal shifts, mixed float/integer masks, exact aliasing and separate destination layouts. C and Rake callers check the boundary, and all five CPU-profile interpreters check the same source. Cross-lane traversal operations remain WIP. Embedded kernel assembly restores the text section before slow functions, so platform startup adapters remain executable code. - WebAssembly partial stores keep their scalar remainder opaque to clang, retaining integer arithmetic on a rack before lane stores. This applies in both addressing modes. Unsigned comparisons at the high-bit boundary select their packed sign-mask sequence explicitly. Independent expected masks and the strict final-object verifier check both changes.
- WebAssembly uniform choices broadcast a Boolean mask and select rack bits. This retains the rack choice through partial stores. Clang previously moved scalar choices after a one-lane extraction, which the strict object verifier rejected. Scalar expected values check both nested choices and every tail remainder, including unchanged output beyond the count.
- Native float streams take
i32,u32andbooluniforms alongside floats. Integer arguments, descriptors, counts and output pointers follow their C register counter independently of float arguments. Direct uniform comparisons and Boolean choices retain protected full-rack and tail semantics. Independent C checks interleaved argument types, signed and unsigned boundaries, unused Boolean argument bits, eight persistent uniforms, in-place updates and separate destinations ending at guard pages. Stack arguments remain work in progress. - Boolean uniforms choose whole racks on all four physical profiles, including results of
allandany. A packed broadcast and sign extension turn the Boolean into a lane mask, with the usual protection for untaken floating-point work. Native kernel arguments accept Cboolthrough the integer argument slots and retain only its value bit. Independent C checks both choices, every reduced lane mask, fused expressions and nested through masks. Assembly callers deliberately set unused upper argument bits. Slow callers check Boolean arguments, literals, results and inlined calls. Marked literals<true>and<false>are Boolean uniforms, including forms with spaces inside the brackets. The lexer previously treated the compact forms as variable references, unlike the Tree-sitter grammar. - Direct uniform
i32andu32comparisons choose whole racks on all four physical profiles. Broadcast operands use the existing typed vector comparisons, with mask protection for untaken floating-point work. Independent C checks all six predicates, signed and unsigned boundaries, literals on either side and retained input racks. Nested conditions inside through masks check exception flags as well as results. Slow callers and WebAssembly programs exercise the same source forms. - Native register kernels take
i32andu32uniforms through the platform C integer argument registers, independently of SIMD argument slots. Their entry imports and vector broadcasts preserve all 32 bits. Verification permits only the declared entry transfers. Independent C checks high-bit values, interleaved argument classes, every integer argument register, eight SIMD arguments and retained racks. Slow callers also use these uniforms and signed 32-bit results. Integer bitwise built-ins and extrema broadcast marked uniforms in either operand, including two uniforms. Typed calls check integer literals against the parameter's range, and explicitly typed rack bindings broadcast their uniforms. Native scalar integer constants stay in vector registers until the C return transfer. Executable semantics and WebAssembly checks cover unsigned boundary bits and wrapping uniform arithmetic. u32sminandmaxcompile on SSE2, AVX2, AVX-512F, NEON and WebAssembly. SSE2 compares sign-bit-biased copies, then selects the original lane bits with three allocated temporary registers. AVX2 and AVX-512F usevpminudandvpmaxud, NEON usesuminandumax, and WebAssembly usesi32x4.min_uandi32x4.max_u. Independent C checks unsigned boundaries, nested clamps, masked selection and retained inputs. Interpreter and WebAssembly goldens include full-width unsigned literals.u32ssupports all six comparisons on SSE2, AVX2, AVX-512F, NEON and WebAssembly. Unsigned rack and uniform types retain their signedness in the IR and interpreter. SSE2 and AVX2 compare sign-bit-biased copies, AVX-512F usesvpcmpud, and NEON usescmhiandcmhs. Independent C checks high-bit boundaries, operand order, mask reductions and live inputs. Interpreter goldens and a whole-program WebAssembly fixture check unsigned ordering and broadcasts.- Static
i32sandu32sshuffles compile on all four physical profiles. They share the float shuffle's bit-preserving selection, allocation and final-object checks. Independent C checks one- and two-rack permutations, repeated lanes, whole-register selection and still-live or aliased inputs. The interpreter also selects 32-bit integer lanes, with hand-specified boundary-bit cases at four, eight and sixteen lanes. Native stream shuffles remain WIP until their tail participation contract is defined. - Native signed
i32sabscompiles on all four physical profiles, retaining the wrapping −2³¹ result. SSE2 uses packed sign extension, XOR and subtraction with an allocated temporary register. AVX2, AVX-512F and NEON use direct full-width absolute-value instructions. Independent widened C arithmetic checks exact bits, retained inputs, nested calls and masked selection. Final-object fixtures refuse narrowed and memory forms, and SSSE3 instructions in the SSE2 profile. - Native
i32sandu32sbit_andnot(a, b)uses a full-width packed instruction on every physical profile, preservinga & ~b. The x86 emitter reverses instruction operands, and SSE2 preserves inputs across destructive register reuse. Independent C checks exact bits, equal operands, literal masks and masked selection. Final-object fixtures refuse narrowed, scalar and memory forms. - Integer rack mask literals use their scalar element's range check. Valid
u32smasks up to 4294967295 are accepted, including a literal as the first operand ofbit_andnot. Native lowering preserves those high-bit masks in its 32-bit representation. An untyped integer literal retains the signedi32range. - Native
i32sandu32sbit shifts accept literal counts from 0 to 31 on all four physical profiles. Full-width packed instructions implement left, logical-right and signed-right shifts. A zero count leaves the rack unchanged. Independent C checks every count, sign extension, masked use and retained inputs. Final-object checks reject narrowed or scalar shifts and out-of-range immediates. Runtime counts use the normalisation sequence above. - Native signed
i32sminandmaxcompile on all four physical profiles. SSE2 uses packed comparison and logical selection with an allocated mask register. AVX2, AVX-512F and NEON use direct full-width extrema instructions. Independent C checks signed limits, nested clamps, masked branches and retained inputs. Final-object checks reject narrower extrema and SSE4.1 instructions in the SSE2 profile. - Native
i32sandu32smultiplication wraps to the low 32 bits on all four physical profiles. SSE2 uses two packed multiplies and four shuffles with allocated vector temporaries. AVX2, AVX-512F and NEON select one full-width multiply instruction. Independent C checks overflow, lane order, masked products and destructive input reuse. Final-object checks refuse MMX, narrowed vectors, scalar work and memory multiply operands. - Native signed
i32snegation uses full-width packed zero/subtract on SSE2, AVX2 and AVX-512F, andneg .4son NEON. The independent C oracle checks wrapping bits, retained inputs and lane-masked selection, including the minimum signed value. The final-object verifier rejects narrower NEON negation forms. - Native
i32sandu32ssupport wrapping add/subtract and bitwise AND/OR/XOR on SSE2, AVX2, AVX-512F and NEON. Signedi32scomparisons produce masks for selection and mask reductions. Mixed integer/float selection preserves the vector C ABI. An independent C oracle checks overflow bits, signed extremes, tines/gaps and still-live inputs. Assembled negative fixtures reject narrowed vectors, scalar work and integer memory operations. Other integer operations remain WIP. - Direct uniform
f32comparisons support whole-rack conditional expressions on all four physical profiles. Broadcast operands use ordered vector comparisons and selection, with benign operands in untaken branches. Independent C checks the six predicates, NaNs, signed zeros, subnormals, mixed scalar/vector arguments and nesting inside through masks. Stream checks cover guarded tails, in-place output and Rake callers. - Float comparison masks support
all,anyandbitmaskon all four physical profiles. Full-width bitwise reductions use one checked temporary vector register, then transfer the completed Boolean or bitset through the scalar C return ABI. Native slow callers can receiveboolandu32results from uniform-f32kernels. Independent C checks every lane-mask pattern and quiet-NaN gaps. Final-object fixtures reject intermediate or incorrect scalar transfers and scalar arithmetic inside a kernel. - Static
f32sshuffles compile on all four physical profiles, selecting lanes from one or two racks across their full width. Register-only sequences preserve exact lane bits, with scratch registers included in the no-spill allocation check. Independent C checks permutations, repeated indices, signalling NaNs and still-live inputs. Final-object fixtures reject unselected permutations and memory stores. Native stream shuffles remain work in progress pending their partial-rack contract. - Literal-index
f32sinsertion compiles on all four physical profiles. It replaces one lane's bits from a uniform argument, literal or extracted value, preserving every other lane. The scan pipeline shares its insertion sequence. Independent C checks the mixed scalar/vector ABI, exact bits, signalling NaNs and preservation of still-live inputs. Native streams still reject extraction and insertion pending their tail participation contract. - Literal-index
f32sextraction compiles on SSE2, AVX2, AVX-512F and NEON. The index bounds follow the selected profile's four, eight or sixteen lanes. Full-width vector transfers preserve the selected bits and return through the scalar C ABI. Independent C checks every lane, signalling NaNs, rebroadcasts and preservation of a still-live rack. - NEON compiles the four float reductions and inclusive scans as three ordered steps. Intermediate values stay in complete vector registers, and strict extrema canonicalise NaNs. The object verifier permits lane broadcasts and prefix insertion only for selected cross-lane functions. The shared scalar oracle checks order, signed zeros and NaNs at every position on all physical profiles.
- Native
f32sfloor,ceil,truncand ties-to-evennearestcompile on SSE2, AVX2, AVX-512F and NEON. SSE2 uses vector conversions and masks, with five temporary registers checked by the no-spill allocator. Other profiles use integral-rounding instructions. Independent binary32 checks cover result bits, active signalling NaNs, inactive lanes and guarded tails. The interpreter'snearestnow preserves negative zero. - Native
f32sminandmaxpreserve NaNs and signed zero on SSE2, AVX2, AVX-512F and NEON. The x86 lowering reuses the strict reduction sequence, with five temporary vector registers included in allocation pressure. NEON uses full-widthfminandfmax. An independent bit-ordering oracle checks values, signalling-NaN exceptions and masked gaps, and guard pages check every stream tail. Native floating-point control settings are now explicit caller obligations. - Native
f32sabsolute values compile on SSE2, AVX2, AVX-512F and NEON as full-rack bitwise operations. An independent IEEE-754 bit oracle checks signed zeros, subnormals, infinities and NaNs, including signalling NaNs under masks. Traversal checks cover every tail and C and Rake callers. - Native
f32traversals accepti32counts as well asi64counts. The selected entry code sign-extends a 32-bit C argument before loop guards and pointer access. Independent C checks supply unspecified upper register bits for every tail and for zero and negative counts on all four profiles. - Native
f32traversals can write one column in a separate mutable destination stack. Its descriptor follows the count in the platform C ABI, and the selected pointer uses the destination's own record layout. C checks values and exact aliases against independent scalar results, with guard pages for every partial rack. Rake callers exercise the same boundary. - Native
f32traversals can update one column of their mutable input stack, with the same selected vector loop and guarded tail as stream output. All read columns are loaded before the update. Independent C checks the mutable descriptor ABI, in-place results, an unread destination and untouched independent columns, alongside Rake callers on every CPU profile. - Native
f32streams accept up to eight uniformf32arguments after their count. Their platform C argument registers are copied into caller-clobbered slots and kept live across full and partial racks by the no-spill allocator. Independent C checks scale, bias and threshold arguments, quiet NaNs and all eight register slots, alongside Rake callers on every physical CPU profile. - Extend native
f32streams to NEON: four-lane vector transfers and count-guarded partial-rack memory operations, with vector arithmetic in the tail. The final ELF function extent is checked against the selected assembly. Independent C and guard pages check one to four columns, exact in-place output and million-element results on all four CPU profiles. NEON ordered comparisons preserve quiet-NaN behaviour through vector operand sanitisation. The numerical oracle checks all six operators. - Correct the safe-root check-only command to include its million-element comparison before returning. Earlier check-only runs covered short counts and guarded tails.
- Extend native
f32streams to SSE2: four-lane racks with guarded memory transfers for partial racks and vector arithmetic throughout. Independent C and guard-page checks pass on SSE2, AVX2 and AVX-512F. SSE2 ordered comparisons now preserve quiet-NaN behaviour through explicit vector sanitisation, with additional mask-register pressure checked by the allocator. - Extend native
f32streams to AVX-512F: sixteen-lane racks, masked loads and stores, and inactive-input sanitisation. Independent C and guard-page checks exercise one, two and four columns on the AVX2 and AVX-512 stream profiles, including every tail remainder and exact in-place output. - Place slow-local aggregate frames using their actual C size and alignment, including header-backed records nested in arrays and Rake records. Small frames retain aligned host-stack storage. Arena padding is reclaimed on return, and independent small-stack checks cover 64-byte-aligned C objects.
- Native large-local arenas allocate on entry and free on the outermost framed return. Only a pointer and cursor occupy TLS, allowing host runtimes to use small worker stacks. Independent pthread checks cover recursive locals, C callback re-entry, returned records and repeated calls.
- Define typed global tines with explicit rack and uniform parameters, and apply them in local tines, through blocks and sweeps.
gapscomplements a tine or composed mask, including unordered floating-point lanes. Through fallbacks are optional: withoutelse, reading an undefined lane is a compile error. Sweeps may omit_when Boolean mask coverage is provably total. Physical and WebAssembly lowering retain their existing predication and inactive-input safety obligations. - Support opaque header-backed C typedefs through typed pointers, including callback signatures. Reject by-value opaque objects and their construction.
- Compile the AVX2
f32stream traversal subset as complete Rake-selected assembly, with masked tails and exact final-function verification. - Rename
crunchtoscratch, without retaining the old keyword.packnow defines one scalar record.stack PackTypedescribes its columnar collection, constructed asstack PackType { field: column }. Runs consume those columns a rack at a time. The compiler, grammar, examples and tutorial use the same hierarchy. - Header-backed C unions have overlapping storage, one-member literals, field access, borrowed arguments and by-value C calls/results. Lowercase C typedefs and primitive-spelled members retain their header identifiers. The platform compiler supplies layout and ABI. Independent C checks cover nested structs, arrays and callbacks on CPU profiles and WebAssembly. Interpreted union storage and Rake-owned union layouts remain WIP.
- Slow pointers distinguish writable
ptr Tfrom read-onlyptr const T. C declarations retain their pointee qualifiers, including callback signatures. Address-taking, view borrowing and opaque pointer casts preserve read-only access. The interpreter now retains location identity through scalar and aggregate assignment, so pointers observe later writes. Taking a checked element's address also checks bounds in the interpreter. - Slow code supports typed C function pointers, noncapturing callbacks and opaque
ptr ()contexts.addr(function)forms a callback, and data pointers can be erased or restored throughbitcast(ptr T, pointer). Native and WebAssembly calls use the platform C ABI, with null indirect calls trapping. - A process entry may take
argc: i32, argv: ptr ptr u8. Native C and WASI startup pass the count, zero-terminated byte strings and trailing null pointer through a compiler-owned typed adapter. The interpreter accepts arguments after--and uses the input path asargv[0]. - Native slow-only programs emit platform C and compile to x86-64 or AArch64 objects. Their public slow functions use the platform C ABI, with header-backed struct imports, pointer parameters and struct returns. Native runs and packs remain work in progress.
- Native mixed programs embed Rake-selected register assembly in their C unit. Slow callers use uniform
f32parameters andf32results, and both object modes check the kernels in the final compiled object. Other scalar kernel boundaries remain work in progress. - The interpreter accepts an explicit target profile for reference
f32rack widths. The browser emits native mixed C and interprets it at that profile's width; it does not execute native machine code. - Native slow frames are thread-local; independent host threads can call exported functions with large recursive locals. Module state stays shared.
- Pointer fields in C struct declarations are checked against
sizeof(void *), rather than assuming a wasm32 pointer. - Slow scalar bit casts use an alias-safe byte copy, shared by the native and WebAssembly C paths.
0.6.0-beta
- SSE2 and AVX-512F compile float scratches and rakes to verified vector kernels alongside AVX2 and NEON. Rack widths are four, sixteen, eight and four binary32 lanes respectively. Native pack traversal and whole programs remain work in progress.
- Native target detection selects AVX-512F, AVX2 with FMA, or SSE2 on x86. SSE2 rejects explicit single-rounded
fmabecause that instruction is absent from its ISA. The AVX-512 profile requires AVX-512F only. - Scratches end with their result expression, without
return. Tines usemeansinstead ofwhen. These are breaking source-language changes. - The browser tutorial offers all five implemented vector profiles for kernel inspection. Execution and whole-program lessons use WebAssembly.
- Tree-sitter preserves a function body across comment-only lines, including comments beginning at the left margin.
- The introduction explains the notation in sequence, with compound masks and a diagram of a 600-record pack. The reference page is now titled Primitives, operations, and targets.
0.5.0-beta
- Rakes end with
sweep:, withoutreturn. This is the rake's result form.returnremains the result of a scratch and an exit from a slow function. Compiler examples, Tree-sitter and the browser tutorial use the new form. slow { ... }is a scoped scalar escape in runs and slow functions. It supports scalar loops, calls, state and memory views, nesting and a scalar tail result. Rack values can't cross the boundary. WebAssembly object verification permits only the helper calls selected by these explicit blocks, and keeps the surrounding vector checks.- Tree-sitter and the browser tutorial recognise and teach slow blocks.
0.4.0-beta
Changes since 0.3.0.
Language and compiler
- Named target profiles fix rack widths.
x86-avx2andaarch64-neoncompilef32sscratches and rakes to assembly, andwasm-simd128compiles scratches, rakes, runs and whole programs to C with onewasm_simd128.hintrinsic for each selected instruction. rakec --print-capabilitiesreports each language feature's status, and each profile rejects what it can't compile with the source location and the missing operation.- Rakes have defined inactive lanes on every profile: benign operands on x86 and AArch64, and none needed on WebAssembly, which has no floating-point exception state. A sweep needs one final
_arm, and its arms' order is its priority. - Fused bindings are pure contiguous data flow. Their identifiers are substituted before fused multiply-adds are formed on AVX2 and NEON, and
fmakeeps its single rounding. A fused stage may now introduce a constant or choose by a mask, and fused bindings inside a through block form regions of their own. - Pure instructions that nothing uses are removed before instruction selection, so a scratch's
repeatcompiles on AVX2 and NEON. wasm-simd128adds:- integer racks,
u8s,i16s,i32s,u32s,i64sandu64s, with wrapping arithmetic, comparisons,minandmax, bitwise operations and bit shifts, dot,narrow,widen_low,widen_high,to_f32andto_i32,- float
min,max,absand rounding, - one- and two-rack shuffles,
bitmask,extract,insert,allandany, and - strict reductions and scans, which AVX2 also compiles.
- integer racks,
exp,log,log2andtanhare fixed sequences of binary32 operations, shared by the interpreter, slow code and racks.- Conditional expressions choose by a mask on every profile, and by a uniform on
wasm-simd128. - A uniform must be marked where it meets a rack:
values * <factor>, notvalues * factor. The type checker reports the unmarked name. wrapandbitcastare conversions only before a parenthesis, so they remain usable as identifiers. A scratch, rake or run can't be named after a C keyword, since each becomes a C function of its own name.- A rake's through block and its sweep compile to one select, since a select whose arm selects on the same mask takes that arm directly.
rakec --print-capabilitiesmarksletexpressions,iscomparisons and expression predicates unavailable, since no source syntax reaches them.!=on floats is ordered everywhere: it is false when either operand is NaN, for scalars as for racks.- Number literals are unsigned, and a minus before one makes it negative, so
n-1subtracts. - Runs on
wasm-simd128traverse packs, with tails that load and store lane by lane and never touch an element past the count. They have nested traversals, counted loops,repeatand uniformif, rack locations and arrays of them, checked and unchecked loads, stores and gathers, and scratches and rakes inlined with their masks. Rake forms each loop's addresses and checks bounds once before the loop. Runs have a published wasm32 C boundary. - A run's uniform arithmetic may use reductions, extractions, bitmasks and scratch calls, each computed as a uniform of its own.
- Slow code on
wasm-simd128: records, arrays, views, pointers, module state, embedded files, constants, control flow and calls to C, compiled with the runs into one C file withint main(void). Aggregates over 256 bytes live in a frame stack in linear memory instead of the 64 KiB C stack, which they overflowed without a trap. rakec --interpretruns a program'smainin Rake's executable semantics. It reports a program withoutmaininstead of failing, and broadcasts a uniform stored where a rack goes, as the compiled code does.- The compiler library exposes the front end, interpreter, C and assembly emitters, and a bounded lane trace to the browser through
js_of_ocaml. The browser build runs in a worker and returns source diagnostics without requiring a server-side compiler. wasm-simd128-relaxedis an opt-in profile withrelaxed_madd,relaxed_nmadd,relaxed_minandrelaxed_max.
Verification
--verify-nativedisassembles every object. On x86 and AArch64 it rejects calls, stack use, split or scalarised racks, instructions outside the profile's list and a wrong count of fused multiply-adds, and it accepts the two-byte nop that pads an AVX2 function. Onwasm-simd128a scratch may hold only register instructions, including the scalar arithmetic clang makes of splat arithmetic, and a run only the vector instructions its source selects or their documented equivalents.test/program_test.shcompares every program fixture in the interpreter and as WebAssembly under wasmtime, andtest/abi_test.shcalls runs from C.test/neon_backend_test.shcompares NEON results under QEMU.- The compiler's parser and the Tree-sitter grammar are compared over the test corpus. The grammar parses indentation with an external scanner, as the compiler does.
- Every Rake example in the README, the documentation and the website is compiled by
tools/check_documentation_examples.sh, and most are run.
Website and documentation
- rake-lang.org now publishes the compiler documentation as a checked set of reference pages with Tree-sitter highlighting.
- The playground is an interactive twelve-lesson tutorial. It runs the same compiler and interpreter in the browser, shows rack values in the Lanes tab, and displays generated C or assembly beside source diagnostics.
Removed
- The experimental MLIR and LLVM lowering and its command-line modes.
- The toolchain-generated memref C wrapper for runs.
sectionYPositions = computeSectionYPositions($el), 10)"
x-init="setTimeout(() => sectionYPositions = computeSectionYPositions($el), 10)"
>
On This Page