package ppx_windtrap
Install
dune-project
Dependency
Authors
Maintainers
Sources
sha256=2e61a86f8c97502c1f8a593c28e255d591a44b1533c6eed8cd0c847afa3d772a
sha512=060d0a47d926d420e0f96f4912c2690d148d6196bcbb701b2463a4fb9ffade5c4ff4d98d703f4b80b2b8ad1bed848c68f1d613e3ba745f9cd5e7d313e049bdc3
doc/CHANGES.html
Changelog
v0.2.0 2026-09-26
Windtrap 0.2.0 is a complete rewrite of Windtrap. It keeps the shape of 0.1.0's API and rebuilds it on a stronger design: one runner for every kind of test, one seed for every generated value in a run, and failures kept as data that the report renders. It also fixes many correctness problems, listed below.
The main additions are stateful testing and mutation testing.
- A stateful test generates sequences of calls and runs each call on the system under test and on a reference, a model written for the test or another implementation. It fails at the first call where the system does not return or raise what the reference does. When a sequence fails, windtrap removes calls and simplifies their arguments for as long as the sequence keeps failing, so the report shows a short sequence that reproduces the bug. It follows Monolith and qcheck-stm, whose parallel mode it shares: with
~domains, the middle of each sequence runs on several domains at once. - Mutation testing makes small changes to the code under test, such as turning
<into<=or&&into||, and reports each change that no test fails on. Each such change points at a gap in the suite.
Windtrap now runs five kinds of test:
- unit tests, with assertions;
- expect tests, inline in a library as with ppx_expect and ppx_inline_test, or in any test with
expect; - snapshot tests, now part of the expect API through
expect_file; - property tests, as with QCheck;
- stateful tests, as with qcheck-stm and Monolith.
It measures a suite in two ways: coverage, as with Bisect_ppx, and mutation testing.
The core of a suite reads as before: test, group, equal and witnesses such as int and list keep their meaning, and so do let%expect_test and [%expect]. Around that core, 0.2.0 makes substantial changes, improvements and additions to the API. The API is flat, so lib/windtrap.mli documents all of it in one file.
See doc/manual/migrating-from-0.1.md for how to migrate from 0.1.0.
Changes that keep a 0.1 suite compiling
These changes compile without an error and change what a 0.1 suite does. Each one is also listed under its area below.
- A test tagged
"disabled"runs, where 0.1 skipped it.skip ()in its body, or--exclude-tag disabledon the command line, leaves it out. - The global
Randomstate is seeded from the path of each test, where 0.1 started every test fromRandom.init 137. - Under
float epsandfloat_rel, NaN equals nothing, where 0.1 found NaN equal to NaN, andfloat_relno longer finds an infinity equal to every float. - A mismatched
expector[%expect]records the failure and returns, where 0.1 raised, so the rest of the test runs. - A call to
exitinside a test fails that test, and the run continues. output ()fails the test under--streamand raisesInvalid_argumentoutside a test, where 0.1 returned"".- A library's inline tests run only in that library's inline runner, never in a
(test)executable that links the library. - A command line that does not parse exits 2, where 0.1 exited 1, and a selection that keeps no test exits 2, where 0.1 exited 0.
-fand-eadd up when repeated, where 0.1 kept the last of each.CIandGITHUB_ACTIONScount as unset when they are empty or0,false,no,noroff.--junit PATHtreats aPATHwithout.xmlas a directory and writes<suite>.xmlin it.- Coverage counts a call that raises as uncovered, so a percentage cannot be compared with a 0.1 one.
WINDTRAP_COVERAGE_FILEnames the dump file itself, where 0.1 used it as a prefix for generated file names.WINDTRAP_UPDATE,WINDTRAP_COVERAGE_LOGand theinline-test=dropcookie are not read.
Declaring tests
- (breaking)
run : ?argv:string array -> string -> test list -> intreturns the exit code instead of callingexit; a suite ends onlet () = exit (run "mylib" tests)(see Reading the exit code). - (breaking)
run's arguments~quick,~bail,~fail_fast,~filter,~exclude,~tags,~exclude_tags,~failed,~list_only,~seed,~timeout,~prop_count,~format,~junit,~stream,~output_dir,~updateand~snapshot_dirare removed, and so istype format; the flags, their mirrors or a synthetic~argvconfigure a run. - (breaking) A test body returns
unit, where 0.1 ignored the body's result; a body that computes a value ends on an assertion or onignore. - (breaking)
?posand?hereare replaced by?__POS__on every constructor and every verb, andtype hereis removed;~__POS__passes the call's position. - (breaking)
Tagis removed; tags are astring list, as in~tags:[ "net" ], and a group's tags are added to those of every test under it. - (breaking)
slow name fnistestwith the tag"slow", which--exclude-tag slowleaves out; a test under a slow group can no longer be marked quick. - (breaking) A test tagged
"disabled"runs, where 0.1 skipped it, and no tag has a meaning for selection;skip ()in the body skips the test, and--exclude-tag disabledleaves it out. - (breaking)
group's~setup,~teardown,~before_eachand~after_eachare removed;bracketandscopedgive each test a resource, andfixtureshares one across the run (see Giving each test its own resource). - (breaking)
test,group,slow,cases,bracketandscopedraiseInvalid_argumentwhen applied if~timeoutis not finite and positive or~retriesis negative, where 0.1 ran~timeout:0.as a one-second alarm. - (breaking)
cases ~name base inputs fntakes a naming function in place of a witness; each child is the testname inputunder the groupbase. - (breaking)
ftestandfgroupare removed;focus tfocuses any test or group, and the tests outside the focus have no row and no count (see Focusing on one test). - (breaking)
runrefuses a suite that holds afocusunderCI, printswindtrap: focused tests committed (focus at <file:line>); remove focus to run under CIand returns 1. runraisesInvalid_argumentwhen called while a run executes, as from a test body; two runs in a row in one process are allowed.- Without
~__POS__, a declaration or an assertion is located from the call stack, which needs-g, dune's default (see Locating a failing assertion). - A
caseschild is a test that-fselects alone, and two inputs with one name makerunrefuse the suite. test,group,slow,cases,bracketandscopedtake?__POS__,?tags,?timeoutand?retries; on a group,timeoutandretriesapply to each test under it that declares none, as ingroup ~timeout:30. "integration" tests.xfail ?reason tmarks a test or group as expected to fail; its failure leaves the exit code,-xand--failedalone, and a pass fails withexpected to fail (<reason>), but the test passed(see Keeping a known bug in the suite).
Assertions
- (breaking) The value
testable ~pp ?equal ?gen ?check (),Testable.gen,Testable.with_gen,Testable.checkandTestable.check_resultare removed;Testable.make ~pp ~equaltakes a required equality and no(), andTestable.structural ~ppcompares withStdlib.( = ). - (breaking)
of_equalandcontramapare removed from the top level;Testable.of_equalandTestable.contramapreplace them. - (breaking) The witnesses
seq,lazy_t,small_intandnatare removed;Testable.contramap List.of_seq (list t)andTestable.contramap Lazy.force treplace the first two, andintthe others. - (breaking)
Ppis removed from the interface; a printer has type'a printer, which isFormat.formatter -> 'a -> unit. - (breaking)
some,ok,errorandno_raiseare removed;some t e visequal (option t) (Some e) v,ok t e risequal t e (require_ok r),error t e risequal t e (require_error r), andno_raise fnisfn (). - (breaking)
raises_invalid_arg m fnandraises_failure m fnare removed;raises (Invalid_argument m) fnandraises (Failure m) fnreplace them. - (breaking) Under
float epsandfloat_rel, NaN equals nothing, where 0.1 found NaN equal to NaN. - (breaking)
float epsraisesInvalid_argumentunlesseps > 0.; exact equality isfloat_exact. - (breaking)
float_relraisesInvalid_argumentwhen a tolerance is negative or NaN, or when both are zero. float_relno longer finds an infinity equal to every float.float epsandfloat_relprint asfloat_exactdoes, where 0.1 printed six significant digits, so a failure never shows two unequal floats alike.Testable.of_equalprints<abstract>, where 0.1 printed<opaque>.slistprints both sides sorted, where 0.1 printed them as given.raisesandraises_matchlet askip, a timeout, an interceptedexitandassumepass through; 0.1'sraisesreported a skip inside it asWrong exception raised, and araises_matchthat accepted any exception swallowed it.raisesandraises_matchreport a wrong exception with the raised one's backtrace.- When
raisesexpects anInvalid_argument,FailureorSys_errorand gets the same constructor with another message, the failure diffs the two messages. - A failure prints
expectedandactualand marks what changed for every witness, from the two printed values: the changed spans of one-line values, and a unified diff with@@hunks for multi-line ones (see Comparing two values). - A changed span is bold in its side's colour, and a
~line marks it unless the report is coloured on a terminal. - Two unequal values that print alike are reported as
both sides render as: <v>. ~msgprints a line above the values, where 0.1 printed it in place of the headline; the headlines, such asValues are not equal, are removed.- A failure prints its location as
file:line, then the source line of the call. - An uncaught exception prints as
uncaught exception:with its backtrace, whichrunrecords withoutOCAMLRUNPARAM=b. - Exception names print without dune's
Dune__exe__prefix. float_exactcompares floats bit for bit, orders-0.below0.as its equality tells them apart, and prints the shortest decimal that round-trips.textis a string witness printed verbatim, so two multi-line texts fail with a line diff.less,at_most,greaterandat_leastcompare under the witness's order, which the base witnesses carry andTestable.with_comparegives; a witness without an order raisesInvalid_argument.require_some,require_ok,require_errorandrequire_matchassert a shape and return its payload (see Unwrapping an option or a result).is_none,is_okandis_errortake?ppand print the payload they did not want, as<abstract>without it.starts_with ~affix,ends_with ~affix,contains ~sub,not_contains ~subandin_order ~subssearch a string, and a failure says where the needle was or was not found in an excerpt of the haystack.satisfies ?claim t pred vassertspred v, andmem t x xsasserts thatxsholdsx; both print the values witht's printer.Exn.invalid_arg,Exn.failureandExn.sys_errorare predicates forraises_match, each with?substring.Lawasserts seventeen textbook laws, such asLaw.associative w op (a, b, c), and a failure names the law, states its equation and prints every term it computed;prop "…" Gen.(triple g g g) (Law.associative w op)checks one over drawn values (see Stating a textbook law).
Property testing
- (breaking)
prop ?__POS__ ?tags ?timeout ?count ?max_discard ?examples name gen lawtakes aGen.tand a law'a -> unitthat asserts with the verbs; a 0.1 boolean law becomesfun l -> is_true (law l). - (breaking)
prop',prop2,prop3andprop4are removed;prop'isprop, and several inputs are one tuple fromGen.pair,Gen.tripleorGen.quad. - (breaking)
~configis removed;?count,?max_discardand--seedreplace it. - (breaking)
Gen.tis abstract and built from the combinators alone;Gen.make_primitive,Gen.no_shrink,Gen.add_shrink_invariantandGen.findare removed. - (breaking)
Gen.oneof,Gen.oneoflandGen.pureare renamedGen.one_of,Gen.of_listandGen.constant;Gen.list_size sg gisGen.list ~size:sg g, andGen.string_size sg cgisGen.string_of ~size:sg cg. Gen.of_list ~ppandGen.constant ~pptake the printer of the values they list, so a list of chosen values prints withoutGen.with_pp.- (breaking)
Gen.sizedis removed;Gen.bind Gen.nat freplacesGen.sized f. - (breaking)
( >>= ),( >|= )andGen.apare removed;let*,let+andand+replace them. - (breaking)
Gen.fixandGen.delayare removed; a recursive type is generated by alet recgenerator over a depth (see Generating a recursive type). - (breaking)
Gen.int32_rangeandGen.int64_rangeare removed;Gen.map Int32.of_int (Gen.int_range lo hi)replaces the first, and likewise forInt64. - (breaking) The
?originofGen.int_range,Gen.float_rangeandGen.char_rangeis removed; a range shrinks toward its point closest to 0, or closest to'a'forGen.char_range. - (breaking) The
?ratioofGen.option,Gen.resultandGen.eitheris removed;Gen.frequencyweighs choices. - (breaking)
Gen.floatdraws finite floats only, where 0.1 also drew NaN and infinities. - (breaking)
Gen.float_range low highincludeshigh, where 0.1 excluded it, and sampling it raisesInvalid_argumentwhen a bound is not finite. - (breaking)
--seedandWINDTRAP_SEEDtake the tokens1:and 16 lowercase hexadecimal digits; any other spelling is a usage error. - (breaking)
cover ~label ~at_least condiscover label cond, which fails the property withnever covered: "label"when no passing case markedlabel. Coverage is judged once every case has run, so a property that fails on a case or gives up lists no label as never covered. - A run has one root seed, and every case derives from it, the test's path and its index; the header or the summary prints
(seed s1:…)when a property is selected. - A failing property's block reads
counterexample (case N, shrunk K steps): <value>and shows the law's own failure with its diff. A report whose failures include a property or a stateful test has onereplay: <command> --seed <token>line above the summary, which reruns the run's selection on the values it drew. - A counterexample prints with its generator's printer, and a value computed by
Gen.maporGen.bindprints ascomputed from <pre-image>. - Shrinking runs the law at most 10,000 times, accepted and rejected candidates alike, where 0.1's
max_shrinkdefaulted to 100. - A shrink cut by its budget of law runs or by a timeout says
counterexample may not be minimal. - A
skipor timeout inside a law is no longer shrunk as a counterexample. - A property gives up once more than
max_discardcases are discarded, twice the effective count by default. collectandclassifylabels print as a table in a failing property's block, and under-vfor a passing one.- A property carries the tag
"prop", and a stateful test"prop"and"stateful". ~examples:[ v1; v2 ]runs the law on those inputs first on every run, unshrunk (see Keeping a counterexample as a regression).Gen.such_that p gendraws untilpholds, at most 100 times, and then discards the case.Gen.with_pp pp gengives a generator its printer.Gen.arraytakes?size, andGen.bytes_ofgeneratesbytes.Gen.list,Gen.array,Gen.string_ofandGen.bytes_ofwithout~sizedraw a length below 64 and about 5 on average;~sizesets a longer one.Gen.int,Gen.int_range,Gen.int32,Gen.int64andGen.nativeintdraw a corner case with probability 0.1: a range's bounds, its point closest to 0 and that point's neighbours, or a type's 0, 1, -1 and extremes.
Stateful testing
stateful name commandsruns generated programs of calls on a system and on a reference, a model written for the test or another implementation, which judges each outcome of the system, and fails at the first call whose outcomes differ (see Writing a stateful test).command name signature reference systemis one operation of the API; a signature takesg @-> …for an argument drawn fromg,t ^-> …for a value oftthat an earlier call made, and ends inreturns w,makes torchooses w.abstract prefixdeclares a type whose values only calls make, namedq1,q2underabstract "q";?ppprints a value's reference side in a failing program,?invariant r sruns on every value after every call, and?release sreleases a system side when a program ends (see Releasing what a program made).- An exception is an outcome: two exceptions are equal when their constructor names match without the module path, and their payloads are not compared (see Checking the exceptions an operation raises).
chooses wcompares an outcome the API leaves open: the reference receives the system's outcome and returns or raises the one it accepts (see Checking an outcome the API leaves open).?preis asked of the reference's arguments when the program runs, and a call it refuses is skipped on both sides and absent from the report.- A stateful test runs
?countprograms (default 100) of at most?stepscalls (default 20). - A failing program shrinks by removing calls and shrinking arguments, and prints as a table of the calls its failing run executed, with the reference side of each argument before the call when its type has a
~pp(see Reading a failing program). - A verb's failure,
Assert_failureorMatch_failureis never an outcome: in a system function it fails the case at that call, in a reference function, like anything a~preraises, it breaks the reference, whose failure shrinks apart from the system's and prints asreference of call N of N(~pre of call N of Nfor a~pre), and in achoosesreference it is the system's mismatch (see Reading a failure of the model). assumeorrejectin a command fails the case withassume or reject in a command; a call's legality is its ~pre.- A drawn argument whose generator has no printer fails the test with
push: argument 2 has no printer; attach one with Gen.with_pp. stateful ~domains:nruns the middle of each program onndomains at once, 50 times, and fails when no order of the calls replayed on the reference gives the outcomes the system gave; the test carries the tagparallel(see Testing on several domains).- On several domains a command that makes a value or has a
~preruns only before the parallel calls;stateful ~domainsabove 1 raisesInvalid_argumentwhen every command makes a value or has a~pre. - A failing program on several domains prints a
domainand aresultcolumn, thenno order of the calls gives these resultsand the closest order, asthe closest order, 2 then 3, differs at call 4: length q1. - On several domains a reference whose replay gives a call before the parallel calls another outcome than the run breaks with
a replay of the reference differs from this run; the reference must behave the same from run to run. - On several domains a replay draws the same programs but not the same schedules, so it may pass.
- A stateful test on several domains takes no retries, a group's included.
- Under
--mutateand--arma stateful test on several domains runs each program once on the test's domain, so a kill does not depend on a schedule. - A stateful test on several domains whose domains cannot be spawned fails with
cannot spawn a worker domain: <message>. - A stateful test fails with
never called: "pop" (over 100 passing cases); a call runs only where its arguments resolve and its ~pre holdswhen a command that a passing case could draw was called by none; a command listed twice is one command, and~count:0judges nothing (see Writing a stateful test).
Baselines and expect tests
- (breaking)
snapshot,snapshot_ppandsnapshotfare removed, and no file lives under__snapshots__;expect s @@ __POS_OF__ {|…|}orexpect_file s "test/name.expected"replaces them. - (breaking)
expectandexpect_exacttake the produced text and a__POS_OF__literal, as inexpect (output ()) @@ __POS_OF__ {|…|}; a failure is located at the literal, and a correcting run rewrites it. - (breaking)
captureandcapture_exactare removed; a test comparesoutput ()with any verb. - (breaking)
output ()fails the test under--streamwiththis test requires capture; rerun without --stream, and raisesInvalid_argumentoutside a test, where 0.1 returned""in both cases. - (breaking) A mismatched expectation or
[%expect]node records the test's failure and returns, where 0.1 raised; one run reports every stale expectation of a test, and one correcting run corrects them. - (breaking)
WINDTRAP_UPDATEis not read; a(test)stanza runs%{test} --corrected, which writes<file>.correctedbeside each file, anddiff?s them fordune promote(see Running expectations under dune). - (breaking)
WINDTRAP_SNAPSHOT_DIR,WINDTRAP_SNAPSHOT_DIFF_CONTEXT,WINDTRAP_SNAPSHOT_MAX_BYTESandWINDTRAP_SNAPSHOT_REPORTare not read. - (breaking)
-uis refused underCI(exit 1), and-uwith--correctedis a usage error. - (breaking)
[%%run_tests]is removed, and inline tests in an(executable)no longer run at exit; the tests belong in a library with(inline_tests). - (breaking) An
(executable)that registers inline tests exits 2 withwindtrap: registered inline tests were never driven. - (breaking) A library's inline tests run only in that library's inline runner, never in a
(test)executable or program that links the library. - (breaking)
[%expect],[%expect_exact]and[%expect.output]compile only inside alet%expect_testbody. - (breaking) The ppx_expect forms windtrap lacks are compile errors, as
[%expect.unreachable]was in 0.1; an attribute such as[@@expect.uncaught_exn], which 0.1 dropped silently, fails withattribute expect.uncaught_exn is not supported by ppx_windtrap; catch and print the exception before an [%expect]. - (breaking) The cookie
inline-test=dropis not read; dune'sinline_testscookie drops the test forms. -urewrites each literal and eachexpect_filefile in place, and every accepted test passes.output ()and the captured log include what C code wrote withprintf, since capture flushes the Cstdoutandstderrbuffers before reading.- A correction is kept only for a test whose every failure is a baseline mismatch and whose source is unchanged since the build; the block says why when none is kept.
- A baseline failure opens on
expect: mismatch,expect: no baselineorexpect_file "<path>": no baseline. Under dune its block ends onaccept: dune promote <file>; a run by hand ends on oneaccept: <command> -uline above the summary, over the run's selection. - A correcting run lists its files under
corrections (N):. - A correction keeps its literal's delimiter and lays out a multi-line text like ppx_expect, and a bare
[%expect]keeps its node. - An
expect_exactcorrection whose text holds a CR is a quoted literal with the CR written\r, so it passes once promoted. - Output a
let%expect_testbody writes after its last node fails the test asexpect: no baseline, and its correction appends;and an[%expect]node that holds it, as ppx_expect does; blank output passes (see Writing expect tests inside a library). - An
[%expect]or[%expect_exact]node that a run of its test never reaches fails the test withthe body returned without reaching this node, and the test keeps no correction; each functor instance is judged alone, and anExpect_test_config.runthat never calls the body fails at the body's first node. - The ppx_expect conformance corpus corrects
negative-tests/trailing.mlbyte for byte as upstream (15 corrections, 8 byte-identical), and the correctednegative-tests/escaped_strings.mlpasses (seetest/conformance/RESULTS.md). - The inline runner runs each partition as
run --correctedunder the suite<lib>/<file>, and exits 1 when a test failed outside its kept corrections. - A test name repeated in one scope, as in a functor applied twice, gets the suffix
(2). module%testkeeps the module's attributes.expect_file actual pathcomparesactualwith the file atpath, relative to the project root; a missing file is a mismatch that-ucorrects by writing the file, and under--correctedit fails the run (see Keeping a baseline in a file).- The project root is
WINDTRAP_PROJECT_ROOT, else the parent of dune's build directory, else the working directory; no marker file is consulted. - On Windows a backslash in the project root separates as
/does, so reports print paths relative to the root. Expect_test_config, fromppx_windtrap.config, wraps expect-test bodies withrunand rewrites captured output withsanitize, and a local module of that name overrides it (see Masking what changes between runs).
Resources and process state
- (breaking)
fixture ?teardown createis acquired by the first call inside a test and released after the last test of the run; calling it outside a test raisesInvalid_argument(see Sharing one resource across the run). - (breaking) A call to
exitinside a test fails that test withthe test called exit and was intercepted; a test must return or raise, never exit the process, and the run continues. - (breaking) While a test runs, the global
Randomstate is seeded from the test's path, where 0.1 started every test fromRandom.init 137. - What
bracket'steardownraises is a[teardown]failure listed beside the body's, where 0.1 replaced the body's failure withFun.Finally_raised. - What
bracket'ssetupraises is a[setup]failure. - A
Stack_overflowfails its test, and onlySys.BreakandOut_of_memoryend the run. scoped scope name fnis the test whose body receives the resource a scope such asIn_channel.with_open_text pathprovides; a scope that swallows the body's failure cannot pass the test.subtest name fnruns a named part of the current test and records its failure while the rest runs. Inside a property's law or a stateful test's function, its failure fails the case, which shrinks, and the report names the subtest under the counterexample.current_test ()returns the running test's path.temp_dir ()andtemp_file ()make scratch paths the runner removes when the attempt ends.setenvandchdirchange the environment and the working directory for the running test, and the runner restores both when the attempt ends (see Using files, variables and a working directory).
Running tests
- (breaking) A command line that does not parse prints one
windtrap:line and the usage and exits 2, where 0.1 exited 1. - (breaking)
--format,WINDTRAP_FORMAT, the TAP report and the dot reporter are removed; the report has two verbosities, the default and-v. - (breaking) A test the selection leaves out is no longer reported as skipped.
- (breaking) A selection that keeps no test prints
<suite>: no tests ran: filter "servr" matched none of 15 tests.and exits 2, where 0.1 exited 0. - (breaking)
-fand-erepeat, where 0.1 kept the last of each, and a test is kept when its path holds an-fpattern and no-epattern. - (breaking) Every bare argument, and every argument after
--, is an-fpattern. - (breaking)
-q/--quickand--bail Nare removed;--exclude-tag slowleaves out the slow tests, and-xstops at the first failure and counts the rest asN not run. - (breaking)
-lprints the paths of the selection in declaration order, where 0.1 printed every path sorted. - (breaking)
-lmakes the checks a run makes, so duplicate paths or afocusunderCIexit 1. - (breaking)
--failedkeeps a record per suite, and with nothing recorded it printswindtrap: no recorded failures match the current suiteand exits 2. - (breaking) A malformed mirror is a usage error naming the variable, where 0.1 ignored a bad value.
- (breaking)
--junit PATHwrites the file beside the terminal report, and treats aPATHwithout.xmlas a directory that gets<suite>.xml. - (breaking)
CIandGITHUB_ACTIONScount as unset when empty or0,false,no,noroff. - A run never fails because it cannot read or write the
--failedrecord, and a record it cannot read counts as empty. - An unknown long option names the nearest flag, as in
windtrap: unknown option '--juint'; did you mean '--junit'?. - A run with nothing to show prints one line, such as
storage: 12 passed, 2 skipped, 1 expected failure in 2.8ms., whose duration is wall-clock time (see Running every suite). - A selection given by mirrors alone that keeps no test of a suite exits 0 (see Passing flags to dune runtest).
- A failure's block prints when its test ends.
- When a failing test wrote output after its last
output ()call, its block closes on the last 10 lines of that output andfull log: <path>, where 0.1 showed no captured output. -vprints one row per test, with its status (PASS,FAIL,SKIPwith its reason, orXFAIL), path and duration, and no group header lines.- On a terminal a dim line such as
[3/15] <path>…names the running test, with or without-v. --slow-threshold SECONDS(default 1) lists underslow testsevery test not taggedslowthat ran that long;Slowest tests:is removed.- A test that passes on a retry is listed under
flaky testsand counted as(N flaky). - A mirror set to the empty string counts as unset.
WINDTRAP_TAIL_ERRORSandWINDTRAP_COLUMNS, which changed nothing in 0.1, are not read.--color autohonoursNO_COLORandTERM=dumb, and--colortakes any case.- A JUnit file that cannot be written prints
windtrap: warning: could not write JUnit report to <file>: <reason>and leaves the exit code alone. - The JUnit file names the suite, carries each failure's text, and leaves deselected tests out.
- Each test's captured output is kept in
<log dir>/<suite>/<groups>/<test>.output, overwritten by each run, with no run directories,latestlinks orTest output saved toline. - Under GitHub Actions each failure is one percent-encoded
::errorannotation after the group, and the summary is the last line. - SIGINT, SIGTERM and SIGHUP print
windtrap: interrupted in <path>and the summary, release the fixtures, and end the process by the same signal. - A timeout is measured to the fraction of a second, where 0.1 rounded it up to whole seconds, and fails with
timed out after 0.5s. - Runner messages print on stderr behind
windtrap:. - A control byte in a name, value or captured line prints as
\xNN. --shard K/Nkeeps bucket K of N of the selection, by a hash of each test path.WINDTRAP_VERBOSE,WINDTRAP_JUNIT,WINDTRAP_OUTPUT,WINDTRAP_SHARDandWINDTRAP_SLOW_THRESHOLDare new mirrors.--helplists each flag with its mirror.- A call on another domain that outlives its test's limit by one more limit fails the test as timed out and ends the run after it, with
run stopped after <test>: a call on another domain outlived the test's limitabove the summary; on Windows, where no limit is enforced, it hangs the run. output,expect,expect_exact,expect_file,collect,classify,cover, a fixture's accessor and the functions of the running test raise a failure when called from a domain other than the one that calledrun, which fails the running test when it reaches the test's domain, as throughDomain.join.
Coverage
- (breaking)
(instrumentation (backend ppx_windtrap))and--instrument-with ppx_windtrapare nowppx_windtrap.coverage, and dune rejectsppx_windtrapas a backend (see Instrumenting a library). - (breaking) A test run prints no coverage line;
windtrap coveragereports coverage. - (breaking) Coverage counts the entries of bodies and branches and the returns of calls, so a call that raises stays uncovered; a percentage cannot be compared with a 0.1 one.
- (breaking) The
windtrap coverageflags--summary-only,-C/--context,--skip-covered,--coverage-path,--source-pathand-jare removed; a positionalPATHreplaces--coverage-path, and--jsonreplaces-j. - (breaking)
--jsondropssource_availableanduncovered_offsets. - (breaking) An instrumented executable writes its dumps under
<build dir>/_coverage/windtrap-<hash>/, and the runner's-ono longer moves them (see Where the dumps are). - (breaking) A missing
PATHor an unreadable, corrupt or 0.1 dump makeswindtrap coverageexit 1, and a usage error exits 2 behindwindtrap:. - (breaking)
WINDTRAP_COVERAGE_FILEnames the dump itself, where 0.1 took it as the prefix of generated file names. - (breaking)
WINDTRAP_COVERAGE_LOGis not read; a dump that cannot be written is onewindtrap: warning:line on stderr. - An executable's first dump after a rebuild removes its older dumps.
windtrap coveragefinds the dumps from any subdirectory.- A dump whose executable was deleted or rebuilt is excluded with a line on stderr.
- The report puts each file's uncovered ranges on its row and ends on
coverage: 71.4% (312/437 points);-ukeeps the table and adds the uncovered source. --min PCTexits 1 belowPCT, and the last line then readscoverage: 71.4% (312/437 points), minimum 80%: FAILED.--lcovprints an LCOV tracefile, which genhtml, Codecov, Coveralls and editor gutters read; GitLab's merge-request view reads neither format windtrap writes.--expect PATHexits 1 unless every source underPATHhas coverage data, and--do-not-expect PATHexempts a file or directory from it.- On Windows
--expectnames each source without coverage data once, spelled with/. [@@coverage off]also excludes a module binding.
Mutation testing
(instrumentation (backend ppx_windtrap.mutate))with--instrument-with ppx_windtrap.mutatecompiles every mutant of a library into the build, each off until armed (see Instrumenting a library for mutation).- A mutant negates a condition, moves a comparison by one, swaps
&&and||, or swaps+and-, and is named<file>:<line>:<col>:<rewrite>. [@mutate off "reason"]and its[@@…]and[@@@…]forms dismiss equivalent mutants; a bare[@mutate off]dismisses every mutant of its expression, in later releases too, and a payload that names a rewrite is refused (see Dismissing an equivalent mutant).--mutate[=PREFIX,…]runs the suite, then each reached mutant in a child process with the tests that reached it, prints each survivor with those tests, and exits 0 (see Testing a suite's mutants).--arm IDruns the suite with one mutant armed and says whether it was killed.windtrap mutantsmerges the verdict files of the executables and exits 1 when a mutant survived every executable that reached it.--mutateis refused (exit 1) on Windows, in a process that spawned a domain, when the dry run fails or has no mutant in scope, and when the determinism probe disagrees with the dry run.- A mutant whose child passed without evaluating its site, as when the dry run cached the site's result, is not evaluated and no survivor; the report lists it under
not evaluatedwith thearm:command that tests it in a new process, and the merge keeps it not evaluated unless an executable killed it (see What a mutation run runs). - A site that no test reaches and that module initialization or a fixture release evaluates is listed under
evaluated outside tests, apart from the lines never reached (see What a mutation run runs).
Packages and libraries
- (breaking)
windtrap.prop(Windtrap_prop, withArbitraryandProp.check),windtrap.myersandwindtrap.clockare removed;windtrapholdsGenandprop. - (breaking)
windtrap.coverage(Windtrap_coverage) is replaced bywindtrap.runtime, which depends on the stdlib alone. - (breaking)
Windtrap.Ppx_runtimeis removed, andppx_windtrap.runtimeholds the inline-test runtime. ppx_windtrapdeclareswindtrapas a runtime library, so a stanza preprocessed with it need not listwindtrapin(libraries …).ppx_windtrap.coverageandppx_windtrap.mutateare the instrumentation backends, andppx_windtrap.configholdsExpect_test_config.- The
windtrapbinary has the commandscoverageandmutants, and an unknown command exits 2. - Linking windtrap no longer adds the top-level modules
ClockandMyersto an executable.
v0.1.0 2026-02-13
Windtrap is an all-in-one OCaml testing framework that unifies unit tests, property-based tests, snapshot tests, and expect tests under a single API. Instead of juggling multiple testing libraries, Windtrap gives you one cohesive package with a PPX for inline expect tests (ppx_windtrap).
- Unit tests with combinators, tags, skip, brackets, and timeouts.
- Property-based testing with configurable seeds and shrinking.
- Snapshot testing with automatic file management and diffing.
- Inline expect tests via
ppx_windtrapwith automatic correction. - CLI test runner with filtering, verbosity, and color support.
- Test coverage reporting with
bisect_ppxintegration.