Repository navigation
Supporting caller-allocated, callee-written buffer without views #68
Description
Activity
giving the callee a permanent window into the caller's linear memory,
Views and
ArrayBuffers are already detachable -- perhaps we can reuse this mechanism to avoid making the views permanent? (They will already be detached on memory growth, which is something we probably also need to think about, but which I haven't done yet).
Regarding
list-to-memoryand its proposed lambda: I want to make sure I am understanding this right, the lambda is mapping a list length to a pointer? To a byte length? Is maybe the genericlist-to-memoryversion missing thearg.get $ptrinstruction that exists in the non-genericlist-to-pre-allocated-memoryversion?ArrayBuffers and typed array views are, but not necessarily a futureslicerefthat gets added as a first-class wasm reference type from which bytes can be loaded and stored. The cost of detachability is constantly needing to probe "is the underlying buffer detached?".
Oh whoops, sorry, I wrote the wrong arg, it should've been
arg.get $ptras you said.Ok so just to be 100% clear: in
list-to-memory f,fis called just once and maps a byte length to a pointer (possibly pre-allocated, possibly malloced) where the bytes can be copied into?Yup! That's the idea at least. Currently in the explainer,
string-to-memorytakes anfthat maps byte length to a pointer, but thefis restricted to be a wasm exported function name, so you could consider this to be a generalization of that.Reacted by Nick FitzgeraldI realized a problem with the above approach: if what we're trying to do is optimize calls into a native runtime (e.g., a native impl of the
read()call mentioned above), it relies on some rather magic compiler/runtime optimizations to effectively hoist thelist-to-memoryso that it happens during the native call (the point at which the native code needs to write its outgoing bytes into something). This seems like a complicated, fragile and incomplete optimization (b/c any intervening effect can break it).Thinking once again about the "slices" approach: one option is to say that the
sliceinterface type is never meant to flow out to core wasm as a first-classslicerefvalue; but, rather,slices are meant to stay encapsulated within adapter code (just likestring) which forces them to only be written into during the adapted call, no different than other interface types.Now there is still the problem that if the other side of a
slice-taking call is JS, that JS gets a persistent typed array view, but maybe that's just "ok"; as long as the long-term shared-nothing wasm-to-wasm precedent is established...There is a use case that does not seem to be addressed by this: using buffers to communicate between modules outside the call to set it up. This seems to be important for graphics buffers where one module is writing to the shared buffer and another reading from it.
True, but I think that use case is best addressed by either multi-memory + memory imports or, to support dynamic acquisition of buffers, a first-class core-wasm "slice reference" that can be used as a dynamic operand to the load/store ops (which will likely be added to wasm at some point in time). I think the only thing such a use case asks from interface types is that interface types not require a shared-nothing boundary (thereby disallowing the memory import or slice reference), which is currently the plan.
For an API like the C standard library
read:where the caller allocates a buffer which the callee (partially) fills in, to express the signature in wasm, ideally we want to avoid
bufbeing ani32argument since (1) this implies linear memory (preventing GC buffers), (2) it isn't multi-memory future-compatible.One solution is to have
readtake a "slice"/"view" parameter (analogous to a Web IDLUint8Array). However, views on linear memory somewhat break encapsulation, by giving the callee a permanent window into the caller's linear memory, leading to use-after-free types of bugs, so I think it's worthwhile to ask whether we can support use cases likereadwith value/copy semantics, without loss of performance.Setting aside the issue of how to signal failure (probably via variant/option return type), I think a natural interface-typed signature would be:
where the API's contract is that the returned list's length is <=
$nbyte. The question is how we can have this interface, but achieve the same performance, in particular, allowing the caller to supply a fixed-size buffer.One idea would be to introduce a variant of the
*-to-memoryinstructions which, instead of calling an exported function to allocate linear memory, would take a (pointer, length)i32pair as operands from the stack (trapping if the required space is greater than the length). E.g.:However, there will be many
*-to-memoryinstructions, and it feels wrong to have to add a*-to-preallocated-memoryvariant for each.An alternative is to generalize the
*-to-memoryfunctions to, instead of only calling exports, to also be able to call any helper function (as described in #65), and for it to be possible for helper functions to be defined inline as unnamed (lambda) functions, so that they can reference the enclosing scope. This would allow the above example to be equivalently expressed as:This generalization would also allow interesting hybrid schemes, e.g., wherein a caller-supplied buffer was used opportunistically, falling back to malloc in too-big cases, which is another common C++ optimization pattern.
With the understanding that helper functions are always designed to be inlined at the callsite (which is always statically determinable), then these lambda functions should compile into direct stack access, with no worse perf than
list-to-preallocated-memory.Thoughts?