Internal API

Note

This is the documentation of Enzymes's internal API. The internal API is not subject to semantic versioning and may change at any time and without deprecation.

Enzyme.Compiler.EnzymeError — Type
EnzymeError

Common supertype for Enzyme-specific errors.

This type is made public so that downstream packages can add custom error hints for the most common exceptions thrown by Enzyme.

source
Enzyme.Compiler.JuliaValueTable — Type
JuliaValueTable

The Julia values one module refers to by name, the part of a compilation's tables that module needs once the compilation is over:

  • slots: the value, and its address, behind each symbolic julia.constgv slot, keyed by the name of the slot (see make_slots_symbolic!);
  • inserted: the value each ejl_<key> global Enzyme inserted stands for, keyed by key (see insert_julia_value!).

julia_value_table takes it out of the context. _thunk resolves the module it compiled with it, and autodiff_cache keeps it next to the bitcode of the thunk, so that a later compilation importing the bitcode (import_cached_autodiff!) can take the values up into its own tables. Holding the values keeps them rooted for as long as the bitcode refers to them by name.

source
Enzyme.Compiler.PLTGot — Type
PLTGot

What the got of a PLT stub (jlplt_<name>_<n>_got) resolves to, read out of the stub before check_ir! walks any function (see plt_gots): lib and sym, what the stub looks up (lib a library name or handle, or nothing for a function of Julia's runtime, called by name). Only that is read ahead: it is what walking the stub folds away. The stub itself and the global it stores the address to are read where the got is rewritten, as rewriting another got may have deleted a global they share.

source
Enzyme.Compiler.EmitTypeNames — Constant
EmitTypeNames[] = true

Also write the printed Julia type next to each type Enzyme records in the IR (the enzymejl_parmtype_str attribute, and the enzymejl_source_type_<T> and enzymejl_allocart_name metadata). They help reading IR dumps.

source
Enzyme.Compiler.abs_ntuple_type — Method
abs_ntuple_type(arg::LLVM.Value, enzyme_ctx::Union{EnzymeContext, Nothing}) -> Union{Nothing, Tuple{LLVM.Value, Any}}

If arg is a call julia.enzyme.ntuple_type(T, count) with T known statically, return (count, T); otherwise nothing.

source
Enzyme.Compiler.arg_kind — Method
arg_kind(T) -> Symbol

Classify an argument of type T as get_specsig_function does: :ghost (omitted), :boxed (a tracked pointer), :byref (an aggregate passed through a derived pointer, plus an inline-roots pointer when it holds tracked pointers), or :byval (a scalar in a register).

source
Enzyme.Compiler.bake_inserted_values! — Method
bake_inserted_values!(mod, inserted)

Write the address of each Julia value the compilation inserted (inserted, see insert_julia_value!) into mod, which refers to it as ejl_<key>, as the module is linked into code that runs. The names stay local to the module and its table until then, so there is nothing to keep of them in a global; the well-known names (JuliaGlobalNameMap, JuliaEnzymeNameMap) the JIT resolves.

source
Enzyme.Compiler.bake_julia_value_globals! — Method
bake_julia_value_globals!(mod, inserted)

Replace each ejl_<key> global of the device module mod, which stands for the Julia value JuliaGlobalNameMap[key], JuliaEnzymeNameMap[key] or inserted[key] (the values the compilation inserted, see unsafe_to_llvm), with the address of that value. On the host the JIT resolves the well-known names, and the module the values were inserted into gets their addresses when it is linked (bake_inserted_values!); GPUCompiler 1.x resolves nothing in device code, so the address is written in when the derivative is handed to the kernel that requested it.

source
Enzyme.Compiler.bakes_julia_values — Method
bakes_julia_values(job)

Whether the back-end of job lets the host addresses of Julia values into the code it compiles. GPUCompiler 1.x always does; 2.x asks the back-end's relocation_lowering, which keeps them symbolic on :patch and :table.

The strategy is asked of a kernel = true flavour of the job: a deferred derivative is linked into the kernel that requested it, and a back-end whose strategy depends on kernel (Metal answers :table for kernels only) would otherwise report :bake for the non-kernel primal job Enzyme holds and let host addresses into a persisted kernel.

source
Enzyme.Compiler.box_inline_union! — Method
box_inline_union!(B, alloctx, val, offset, UT) -> LLVM.Value

val is an aggregate that holds an isbits Union field of type UT inline at byte offset: the payload, then a selector byte with the 0-based index of the member in Base.uniontypes order. Return the selected member boxed, as a tracked pointer; a singleton member is its instance. The tape slot of a Union tape holds the union boxed, as jl_type_to_llvm declares it, and the reverse rule takes such a tape boxed too.

source
Enzyme.Compiler.byref_alloca_type — Method
byref_alloca_type(inst::LLVM.AllocaInst, DL)

The Julia type of the untyped alloca inst (iN or [N x iM]), if it is passed (possibly through casts) by reference (BITS_REF) as an argument of calls whose parameter type is known (enzymejl_parmtype), all agreeing on one concrete immutable type of the alloca's size. Otherwise nothing.

source
Enzyme.Compiler.call_convention — Method
call_convention(mi::MethodInstance, ci::CodeInstance) -> Symbol

Decide from the inlining annotation how the function mi reaches the calling module. :inline means: emit its IR into the calling module and always-inline it. :call means: call its natively compiled entry point. An @inline method gets :inline, and a @noinline method gets :call. An unannotated method gets :inline exactly when Julia would inline it, as recorded on its native CodeInstance ci (see codeinst): the native compiler keeps the source of a method only when the method is inlineable. A method with a constant result gets :inline whatever its annotation, because Julia compiles no code for it.

source
Enzyme.Compiler.call_same_with_inverted_arg_if_active! — Method

Helper function for llvm-level rule generation. Will call the same function (and optional postprocessing), if the argument at index cmpidx isn't active. This takes into account runtime activity as a reason the value may not be active.

If postprocess_const is set, the original function will always be called, but the postprocessing will be conditionally gated as follows.

If the relevant input is active (and verified by runtime activity), postprocess(B, result, args) will run as normal Otherwise postprocess_const(B, result, args) will run

source
Enzyme.Compiler.check_emitted_specsig — Method
check_emitted_specsig(mod::LLVM.Module, llvmf::LLVM.Function, mi::MethodInstance, RT::Type)

Compare the signature of llvmf, which Julia's codegen just emitted into mod for the MethodInstance mi with return type RT, against the one specsig derives, and throw an AssertionError on a difference (see check_specsig). Only invoke_codegen! relies on the derivation, but prepare_llvm checks every function Julia emits, so the derivation meets every signature shape the differentiated code contains and a mismatch is reported where it arises rather than as corruption at the first native call.

source
Enzyme.Compiler.check_specsig — Method
check_specsig(llvmf, mi, RT)

Throw an AssertionError when the signature of the emitted function llvmf differs from the one specsig derives for mi and RT. The comparison covers the return type, the parameter types, and the sret and swiftself attributes. The other derived attributes only aid optimization and are skipped.

nested_codegen! compiles without gcstack_arg, so the derivation takes the presence of pgcstack from has_gcstack_arg.

source
Enzyme.Compiler.clear_caches! — Method
clear_caches!()

Empty every cache Enzyme keeps in a global, dropping what this session compiled, looked up and rooted along with them.

Nearly all of it means something only to the session that filled it. A CompileResult holds the address the JIT gave a thunk, captured_constants roots objects because their addresses were written into that code, the rule and activity memos are keyed on world ages, and the jl_load_and_lookup handles are ones this process opened. Anything outliving the session must not carry them, which is why Enzyme's precompile workload ends with this call: what it left behind would otherwise be serialized into Enzyme's package image and inherited, dead, by every session that loads it.

This is meant for the end of precompilation and not for a live session. It hands back the thunks the JIT compiled and unroots the objects their code refers to by address, so a thunk still held anywhere is left pointing at objects that may now be collected.

Caches filled by __init__ rather than by compiling are left alone: they are rebuilt per session and so never reach an image.

source
Enzyme.Compiler.codeinst — Method
codeinst(mi::MethodInstance, world::UInt) -> Union{Nothing, CodeInstance}

Infer and compile the function mi at world with Julia's native interpreter and JIT, exactly as an ordinary call of it would, and return its CodeInstance. Return nothing when inference fails. The result is the same CodeInstance ordinary callers use, so the function is compiled at most once per process. See invoke_codegen! for why the native interpreter, and not EnzymeInterpreter, infers the functions that are called natively.

source
Enzyme.Compiler.copy_abi_attrs! — Method
copy_abi_attrs!(call::LLVM.CallInst, fn::LLVM.Function)

Copy the zeroext, signext and swiftself parameter attributes of fn to the call site call. restore_lookups replaces the callee with a constant address, after which only the call site's attributes describe the convention the callee expects: the extension of a small integer argument, and the register pgcstack goes in. A target without the swift calling convention takes pgcstack as a plain parameter, which carries no attribute to copy.

source
Enzyme.Compiler.current_pgcstack — Method
current_pgcstack() -> Ptr{Cvoid}

Give the pgcstack of the running task, for an llvmcall that takes it as an argument (see use_gcstack_arg!).

Julia computes current_task() from the pgcstack of the function that this is inlined into (on 1.13 that is the "gcstack" argument of the function). So the function gets no julia.get_pgcstack call from this. An llvmcall of julia.get_pgcstack would not do: Julia inlines it where it is called, which is the problem that use_gcstack_arg! avoids.

source
Enzyme.Compiler.decay_args_readonly — Method
decay_args_readonly(st, fop, args)

Whether the call st merely reads through each of the argument positions in args. That is the case when the position itself is marked readonly / readnone, or when the callee as a whole only reads memory – which is how plain libcalls such as memcmp are annotated. Positions that carry an sret-like marker are never treated as read-only, since the callee writes the object back through them.

source
Enzyme.Compiler.declare_native! — Method
declare_native!(mod, mi, RT, specptr, name, world)

Declare the natively compiled function mi, with return type RT, in mod, with the signature specsig derives. Store the entry point specptr in the enzymejl_needs_restoration attribute. restore_lookups turns that attribute into the address once the calling module is final.

source
Enzyme.Compiler.declare_ntuple_type! — Method
declare_ntuple_type!(mod::LLVM.Module)

Declare julia.enzyme.ntuple_type(T, count), which returns NTuple{count, T}. It stays a bare declaration while Enzyme differentiates, including in the module kept for nested differentiation, so that no optimization drops its constant T argument and abs_ntuple_type can always recover T. define_ntuple_type! gives it a body after differentiation.

source
Enzyme.Compiler.define_ntuple_type! — Method
define_ntuple_type!(mod::LLVM.Module)

Give julia.enzyme.ntuple_type, if mod declares it, an alwaysinline body. For count <= NTUPLE_TYPE_STACK_SIZE it fills a stack array with T and calls jl_apply_tuple_type_v; otherwise it calls jl_apply_tuple_type(jl_svec_fill(count, T)). Both find the interned type several times faster than jl_f_apply_type(NTuple, count, T), which first instantiates the NTuple UnionAll. Call this once differentiation is done, after the module for nested differentiation has been saved.

source
Enzyme.Compiler.dereferenceable_root_ptr — Method
dereferenceable_root_ptr(v) -> Bool

Whether a pointer-sized load from v is safe even where the program would not have loaded from it: v points into an alloca, or at a constant offset into an argument whose dereferenceable attribute covers the load.

source
Enzyme.Compiler.deserves_argbox — Method
deserves_argbox(T) -> Bool

Say if Julia's codegen passes a value of type T boxed, as a tracked jl_value_t*, rather than unboxed on the stack (deserves_argbox in codegen.cpp). Only a concrete immutable type that is a singleton, or that Julia can allocate inline (jl_datatype_isinlinealloc), goes unboxed. A struct with uninitialized fields, or one whose field layout the GC cannot describe inline, is boxed even though jl_type_to_llvm gives it an LLVM struct type.

source
Enzyme.Compiler.ejl_value — Method
ejl_value(key, inserted)::Union{Some{Any}, Nothing}

The Julia value the global ejl_<key> stands for: a well-known Julia global (JuliaGlobalNameMap), one Enzyme knows when it loads (JuliaEnzymeNameMap), or one the compilation inserted (inserted, see insert_julia_value!); nothing if none. A load folded through a binding (Julia 1.10) stands for the binding's value.

source
Enzyme.Compiler.emit_ntuple_type! — Method
emit_ntuple_type!(B, count, T) -> LLVM.Value

Emit the type NTuple{count, T} for a runtime count::Int as a call to julia.enzyme.ntuple_type (see declare_ntuple_type! and define_ntuple_type!).

source
Enzyme.Compiler.emit_unresolved_llvm — Method
emit_unresolved_llvm(job)

GPUCompiler.emit_llvm(job), without resolving the references to Julia values on GPUCompiler 2.x: what it returns for a job compiled on behalf of another, also for a toplevel one, so that nothing is baked into the module before Enzyme is done with it. The relocation records say what each slot holds; link_julia_values! resolves them once the module is linked. GPUCompiler 1.x always resolves; make_slots_symbolic! undoes it.

source
Enzyme.Compiler.extract_roots_from_value! — Method
extract_roots_from_value!(builder, sret, roots)

Store the GC-tracked fields of the value sret into the returnRoots array roots, which must have room for CountTrackedPointers(value_type(sret)).count entries.

This is the split recombine_value! undoes: a caller that needs to hand a value on through the sret/returnRoots convention writes the tracked pointers here and the inline data into the sret buffer separately.

source
Enzyme.Compiler.fixup_1p12_sret! — Method
fixup_1p12_sret!(f::LLVM.Function)

Rewrite the untyped store of a return value that needs both an sret buffer and a returnRoots array into field-wise stores that skip the GC-tracked slots.

Julia 1.12 changed the convention for such a return: the callee writes the tracked pointers only into returnRoots and leaves the corresponding slots of the sret buffer undefined, so a caller has to recombine the two halves (which is what recombine_value! does). Codegen spells the write of the remaining, inline data as a single untyped llvm.memcpy out of an [N x i64] alloca, which also drags the undefined bytes into the tracked slots – and Enzyme would then take those for live jlvalues. Replacing the memcpy with stores of just the untracked fields keeps the tracked slots alone, matching what the convention promises the caller.

The rewrite is keyed on the ABI actually present in the IR rather than on a version bound: it only fires for a memcpy whose destination parameter carries an sret attribute for exactly RT.

source
Enzyme.Compiler.gcstack_arg_index — Method
gcstack_arg_index(fn::LLVM.Function) -> Int

Give the index of the pgcstack parameter of fn, or 0 when fn has none.

Julia's codegen marks that parameter swiftself where the target supports the swift calling convention, and turns the convention off where it does not (jl_codegen_output_t::use_swiftcc, false on RISC-V). Julia 1.13 also gives the parameter a gcstack string attribute, which specsig writes on every version. Hence recognize either mark.

source
Enzyme.Compiler.import_cached_autodiff! — Method
import_cached_autodiff!(mod, ptr, FT)

Make the derivative thunk compiled at ptr callable from mod, and return the function to call in its place.

cached_compilation keeps the bitcode of every thunk it compiles in autodiff_cache, so that a later compilation which sees a call to the thunk's address can use its IR instead of calling an opaque pointer. That blob is a fixed serialization: every module it is linked into gets the very same symbol names for the Julia functions it carries. compile_unhooked links all the modules nested_codegen! emitted into one module, so two of them importing the same blob define those symbols twice and LLVM.link! rejects the second one ("symbol multiply defined", EnzymeAD/Enzyme.jl#2788). The names collide even though nothing else does, because they were minted once, when the thunk was compiled, and are replayed verbatim on every import.

Import the blob once per compilation instead. The first module to need it gets the definitions; every later one gets a declaration of the entry, which the final link binds to that single definition. The entry stays externally visible for as long as those declarations do: compile_unhooked internalizes it again once every module is linked (see internalize_imported_thunks!), so it is inlined and discarded just as a lone import was.

source
Enzyme.Compiler.instantiate_annotation — Method
instantiate_annotation(A, rt, width)

Fill in the free parameters of a (possibly partially applied) activity annotation A with element type rt and batch width width.

The batch annotations take a second parameter carrying the batch width. Applying only A{rt} to them leaves that parameter free, and a subsequent A{rt} binds the element type to it, yielding an annotation whose batch_size is a type rather than the width. Filling both explicitly keeps batch_size(A) == width, which the shadow-return ABI in create_abi_wrapper and enzyme_call asserts.

source
Enzyme.Compiler.internalize_imported_thunks! — Method
internalize_imported_thunks!(mod)

Undo the external linkage import_cached_autodiff! gave the entry of every cached thunk this compilation imported.

The entry is externally visible only so that the modules which did not import the blob can declare it and have the final link bind them to the one definition. Once mod holds every such module, nothing outside it refers to the entry any more, and leaving it visible would keep a copy of the thunk that post_optimize! may neither inline nor drop.

source
Enzyme.Compiler.invoke_codegen! — Function
invoke_codegen!(mode, mod, funcspec, alwaysinline = false)

Make the rule funcspec callable from mod.

Rules with the :inline convention (see call_convention) are emitted into mod by nested_codegen! and marked always-inline. Like all code emitted into mod, they are inferred by EnzymeInterpreter.

Rules with the :call convention are not emitted at all. Julia's native interpreter infers the rule, exactly as an ordinary call of it would, and Julia's JIT compiles the resulting CodeInstance and its callees (see codeinst). declare_native! declares the specialized entry point in mod, and restore_lookups binds the entry's address. Such a rule is compiled once for the whole process, shares its CodeInstances and native code with ordinary callers, and is called directly through its specsig, without boxing or dispatch. Callers build its arguments as for an emitted rule. A rule that Julia compiled without a specialized entry point, or for another MethodInstance, is an error (see native_codeinst).

The two interpreters must not mix. Code inferred by the native interpreter has Julia's semantics: no call in it is marked for a rule, within_autodiff() is false, and no Enzyme intrinsic (such as ignore_derivatives) is left for the Enzyme pipeline to lower. Such code is legal to call, because a rule body is primal code that Enzyme never differentiates. It is not legal to emit into mod, where check_ir, the rule handlers and the Enzyme pipeline expect EnzymeInterpreter output and its enzymejl_* attributes. Hence a natively called rule reaches mod only as a declaration bound to an address: it has no body that LLVM could inline. Conversely, code inferred by EnzymeInterpreter must not be handed to Julia's JIT: within_autodiff() is true there, so a rule body that calls ignore_derivatives emits a call to __enzyme_ignore_derivatives, which only the Enzyme pipeline resolves.

Nested differentiation is the exception to "never differentiates". When a derivative is differentiated again, its rule calls must be differentiated through, so they need a body. materialize_native_invokes! gives the declarations one before the outer differentiation runs.

Every rule is emitted by nested_codegen! where the call ABI does not exist (see native_invoke_available).

source
Enzyme.Compiler.is_mutable_array — Method
is_mutable_array(T::Type)::Bool

Return true if T refers to mutable memory of elements of type eltype(T), as an Array does. The activity of T then comes from eltype(T). A package can add methods for its array or pointer types, for example in a package extension. Enzyme calls this function in the world of the compilation.

source
Enzyme.Compiler.is_readonly — Method
is_readonly(attr::LLVM.Attribute)::Bool

Whether attr on its own establishes that the function or argument position it is attached to is only read from. That is readonly / readnone, and on LLVM 16+ a memory effect whose modref is read-only.

source
Enzyme.Compiler.jit_gcstack_arg — Method
jit_gcstack_arg() -> Bool

Say if Julia's JIT passes pgcstack as an argument to compiled code. The JIT compiles with jl_default_cgparams, so read its gcstack_arg field. The Base.CodegenParams mirror gives the field offset. The value is a process constant, so read it once and cache it.

source
Enzyme.Compiler.jit_uses_swiftcc — Method
jit_uses_swiftcc() -> Bool

Say if Julia's codegen uses the swift calling convention on this target. LLVM does not support the convention on RISC-V, so Julia turns it off there (jl_codegen_output_t::use_swiftcc). With the convention on, and jit_gcstack_arg set, the pgcstack parameter gets the swiftself attribute; without it, pgcstack keeps its place in the signature but carries no attribute and the function keeps the C calling convention.

source
Enzyme.Compiler.julia_value_of_slot — Method
julia_value_of_slot(gv, enzyme_ctx)

The Julia value a load of the global slot gv yields, as the compilation in flight knows it (slot_value: what codegen reported, or a module compiled earlier brought along), or the value of the box GPUCompiler 2.x materialized for it in device code (materialized_box_value); nothing if neither.

This is the preferred source: it is what codegen itself said the slot refers to, and it does not depend on the address of the value having been written into the IR, which Enzyme removes until the module is linked (see make_slots_symbolic!). For a slot with no record the caller falls back to decoding the initializer with slot_initializer_address.

source
Enzyme.Compiler.julia_value_table — Method
julia_value_table(ctx, mod)::JuliaValueTable

The entries of the tables of ctx for what mod refers to by name: its slots that are still symbolic, and the ejl_ globals Enzyme inserted into it.

source
Enzyme.Compiler.legalize_readonly_decay! — Method
legalize_readonly_decay!(st, inst, args)

Rewrite the argument positions args of the read-only call st, all of which are the illegal addrspace(10) -> addrspace(0) cast inst, to use a legally derived pointer instead: decay to addrspace 11 and go through julia.pointer_from_objref, with the object gc-preserved across the call. Julia's late GC lowering understands that form; the raw cast it does not.

source
Enzyme.Compiler.link_julia_values! — Method
link_julia_values!(mod, meta)

Resolve the Julia values the toplevel module mod refers to, as it is linked into what runs: on GPUCompiler 2.x by baking its relocation records (meta.relocations), which also cover what is not a slot of a Julia value, and from the module's table (meta.value_table): its slots (resolve_slots!) and the values the compilation inserted (bake_inserted_values!).

source
Enzyme.Compiler.link_split_existing! — Method
link_split_existing!(mod::LLVM.Module, newmod::LLVM.Module)

Link newmod into mod like LLVM.link!(mod, newmod), but set LLVMInternalLinkage on any function defined in both modules before linking. This allows LLVM's linker to natively internalize and resolve duplicate definitions without string comparisons or linker collisions.

source
Enzyme.Compiler.make_slots_symbolic! — Method
make_slots_symbolic!(mod, ctx)

Drop the address GPUCompiler 1.x writes into each julia.constgv slot of mod whose value is recorded in the table of ctx, leaving the slot a declaration: the shape GPUCompiler 2.x emits for a job compiled on behalf of another, with the value known only by name. Enzyme works on the module through the table of Julia values, and resolve_slots! writes the addresses back in once the module is linked into what runs: late, as GPUCompiler 2.x resolves its relocations.

source
Enzyme.Compiler.mark_load_dereferenceable! — Method
mark_load_dereferenceable!(inst::LLVM.LoadInst, @nospecialize(source_typ), byref)

Julia only marks a loaded pointer to a heap object as dereferenceable when it comes from a mutable struct's field; a field of an immutable struct gets no such metadata. Enzyme's abs_typeof knows the exact Julia type of the loaded value, so if inst loads a tracked pointer to an object of the concrete type source_typ (byref == MUT_REF), mark it dereferenceable_or_null for that type's size, as Julia's codegen does in the mutable case. Whether it is also non-null is left to Julia's own nonnull, which it emits exactly when the field cannot be #undef; with both, LLVM may speculate loads through the pointer, which lets LICM hoist e.g. an array's Memory pointer out of loops it is not guaranteed to be loaded in, where Enzyme otherwise has to cache it per iteration.

source
Enzyme.Compiler.mark_loads_dereferenceable! — Method
mark_loads_dereferenceable!(fn::LLVM.Function, enzyme_ctx)::Bool

Apply mark_load_dereferenceable! to every load of a tracked pointer in fn whose Julia type abs_typeof can determine. Runs as a pass right before the loop passes of the early pipeline, since earlier passes recreate such loads without their metadata.

source
Enzyme.Compiler.mark_symbolic_slot! — Method
mark_symbolic_slot!(gv)

Tell Enzyme what the initializer of the julia.constgv slot gv said before it became a declaration: the slot never changes (constant), and as it only ever holds the address of a Julia value, it is inactive. Activity analysis inferred that from a constant initializer; a declaration gives it nothing to go by, and an active-looking slot would need a shadow global. Type analysis is kept from accumulating information on the slot, which every load of it shares (as for the globals unsafe_to_llvm inserts); a load LLVM folded into the address gave it none either.

source
Enzyme.Compiler.materialize_native_invokes! — Method
materialize_native_invokes!(mode, mod)

Give every natively called function that mod declares a body, so that the differentiation of mod can differentiate through it.

A derivative that is differentiated again reaches the outer primal module either linked in as a deferred job, or embedded through enzyme_call. Either way it calls its rules through the declarations declare_native! made, which restore_lookups leaves symbolic until the module is compiled (see _thunk). Such a declaration is an opaque call to the outer differentiation. This emits the rule with nested_codegen!, as an :inline rule, and defines the declaration as an always-inline wrapper that forwards to it. The wrapper drops the pgcstack parameter when the emitted rule takes none.

source
Enzyme.Compiler.materialized_box_value — Method
materialized_box_value(ctx, gv)::Union{Some{Any}, Nothing}

The value of the box the slot gv points into, if GPUCompiler 2.x materialized one for it: in device code it gives an isbits value a box in the module instead of the address of a host one, {[padding,] header, bytes}, and points the initializer of the slot at its bytes. The type is the header's small type tag, or, for a type that has none, the target of the relocation record of the header in the modules ctx emitted. The value only lives in the module: analysis can read it, but no host address stands for it, so a load of the slot is never folded.

source
Enzyme.Compiler.memtransfer_truetype — Method
memtransfer_truetype(world, ptr, sz, enzyme_ctx)

The enzyme_truetype metadata for a memcpy/memmove/memset of sz bytes at ptr, or nothing if ptr cannot be traced back to a Julia object of known concrete type. The type is read off the Julia layout, so it is not limited to the offsets Enzyme's type analysis keeps.

source
Enzyme.Compiler.module_targets_host — Method
module_targets_host(mod) -> Bool

Say if mod targets the machine this process runs on. Compare the architecture component of the module's target triple with the host triple (Sys.MACHINE) and with Sys.ARCH. The two spellings can differ: Darwin writes arm64 where Sys.ARCH says aarch64.

source
Enzyme.Compiler.native_codeinst — Method
native_codeinst(mod, mi, world) -> Union{Nothing, Tuple{CodeInstance, Ptr{Cvoid}}}

Return the native CodeInstance of the function mi and its specialized entry point when invoke_codegen! binds the function, called from mod, to that code. Return nothing when it emits the function with nested_codegen! instead: where native calls are unavailable (see native_invoke_available), and for the :inline convention.

Throw a CallingConventionMismatchError when Julia compiled the function for a MethodInstance other than mi, or without a specialized entry point (the boxed jl_fptr_args ABI, which Julia picks when every argument is boxed and so is the return). The declaration is derived from mi.specTypes, so the first would bind it to code with another signature. Every rule takes an annotation, an immutable struct passed unboxed, so Julia always gives a rule a specialized entry point. Neither case is known to occur; an error reports the broken assumption instead of hiding it.

source
Enzyme.Compiler.native_invoke_available — Method
native_invoke_available(mod::LLVM.Module) -> Bool

Say if functions called from mod may be bound to natively compiled code (see invoke_codegen!). That binds a process address, so not during precompilation, and only for modules that target the host.

source
Enzyme.Compiler.native_return_type — Method
native_return_type(mod::LLVM.Module, mi::MethodInstance, world::UInt) -> Union{Nothing, Type}

Return the return type of the function mi as its natively compiled code has it, when invoke_codegen! binds the function, called from mod, to that code. Return nothing when it emits the function with nested_codegen! instead. The rule handlers derive the tape type and check the returned derivatives against this type, so it must come from the inference that produced the code they call.

source
Enzyme.Compiler.nullify_rooted_values! — Method
nullify_rooted_values!(builder, sret)

Return the value sret with every GC-tracked field replaced by a null reference.

Used where only the inline data of a value is wanted, and the tracked fields are either held elsewhere (in a returnRoots array, see extract_roots_from_value!) or known not to be needed, so that leaving the original pointers in place would root objects that must not be kept alive.

source
Enzyme.Compiler.plt_gots — Method
plt_gots(mod)::Dict{String, PLTGot}

What every PLT got of mod resolves to, keyed by the name of the got. Rewriting a load of a got reads the library and symbol out of the ijl_load_and_lookup call in its stub, and walking the stub folds that call away; reading them all before the walk makes the order in which check_ir! walks the functions irrelevant (GPUCompiler 1.x happens to list stubs after their users, 2.x before).

source
Enzyme.Compiler.recombine_value! — Method
recombine_value!(builder, sret, roots; must_cache=false)

Rebuild a whole return value from the two halves the sret/returnRoots calling convention splits it into.

A callee returning a type with both GC-tracked and inline fields writes the tracked pointers into the returnRoots array and the remaining data into the sret buffer, leaving the tracked slots of sret undefined. sret here is the value already loaded out of that buffer and roots the pointer to the root array; the tracked fields are loaded from roots and inserted back into their slots, and the completed value is returned. This is the inverse of extract_roots_from_value!; see recombine_value_ptr! for the variant that takes sret as a pointer.

must_cache marks the loads from roots as must-cache, for a caller that needs the recombined value to survive into the reverse pass.

source
Enzyme.Compiler.recombine_value_ptr! — Method
recombine_value_ptr!(builder, jltype, sret, roots; must_cache=false)

Like recombine_value!, but loads the inline half out of the sret buffer rather than taking it as an already-loaded value.

Both sret and roots are pointers; a fresh jltype value is built by loading the untracked fields from sret and the tracked ones from roots.

source
Enzyme.Compiler.record_julia_values! — Method
record_julia_values!(ctx, job, meta)

Make the Julia value each global slot of a freshly emitted module refers to known to ctx.

That is the authoritative answer to "which object is this global?", and unlike an address decoded from an initializer it does not depend on the address having been written into the IR, which no back-end has before Enzyme is done (see emit_unresolved_llvm and make_slots_symbolic!). GPUCompiler 1.x reports the address of the object behind each slot as gv_to_value, keyed by the name of the slot, which is copied into the tables of ctx. GPUCompiler 2.x reports a relocation record per slot, keyed the same way, whose target is the value: ctx keeps the records themselves, and whether the back-end of job lets host addresses into the module (see bakes_julia_values). Either way absint and abs_typeof read the value through slot_value, resolve_slots! writes the address back in from slot_address, and the context keeps the values rooted for the duration of the compilation.

source
Enzyme.Compiler.record_symbolic_slots! — Method
record_symbolic_slots!(mod, relocs, ctx)

Give every slot of mod that is still a declaration, and whose value the table of ctx knows, a relocation record in relocs, unless it has one. The slots of the primal module come with records; those of a module linked in later (nested_codegen!, an imported thunk) do not, and whoever resolves the records must see them too.

source
Enzyme.Compiler.relocate_julia_value_globals! — Method
relocate_julia_value_globals!(mod, relocs, inserted)

Hand the Julia values Enzyme refers to in the device module mod to GPUCompiler 2.x's relocation machinery.

Enzyme refers to a Julia value through a global ejl_<key> whose address is the value (JuliaGlobalNameMap[key], JuliaEnzymeNameMap[key], or inserted[key] for the values the compilation inserted, see unsafe_to_llvm). On the host the JIT resolves the well-known ones and the module gets the addresses of the inserted ones when it is linked (bake_inserted_values!); nothing does in device code. Replace the global with what codegen emits for a Julia value: a load of a word-sized slot, here ejl_slot_<key>, with a relocation record for the value. The job that requested the derivative then lowers it with its own strategy, as it does Julia's own slots: bakes the address in, or leaves it for its loader.

source
Enzyme.Compiler.resolve_slots! — Method
resolve_slots!(mod, table::JuliaValueTable)

Write the address of its value back into every julia.constgv slot of mod that make_slots_symbolic! left a declaration, from the module's table (see julia_value_table). This is the resolver for GPUCompiler 1.x, which resolves nothing itself: it runs where the module is linked into code that runs, so that everything before sees the slots as names.

source
Enzyme.Compiler.restore_native_invokes! — Method
restore_native_invokes!(mod::LLVM.Module)

Bind the declarations of natively called functions in mod to their entry addresses. restore_lookups(mod; native_invokes = false) skips them, so that the module can still be differentiated again while they are symbolic.

source
Enzyme.Compiler.rewrite_abi_converter_calls! — Method
rewrite_abi_converter_calls!(mod::LLVM.Module)

Julia 1.12+ lowers @cfunction to a world-age-guarded dispatch site that calls jl_get_abi_converter to obtain a callable pointer for the target in the current world. That runtime resolver picks a code instance from the native JIT cache, which is the wrong one for GPUCompiler-emitted code: both its owner (ci->owner) and its ABI (compiled with gcstack_arg = true, i.e. pgcstack in the swiftself register) differ from this module (gcstack_arg = false), so the raw specsig pointer handed back is called with a mismatched ABI and reads a garbage GC stack out of the swiftself register (#3284).

Rewrite each such site to act like jl_apply_generic instead: Julia's codegen already emitted an in-module unspecialized dispatcher thunk for every site (stored in the cfuncdata global) that boxes the arguments and dispatches via jl_apply_generic. That thunk was compiled with this module's own ABI and resolves the callee from the correct cache in the current world, so replace every jl_get_abi_converter call with it and let the world-age guard fold away.

source
Enzyme.Compiler.root_load_in — Method
root_load_in(inst, ptr, bb, T_prjlvalue)

Whether inst is a plain (non-volatile, non-atomic) load of a tracked pointer from ptr, placed in the block bb.

source
Enzyme.Compiler.root_path_may_write — Method
root_path_may_write(inst) -> Bool

Whether inst may write memory. Unlike mayWriteToMemory, which only reads the attributes of the call site, a call is also known not to write when its callee is read-only, as for intrinsics such as llvm.smax that LLVM leaves between a phi and the loads it moved into a successor.

source
Enzyme.Compiler.roots_follow — Method
roots_follow(args, ai, removedRoots) -> Bool

Whether the classified argument after args[ai] carries the inline roots of args[ai] and is folded back into it (removedRoots), so that args[ai] is handled through its data pointer rather than loaded whole.

source
Enzyme.Compiler.slot_address — Method
slot_address(ctx, name)::Union{Ptr{Cvoid}, Nothing}

The address resolve_slots! writes into the slot called name, or nothing if the compilation ctx has none: it does not know the slot, or the back-end keeps it symbolic (GPUCompiler 2.x's :patch and :table), in which case a host address must not go into the module. On GPUCompiler 2.x the address is that of the instance GPUCompiler roots, which gives an isbits value a box whose address stays valid.

source
Enzyme.Compiler.slot_object_address — Method
slot_object_address(gv, enzyme_ctx) -> Union{LLVM.Value, Nothing}

The address of the object the julia.constgv slot gv holds, as a constant, or nothing if it is not known. It is only read from at compile time, never written into the IR: what a fold inserts is a named global for the object, which the JIT or GPUCompiler resolves.

source
Enzyme.Compiler.slot_value — Method
slot_value(ctx, name)::Union{Some{Any}, Nothing}

The Julia value the slot called name refers to, as the compilation ctx knows it: from its tables (GPUCompiler 1.x's gv_to_value, or a module compiled earlier, see merge_julia_value_table!), or from the relocation records of the modules it emitted (GPUCompiler 2.x); nothing if neither knows the slot.

source
Enzyme.Compiler.specsig — Method
specsig(mi, RT; gcstack_arg = jit_gcstack_arg()) -> (retty, params, param_attrs)

Derive the signature Julia's codegen (get_specsig_function) gives the specialized entry of the function mi with return type RT. The parameter order is [sret][return roots][pgcstack] args....

The return:

  • A return that deserves an sret becomes a leading pointer parameter with the sret attribute. On 1.12+ the buffer holds the layout with the tracked pointers stripped, and the tracked pointers go through the return-roots parameter.
  • A small Union return gets its byte buffer as the leading pointer parameter, and the function returns {boxed-or-null, selector}.
  • A boxed return is a tracked pointer.
  • Every other return keeps the value's LLVM type.

The arguments:

  • Ghost and Type{T} arguments are omitted.
  • Boxed arguments are tracked pointers.
  • Aggregates go by reference in the derived address space. On 1.12+ a pointer to their inline roots follows.
  • Scalars go by value.

With gcstack_arg set, pgcstack follows the return parameters, with the swiftself attribute where jit_uses_swiftcc holds. It always carries the gcstack string attribute, as Julia 1.13 writes, so that gcstack_arg_index finds it on a target without the convention too. Pass false to model a function compiled without pgcstack, as nested_codegen! emits.

classify_arguments, get_return_info and enzyme_custom_setup_args follow the same rules, so the derived signature matches the arguments the rule handlers in customrules.jl pass.

source
Enzyme.Compiler.split_value_into! — Method
split_value_into!(builder, val, sret, roots)

Store the value val through the sret/returnRoots convention: its GC-tracked fields into the roots array and every other field into the sret buffer, leaving the tracked slots of the buffer alone, as a caller reads them from roots only. The inverse of recombine_value_ptr!.

source
Enzyme.Compiler.split_value_size — Method
split_value_size(dl, T) -> Int

Size in bytes of the data half of a value of LLVM type T when Julia keeps it on the stack split into inline data and GC roots (split_value_size in cgutils.cpp).

Up to Julia 1.13.0 the data half kept the full layout, with the tracked slots left undefined. Since 1.13.1 (JuliaLang/julia#60388) codegen shrink-wraps it: pointer words at the very end of the layout are dropped, while interior pointer slots stay as padding so the offsets of the remaining fields are unchanged. The buffer behind the by-reference pointer of an argument with inline roots, and the stack copy a new makes of such a value, are only this many bytes; a callee may not describe or touch more. An sret buffer is not affected: its parameter is still declared dereferenceable for the full layout and callers allocate it whole, only the memcpy that fills it may stop short (see fixup_1p12_sret!).

source
Enzyme.Compiler.unfold_root_phi_loads! — Method
unfold_root_phi_loads!(f::LLVM.Function) -> Bool

Turn a load through a phi of root-array pointers back into a phi of loads.

Julia 1.13.1+ keeps the inline GC roots of a split value lazily, as a pointer into whatever roots array already holds them (jl_gc_roots_t), so the roots a phi merges may be loaded right in its predecessors: from the roots argument of the function on one edge, from a local roots alloca on another. LLVM's instcombine then folds that phi of loads into one load of a phi ptr of the arrays. Enzyme promotes every alloca of the augmented primal to a heap allocation and cannot push that address-space change through a phi whose other operands are not allocas ("Illegal address space propagation"). Sinking the load back into the predecessors gives the alloca only plain loads again.

Only rewrite what is certainly equivalent: a phi ptr in address space 0 with an alloca among its incoming values, whose users are all non-atomic loads of tracked pointers, in its own block or in a successor whose only predecessor it is, that no memory write precedes. Each load is re-created before the terminator of every predecessor. That load is speculative on an edge whose predecessor has other successors, and on every edge when the original load is in the successor, so it is then only done for a pointer that is always dereferenceable: an alloca, or an argument whose dereferenceable bytes cover it.

source
Enzyme.Compiler.use_gcstack_arg! — Method
use_gcstack_arg!(f::LLVM.Function, arg::LLVM.Argument)

Make the llvmcall f use its argument arg as the pgcstack of the code it inlines. The caller gives this argument with current_pgcstack, as a Ptr{Cvoid}.

Julia inlines an llvmcall into its caller, together with the alwaysinline functions that the llvmcall calls. A julia.get_pgcstack call in that code thus goes into the middle of the caller. On 1.13 the caller takes its pgcstack as its own "gcstack" argument. But the GC lowering prefers a julia.get_pgcstack call in the entry block to that argument, and pushes the GC frame of the caller only after the call. Then no safepoint before the call has the roots of the caller.

Thus inline the alwaysinline functions into f here, and replace every julia.get_pgcstack call in f with arg, as LowerPTLS does for a function that takes its pgcstack as an argument.

source
Enzyme.Compiler.Interpreter.codeinst_entry — Method
codeinst_entry(ci::CodeInstance) -> (specptr, invoke)

Compile ci if needed and return its two entry points. specptr is the specialized-signature entry, or C_NULL when ci only got a boxed jl_fptr_args entry. invoke is the boxed invoke(F, args, nargs, ci) entry, or C_NULL when ci could not be compiled.

source
Enzyme.Compiler.Interpreter.enzyme_call_kind — Method
enzyme_call_kind(interp, specTypes) -> Union{Nothing, Symbol}

Say how Enzyme handles a call of the signature specTypes, or nothing when it handles it like any other call:

  • :primitive: Enzyme differentiates the call itself (is_primitive_func).
  • :alwaysinline: the inliner must always inline it (is_alwaysinline_func).
  • :inactive, :frule, :rrule: the signature has a rule of that kind, and interp has that kind of rule enabled.

Every place that decides whether a call needs Enzyme's handling asks this, so that a call reaches the pipeline the same way whichever path inference took to it: FutureCallinfoByType for an ordinary call, and invoked_ci_needs_rule for invoke(f, ci, args...).

source