Skip to content

Multimode batched evaluation of factorized CC (cost model + placement + dry-run) - #583

Open
evaleev wants to merge 248 commits into
masterfrom
evaleev/feature/multimode-batched-eval
Open

Multimode batched evaluation of factorized CC (cost model + placement + dry-run)#583
evaleev wants to merge 248 commits into
masterfrom
evaleev/feature/multimode-batched-eval

Conversation

@evaleev

@evaleev evaleev commented Jul 29, 2026

Copy link
Copy Markdown
Member

Multimode batched evaluation of factorized coupled-cluster equations

Adds cost-model-driven multimode batching to the SeQuant evaluator so
large-system CSV/PNO-CC residuals can be evaluated without forming their
largest transients whole. Six squashed commits (dry-run backend, optimizer,
evaluator, supporting core, tests, docs).

What it does

  • Batch modes with two kinds: external (occ / PNO pair, free on the result
    -> scattered into disjoint slices) and contracted (DF aux, summed ->
    accumulated). Loops nest external-outside-contracted.
  • Perf-first cost model (DenseTimeSpace): minimizes flops with
    peak_threshold as a ceiling; role-split (contracted/external) batchability;
    order-aware placement over the combined nest.
  • Batched evaluator: cache scope chain with fall-through; slice-on-use
    (a cached intermediate fetched from an outer scope is sliced to the current
    block) decouples correctness from placement; per-level placement driven by
    a per-canonical lifetime mask (cross-occurrence proto-aware meet) unioned
    with contracted residency; iterative (stack-safe) tree traversal.
  • Dry-run cost-profile backend predicts peak/flops/exec over the same IR.

C60 PNO-CCSD dry run (55-term residual, aux K@256, occ@8, 100 GB budget)

batching dry-run peak schedule flops (DP cost) modeled roofline time
none 38 897 GB baseline 1.06e17
contracted-aux only 6 047 GB (6.4x) identical schedule 2.94e16
ext-occ + contracted-aux 443.6 GB (87.7x) identical schedule 1.13e17

The DP selects the same factorization regardless of what is batchable
(flops are unchanged); batching only slices modes to lower the peak. The
roofline-time column moves because a giant intermediate executed whole is
memory-bound (machine_balance x traffic) but compute-bound when sliced -- the
cache-blocking win of the same schedule, not a cheaper one. Recompute overhead
(avoidable_time) is 1.8% -> 6.5% -> 39.8% as slicing gets more aggressive.

Validation

  • Units: [eval] 449, [lifetime_mask] 76, [optimize] 628 assertions green;
    OFF (order-blind) path byte-identical.
  • MPQC he10 CSV-CCk on this stack: batched (371 external scatter + 962 aux group
    events) matches unbatched to < 1e-9, within the 1e-7 precision, no aborts.

Follow-ups (non-blocking, from the final review): dedup the proto-expansion
helper; add a real-forest hidden-tag hash-regression test; revisit the
stamp_lifetime_masks const_cast.

evaleev added 6 commits July 29, 2026 14:39
CostProfile (peak/flops/exec) over the factorized IR via a zero-data dry-run
evaluation, driving the batched cost model's predictions.
Perf-first (DenseTimeSpace) objective with peak_threshold as a ceiling;
role-split (contracted/external) batchability; order-aware placement over the
combined nest; per-node batch annotations consumed by the evaluator.
External-mode scatter + contracted accumulate; cache scope chain with
fall-through; slice-on-use; per-level placement driven by a per-canonical
lifetime mask (cross-occurrence meet) unioned with contracted residency;
iterative (stack-safe) tree traversal.
Index-space occupancy predicates robust to non-physical spaces; logger;
is_valid accepts Power; convention.
evaleev added 23 commits July 29, 2026 23:06
The multimode-batched-eval tests had only run under Release/IGNORE, so several
Debug-only asserts (and one ASan bug) were masked. Fix them so the suite is
green under Debug (SEQUANT_ASSERT_BEHAVIOR=ABORT + AddressSanitizer):

- Covariant tensor forms in the synthetic batched-eval tests: contracted
  indices had been placed in the same bra/ket slot, tripping create_graph's
  strict-braket invariant. Reorient contractions Einstein-properly and move
  Hadamard/external indices to the aux slot (test_lifetime_mask, test_eval_ta,
  test_eval_dryrun). Physical-tensor tests were unaffected (Symm braket).
- test_eval_dryrun: a rank-4 CSV composite used a duplicate proto index; use a
  distinct fourth occ index.
- test_eval_ta: rand_tensor_yield now sizes the m (mu~) space; and copy the
  compared tiles by value in shape_spike_ToT_inner_contraction_to_flat_T (it
  bound a reference into a temporary Future -> ASan stack-use-after-scope).
- test_cache_manager: the batch-axis veto is phase-2 -- a node carrying a batch
  mode free on its own result is batch-variant (External or Contracted) and
  correctly refused run-scope caching; update the stale case-3 expectation.
- cost_model: seeded_root_peak_batched must admit the seed via the external-role
  predicate too. The seed is an external mode, so build_context's role filter
  gates it through is_batchable_external_index; overriding only the
  contracted-role predicate dropped the seed and tripped the k_seed assert.
- Hide ([.]) the all-C60-terms perf-first cost diagnostic: it runs optimize()
  on every summand (tens of minutes in Debug) and has no correctness checks.
- Add a hidden ([.]) water-20 occ-batching overcompute dry-run diagnostic.
order_aware_recompute=false selects the set-keyed DP that ignores the
per-batch-block replay recompute, which under-costs every batched schedule and
is never the more realistic default. Default it to true in BatchPolicy and
CostParams (MPQC and other callers inherit it via optimize()).

Pin it false in the two cases that specifically characterize the legacy
set-keyed behavior: reconstruct_batched_modes_emits_external_per_node (its own
comment documents the order-aware-off emit_external regime) and the C60
objective-determines-factorization case (peak-first forms the fully-sliceable
4-PAO only under the set-keyed peak model; under the realistic resident-scan
model the contrast collapses -- peak-first also avoids it and perf-first flops
then exceed peak-first because the recompute is charged).
OptimizeOptions::inner_pow and PeakBatchedModel::inner_pow already have no
default (empty + composite indices -> inner_aware_volume throws, so the old
silent mis-sizing fallback cannot recur). But 9 OptimizeOptions{...} designated
initializers omitted inner_pow, which g++ -Wextra -Werror flags as
missing-field-initializers (clang does not, so macOS CI and local clang builds
missed it) -- breaking every Linux Debug job.

Add an explicit .inner_pow = {} (composite-free no-op) at all 9 sites
(optimize.cpp compatibility_opts + 8 in test_optimize). Also finish the
removal that had missed CostParams::inner_pow: drop its stale = {} default and
its stale "sized by idxsz (k=1)" fallback comment so all three inner_pow fields
are uniformly no-default (all 21 CostParams{...} sites already set it).
The static per-node walk in cost_profile() prices each node once, so its
flops/exec are order- and batching-blind and never reflect the per-occ-block
REPLAY recompute the batched evaluator does at runtime -- the reason a dry-run
could not predict occ-batching being slower than aux-only.

Split the reported cost:
- Rename CostProfile::{flops,exec_cost,n_ops} -> model_{flops,exec,n_ops} (the
  static DP-model quantities, unchanged).
- Add dryrun_{flops,exec,n_ops}, tallied from the existing Trace::On replay: an
  optional CostSink is attached to the shared dry-run CostModel, and every
  actual product-op execution (DryRunOps::prod) folds its own SLICED-extent
  flops/exec (the same numbers already computed for the per-op OpCost log) into
  it. Because a sliced occ-dependent op run N times does ~1/N work per pass, its
  sliced-cost sum is work-neutral; only occ-INDEPENDENT work re-executed at full
  size once per block inflates -- so dryrun_* isolates exactly the recompute.

The sink is opt-in (nullptr default) and lives on the dry-run CostModel, which
is constructed only inside cost_profile(); eval.hpp / make_evaluator are
untouched, so mpqc's production evaluator path is byte-identical.

water-20 [.][dryrun-water20-overcompute], heavy occ batching: model_flops
ratio 1.0 (flat), dryrun_flops/exec ratio ~1.98, dryrun_n_ops ~55x -- the
recompute is now visible where the model walk saw nothing.
Extend [.][dryrun-water20-overcompute] to three configs -- aux-only, occ+aux
order_aware=false (MPQC production / root-level forest seed), occ+aux
order_aware=true (node-level placement) -- and report the recompute-aware
dryrun_{exec,n_ops} per config plus the OA-true-vs-false ratio.

Diagnoses the water-20 order_aware=true slowdown: under occ batching,
order_aware=true does ~2x the dryrun_exec and ~12x the op-executions of
order_aware=false, and (via its resident-scan peak model reporting higher
peaks) also tips more terms over peak_threshold so occ batching engages where
order_aware=false leaves them un-batched -- a double hit that matches the
observed runtime regression.
Flipping the default to true (previous commit on this branch) turned on the
order-aware cost model AND, together with batch_spectator_indices, node-level
external placement. On water-20 that path is a runtime regression -- the
water-20 dryrun diagnostic measures ~2x the traffic and ~12x the op-executions
of the root-level forest seed for the same term -- and, worse, was reported on
Owl (job 649250) to produce a WRONG schedule: incorrect PNO-CCSD iteration
energies and malformed eval ops. Restore the known-good default (false =
legacy set-keyed DP + root-level forest seed) until node-level placement is
root-caused and fixed. The two tests that pin order_aware off explicitly keep
passing (the pin now merely matches the default).
order_aware_recompute conflated two orthogonal concerns: the order-aware
recompute COST MODEL (which factorization the DP selects) and the node-level
external-mode EMISSION placement (per-node External stamps vs the root-level
forest seed). node_level_placement was defined as order_aware_recompute &&
batch_spectator_indices, so enabling the more-realistic cost model forced the
node-level emission along with it.

A water-8 A/B (holding the cost model fixed, toggling only emission) shows the
regression is the EMISSION, not the cost model: node-level placement runs ~6x
slower (~124 vs ~20 s/iter) and emits ~8x more batch scopes than the root-level
seed, because it nests a batch scope at every carrying node and the batched
evaluator replays each. The order-aware cost model with root-seed emission is
correct and cheap. Node-level placement also produces a wrong residual on
water-20 (size-dependent; not reproduced at water-8/he10).

Split node_level_placement into its own BatchPolicy/CostParams/CostModel flag
(default false), threaded alongside order_aware_recompute. order_aware_recompute
now drives only selection; node_level_placement drives only emission. Both
default false, so this is behavior-neutral. A gated SEQUANT_NODE_LEVEL_PLACEMENT
env knob forces the placement for A/B diagnostics without recompiling a flag
through the caller.

Tests that engaged node-level placement via order_aware_recompute=true now set
node_level_placement=true explicitly; the node-level correctness sweep now
drives the emission via its sweep variable.
Now that node-level emission is separately gated by node_level_placement
(default off), the order-aware recompute cost model is selection-only and safe
to default on: it charges recompute realistically and picks better-batching
factorizations, while emission stays the correct, cheap root-level forest seed.
Verified no churn across the [optimize] suite (628 assertions) and the batched
emission tests. The internal CostModel/oracle-helper member defaults stay false
(documented seeded-probe reason); the public BatchPolicy/CostParams path
threads the true default through.
…le recompute

Add a batched-evaluation schedule-visualizer pipeline and make avoidable
recompute a first-class per-node output of cost_profile().

- schedule_dump.hpp (new): per-term IR schedule-record emitter
  (schedule_ir_json) and a shared cost_op_signature() join key (result index
  labels + the sorted operand pair). One definition, used by three producers
  so a DAG node, its runtime Build event, and cost_profile's per-node number
  all carry the identical key -- no signature is reconstructed downstream.
- eval.hpp: runtime schedule-dump hooks (SCHEDULE_RUN_EVENT / RUN_GROUP),
  each internal Build event stamped with the same signature and per-loop
  dependent-mode flags, all gated by SEQUANT_SCHED_DUMP (production path
  byte-identical when unset).
- CostProfile gains per-node avoidable_nodes {label,count,exec,flops} plus
  avoidable_exec / avoidable_ops and avoidable_time(). DryRunOps::prod tallies
  each build's necessary = product of block counts of its touched (dependent)
  modes, so builds - necessary is the avoidable recompute (a node rebuilt once
  per block of a mode it does not touch). The cost model is dense, so necessary
  is exact and no empirical correction is needed. avoidable_nodes_from_sink()
  is the shared rollup, reused by the schedule-dump test to emit these numbers.
- Consolidate the two dryrun avoidable witnesses (occ-veto, extmode) onto
  cost_profile's structural per-node rollup, replacing the BatchGroup/BatchIter
  trace-parse string-match reconstruction with a read of cp.avoidable_* (the
  structural signature is relabeling-proof; the trace is still parsed only for
  the scatter/group Begin markers cost_profile does not expose).
…uckets

The per-node avoidable-recompute rollup keyed only by the label signature
(result+operand indices), so a node built at many slice sizes landed in one
bucket. Because the exec model is roofline-like (nonlinear in slice size),
those builds span orders of magnitude (a 257x spread on the C60 external-occ
arm); pricing avoidable as count * a single per-build exec then exceeded the
whole replay's exec -- the impossible avoidable_time > 100%.

Fix: key the sink by the EXTENDED signature (label + the touched modes'
realized extents), so every build in a bucket ran at the same slice-context
and hence the same roofline cost. Within a cost-homogeneous bucket
avoidable_exec = total_exec * (builds - necessary)/builds is exact; the
buckets of one DAG node aggregate back to a per-label AvoidableNode (what the
visualizer joins on by hash->sig).

Verified: buckets are homogeneous (count*last == total*frac per bucket), all
arms bounded <= 100% (C60 external-occ 390% -> 0.005%), and the giant term
reads 77%, matching an independent dependent-mode analysis (78%). NodeCost now
accumulates total_exec/total_flops; min/max/last are kept only for the
SEQUANT_AVOIDABLE_DEBUG homogeneity dump.
Redefine the per-value avoidable-recompute metric: it is now measured in FLOPs
against the batching-free (unlimited-memory) ideal -- the arithmetic the batched
replay repeats beyond building each value once at full extent -- rather than in
roofline exec against a within-scheme "necessary" reference.

Two problems with the exec-weighted metric drove this:
 - roofline exec is nonlinear in slice size, so one value's differently-sized
   builds spanned ~257x; pricing avoidable as count x a single per-build exec
   exceeded the whole replay's exec (avoidable_time > 100%). The prior
   cost-homogeneous slice-context bucketing removed the >100% but stayed
   exec-weighted;
 - referencing "necessary = distinct slices the scheme produces" is circular --
   it scores an un-hoisted value's per-block rebuilds as necessary, so it cannot
   see the recompute hoisting exists to avoid.

FLOPs is linear in extents, hence additive across slices: disjoint per-block
slices that tile a value sum to exactly full_flops (0 avoidable), while a value
rebuilt full per block sums to N*full ((N-1)*full avoidable). So avoidable =
max(0, total_flops - full_flops) per value, bounded in [0, dryrun_flops] by
construction, needs no slice-context bucketing, and answers "what does batching
cost vs. infinite memory". NodeCost drops to {builds, total_flops, full_flops};
CostProfile.avoidable_exec -> avoidable_flops, avoidable_time() = avoidable_flops
/ dryrun_flops.

Witnesses re-baselined (nterms=55, FLOPs): occ-veto 1.8/6.5/15.4%; the extmode
witness shows external-occ (~1.95%) and contracted-occ (~1.97%) essentially
equal -- external-mode batching is NOT a recompute fix on the C60 forest (its
original conclusion, now with honest magnitudes). The [cost_profile] giant reads
77.8%, matching an independent dependent-mode analysis (78%).
The batch-variant caching veto in cache_manager had two disjuncts: (a) a node
whose own batched_here() carries a Contracted, batchable mode FREE in its own
result, and (b) a non-empty cross-occurrence lifetime mask. Disjunct (a) is
structurally dead: a Contracted mode is summed AT the node, so it can never be
free in that node's result, and post-role-split a free index is stamped
External, never Contracted -- so the condition never holds (the occ-veto test's
[veto-reach] probe read 0, structurally, not by forest accident). It only ever
guarded a malformed emission.

Remove disjunct (a) and the `is_batchable_contracted_index` parameter from the
cache_manager factory (the only functional caller, build_dryrun_cache, drops the
arg; no production caller passed it), the `CacheConfig::is_batchable_index`
field and the cost_profile() overwrite that fed it, and the now-obsolete
[veto-reach]/[veto-hazard] probes plus the disjunct-(a) sub-test. Disjunct (b)
(the cross-occurrence lifetime-mask veto -- the load-bearing F1 correctness
guard) is unchanged.

Behavior-preserving: [cache_manager] (200), [lifetime_mask] (76), [dryrun], and
[eval] (TA production path, 457) all green. Does NOT touch
BatchPolicy::is_batchable_contracted_index (the batching decision) or
BatchPolicy::is_batchable_index() (the eval accept union) -- both stay.
Design note reframing batched-eval cache placement as register allocation.
Three identities -- value (hash), instance (use-site), cell (a materialized copy
serving a subset of instances). A value's instances partition into cells; perfect
CSE = one cell/value, no CSE = one cell/instance, and the peak budget chooses the
granularity in between (a partial un-CSE / materialization DAG). Cell identity =
(value, home-scope, split-index): home-scope is the loop level (the axis batching
adds), split-index names a same-scope peak split (the RA live-range-split rename).

Placement is register allocation + loop-invariant code motion + rematerialization:
hoisting a shared value lengthens its live range (peak) to save recompute; slicing
adds partial-hoist granularity. Objective: minimize recompute (the rational,
batching-aware reuse count W-1 times build cost) subject to the whole-forest peak
profile <= peak_threshold. Peak is a placement (post-CSE, whole-forest) constraint,
not a factorizer one; cost_profile()'s replay peak is the detection safety net, and
a peak that survives full splitting is factorization-inherent.

Includes a prior-art section (rematerialization/checkpointing -- Checkmate;
electronic-structure space-time tradeoff -- Cociorva/Sadayappan PLDI 2002; register
allocation; pebble games), four worked cases, and open items (group-scoped cache
keying, the greedy split move, per-placement footprint, W's fixed point).
O1 (cell keying) resolved as a router + dumb stores, not a wider cache key. Add
§7a "Runtime realization": one value-keyed store per (home-scope, split-index)
-- the cache stays TreeNode-keyed unchanged -- plus an explicit router
{value, use-site} -> (home-scope, split-index) that is the placement pass's
output and replaces the implicit parent_ fall-through search. Reads route via
the map then reuse the EXISTING Enter-stage slicer, (use-scope - home-scope)
INTERSECT carried(N), fed the home scope directly instead of via hops; default
{value} -> (home, 0) is byte-identical. Standardize terminology on "home scope"
(= the code's "lifetime scope" = store scope; consumer's is "use scope"). Update
§4, §9, and O1 accordingly; residual O1 sub-items are the use-site/occurrence id,
the parent_/hops audit, and the naming standardization.
Add §7b: the placement pass as a register-allocation spill loop. Seed = perfect
CSE (recompute-minimal, peak-maximal); walk up the recompute axis to walk down
peak until peak <= threshold. Objective and constraint are exactly
cost_profile()'s avoidable_flops and peak_bytes -- no new measurement. Moves:
SHRINK (slice a carried mode a cell holds full -- the existing external-slice /
node_level_placement, now driven off the true whole-forest peak) and EVICT
(delay/un-hoist an invariant cell held idle, or split a long-lived cell's
instances into short-lived groups -- the new CSE-aware move the per-term DP
cannot see). Greedy: candidates = cells alive at the binding peak point, prefer
free shrinks then max ΔPeak/ΔRecompute (the spill metric), apply, incrementally
re-cost, repeat; terminate on fit or on a factorization-inherent peak. Residual
sub-items O2a (incremental profile update), O2b (per-move estimator/lookahead),
O2c (subsume vs run-after the DP external-slice pass).
Clarify §7b: O2 runs after the per-term min-time factorizer and takes the
factorization AND batch-loop assignments (batched_here) as FIXED, deciding only
the whole-forest eval/placement strategy (home-scope + router); it never adds,
removes, or re-assigns a batch loop. Reframe "shrink" from "slice a carried
mode" to "re-home a cell into an EXISTING carried loop" -- a placement choice on
the fixed nest, not a batching change; deciding to batch an un-batched mode
(adding a loop) is the factorizer's lever. Split the termination boundary into
two non-O2 failure modes: factorization-inherent (a single intermediate > budget)
vs. re-batch-needed (fixed batching left placement too little room, e.g. a shared
cell needing slicing on a mode no single term batched) -- both detected via
peak_bytes and fed back, giving the structure factorize+batch -> O2 place -> if
infeasible re-batch.
Add §7c. Cell footprint is home-relative: a carried mode is sliced (block extent)
iff its fixed batch loop encloses the cell's home, else held full -- the existing
moment-aware memsize with home-relative extent overrides, so O2's shrink ΔPeak is
just the footprint delta. The peak profile is max weighted-interval overlap: each
cell is a [first-use, last-use] interval (from the router's use-sites + the static
schedule order) weighted by footprint; peak = max over static points of the sum of
live cells' footprints (a sweep line), and the argmax is O2's binding peak point.
Because it SUMS co-resident live cells it corrects today's peak_bytes =
max(scratch, cache) under-count (a lower bound per §1); the replay stays the oracle
(must sum, not max, co-residency). The weighted-interval form updates incrementally
under an O2 move (feeds O2a). Residual O3a-c: the sweep structure, the summed-
co-residency replay oracle, composite/proto sizing.
Add §7d. Define home_scope(value) = deepest scope enclosing the loops of
(sliced_modes ∪ demoted_external_modes). sliced_modes is the cross-occurrence
meet (max-reuse upper bound); the demotion fold adds the External batched_here
stamps the meet demoted (has_demoted_external) -- occurrences bind them to
incompatible blocks, so the value can't be a single full value above those loops
and its home must be inside them. The fold is exactly what unifies the current
cache-veto-vs-has_demoted_external disagreement into one authority both the cache
and the runtime read. Per-block is temporal (one external-loop-homed cell
re-instantiated per iteration), so no split-index -- that stays reserved for O2's
peak-driven same-scope splits. Structural and computed from the meet before O2,
which only lowers homes further for peak; consistent with W (the demoted mode is
free tiling). Residual O5a-b: confirm the exact signal / edge cases and the seed
router construction. Also tie O6 to §7b's two failure modes.
O4 (W's computation order) is not a fixed point: W is a function of the current
placement, well-defined at the home_scope seed and re-costed incrementally per O2
move -- seed-then-refine, subsumed by §7b/§7d. O6 (feedback) scoped to a minimal
detect-and-report step (surface the binding cell + failure mode so a schedule
fails loudly, not silent OOM), with the re-batch/re-factorize hint as a follow-on
that the detect step precedes. All major open items (O1-O6) now designed or
resolved; the spec is design-complete.
Phased plan for the placement-as-register-allocation design. Phase 1 (detailed,
bite-sized TDD) corrects cost_profile()'s peak_bytes from max(scratch, cache)
hwmarks to the instant-resolved co-resident SUM across the scope chain (spec
7c/O3b) -- adds CacheManager current_residency()/chain_residency(), threads the
chain sum into note_working_set, simplifies the fold, and re-baselines the
documented-RED peak figures from measurement. Phases 2-5 (router+home_scope seed,
static peak sweep, the O2 greedy, feedback) are a roadmap, each a future plan.
Global constraints: no en-dashes, clang-format, byte-identical perfect-CSE
default, replay stays the peak oracle.
evaleev added 30 commits August 22, 2026 20:34
External-axis batching sized its scatter destination and enumerated its
batch chunks by borrowing an axis's tiling from whatever array in the DAG
happened to carry it (the "carrier"): Result::pre_sized_zeros_over_mode and
mode_batches, plus outer_mode_tiling to read a carrier's outer tiling
type-agnostically. That leaked a backend artifact (tiling) into the neutral
eval layer and forced the executor to find a carrier, work out which of its
modes carried the axis, and reconcile the carrier's Result type against the
destination's -- accidental complexity, and the source of a scatter crash when
a flat destination borrowed from a nested-tensor carrier.

Replace it with two backend-provided closures on a BackendArrayOps, wired onto
the CacheManager (parent-walking, like placement_router) so the forest,
whole-scope, and ordered (DAG) executors all build identical arrays from one
source:
  - make_zeros(index list) -> a zero destination shaped by the descriptor
  - axis_batches(axis, target) -> the axis's batch chunks
Both are realized backend-side -- make_ta_array_ops (TiledArray) and
make_dryrun_array_ops (dry-run cost backend, wired into meter()) -- from the
spaces alone; nothing is read out of the DAG. Tiling never appears in the
neutral layer.

Remove Result::mode_batches / pre_sized_zeros_over_mode / outer_mode_tiling
(base + TA + dry-run) and the carrier fallbacks in all three executors, which
now assert that a backend is wired. Forest scratch caches inherit the array
ops explicitly, since make_batched_scratch takes the real cache by const-ref
and so cannot set a parent for the array-ops walk.

Tests supply their own ops the same way: rand_tensor_yield::array_ops() from
its own extents, make_dryrun_array_ops from the CostModel. Unit tests of the
removed primitives are redirected to the survivors (mode_batches_of_trange1,
axis_batches/make_zeros, direct zero-ToT construction).
ordered_axis_leaf (ordered_executor) and the static
DryRunOps::pre_sized_zeros_over_mode (dry-run backend) were only ever called
from the carrier fallbacks and the removed pre_sized virtual, both gone now.
Also refresh ordered_n_blocks' doc to reflect axis_batches. No behavior change.
Add populate_cell_mode_to_level(ordered, rich), which walks an
OrderedSchedule's realized ScopeBlock tree and, for every BuildStep it
reaches, derives the value's ModeToLevel from the enclosing non-root
blocks on its root-to-block path (their axis/level give the ordered
loop axes and matching DagScopeLevels) against the cell's own carried
(canon_indices), via mode_to_level_from_signature -- the same
position-in-result substrate slicing_signature/index_position use.
Matching is by exact Index identity, not axis TYPE alone: a ScopeBlock
realizes one axis TYPE via ONE representative Index, so a cell
enclosed by that loop but carrying a DIFFERENT physical index of the
same type correctly gets no level for it (the water-8 DF-leaf
over-slice scenario the design targets).

Escaped (ScopeBlock::outputs) values and forest leaves are never
BuildSteps and are left at mode_to_level's default-empty value, which
is correct: neither is ever read as a home-resident operand under
runtime slicing. A value_id reached via more than one BuildStep site
must agree on the same result (SEQUANT_ASSERT) -- divergence there
would mean the scheduler failed to split a physically-divergent value.

Wire it into scope_executor.hpp's ordered-scheduler path: rich is now
built non-const so populate_cell_mode_to_level can write back into it
right after build_ordered_schedule returns; build_ordered_schedule
itself only reads rich, so the reordering is behavior-preserving.

Test: extends the 2-axis occ-outer/aux-inner [sp2-noninner] fixture
with two more roots, each wrapping a local intermediate (M3{;i_3},
Y{;i_5}) that carries exactly one occ mode and no aux dependence --
both land as plain BuildSteps directly inside the occ ScopeBlock. M3
uses the same physical index the block's canonical chain picked as its
axis and gets a level; Y uses a different physical occ index and,
despite being enclosed by the exact same block, gets none.
…to_level cache seam

CacheManager::BatchContext's entry becomes a named struct (BatchContextEntry:
axis, level, range, exact_axis) instead of a bare pair, additively carrying
each realized batch loop's DagScopeLevel alongside the existing exact-axis
element range. Every push site (ordered executor, whole-scope executor,
forest evaluator) now fills level (block.level for the ordered path; a
synthesized depth/space/0 for forest/whole-scope pushes, which still resolve
by exact_axis) and every reader is updated to name .axis/.range instead of
.first/.second. Resolution itself is unchanged (still index_position(nd,
axis) plus the space-mapped fallback), so this is pure plumbing.

Also adds a non-owning node->ModeToLevel cache seam (cell_mode_to_level_,
mirroring array_ops_/placement_router_), populated by the ordered-executor
dispatch entry point from RichSchedule::cells and wired onto the cache for
the call's duration via an RAII guard. Nothing reads the seam yet; a later
task's slice-on-use resolution will.
…), not just BuildSteps

populate_cell_mode_to_level previously walked only BuildStep home paths and
derived each value's mode_to_level from the blocks enclosing where it is
BUILT. That missed every value that is SLICED but not built via a BuildStep:
forest LEAVES (an input DF tensor fetched under -- and sliced by -- the aux
loop it carries) and ESCAPED values (ScopeBlock::outputs, evaluated inline).
Their maps stayed default-empty, so the equivalence gate in slice_to_use
fired on an exact index_position match with no map entry (e.g. the water-20
aux DF leaf carrying [mu-tilde, i, Kappa], sliced at position 2 on Kappa).

A value is sliced at runtime by an enclosing loop L exactly when it is
FETCHED as an operand inside L's block and its canonical result carries L's
block axis. Rebuild the map from that fact:

  1. populate_build_scope_walk records, for every produced or escaped value,
     the enclosing loop blocks (axis + level) it reads its operands inside,
     reading each level straight off the block tree -- so a forced-split
     sibling's ordinal is the REAL one the runtime pushes, never guessed.
  2. ordered_schedule_dep_graph gives every operand edge (leaves included).
  3. For each consumer P and operand W, stamp P's enclosing blocks onto W's
     carried mode that exactly equals each block axis. Matching against the
     block's representative axis (not W's own occurrence label) is
     runtime-exact across relabelings and keeps the over-slice fix: a value
     whose carried mode is a different physical index of the loop's space
     gets no level for it. Divergent stamps across fetch sites (an unsplit
     value) trip the same consistency SEQUANT_ASSERT as before.

The Task-4 map tests ([dag-scope], [sp2-noninner]) pass unchanged.
Cross-check, in slice_to_use, that the DAG-scope node->mode_to_level map
(cache_manager seam) agrees with today's exact index_position resolution:
wherever the old path found an exact match for a batch-context axis on a
node, mode_to_level->mode_of(that level) must return the same carried
position (SEQUANT_ASSERT). Where the old path needed the space-mapped
fallback instead (the over-slice-prone slot), the map is the CORRECTION --
logged as [SLICE-CORRECTION] under SEQUANT_UT_SLICE_DIAG, not asserted.

Debug-only: the map is never used to slice here (that is a later task);
the gate is inert when the seam is unwired (forest / whole-scope / hand-
built contexts). With populate_cell_mode_to_level now covering leaves and
escapes, the exact arm no longer fires: [ordered-executor], [scope-executor],
[dag-scope], [sp2-noninner], [slice-on-use] show zero exact-arm assertions
and zero [SLICE-CORRECTION] lines; the b3 water-20 aux-only repro no longer
aborts on the gate.
slice_to_use's shared entry-resolution closure (eval.hpp) now slices by a
per-path p_new: exact_axis present (forest/whole-scope, which push each
member's own physical axis) resolves the same intra-tree exact match as
before; the ordered push style (exact_axis == nullopt) resolves through the
schedule's node->mode_to_level map instead of the old exact+space-fallback
guess. The Task-5 equivalence assert stays active as a transitional check.

Exposed a wiring gap: only the eval::evaluate dispatch wrapper pre-wired the
node->ModeToLevel cache seam, so any direct caller of evaluate_ordered_schedule
/ evaluate_ordered_multiroot (several unit tests) ran with an unwired seam and
silently left ordered-push entries unsliced under the flip. Factored
populate_cell_mode_to_level's construction into a const-rich-compatible
detail::compute_cell_mode_to_level core and self-wired the seam inside
detail::run_ordered_schedule_pre_results (the shared core both entry points
delegate to), removing the now-redundant wiring from the dispatch wrapper.

Added a debug assert that (level.depth, level.space, level.ordinal) names one
representative axis GLOBALLY across the realized block tree, not just among a
block's own direct siblings (well_formed's existing check) -- the invariant
mode_to_level's position resolution relies on.

Updated three hand-built BatchContextEntry test fixtures (test_eval_ta.cpp,
test_placement_router.cpp, test_placement_remat.cpp) that simulate
forest/whole-scope exact-axis contexts but had left exact_axis unset from
their Task-6-era construction; they now set it to match their documented
exact-axis intent.

Aux-suite and tripwire failures are byte-identical to pre-flip baseline
(4 and 5 pre-existing failures respectively, same values); [dag-scope] and
[sp2-noninner] pass unchanged.
Reverts the uncommitted member-anchored mode_to_level WIP (dag_scope.hpp,
eval.hpp, member_axis.hpp, ordered_executor.hpp, ordered_schedule.hpp,
test_ordered_schedule.cpp) back to ddc360e content -- that approach failed
on symmetric CSE intermediates (I(i1,i2) == I(i2,i1)), which a per-cell
(cell, loop) -> mode map cannot represent when different use-sites bind the
loop to different physical slots of the same shared value. The env-gated
[PROD] diagnostic in backends/tiledarray/result.hpp is kept (unstaged) as a
validation tool for the follow-on tasks.

Adds tests/unit/test_sliced_canonical_layout.cpp: a minimal probe using a
Symmetry::Symm 2-occ-index intermediate shared by two use-site contractions
that bind the batched loop to different physical slots (slot 1 in one, slot
0 in the other). It pins today's baseline: occurrence_key correctly folds
the two occurrences to one router key, yet index_position shows the loop
sits at a different physical slot per occurrence for the same value -- the
exact ambiguity the loop-coloring design
(doc/dev/specs/2026-08-23-sliced-value-canonical-layout-loop-coloring-design.md)
resolves by coloring sliced slots with their DAG-scope loop identity.
The Task-7-P2 consumer-aware seam (LoopColoredSliceSeam::by_hash_consumer +
SlicedModeAssignment::occ_facts) replaced the old per-cell
node->ModeToLevel map on the ordered path, and the transitional
equivalence assert was already retired. Delete the whole dead chain:
ValueCell::mode_to_level, compute_cell_mode_to_level /
populate_cell_mode_to_level, the cache-side plumbing
(cell_mode_to_level_/set_cell_mode_to_level/cell_mode_to_level/
mode_to_level_of), the executor's self-wiring block and its
CellModeToLevelGuard, the eval.hpp m2l equivalence-gate diagnostic and its
two inherit sites, and the now-obsolete populate_cell_mode_to_level unit
test. The ModeToLevel TYPE and mode_to_level_from_signature (used by
compute_sliced_mode_assignment and slicing_signature.hpp) are unrelated
and stay.

Also convert the manual current_consumer save/restore around
run_ordered_contracted_block's two evaluate_impl call sites into an RAII
guard (CurrentConsumerGuard, mirroring LoopColoredSliceSeamGuard): the
manual restore leaked on a throwing evaluate_impl. Behavior on the
non-throwing path is unchanged (byte-identical test results).
… deadlock)

Batched slice-on-use resolved a canonical Index label against the fetched
node, but a divergent (relabeled) CSE occurrence is a different index-frame
that only shares SLOTS with the value's canonical frame, so the label match
failed (unsliced) or mis-resolved -> a shared occ index served full (32) in
one operand and sliced (16) in the other -> the ToT einsum's tiled-range
TA_ASSERT is elided in Release so the DistEval deadlocked.

Fix, in-frame end to end:
- compute_sliced_mode_assignment stamps each occurrence in its OWN frame
  (iterate occ.carried, not the value's first-occurrence w.carried).
- LoopColoredSliceSeam by_hash/by_hash_consumer and occ_facts carry a physical
  POSITION computed in each occurrence's own frame at schedule time; mode_of
  returns a position; slice_to_use uses it directly -- no index_position, no
  cross-frame label match, no first-match guess.
- Remove space_mapped_slicing (it guessed a mode by space); the resident-reads
  guard, the [homed] escape-output trace line, and the truthful sliced=/scope=
  and tiled-range traces stay.

Fixes the water-8 occ+aux 5-member (L*R) deadlock. A sibling-occ-loop case
(two forced-split occ loops sharing one canonical axis) still mis-binds a
mode's loop by SPACE in the assignment -- documented in
doc/dev/specs/2026-08-24-frame-correct-slot-slicing-design.md; the fix is the
next step. Env-gated investigation diagnostics (SEQUANT_UT_SMA_DIAG/HOME_DIAG/
PROD_TR, trange_annot) are retained for that pass and reverted at finalize.
Two TEST_CASEs assert the pre-2026-08-23 label seam (mode_of -> optional<Index>,
by_hash keyed by Index) and no longer compile against the position-based seam,
which blocks linking unit_tests-sequant. Guard them with #if 0 and a marker;
the loop-open-vs-sliced-mask plan Task 4 rewrites the seam and these cases
against the final per-occurrence positional semantics.
Add EvalExpr::batch_loops_opened_here() -- the subset of batched_here() for
which a node is the loop-OPEN site (outermost node introducing the physical
batch loop), distinct from the per-node sliced mask the DP stamps on every
carrying node. Carried on NodeBatchAnnotation::opened_here and applied by
binarize alongside set_batched_here. Lets a consumer reconstructing the
enclosing-loop nest (peak_profile's ectx) count each physical loop once
instead of once-per-carrying-node. Default empty; OFF path byte-identical.

Task 1 of the loop-open-vs-sliced-mask plan.
reconstruct_batched_modes now fills NodeBatchAnnotation::opened_here -- the
subset of a node's axes for which THIS node introduces the physical batch loop:
 - external, root-seed path: opened at the term root only (an external mode is
   on the final result, so the root is its outermost carrier);
 - external, node-level path: the injection site (new opened_at_node, NOT the
   D1-propagated carrying descendants that placed_at_node also records);
 - contracted: at the unique contraction node (aprime), mirroring the emitted
   Contracted axes.

Lets peak_profile build its enclosing-loop nest (ectx) from opens so one
physical loop counts once instead of once-per-carrying-node. axes emit
unchanged; OFF path byte-identical. Observable witness is the w8 ectx (Task 3);
the optimize-suite external-emission unit tests are pre-existing-red on this
branch (orthogonal calibration drift, surfaced now that the suite links again).

Task 2 of the loop-open-vs-sliced-mask plan.
peak_profile's visit accumulated n->batched_here() into the enclosing-loop
context, but the DP stamps an external mode's sliced mask on EVERY carrying
node, so one physical batch loop piled up once-per-carrying-node -- ectx became
the same occ index repeated (i i i ...) and no longer matched the DAG scope's
one-loop-one-level de-duplication. Read n->batch_loops_opened_here() instead,
which names each physical loop once. own_modes likewise moves to opens, matching
its documented intent ('its own node, not an ancestor's').

Verified on w8 occ+aux (SEQUANT_UT_SMA_DIAG): |ectx| drops from 6/7/3/4 to 2/1,
no duplicates; single-external occurrences now align |ectx|==|scope|==1 and
two-external ones show the two distinct loops cleanly, so dedup(ectx) & carried
yields clean per-occurrence slice positions for Task 4. OFF path byte-identical.

Task 3 of the loop-open-vs-sliced-mask plan.
…adlock)

Replace compute_sliced_mode_assignment's EXACT + REGIME-2 base_key passes with
a single per-occurrence positional rule. For each occurrence occ of value W
consumed by C, the loops the runtime crosses fetching W are C's enclosing DAG
blocks (build_scope), outermost first; dedup(occ.ectx) -- now one entry per
physical loop (Task 3) -- names those loops in W's OWN frame, outermost first.
Pair by nest position: block scope[k] slices occ mode nest[k], at physical
position index-in(occ.carried). All in occ's own frame -- no base_key, no
cross-frame label match, no first-match guess. Divergent (relabeled) and
symmetric occurrences (one value sliced on different positions by different
consumers) are handled uniformly by the consumer-keyed occ_facts; by_value is
left empty (the ordered executor always fetches under a tracked consumer).

The base_key block-find collapsed sibling occ loops (all occ share base_key
'i'), mis-slicing divergent occurrences -> two operands of one contraction
disagreed on a shared occ mode's extent -> ToTxToT DistEval deadlock (TA_ASSERT
elided in Release). w8 occ+aux ordered now completes: CSV-CCk Energy
-1.602851115435274 vs reference -1.6028511154353591 (8.5e-14, FP summation
order; lossless). SMA diagnostics retained (inert); removed in Task 5.

Task 4 of the loop-open-vs-sliced-mask plan.
…-leaf assignment case

peak_profile now builds OccurrenceRec::ectx from batch_loops_opened_here (Task
3), so fixtures that manually stamp batched_here must also stamp the loop-open:
 - test_eval_ta forest-descent equivalence cases (one Contracted block, one
   External block, reset-in-Contracted) and the orderedsched_stamp helper: aux
   Contracted opens at its contraction node, occ External opens at the root only;
 - test_eval_dryrun whole-scope outer-homed-aux case: both loops open at the root.

Guard 'compute_sliced_mode_assignment: a DF-leaf's aux+occ modes map ...': it
asserts the value-keyed loop_of/by_value API, which the per-occurrence occ_facts
(consumer-keyed) seam supersedes -- by_value is now legitimately empty and
loop_of returns nullopt. To be rewritten against occ_facts / by_hash_consumer.

Restores the eval unit suite to its pre-existing baseline (0 new regressions);
w8 occ+aux ordered stays lossless (-1.6028511154354086).
…ts gap)

Because by_value is now empty (per-occurrence occ_facts is the slice source and
the value-keyed by_hash fallback is unused), a fetch the schedule failed to
attribute a sliced-mode fact to would silently serve the operand UNSLICED and
mismatch its contraction partner -- and TA's tiled-range assert is elided in
Release, so the mismatched ToTxToT DistEval deadlocks instead of erroring.

Add LoopColoredSliceSeam::participates(hash, loop) -- whether the value has ANY
sliced-mode fact under the loop (any consumer, either map) -- and a guard in
slice_to_use: when the value PARTICIPATES in the crossed loop but this fetch got
no fact, throw sequant::Exception instead of proceeding unsliced. A value with
NO fact under the loop is genuinely invariant and correctly left unsliced, so the
guard does not fire there. w8 occ+aux ordered stays lossless
(-1.6028511154355747) -- occ_facts is complete for it, the guard never fires.
Remove the SEQUANT_UT_SMA_DIAG investigation scaffolding from
compute_sliced_mode_assignment: the [SMA-ALIGN]/[SMA-PAIR] traces and the
hardcoded w8 target-hash predicate (_sma_tgt). These were env-gated (inert
without the var), so removal is byte-identical. The general env-gated eval
traces ([HOME-SLICE], [PROD-PRE], trange_annot) are left in place -- they are
not hardcoded to w8 and remain useful for the ongoing multimode-batched-eval
branch.

Task 5 of the loop-open-vs-sliced-mask plan.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant