Type: chore (methodology) · Area: 15 (Legacy pure-R search API), but the result governs the whole rotation
Three finder passes were run over area 15 on 2026-08-05/06 to decide red-team tier policy. They settled one question and left the more useful one open, because the same control was never run.
What was established
| Arm |
Tokens |
Rel. price |
≈ sonnet-equivalents |
Candidates |
sev:high |
sonnet |
136,777 |
1× |
137k |
5 |
0 |
opus (sonnet's findings excluded) |
180,024 |
2× |
360k |
26 |
4 |
fable (nothing excluded) |
162,265 |
4× |
649k |
31 |
3 + 1 cross-area |
Settled: cheap-first is not a saving. The sonnet pass removed no work from the opus pass — opus still read all 2,185 loc and found five times as much — so it was an added pass, not a substituted one. start_tier: sonnet on a never-visited area does not pay.
What was NOT established, and the control that would establish it
The finder sets are not nested in either direction. Fable independently re-found ~20 of opus's 26 and both of sonnet's distinctive findings (which opus had been told to skip) — but it missed five that opus found, including TBRSwap()'s indefinite hang on a trifurcating root. Opus likewise missed ~10 that fable found, including a reproducible SPR() crash (138 hits in an exhaustive n=5–9 sweep).
So "higher tier finds a superset" — the premise the whole escalation ladder rests on — is false as stated. What the data actually shows is that any competent independent pass finds a partly-different set.
That leaves the operative question unanswered: is fable's marginal contribution due to capability, or merely to being another fresh pass?
The missing arm: a second opus finder over area 15's identical scope, with no exclusions — matching the fable arm's conditions exactly, so the two are comparable. Then measure:
- How much of fable's 31 does opus-2 independently recover?
- How much of opus-1's 26 does opus-2 recover? (An estimate of single-pass recall at a fixed tier — interesting in its own right, and never measured here.)
- Does opus-2 find anything neither opus-1 nor fable found?
Cost is ~360k sonnet-equivalents against fable's ~649k. If opus-2 recovers a marginal set comparable to fable's, then the cheap escalation step is a fresh peer agent, not a rung bump, and the skill's existing "re-visit at the same tier with a fresh agent" should extend to: try one more fresh agent at the same rung before paying for the next one.
If opus-2 comes back with substantially less than fable did, that is direct evidence for the rung and the current ladder stands as written.
Why area 15 is the right site, and why this should not wait
It is the only area with three measured passes and a fully enumerated finding set, so opus-2's output can be scored against real ground truth rather than impressions. That property decays as #125–#144 get fixed: once the defects are gone, no later pass can be scored against them.
Not to be confused with
This is not a request for more findings from area 15. If opus-2 files new issues, fine, but the deliverable is the three overlap numbers above and a one-paragraph verdict in dev/red-team/log.md on whether the fable rung earns its price.
Reopening condition
If closed unrun, reopen when a future round proposes a fable escalation on cost grounds — that decision has no evidence behind it until this control exists.
Type: chore (methodology) · Area: 15 (Legacy pure-R search API), but the result governs the whole rotation
Three finder passes were run over area 15 on 2026-08-05/06 to decide red-team tier policy. They settled one question and left the more useful one open, because the same control was never run.
What was established
sev:highsonnetopus(sonnet's findings excluded)fable(nothing excluded)Settled: cheap-first is not a saving. The sonnet pass removed no work from the opus pass — opus still read all 2,185 loc and found five times as much — so it was an added pass, not a substituted one.
start_tier: sonneton a never-visited area does not pay.What was NOT established, and the control that would establish it
The finder sets are not nested in either direction. Fable independently re-found ~20 of opus's 26 and both of sonnet's distinctive findings (which opus had been told to skip) — but it missed five that opus found, including
TBRSwap()'s indefinite hang on a trifurcating root. Opus likewise missed ~10 that fable found, including a reproducibleSPR()crash (138 hits in an exhaustive n=5–9 sweep).So "higher tier finds a superset" — the premise the whole escalation ladder rests on — is false as stated. What the data actually shows is that any competent independent pass finds a partly-different set.
That leaves the operative question unanswered: is fable's marginal contribution due to capability, or merely to being another fresh pass?
The missing arm: a second
opusfinder over area 15's identical scope, with no exclusions — matching the fable arm's conditions exactly, so the two are comparable. Then measure:Cost is ~360k sonnet-equivalents against fable's ~649k. If opus-2 recovers a marginal set comparable to fable's, then the cheap escalation step is a fresh peer agent, not a rung bump, and the skill's existing "re-visit at the same tier with a fresh agent" should extend to: try one more fresh agent at the same rung before paying for the next one.
If opus-2 comes back with substantially less than fable did, that is direct evidence for the rung and the current ladder stands as written.
Why area 15 is the right site, and why this should not wait
It is the only area with three measured passes and a fully enumerated finding set, so opus-2's output can be scored against real ground truth rather than impressions. That property decays as #125–#144 get fixed: once the defects are gone, no later pass can be scored against them.
Not to be confused with
This is not a request for more findings from area 15. If opus-2 files new issues, fine, but the deliverable is the three overlap numbers above and a one-paragraph verdict in
dev/red-team/log.mdon whether the fable rung earns its price.Reopening condition
If closed unrun, reopen when a future round proposes a
fableescalation on cost grounds — that decision has no evidence behind it until this control exists.