Skip to content

chore(red-team): run the missing control - a second opus pass over area 15 with no exclusions - to decide whether the fable rung earns its price #148

Description

@ms609-agent

Type: chore (methodology) · Area: 15 (Legacy pure-R search API), but the result governs the whole rotation

Three finder passes were run over area 15 on 2026-08-05/06 to decide red-team tier policy. They settled one question and left the more useful one open, because the same control was never run.

What was established

Arm Tokens Rel. price ≈ sonnet-equivalents Candidates sev:high
sonnet 136,777 137k 5 0
opus (sonnet's findings excluded) 180,024 360k 26 4
fable (nothing excluded) 162,265 649k 31 3 + 1 cross-area

Settled: cheap-first is not a saving. The sonnet pass removed no work from the opus pass — opus still read all 2,185 loc and found five times as much — so it was an added pass, not a substituted one. start_tier: sonnet on a never-visited area does not pay.

What was NOT established, and the control that would establish it

The finder sets are not nested in either direction. Fable independently re-found ~20 of opus's 26 and both of sonnet's distinctive findings (which opus had been told to skip) — but it missed five that opus found, including TBRSwap()'s indefinite hang on a trifurcating root. Opus likewise missed ~10 that fable found, including a reproducible SPR() crash (138 hits in an exhaustive n=5–9 sweep).

So "higher tier finds a superset" — the premise the whole escalation ladder rests on — is false as stated. What the data actually shows is that any competent independent pass finds a partly-different set.

That leaves the operative question unanswered: is fable's marginal contribution due to capability, or merely to being another fresh pass?

The missing arm: a second opus finder over area 15's identical scope, with no exclusions — matching the fable arm's conditions exactly, so the two are comparable. Then measure:

  1. How much of fable's 31 does opus-2 independently recover?
  2. How much of opus-1's 26 does opus-2 recover? (An estimate of single-pass recall at a fixed tier — interesting in its own right, and never measured here.)
  3. Does opus-2 find anything neither opus-1 nor fable found?

Cost is ~360k sonnet-equivalents against fable's ~649k. If opus-2 recovers a marginal set comparable to fable's, then the cheap escalation step is a fresh peer agent, not a rung bump, and the skill's existing "re-visit at the same tier with a fresh agent" should extend to: try one more fresh agent at the same rung before paying for the next one.

If opus-2 comes back with substantially less than fable did, that is direct evidence for the rung and the current ladder stands as written.

Why area 15 is the right site, and why this should not wait

It is the only area with three measured passes and a fully enumerated finding set, so opus-2's output can be scored against real ground truth rather than impressions. That property decays as #125#144 get fixed: once the defects are gone, no later pass can be scored against them.

Not to be confused with

This is not a request for more findings from area 15. If opus-2 files new issues, fine, but the deliverable is the three overlap numbers above and a one-paragraph verdict in dev/red-team/log.md on whether the fable rung earns its price.

Reopening condition

If closed unrun, reopen when a future round proposes a fable escalation on cost grounds — that decision has no evidence behind it until this control exists.

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:15Red-team focus area 15choreInfrastructure / process work, not a red-team finding

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions