Referee Launch Plan
Ravage has two distinct launch stages. They must not be collapsed into one
claim: the current artifact is a frozen Ravage-only baseline, while the planned
referee launch requires real external-agent rows and repeated runs.
Stage 1: Frozen Ravage Baseline
The current public artifact is one uninterrupted Ravage run across all 104
public XBEN cases under a frozen description-only black-box contract:
- 85 / 104 exact randomized flags
- 81.73% solve rate
- $55.758722 in provider-usage list-price model cost
- $0.655985 per valid flag
- 1,151 model replies accounted for, with zero unmatched attempts
- 16 failures, 1 error, and 2 timeouts retained
- no retries, row replacement, or code changes during execution
Review the evidence bundle.
The reviewer guide links the readable ledger, per-case TSV, machine-readable
report, provenance, audit, and checksum commands. Link public announcements to
that guide, not directly to the raw archive.
XBOW now labels the public suite outdated and saturated
and says its vulnerabilities are present in model training. Randomized flags
protect exact-output scoring, but they do not make the vulnerability patterns
novel again. Stage 1 is therefore an auditable public-benchmark regression and
harness-integrity baseline, not a frontier-capability, production-efficacy, or
head-to-head claim.
What Stage 1 Does Not Establish
- Repeatability: there is one frozen full run, not a repeated sample with
variance or confidence intervals.
- Competitive rank: no external agent has a published row from Ravage’s
harness.
- Unseen-task performance: XBEN and the committed solution traces are public.
- Production pentest quality: flag capture does not measure normal-application
false positives, business impact, remediation, reporting, or organizational
safety.
- Model independence: the result belongs to the complete Ravage plus
gpt-5.4-2026-03-05 configuration.
- Total operating cost: the reported dollars cover provider-usage list-price
model text tokens, not compute, storage, engineering, or review labor.
- Portable speed: the recorded runtime used amd64 containers on an ARM host.
The broader product and evidence caveats are listed under
Honest Limitations.
Stage 1 Launch Copy
We ran Ravage once across all 104 public XBEN cases under a frozen,
description-only black-box contract. It captured 85 exact randomized flags at
$55.76 in provider-usage list-price model cost. Every failure, error, timeout,
model reply, and artifact hash is retained. Because XBEN is now saturated, this
is a reproducible regression baseline—not a claim of unseen-task superiority.
This evidence release can be published now. Do not call it a leaderboard or
say that Ravage benchmarked other agents.
Stage 2: Cross-Agent Referee
The referee launch starts only after Ravage and at least two real external
agents have complete results under one declared protocol. The repository
currently contains the harness and smoke examples, but no public
external-agent matrix.
Before any result is known, freeze and publish:
- benchmark and harness commit SHAs;
- complete target set and input-visibility contract;
- agent names, versions, adapters, models, and reasoning settings;
- time, turn, request, retry, isolation, and cost policies;
- evaluator-owned exact-flag and false-positive rules;
- infrastructure-failure and exclusion policy;
- cost source, including whether it is measured, list-price computed, or
unknown;
- whether each row was maintainer-run or operator-run.
If systems cannot use meaningfully comparable models, budgets, or tool access,
report them in separate cohorts rather than presenting a controlled ranking.
The table compares complete system configurations, not the underlying models
in isolation.
Minimum Public Artifact
A public head-to-head claim needs all of these:
report.json, leaderboard.tsv, and leaderboard.md;
artifacts.sha256 plus immutable code and target provenance;
- per-agent
stdout.log, stderr.log, and agent-result.json;
- complete rows for every selected target, including errors and timeouts;
- valid, invalid, self-reported, false-positive, and out-of-scope counts;
- cost source or an explicit
cost_status: unknown;
- an explanation of every failed, errored, or excluded row;
- repeat runs reported separately rather than merged into a best-of score.
Do not publish a headline score from a local note, screenshot, or
self-reported finding count. Scores come from evaluator-scored artifacts.
Execution Order
- Tag a stable harness version and select comparable external adapters.
- Publish and freeze the protocol before seeing full-matrix results.
- Run a small non-reportable adapter smoke test.
- Freeze the source and execute the complete matrix without source changes.
- Repeat the same complete matrix from the same frozen source.
- Audit all errors and false positives plus a sample of passes.
- Give named maintainers a factual-review window and publish any errata.
- Publish the table only after both runs and every terminal row are explained.
Scoreboard Rules
- Exact evaluator-owned flags count as valid.
- Unknown flags count as invalid.
- Findings without replayable proof count as false positives.
- Out-of-scope findings are reported separately and count against false
positives.
- Unknown cost is not zero cost.
- Cost per valid flag is reported only when cost is known or computable.
- Errored and timed-out rows remain visible.
- A repeat run is a separate observation, not a replacement result.
Stage 2 Launch Copy
Use this only after the required multi-agent artifacts exist:
We ran Ravage, X, and Y under one declared benchmark protocol with
evaluator-scored exact flags, replayable proof requirements, false-positive
accounting, and cost per valid flag. This compares the complete system
configurations—not their underlying models in isolation. Every row links to
its provenance, terminal status, cost accounting, and artifact bundle.
Claim Language
Use precise descriptions:
- evaluator-scored exact randomized flags;
- frozen public-benchmark regression;
- complete artifact bundle;
- declared model, limits, and cost accounting;
- Ravage-only baseline until external rows exist.
Avoid claims that the evidence does not support:
- cheat-proof;
- the only benchmark or the only cost-aware benchmark;
- nobody else reports cost;
- state of the art;
- better than Strix, MAPTA, or another system without a controlled comparison;
- plural “agents” before stage two exists.
XBEN, MAPTA, Strix, BountyBench, HAL, and other evaluation projects already
publish benchmark, evidence, or cost artifacts. Ravage’s prospective
differentiator is the declared web-agent execution contract and interoperable
row-level evidence—not the invention of benchmarks, scoreboards, or cost
tracking.