ravage

Referee Launch Plan

Ravage has two distinct launch stages. They must not be collapsed into one claim: the current artifact is a frozen Ravage-only baseline, while the planned referee launch requires real external-agent rows and repeated runs.

Stage 1: Frozen Ravage Baseline

The current public artifact is one uninterrupted Ravage run across all 104 public XBEN cases under a frozen description-only black-box contract:

Review the evidence bundle. The reviewer guide links the readable ledger, per-case TSV, machine-readable report, provenance, audit, and checksum commands. Link public announcements to that guide, not directly to the raw archive.

XBOW now labels the public suite outdated and saturated and says its vulnerabilities are present in model training. Randomized flags protect exact-output scoring, but they do not make the vulnerability patterns novel again. Stage 1 is therefore an auditable public-benchmark regression and harness-integrity baseline, not a frontier-capability, production-efficacy, or head-to-head claim.

What Stage 1 Does Not Establish

The broader product and evidence caveats are listed under Honest Limitations.

Stage 1 Launch Copy

We ran Ravage once across all 104 public XBEN cases under a frozen,
description-only black-box contract. It captured 85 exact randomized flags at
$55.76 in provider-usage list-price model cost. Every failure, error, timeout,
model reply, and artifact hash is retained. Because XBEN is now saturated, this
is a reproducible regression baseline—not a claim of unseen-task superiority.

This evidence release can be published now. Do not call it a leaderboard or say that Ravage benchmarked other agents.

Stage 2: Cross-Agent Referee

The referee launch starts only after Ravage and at least two real external agents have complete results under one declared protocol. The repository currently contains the harness and smoke examples, but no public external-agent matrix.

Before any result is known, freeze and publish:

If systems cannot use meaningfully comparable models, budgets, or tool access, report them in separate cohorts rather than presenting a controlled ranking. The table compares complete system configurations, not the underlying models in isolation.

Minimum Public Artifact

A public head-to-head claim needs all of these:

Do not publish a headline score from a local note, screenshot, or self-reported finding count. Scores come from evaluator-scored artifacts.

Execution Order

  1. Tag a stable harness version and select comparable external adapters.
  2. Publish and freeze the protocol before seeing full-matrix results.
  3. Run a small non-reportable adapter smoke test.
  4. Freeze the source and execute the complete matrix without source changes.
  5. Repeat the same complete matrix from the same frozen source.
  6. Audit all errors and false positives plus a sample of passes.
  7. Give named maintainers a factual-review window and publish any errata.
  8. Publish the table only after both runs and every terminal row are explained.

Scoreboard Rules

Stage 2 Launch Copy

Use this only after the required multi-agent artifacts exist:

We ran Ravage, X, and Y under one declared benchmark protocol with
evaluator-scored exact flags, replayable proof requirements, false-positive
accounting, and cost per valid flag. This compares the complete system
configurations—not their underlying models in isolation. Every row links to
its provenance, terminal status, cost accounting, and artifact bundle.

Claim Language

Use precise descriptions:

Avoid claims that the evidence does not support:

XBEN, MAPTA, Strix, BountyBench, HAL, and other evaluation projects already publish benchmark, evidence, or cost artifacts. Ravage’s prospective differentiator is the declared web-agent execution contract and interoperable row-level evidence—not the invention of benchmarks, scoreboards, or cost tracking.