Harness benchmark

Same model, same tasks, same machine, same meter — only the harness changes. Haider competes under identical rules and gets no advantage on this page.

connecting Fetching the current snapshot. last update unknown

Live runs

Counted runs of the active round, as they happen. An in-flight request shows unsettled — never a guess, never a zero.

No round is running.

Simulated retail is what a run would have cost you on Haider. Only true upstream cost — what the benchmark paid the provider — enters the ranking.

Leaderboard

Per-model tables are primary; a cross-model average is never the headline. Every harness/model pair is flagged same-vendor or cross-vendor by one rule — Haider and DeepSeek included.

No ranked result yet.

A leaderboard exists only after a completed round. Aborted rounds and public calibration tasks never become rankings.

Published rounds

Immutable once published — every counted and invalid run, CSV and JSON.

Nothing published yet.

Method

The short version. Every claim expands to its full committed rules.

  • Frozen before a round runs — model lane, rate cards, tasks, images, limits, seed, spend caps; all named by digest in a signed manifest.
  • The meter is the truth — a benchmark-owned proxy prices every accepted request; a request that cannot settle voids the whole run. Self-reports never win.
  • One scored outcome — at most three attempts; a wrong answer is never retried into a better score.
  • A winner needs intervals, not bars — sealed-task success with Holm-adjusted 95% CIs decides; ties fall to cost per success, then p90 time; otherwise no reliable difference.
  • Isolated execution — disposable VMs, offline network, hidden verifiers, no run on the revenue server.
  • The conflict is real — Haider owns the meter, the venue, and a contestant; the controls below are the minimum credible set, not a clean bill of health.
What is frozen, what the meter records, what counts as success

Fixed for every counted run

  • The model, resolved to one exact direct-provider lane with fallback disabled.
  • The upstream account rate card and the separately versioned retail simulation policy, both named by digest in the round manifest.
  • The task prompt and repository bundle.
  • Runner image family and hardware limits: a 4 vCPU, 8 GiB VM with the harness cgroup capped at 3 vCPU, 6 GiB, 1,024 processes and 20 GiB writable disk.
  • A 30-minute hard agent timeout.
  • At most 60 accepted model requests per attempt and 180 per logical run.
  • A cost ceiling of $1.00 true upstream usage per attempt and $3.00 per logical run.
  • Offline network policy, except the metering proxy, the run-scoped result API, and exact write-only upload grants through a controlled egress proxy.
  • The verification image, the hidden verifier, and the scoring rubric.
  • Three independent repeats, and no persistent home directory, session, compiler output or cache from a prior run.

Published the moment a round starts

  • The round manifest and its cryptographic signature.
  • Model alias, direct-provider lane policy, proxy build, upstream account-rate digest, and retail-simulation policy digest.
  • Public tasks and their verifiers.
  • Sealed-task metadata and the corpus Merkle root.
  • Harness and adapter versions, checksums, images, sanitized full configs, launch contract, and protocol path.
  • Runner and verifier image digests and the resource and network policy.
  • Metric definitions, the statistical method, the randomization seed, the stop and invalidation rules, and known shim qualification status.

The meter, not the self-report

Benchmark traffic uses direct provider keys and never buys inference from the Haider retail gateway, so a benchmark-owned metering proxy is the source of truth. It records an accepted request before forwarding, drains streaming responses for terminal usage, deduplicates by request identity, and prices exact billable token classes against the frozen account rate card.

A run is publishable only if every accepted request across all of its attempts settles to exactly one usage record or a provider-confirmed zero-bill outcome. A request that cannot settle makes the whole logical run measurement-invalid, even when a later attempt produced a correct patch: a later success cannot repair unknown cost. Proxy tokens and true upstream cost win any disagreement with what a harness reports about itself.

One scored outcome, three attempts at most

A logical run has one scored outcome and at most three operational attempts: the initial attempt plus no more than two automatic retries after a retryable error. Harness crashes, infrastructure faults, provider 5xx responses and timeouts are retryable.

A normal answer that fails verification is a legitimate task failure and is never retried into a better score. Cost per success is true upstream cost and includes failed runs and every automatic retry. Every attempt, reason, and reconciled dollar stays in the audit record.

What counts as success

  • The harness terminates normally before the hard limits.
  • The patch applies cleanly to a fresh copy of the exact base repository.
  • Every mandatory hidden verification gate passes.
  • Forbidden files, baseline tests, lockfiles and repository integrity constraints remain intact unless the task explicitly permits otherwise.
  • Usage reconciliation is complete across every operational attempt.

No human grades outputs. A normal harness exit is not success. A plausible answer is not success. A patch that passes tests after the harness crashed is recorded as would_have_passed and is still not a successful run.

Where the harnesses run

Harnesses never execute on the revenue server. Each attempt gets a fresh disposable VM in a separate cloud project, with a rootless container inside it, one task, a clean home directory and fixed limits. It can reach only a controlled egress proxy, which relays inference to the metering proxy and run-scoped uploads to the result API.

It cannot reach the Haider customer gateway, the production origin, private ranges, metadata endpoints, package registries, or the general internet. Dependencies are baked before the run. The VM is destroyed after its durable result is acknowledged, or reclaimed after failure.

How a winner is decided, and how uncertainty is reported

The decision rules

  • Success rate on the sealed tasks is primary. Every task carries equal weight.
  • Harness A is called more successful than B only when the Holm-adjusted 95 % interval for the paired, task-clustered success-rate difference excludes zero. A point-estimate lead is not a win.
  • If success ties, A is called more efficient only when its 95 % success lower bound is above −3 percentage points and its cost per success is at least 10 % lower with the paired cost-ratio interval entirely below 1.0.
  • If cost also ties, p90 agent wall time under the same 10 %-and-interval rule.
  • Conditional quality is reported, but it never rescues a low-success harness.
  • If no rule is met, the conclusion is no reliable difference, even if one bar is taller.

How uncertainty is reported

  • 10,000 deterministic bootstrap resamples from a published seed, clustered by task so all three repeats move together.
  • Success rate, paired success differences, cost and token per success, and p50/p90 time, cost, turns and tool calls all carry 95 % intervals.
  • Each task's 0/3, 1/3, 2/3 or 3/3 record is published, along with the fraction of tasks whose repeats disagree.
  • Planned pairwise significance claims are corrected with Holm's method within a round.
  • Zero successes makes cost per success and tokens per success infinite — not null, not zero.
  • Cache-hit rate with zero input tokens is N/A, not zero.
  • Three repeats expose stochastic instability; they do not produce precise per-task estimates, and this page will not pretend otherwise.
Conflict-of-interest controls, what cannot happen, what you can reproduce

Conflict-of-interest controls

  • Haider's and every competitor's versions and configs are frozen before sealed task access.
  • Task curation and verifier access are separated from Haider harness development.
  • The corpus Merkle root, analysis plan, schedule seed, exclusions, spend and stop rules, and harness digests are committed before running.
  • Adapters and analysis are open for review.
  • Every attempt, invalidation, and failed verifier stays in the audit export.
  • Maintainers are invited to inspect their redacted config before freeze and to file a public dispute after publication.
  • The first public round's task and invalidation log gets an external review.

What cannot happen

  • The round is not stopped early because Haider is winning or losing. It stops only for a predeclared event: the spend cap, a model identity or rate-card change, sealed-task exposure, metering reconciliation failure, production health degradation, or more than 5 % of attempts ending in retryable infrastructure or provider errors.
  • No post-hoc task deletion, no post-hoc metric deletion, and no new metric that reverses the order.
  • A flaky task is removed from every harness, which creates a new corpus manifest and restarts the statistical round.
  • No post-hoc “optimized Haider” rerun inside the same round. Haider may improve and enter a new, separately versioned round.
  • If Haider loses, the executive result names the winner and gives the measured difference before discussing caveats.
  • Results from different model aliases, model revisions, harness versions, task versions or benchmark envelopes are separate rounds and are never pooled.

What you can reproduce

Third parties can reproduce public cells immediately with their own direct-provider accounts, the published proxy and runner images, and the published rate and policy digests. Exact true-cost reproduction needs the same provider account rate card; exact customer-view reproduction needs the same Haider per-request retail and rounding policy.

Sealed cells become reproducible on release: after 90 days or the next round, whichever comes first, the retired sealed prompts, repositories where licensing permits, hidden verifiers, gold-patch hashes and redacted run traces are published and replaced before another ranked round. Where an upstream alias cannot be version-pinned, the report says so rather than implying bit-for-bit reproducibility.

What this page never shows

Public fields are a typed allowlist; failures publish as a stable code plus one fixed message — never copied output.

  • prompt text
  • repository contents
  • commands
  • tool arguments
  • tool results
  • stdout / stderr
  • patches and diffs
  • traces containing agent-authored source
  • hidden-test names
  • worker or controller hostnames and IPs
  • internal controller paths
  • provider response bodies
  • provider request IDs
  • capability or key IDs
  • customer telemetry