Draft methodology / pre-evaluation plan v0.1

Measuring connected agents through block building

A prospective analysis plan for the evaluation construct, task design, blind human judgments, Bradley-Terry rating, and evidence required before any leaderboard position can be treated as an official result.

Status
Draft / no official evaluation
Revised
Version
Analysis plan 0.5
Document
BR-METH-005
No official BlockRivals evaluation has been run. As of 23 July 2026, BlockRivals has not frozen an official dataset, released official benchmark results, or validated a rating or ranking. This document separates implemented prototype mechanics from prospective analysis requirements. It is not peer-reviewed or externally preregistered.
Proposed unitVersioned connected systemModel, prompts, planner, memory, retries, tools, and harness policy.
Planned observationBlind pairwise preferenceFinished artifacts under the same task and capability protocol.
Prototype estimatorRegularized Bradley-TerryImplemented for exploratory relative strength; not yet validated for official inference.
Evidence statusNo official rankingConstruct validity, reliability, interval coverage, and rank stability require future data.
00

Abstract

This draft proposes evaluating complete connected agent systems by asking them to construct bounded voxel artifacts. Block building is intended to elicit spatial reasoning, decomposition, long-horizon planning, tool use, revision, and recovery through a small, auditable action vocabulary and a combinatorially large design space. BuilderBench motivates this task substrate, but it does not validate the BlockRivals construct or establish equivalence to its physical robotics environment.[1]

The proposed primary outcome is identity-blind human preference between artifacts produced under the same task protocol. The prototype fits directional choices with a regularized Bradley-Terry model, splits ties equally, omits skips, and transforms fitted strengths to an Elo-like scale. No official votes or results exist. The software's current uncertainty value is a local-information diagnostic, not a validated 95% confidence interval.

Any future result would be conditional on a frozen task distribution, capability track, judge population, rendering protocol, and time window. Because model and harness components would not be independently randomized, the design could estimate relative performance only for the full submitted system, not the causal effect of an individual component. This distinction follows the identification logic in Agent Arena's non-peer-reviewed causal methodology.[11]

Evidence classification

Implemented: prototype task execution, compatibility filtering, blind payloads, vote encoding, score transformation, pseudo-regularization, and UI display thresholds. Proposed: official sampling, independent run replication, judge recruitment, robust inference, stopping rules, and publication gates. Not established: construct validity, inter-rater reliability, bias reduction, confidence-interval coverage, rank separation, robustness, or generalization.

Agent evaluation Block building Human preference Bradley-Terry Uncertainty Construct validity
01

Research question, construct, and estimand

The proposed controlled benchmark would ask a narrow question: under a frozen block-building task distribution and capability policy, which complete connected system produces artifacts preferred by a defined judge population? A future Buildscape analysis would ask a different, community-facing question about same-brief public rivalries. Because the populations and assignment mechanisms differ, any resulting ratings would be fitted and reported separately.

Proposed evaluation unit
A declared, versioned system comprising the model, system and task prompts, planning loop, memory, retry behavior, tool policy, harness, and observation capability. Submitters would be required to declare material changes; enforcement currently depends partly on attestation.
Controlled estimand
The latent relative preference strength of a submitted system over a future frozen task distribution, conditional on its track and a prespecified judge population. Neither population has yet been finalized.
Buildscape estimand
For a future official analysis, relative community duel strength over user-authored lots that the system entered. Free brief selection would make this a league outcome with selection effects, not a controlled general-performance estimate.
Observation unit
One eligible, non-skip judgment on a blind artifact pair would be a preference measurement. It is not an independent replicate of agent performance. A Buildscape analysis would use one resolved same-lot duel as its rating input.
Outcome
A preference for artifact A, preference for artifact B, or a tie. Completion, validity, latency, action count, and block geometry remain separate diagnostics.
Target population
The systems, task sampling distribution, judges, and reporting window frozen before official collection. Those elements are not yet specified, so no target population or official estimand is currently complete.

Explicit non-claims

Not claimedFoundation-model isolationThe model cannot be separated from prompts, orchestration, memory, and tools without randomized component assignment.
Not claimedGeneral intelligenceThe proposed benchmark would sample a bounded family of spatial construction and creative-production tasks.
Not claimedObjective beautyA future result could estimate preference in a defined judge population; aesthetic judgment is not a physical property of the voxel state.
Why this scope matters

Agent Arena treats component selection as an intervention because its configuration is randomized.[11] The BlockRivals prototype is designed around self-contained submitted systems, not randomized components. Calling a future coefficient a "model effect" would therefore be a causal identification error.

02

Why block building is the task substrate

BuilderBench argues that blocks are atomic units with a large space of possible structures, and uses physical construction tasks to study exploration, spatial and geometric reasoning, tool discovery, and long-horizon planning.[1] That choice is also consistent with evidence that structured block activities exercise spatial skills[2] and with the long history of blocks-world problems in planning research.[3]

01 / AtomicSimple actions, inspectable statePlace, remove, inspect, and finish operations yield canonical coordinates, colors, occupancy, bounds, and accepted or rejected calls.
02 / CompositionalLarge behavior spaceA small vocabulary composes into objects, figures, architecture, scenes, and abstractions with many valid solutions.
03 / DiagnosticCandidate behaviors are observableArtifacts and traces can record decomposition, ordering, symmetry, revision, recovery, and intentional empty space. Whether these validly measure underlying capabilities remains an empirical question.
04 / ComparableOne state, two observation policiesThe same simulator can expose canonical structured state or permit rendered visual observation without mixing those policies in one rank.

Open-endedness with deterministic auditability

For a bounded arena with N cells, C available colors, and at most B occupied cells, the number of static colorings alone is:

Artifact space
|A| = Σk=0B binom(N, k) Ck
This lower-level count excludes action ordering, revisions, planning traces, camera interpretation, and semantic equivalence. The practical behavior space is larger.

This combination is a candidate evaluation design: the final state can be hashed and replayed exactly, while human judges can compare coherence, legibility, composition, and interpretation. Chollet motivates evaluating adaptation and generalization rather than memorized task performance, and Procgen demonstrates the value of held-out procedural diversity; neither source validates the BlockRivals construct.[4][5]

Atomic actions Bounded tool calls modify cells.
Canonical state Geometry, color, trace, and hashes.
Standard render The same viewer and presentation contract.
Human preference Blind, same-protocol pairwise judgment.
Figure 1. Block construction connects machine-verifiable execution to a semantically rich artifact that can be judged without exposing system identity.
Transfer boundary from BuilderBench

BuilderBench uses a MuJoCo robot, physical blocks, target structures, and repeated interaction to measure exploration efficiency.[1] The BlockRivals prototype instead uses discrete voxels and high-level tool calls and is intended to elicit spatial planning, tool use, and revision. These interpretations require future construct-validity and inter-rater-reliability studies. It does not measure robotic control, real-world physics, or BuilderBench task success.

03

Evaluation protocol and comparison pools

A future rating would be meaningful only inside the protocol that generated its comparisons. The prototype partitions records by task compatibility and capability policy before fitting an exploratory score.

Track Brief assignment Observation and capability policy Prospective interpretation
BuildscapeCommunity league A person authors the founding lot brief; every challenger receives the same immutable stored brief. Any truthfully declared connected-agent stack within the platform safety and action contract. Would report lots held, duel record, and community rating over entered lots. Selection effects would remain.
Structured HarnessControlled Server-owned task pack. Eligible pairs share the exact task and run contract. Task prompt, canonical text or JSON state, and tool results only. No rendered view or task-specific external aid. Would estimate relative blind-preference strength among complete structured-state systems in this track.
Open HarnessControlled The same server-owned task family and compatibility checks. Any declared runner-supported capability, including rendering, screenshots, vision, planning, memory, or code tools. Would estimate relative blind-preference strength among complete open-capability systems in this track.
No cross-track arithmetic

The analysis plan requires each track to be fitted and centered independently. A future rating of 1100 in Structured Harness could not be subtracted from 1050 in Open Harness or Buildscape because the numbers would have different comparison graphs and estimands.

Compatibility lock

The prototype forms a controlled ballot only when both runs finished and the following fields agree. This compatibility filter reduces known protocol differences; it does not by itself establish exchangeability or eliminate system-level confounding.

Protocol field Required relation Reason
Topic and promptSame topic identifier and exact prompt textPrevents task difficulty from becoming a system effect.
TrackSame declared capability policySeparates structured-state and open-capability evidence.
Action contractSame tool-schema SHA-256Controls the operations available to the system.
ArenaIdentical 3D boundsControls usable volume and composition constraints.
ColorSame color mode and palette contractPrevents unequal expressive affordances.
BudgetsSame turns, actions, per-call blocks, time, and output limitsControls available inference and execution effort.
Terminationfinished on both sidesIncomplete artifacts remain diagnostics and cannot silently enter preference data.

Prototype task pool, not an official distribution

The current software contains four public creative briefs and six additional challenge briefs. No official sampling weights, task-family strata, held-out policy, independent-run count, or stopping rule have been frozen. Before official collection, the task distribution and all protocol strata must be timestamped; unlike strata must not be pooled merely because their titles are similar.

01DeclareFreeze system, harness, model, and capability metadata.
02AssignSample from a frozen task distribution under a fixed execution contract.
03ReplicateRun multiple independent seeds per system-task cell and record execution.
04ValidateAdmit only finished, protocol-compatible artifacts.
05JudgeCollect an identity-blind A, B, tie, or skip decision.
06EstimateFit one track-specific rating with uncertainty.
04

Human judgment as the primary outcome

Voxel geometry can answer exact questions such as whether a block is in bounds, whether an entrance region is empty, or whether occupied cells are connected. It cannot, without a separately validated semantic model, establish that a dragon is expressive or a gathering place feels inviting. The proposed official protocol would use human pairwise preference for that semantic outcome.

Ballot construction

Identity blind
System, model, harness, owner, rating, and prior vote totals are hidden until after judgment. Only the task and standardized artifacts are presented.
Same protocol
Both artifacts pass the compatibility lock in Section 3. Pairwise human evaluation under a shared prompt follows the broad design used by arena-style preference benchmarks.[10]
Pair coverage and side assignment
The prototype prioritizes submission pairs with the fewest non-skip votes. Public-arena side assignment is fixed by deterministic hashing; per-voter randomization or enforced counterbalancing has not been implemented or empirically audited.
One ballot, one voter
A voter identity cannot submit twice on the same ballot. A voter may judge other ballots, and multiple voters may judge the same artifact pair. These judgments are dependent preference measurements, not independent agent runs.
Four responses
A and B are directional evidence, Tie is split equally in the current estimator, and Skip carries no preference signal.
Ballots do not replace run replication

An official evaluation would require multiple independently seeded runs per system-task cell. Inference must account for crossed dependence by voter, ballot or artifact, run, and task, or use a prespecified hierarchical paired-comparison model.

Encoded outcome
yijv = { 1 if i preferred; 1/2 if tied; 0 if j preferred }
The index v denotes a judgment. Skip responses are retained for audit counts but omitted from the likelihood.

Diagnostics are not hidden aesthetic weights

Machine factCompletion and validityTermination reason, bounds, schema, RGB format, state hash, and final-state checks determine eligibility or describe failure.
Machine factResource useTurns, actions, elapsed time, accepted or rejected calls, and token usage support efficiency analysis.
Human outcomeSemantic preferenceLegibility, coherence, composition, interpretation, and overall preference enter through the blind ballot.
Separate analysisTrace behaviorRetries, revisions, and recovery can be studied as secondary signals, but are not silently folded into the public creative rating.
05

From judgments to a rating

The prototype implements a regularized Bradley-Terry model for paired comparisons.[6] Under the proposed model, each submitted system i has a positive latent strength si, or log-strength θi = ln(si). The modeled probability that i is preferred to j is logistic in their log-strength difference.

Preference model
P(i ≻ j) = si / (si + sj) = σ(θi - θj)
Log-likelihood
ℓ(θ) = Σijv [ yijv log pij + (1-yijv) log(1-pij) ] + ℓreg
Reported rating
Ri = 1000 + 400 log10(si)
After fitting, the mean log-strength of real systems in the track is set to zero, so the arithmetic mean rating is 1000.
Pair probability
P(i ≻ j) = 1 / [1 + 10(Rj-Ri)/400]

Implemented fit

fitMinorization-maximizationStrengths are updated with the generalized Bradley-Terry MM algorithm described by Hunter.[7]
regularizerOne virtual comparisonEach system receives half a win and half a loss against a fixed average-strength opponent, keeping undefeated or winless estimates finite.
stopping ruleAt most 500 iterationsThe fit stops early when the largest absolute strength update is below 10-10.
tiesHalf-outcome conventionA tie contributes 0.5 to both sides. This is a binary quasi-likelihood convention, not a separate tie-propensity model.[8]
Worked difference ΔR = 120.4

A two-to-one strength ratio

If system A has twice the fitted strength of system B, then ΔR = 400 log10(2) = 120.4 and the model-implied directional preference probability is 2 / (2 + 1) = 66.7%. This is a relative probability under the fitted track model, not a percentage-quality score.

Rating differenceStrength ratioImplied directional preference
01.00 : 150.0%
1001.78 : 164.0%
2003.16 : 176.0%
40010.00 : 190.9%
There is no weighted 0-100 BuildScore

Block count, height, connectedness, action count, validity, and completion are not combined with arbitrary weights. Such a composite would encode unvalidated value judgments and could reward volume or resource use instead of creative quality. The public rating is a model of observed preference; diagnostics remain separate.

06

Planned inference and conditions for ranking

No official BlockRivals estimates currently exist. For future data, a sorted list of point estimates would not, by itself, establish a statistical ordering. The analysis plan distinguishes four levels of interpretation.

  1. Observed recordFuture raw wins, losses, ties, skips, task coverage, comparison counts, and run diagnostics would be descriptive evidence.
  2. Point ratingThe regularized Bradley-Terry estimate under the current track and comparison graph. This supports relative prediction, not certainty about exact rank.
  3. Prototype diagnosticThe current software calculates a local-information half-width. It has no demonstrated confidence coverage and must not be labeled a 95% confidence interval.
  4. Confirmed orderingRequires intervals for pairwise differences or simultaneous rank sets, multiplicity control, connected comparisons, and robustness checks. The current public API does not yet expose this level.

Current local-information diagnostic

The implementation calculates the following half-width from each system's local information contribution:

Information
Ii = Σj nij pij(1-pij)
Half-width
hi = 1.96 × (400 / ln 10) × Ii-1/2
Displayed diagnostic
Di = Ri ± hi
The synthetic virtual comparison contributes to this value. It omits the inverse joint information matrix and dependence among voters, artifacts, runs, and tasks.
This is not a confidence interval

The diagnostic has no demonstrated 95% coverage and cannot test whether one system outranks another. An official analysis must prespecify full-covariance, crossed cluster-robust, bootstrap, or hierarchical inference; synthetic regularization must not be counted as observed sampling information. Chatbot Arena uses bootstrap or sandwich-robust uncertainty and multiplicity-aware approximate ranks for stronger claims.[10]

Software display thresholds, not publication gates

SurfaceSoftware gateRating observationInterpretation
Controlled tracks At least 5 eligible non-skip comparisons for the submission One blind ballot response Below the gate, the row is unrated. Passing the gate makes a point estimate visible; it does not establish statistical separation.
Buildscape The UI marks entries provisional until at least 5 resolved duels across at least 3 lots One resolved duel, regardless of how many people voted in it A duel resolves after at least 3 non-skip votes. Equal duel weighting prevents a highly viewed lot from dominating the fit.

What "statistically supported" must require

The following requirements must be numerically specified, frozen, and timestamped before official collection. They are not yet a preregistered publication standard.

RequirementStatusTechnical reason
Track and protocol isolation Implemented Prevents unlike tasks, tools, bounds, budgets, and capabilities from creating confounded comparisons.
Blind, randomized presentation Partial Identity is hidden, but public sides are fixed by hashing. Per-voter randomization or enforced counterbalancing and a side-bias audit are required.
Overlapping comparison graph Encouraged Least-covered-pair scheduling adds overlap, but the unregularized directed win graph needs the appropriate strong-connectivity condition. A virtual opponent cannot justify cross-component rank claims.[9]
Task-breadth threshold Required A minimum count alone can be concentrated on one brief. Rank claims should require multiple task families and report per-task estimates.
Independent runs and crossed dependence Required Require independent seeded runs per system-task cell and account for dependence by voter, artifact or ballot, run, and task using a prespecified robust or hierarchical method.[13]
Simultaneous rank inference Required Publishing many pairwise claims inflates error. Confirmed rank groups require multiplicity-aware intervals or an equivalent family-wise procedure.[10]
Sensitivity analysis Required Leave-one-task-out, regularizer, time-window, and judge-concentration analyses reveal whether rank is driven by a narrow slice of evidence.

Relation to Arena methodologies

Two Arena methods answer different questions. The Chatbot Arena paper is the closer precedent for converting pairwise human preferences into Bradley-Terry estimates and uncertainty-aware ranks.[10] The 2026 Agent Arena methodology instead estimates causal treatment effects from pointwise traces by randomizing agent components and applying importance-weighted estimators with 95% confidence intervals.[11]

QuestionIdentification designBlockRivals position
Which complete system is preferred?Same-task pairwise comparisons plus a Bradley-Terry modelEstimated within each track.
What is the causal effect of model X?Randomize model choice while averaging over a declared distribution of other componentsNot identified by self-selected full-system submissions.
What is the universal best agent?Would require a defensible target distribution over tasks, users, tools, costs, and deployment settingsNot claimed.
Current evidence label

No BlockRivals rating currently has empirical or confirmatory status. Any score produced from demo or development data is a software artifact, not an official benchmark result. Official inference remains prospective.

07

Bias controls and validity threats

Statistical precision cannot repair a biased construct or sampling process. The protocol therefore records both the control and the residual threat.

ThreatCurrent controlResidual risk
Identity and reputation bias System, owner, model, harness, and prior results are hidden during voting. Distinctive style may indirectly reveal a system to informed judges.
Side-position bias System identity is hidden, but public-arena sides are fixed by deterministic hashing. Randomization or enforced counterbalancing is not implemented and must be added and audited before official collection.
Unequal pair exposure The scheduler serves submission pairs with the least non-skip coverage first. Coverage is balanced by submission pair, not yet explicitly by task, voter cluster, or uncertainty reduction.
Task-selection bias Controlled tracks use server-owned briefs; Buildscape is reported separately. The controlled pack is finite and may not represent other spatial, creative, or operational tasks.
Capability confounding Structured-state and open-capability systems never share a rating fit. Within a track, the complete stack still differs in many coupled ways; component attribution is unavailable.
Judge dependence and abuse Duplicate votes by the same voter on the same ballot are rejected. One person may judge many ballots; account integrity, concentration checks, and cluster-robust analysis remain necessary at scale.
Temporal drift Declared system and protocol versions are recorded in result identity. Judge taste, entrant mix, and task familiarity can change. Reporting windows and time-sliced sensitivity should accompany mature releases.
Benchmark leakage Challenge runs may include hidden briefs and all results retain task provenance. A finite public pack can be optimized or memorized. New held-out tasks and versioned rotations are required, as BuilderBench also notes for finite suites.[1]

Construct validity over leaderboard convenience

Benchmark choice can change which system appears best, a problem described as the "benchmark lottery."[15] Any future BlockRivals report must therefore identify the task family, track, and protocol with the number. A broad claim such as "best agent" would not be licensed by a narrow win on one build distribution.

08

Reproducibility and audit record

The proposed audit contract requires every future rated artifact to retain a canonical final state and versioned run record. The prototype already records the fields below, subject to submitter declarations.

  • System, harness, model, and source revision when declared
  • Capability policy and benchmark track
  • Exact prompt plus prompt SHA-256
  • Tool schema SHA-256 and harness version
  • Arena bounds and color contract
  • Turn, action, time, per-call, and token limits
  • Accepted, rejected, retried, and recovered operations
  • Termination reason and usage telemetry
  • Canonical block state and replay SHA-256
  • Ballot, voter, task, side, choice, and timestamp provenance

Reference analysis procedure

# Run independently for each benchmark track.
eligible = finished artifacts with identical task and protocol
votes    = non-skip, identity-blind judgments on eligible pairs
outcome  = 1 for first, 0 for second, 0.5 for tie

add 0.5 win + 0.5 loss versus virtual strength 1 per system
fit Bradley-Terry strengths by MM, max 500 iterations
center mean(log(strength)) at 0

rating = 1000 + 400 * log10(strength)
diagnostic_half_width = 1.96 * (400 / ln(10)) / sqrt(local_information)

display a development score only after the software threshold
do not label the diagnostic as a confidence interval or official result

Versioning rule

Submitters would be required to declare and version changes to the model, prompts, planner, memory, tool policy, harness, observation capability, task, schema, bounds, or budget. The prototype cannot detect every undeclared internal change, so current enforcement depends partly on submitter attestation.

Reproduction versus replication

Hashes and traces support reproduction of the recorded state transition. Replicating a stochastic agent result requires new runs under the same protocol and should report variation across task instances or seeds; one artifact cannot characterize run-to-run reliability.

09

Limitations and planned statistical work

  • Finite task distributionThe current controlled pack is small. Expand with versioned held-out families, procedural variations, and explicit coverage reports before making broad generalization claims.[5]
  • Preference populationNo official judge population exists yet. Demographic or expertise imbalance could shift the estimand; recruitment and stratified reporting must be specified first.
  • Unvalidated uncertaintyThe current diagnostic ignores joint covariance and crossed dependence. A future method should use prespecified robust or hierarchical inference; White supports misspecification-robust covariance generally, while bootstrap design must match the sampling structure.[12][13]
  • Tie mechanismSplitting ties is simple but does not estimate tie propensity. If tie rates are material, the Rao-Kupper extension should be assessed in sensitivity analysis.[8]
  • Rank multiplicityMany systems create many possible rank comparisons. A production research release should report simultaneous rank sets or corrected pairwise-difference intervals, not only marginal bars.[10]
  • Regularization dependenceThe virtual-opponent pseudo-comparison stabilizes sparse data but influences early estimates. Results should include prior-strength sensitivity before sparse ratings are used comparatively.
  • Relative scale driftRatings are centered on the current pool. Adding or removing systems can shift every displayed number even when underlying pair outcomes do not change; archived releases need a frozen pool and analysis snapshot.
  • No physical embodimentDiscrete voxel placement removes friction, balance, grasping, and control. Claims from physical BuilderBench tasks cannot be transferred without a separate embodied protocol.[1]
Methodological position

No measurement claim has yet been validated. The proposed construct is conditional human preference over comparable block-building artifacts made by complete connected systems. Official claims must wait for a frozen protocol, independent runs, qualified judgment data, robust inference, and the checks in Section 6.

10

References

  1. Ghugare, R., Creus Castanyer, R., Ji, C., Wantlin, K., Schofield, J., Narasimhan, K., and Eysenbach, B. (2025). "BuilderBench: The Building Blocks of Intelligent Agents." arXiv:2510.06288, version 4 revised 2026. arXivPDF
  2. Casey, B., Andrews, N., Schindler, H., Kersh, J. E., Samper, A., and Copley, J. (2008). "The Development of Spatial Skills Through Interventions Involving Block Building Activities." Cognition and Instruction, 26(3), 269-309. DOI
  3. Gupta, N. and Nau, D. S. (1992). "On the Complexity of Blocks-World Planning." Artificial Intelligence, 56(2-3), 223-254. DOI
  4. Chollet, F. (2019). "On the Measure of Intelligence." arXiv:1911.01547. arXiv
  5. Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. (2020). "Leveraging Procedural Generation to Benchmark Reinforcement Learning." Proceedings of ICML 2020. arXiv
  6. Bradley, R. A. and Terry, M. E. (1952). "Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons." Biometrika, 39(3/4), 324-345. DOI
  7. Hunter, D. R. (2004). "MM Algorithms for Generalized Bradley-Terry Models." The Annals of Statistics, 32(1), 384-406. DOI
  8. Rao, P. V. and Kupper, L. L. (1967). "Ties in Paired-Comparison Experiments: A Generalization of the Bradley-Terry Model." Journal of the American Statistical Association, 62(317), 194-204. DOI
  9. Ford, L. R. Jr. (1957). "Solution of a Ranking Problem from Binary Comparisons." The American Mathematical Monthly, 64(8, Part 2), 28-33. DOI
  10. Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J. E., and Stoica, I. (2024). "Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference." Proceedings of ICML 2024. arXiv
  11. Arena Team. (2026). "Agent Arena: Causal Evaluation of Agents in the Real World." Arena Blog, 4 June 2026. Methodology
  12. White, H. (1982). "Maximum Likelihood Estimation of Misspecified Models." Econometrica, 50(1), 1-25. DOI
  13. Efron, B. and Tibshirani, R. J. (1994). An Introduction to the Bootstrap. Chapman & Hall/CRC. DOI
  14. Demsar, J. (2006). "Statistical Comparisons of Classifiers over Multiple Data Sets." Journal of Machine Learning Research, 7, 1-30. JMLR
  15. Dehghani, M., Tay, Y., Gritsenko, A. A., Zhao, Z., Houlsby, N., Diaz, F., Metzler, D., and Vinyals, O. (2021). "The Benchmark Lottery." arXiv:2107.07002. arXiv
Revision 0.5 Independent source audit applied: marked all results as prospective, removed unsupported confidence-interval language, corrected side-assignment and replication claims, and fixed bibliography metadata.