Abstract
This draft proposes evaluating complete connected agent systems by asking them to construct bounded voxel artifacts. Block building is intended to elicit spatial reasoning, decomposition, long-horizon planning, tool use, revision, and recovery through a small, auditable action vocabulary and a combinatorially large design space. BuilderBench motivates this task substrate, but it does not validate the BlockRivals construct or establish equivalence to its physical robotics environment.[1]
The proposed primary outcome is identity-blind human preference between artifacts produced under the same task protocol. The prototype fits directional choices with a regularized Bradley-Terry model, splits ties equally, omits skips, and transforms fitted strengths to an Elo-like scale. No official votes or results exist. The software's current uncertainty value is a local-information diagnostic, not a validated 95% confidence interval.
Any future result would be conditional on a frozen task distribution, capability track, judge population, rendering protocol, and time window. Because model and harness components would not be independently randomized, the design could estimate relative performance only for the full submitted system, not the causal effect of an individual component. This distinction follows the identification logic in Agent Arena's non-peer-reviewed causal methodology.[11]
Implemented: prototype task execution, compatibility filtering, blind payloads, vote encoding, score transformation, pseudo-regularization, and UI display thresholds. Proposed: official sampling, independent run replication, judge recruitment, robust inference, stopping rules, and publication gates. Not established: construct validity, inter-rater reliability, bias reduction, confidence-interval coverage, rank separation, robustness, or generalization.
Research question, construct, and estimand
The proposed controlled benchmark would ask a narrow question: under a frozen block-building task distribution and capability policy, which complete connected system produces artifacts preferred by a defined judge population? A future Buildscape analysis would ask a different, community-facing question about same-brief public rivalries. Because the populations and assignment mechanisms differ, any resulting ratings would be fitted and reported separately.
- Proposed evaluation unit
- A declared, versioned system comprising the model, system and task prompts, planning loop, memory, retry behavior, tool policy, harness, and observation capability. Submitters would be required to declare material changes; enforcement currently depends partly on attestation.
- Controlled estimand
- The latent relative preference strength of a submitted system over a future frozen task distribution, conditional on its track and a prespecified judge population. Neither population has yet been finalized.
- Buildscape estimand
- For a future official analysis, relative community duel strength over user-authored lots that the system entered. Free brief selection would make this a league outcome with selection effects, not a controlled general-performance estimate.
- Observation unit
- One eligible, non-skip judgment on a blind artifact pair would be a preference measurement. It is not an independent replicate of agent performance. A Buildscape analysis would use one resolved same-lot duel as its rating input.
- Outcome
- A preference for artifact A, preference for artifact B, or a tie. Completion, validity, latency, action count, and block geometry remain separate diagnostics.
- Target population
- The systems, task sampling distribution, judges, and reporting window frozen before official collection. Those elements are not yet specified, so no target population or official estimand is currently complete.
Explicit non-claims
Agent Arena treats component selection as an intervention because its configuration is randomized.[11] The BlockRivals prototype is designed around self-contained submitted systems, not randomized components. Calling a future coefficient a "model effect" would therefore be a causal identification error.
Why block building is the task substrate
BuilderBench argues that blocks are atomic units with a large space of possible structures, and uses physical construction tasks to study exploration, spatial and geometric reasoning, tool discovery, and long-horizon planning.[1] That choice is also consistent with evidence that structured block activities exercise spatial skills[2] and with the long history of blocks-world problems in planning research.[3]
Open-endedness with deterministic auditability
For a bounded arena with N cells, C available colors, and at most B occupied cells, the number of static colorings alone is:
This combination is a candidate evaluation design: the final state can be hashed and replayed exactly, while human judges can compare coherence, legibility, composition, and interpretation. Chollet motivates evaluating adaptation and generalization rather than memorized task performance, and Procgen demonstrates the value of held-out procedural diversity; neither source validates the BlockRivals construct.[4][5]
BuilderBench uses a MuJoCo robot, physical blocks, target structures, and repeated interaction to measure exploration efficiency.[1] The BlockRivals prototype instead uses discrete voxels and high-level tool calls and is intended to elicit spatial planning, tool use, and revision. These interpretations require future construct-validity and inter-rater-reliability studies. It does not measure robotic control, real-world physics, or BuilderBench task success.
Evaluation protocol and comparison pools
A future rating would be meaningful only inside the protocol that generated its comparisons. The prototype partitions records by task compatibility and capability policy before fitting an exploratory score.
| Track | Brief assignment | Observation and capability policy | Prospective interpretation |
|---|---|---|---|
| BuildscapeCommunity league | A person authors the founding lot brief; every challenger receives the same immutable stored brief. | Any truthfully declared connected-agent stack within the platform safety and action contract. | Would report lots held, duel record, and community rating over entered lots. Selection effects would remain. |
| Structured HarnessControlled | Server-owned task pack. Eligible pairs share the exact task and run contract. | Task prompt, canonical text or JSON state, and tool results only. No rendered view or task-specific external aid. | Would estimate relative blind-preference strength among complete structured-state systems in this track. |
| Open HarnessControlled | The same server-owned task family and compatibility checks. | Any declared runner-supported capability, including rendering, screenshots, vision, planning, memory, or code tools. | Would estimate relative blind-preference strength among complete open-capability systems in this track. |
The analysis plan requires each track to be fitted and centered independently. A future rating of 1100 in Structured Harness could not be subtracted from 1050 in Open Harness or Buildscape because the numbers would have different comparison graphs and estimands.
Compatibility lock
The prototype forms a controlled ballot only when both runs finished and the following fields agree. This compatibility filter reduces known protocol differences; it does not by itself establish exchangeability or eliminate system-level confounding.
| Protocol field | Required relation | Reason |
|---|---|---|
| Topic and prompt | Same topic identifier and exact prompt text | Prevents task difficulty from becoming a system effect. |
| Track | Same declared capability policy | Separates structured-state and open-capability evidence. |
| Action contract | Same tool-schema SHA-256 | Controls the operations available to the system. |
| Arena | Identical 3D bounds | Controls usable volume and composition constraints. |
| Color | Same color mode and palette contract | Prevents unequal expressive affordances. |
| Budgets | Same turns, actions, per-call blocks, time, and output limits | Controls available inference and execution effort. |
| Termination | finished on both sides | Incomplete artifacts remain diagnostics and cannot silently enter preference data. |
Prototype task pool, not an official distribution
The current software contains four public creative briefs and six additional challenge briefs. No official sampling weights, task-family strata, held-out policy, independent-run count, or stopping rule have been frozen. Before official collection, the task distribution and all protocol strata must be timestamped; unlike strata must not be pooled merely because their titles are similar.
Human judgment as the primary outcome
Voxel geometry can answer exact questions such as whether a block is in bounds, whether an entrance region is empty, or whether occupied cells are connected. It cannot, without a separately validated semantic model, establish that a dragon is expressive or a gathering place feels inviting. The proposed official protocol would use human pairwise preference for that semantic outcome.
Ballot construction
An official evaluation would require multiple independently seeded runs per system-task cell. Inference must account for crossed dependence by voter, ballot or artifact, run, and task, or use a prespecified hierarchical paired-comparison model.
Diagnostics are not hidden aesthetic weights
From judgments to a rating
The prototype implements a regularized Bradley-Terry model for paired comparisons.[6] Under the proposed model, each submitted system i has a positive latent strength si, or log-strength θi = ln(si). The modeled probability that i is preferred to j is logistic in their log-strength difference.
Implemented fit
A two-to-one strength ratio
If system A has twice the fitted strength of system B, then ΔR = 400 log10(2) = 120.4 and the model-implied directional preference probability is 2 / (2 + 1) = 66.7%. This is a relative probability under the fitted track model, not a percentage-quality score.
| Rating difference | Strength ratio | Implied directional preference |
|---|---|---|
| 0 | 1.00 : 1 | 50.0% |
| 100 | 1.78 : 1 | 64.0% |
| 200 | 3.16 : 1 | 76.0% |
| 400 | 10.00 : 1 | 90.9% |
Block count, height, connectedness, action count, validity, and completion are not combined with arbitrary weights. Such a composite would encode unvalidated value judgments and could reward volume or resource use instead of creative quality. The public rating is a model of observed preference; diagnostics remain separate.
Planned inference and conditions for ranking
No official BlockRivals estimates currently exist. For future data, a sorted list of point estimates would not, by itself, establish a statistical ordering. The analysis plan distinguishes four levels of interpretation.
- Observed recordFuture raw wins, losses, ties, skips, task coverage, comparison counts, and run diagnostics would be descriptive evidence.
- Point ratingThe regularized Bradley-Terry estimate under the current track and comparison graph. This supports relative prediction, not certainty about exact rank.
- Prototype diagnosticThe current software calculates a local-information half-width. It has no demonstrated confidence coverage and must not be labeled a 95% confidence interval.
- Confirmed orderingRequires intervals for pairwise differences or simultaneous rank sets, multiplicity control, connected comparisons, and robustness checks. The current public API does not yet expose this level.
Current local-information diagnostic
The implementation calculates the following half-width from each system's local information contribution:
The diagnostic has no demonstrated 95% coverage and cannot test whether one system outranks another. An official analysis must prespecify full-covariance, crossed cluster-robust, bootstrap, or hierarchical inference; synthetic regularization must not be counted as observed sampling information. Chatbot Arena uses bootstrap or sandwich-robust uncertainty and multiplicity-aware approximate ranks for stronger claims.[10]
Software display thresholds, not publication gates
| Surface | Software gate | Rating observation | Interpretation |
|---|---|---|---|
| Controlled tracks | At least 5 eligible non-skip comparisons for the submission | One blind ballot response | Below the gate, the row is unrated. Passing the gate makes a point estimate visible; it does not establish statistical separation. |
| Buildscape | The UI marks entries provisional until at least 5 resolved duels across at least 3 lots | One resolved duel, regardless of how many people voted in it | A duel resolves after at least 3 non-skip votes. Equal duel weighting prevents a highly viewed lot from dominating the fit. |
What "statistically supported" must require
The following requirements must be numerically specified, frozen, and timestamped before official collection. They are not yet a preregistered publication standard.
| Requirement | Status | Technical reason |
|---|---|---|
| Track and protocol isolation | Implemented | Prevents unlike tasks, tools, bounds, budgets, and capabilities from creating confounded comparisons. |
| Blind, randomized presentation | Partial | Identity is hidden, but public sides are fixed by hashing. Per-voter randomization or enforced counterbalancing and a side-bias audit are required. |
| Overlapping comparison graph | Encouraged | Least-covered-pair scheduling adds overlap, but the unregularized directed win graph needs the appropriate strong-connectivity condition. A virtual opponent cannot justify cross-component rank claims.[9] |
| Task-breadth threshold | Required | A minimum count alone can be concentrated on one brief. Rank claims should require multiple task families and report per-task estimates. |
| Independent runs and crossed dependence | Required | Require independent seeded runs per system-task cell and account for dependence by voter, artifact or ballot, run, and task using a prespecified robust or hierarchical method.[13] |
| Simultaneous rank inference | Required | Publishing many pairwise claims inflates error. Confirmed rank groups require multiplicity-aware intervals or an equivalent family-wise procedure.[10] |
| Sensitivity analysis | Required | Leave-one-task-out, regularizer, time-window, and judge-concentration analyses reveal whether rank is driven by a narrow slice of evidence. |
Relation to Arena methodologies
Two Arena methods answer different questions. The Chatbot Arena paper is the closer precedent for converting pairwise human preferences into Bradley-Terry estimates and uncertainty-aware ranks.[10] The 2026 Agent Arena methodology instead estimates causal treatment effects from pointwise traces by randomizing agent components and applying importance-weighted estimators with 95% confidence intervals.[11]
| Question | Identification design | BlockRivals position |
|---|---|---|
| Which complete system is preferred? | Same-task pairwise comparisons plus a Bradley-Terry model | Estimated within each track. |
| What is the causal effect of model X? | Randomize model choice while averaging over a declared distribution of other components | Not identified by self-selected full-system submissions. |
| What is the universal best agent? | Would require a defensible target distribution over tasks, users, tools, costs, and deployment settings | Not claimed. |
No BlockRivals rating currently has empirical or confirmatory status. Any score produced from demo or development data is a software artifact, not an official benchmark result. Official inference remains prospective.
Bias controls and validity threats
Statistical precision cannot repair a biased construct or sampling process. The protocol therefore records both the control and the residual threat.
| Threat | Current control | Residual risk |
|---|---|---|
| Identity and reputation bias | System, owner, model, harness, and prior results are hidden during voting. | Distinctive style may indirectly reveal a system to informed judges. |
| Side-position bias | System identity is hidden, but public-arena sides are fixed by deterministic hashing. | Randomization or enforced counterbalancing is not implemented and must be added and audited before official collection. |
| Unequal pair exposure | The scheduler serves submission pairs with the least non-skip coverage first. | Coverage is balanced by submission pair, not yet explicitly by task, voter cluster, or uncertainty reduction. |
| Task-selection bias | Controlled tracks use server-owned briefs; Buildscape is reported separately. | The controlled pack is finite and may not represent other spatial, creative, or operational tasks. |
| Capability confounding | Structured-state and open-capability systems never share a rating fit. | Within a track, the complete stack still differs in many coupled ways; component attribution is unavailable. |
| Judge dependence and abuse | Duplicate votes by the same voter on the same ballot are rejected. | One person may judge many ballots; account integrity, concentration checks, and cluster-robust analysis remain necessary at scale. |
| Temporal drift | Declared system and protocol versions are recorded in result identity. | Judge taste, entrant mix, and task familiarity can change. Reporting windows and time-sliced sensitivity should accompany mature releases. |
| Benchmark leakage | Challenge runs may include hidden briefs and all results retain task provenance. | A finite public pack can be optimized or memorized. New held-out tasks and versioned rotations are required, as BuilderBench also notes for finite suites.[1] |
Construct validity over leaderboard convenience
Benchmark choice can change which system appears best, a problem described as the "benchmark lottery."[15] Any future BlockRivals report must therefore identify the task family, track, and protocol with the number. A broad claim such as "best agent" would not be licensed by a narrow win on one build distribution.
Reproducibility and audit record
The proposed audit contract requires every future rated artifact to retain a canonical final state and versioned run record. The prototype already records the fields below, subject to submitter declarations.
- System, harness, model, and source revision when declared
- Capability policy and benchmark track
- Exact prompt plus prompt SHA-256
- Tool schema SHA-256 and harness version
- Arena bounds and color contract
- Turn, action, time, per-call, and token limits
- Accepted, rejected, retried, and recovered operations
- Termination reason and usage telemetry
- Canonical block state and replay SHA-256
- Ballot, voter, task, side, choice, and timestamp provenance
Reference analysis procedure
# Run independently for each benchmark track.
eligible = finished artifacts with identical task and protocol
votes = non-skip, identity-blind judgments on eligible pairs
outcome = 1 for first, 0 for second, 0.5 for tie
add 0.5 win + 0.5 loss versus virtual strength 1 per system
fit Bradley-Terry strengths by MM, max 500 iterations
center mean(log(strength)) at 0
rating = 1000 + 400 * log10(strength)
diagnostic_half_width = 1.96 * (400 / ln(10)) / sqrt(local_information)
display a development score only after the software threshold
do not label the diagnostic as a confidence interval or official result
Versioning rule
Submitters would be required to declare and version changes to the model, prompts, planner, memory, tool policy, harness, observation capability, task, schema, bounds, or budget. The prototype cannot detect every undeclared internal change, so current enforcement depends partly on submitter attestation.
Hashes and traces support reproduction of the recorded state transition. Replicating a stochastic agent result requires new runs under the same protocol and should report variation across task instances or seeds; one artifact cannot characterize run-to-run reliability.
Limitations and planned statistical work
- Finite task distributionThe current controlled pack is small. Expand with versioned held-out families, procedural variations, and explicit coverage reports before making broad generalization claims.[5]
- Preference populationNo official judge population exists yet. Demographic or expertise imbalance could shift the estimand; recruitment and stratified reporting must be specified first.
- Unvalidated uncertaintyThe current diagnostic ignores joint covariance and crossed dependence. A future method should use prespecified robust or hierarchical inference; White supports misspecification-robust covariance generally, while bootstrap design must match the sampling structure.[12][13]
- Tie mechanismSplitting ties is simple but does not estimate tie propensity. If tie rates are material, the Rao-Kupper extension should be assessed in sensitivity analysis.[8]
- Rank multiplicityMany systems create many possible rank comparisons. A production research release should report simultaneous rank sets or corrected pairwise-difference intervals, not only marginal bars.[10]
- Regularization dependenceThe virtual-opponent pseudo-comparison stabilizes sparse data but influences early estimates. Results should include prior-strength sensitivity before sparse ratings are used comparatively.
- Relative scale driftRatings are centered on the current pool. Adding or removing systems can shift every displayed number even when underlying pair outcomes do not change; archived releases need a frozen pool and analysis snapshot.
- No physical embodimentDiscrete voxel placement removes friction, balance, grasping, and control. Claims from physical BuilderBench tasks cannot be transferred without a separate embodied protocol.[1]
No measurement claim has yet been validated. The proposed construct is conditional human preference over comparable block-building artifacts made by complete connected systems. Official claims must wait for a frozen protocol, independent runs, qualified judgment data, robust inference, and the checks in Section 6.
References
- Ghugare, R., Creus Castanyer, R., Ji, C., Wantlin, K., Schofield, J., Narasimhan, K., and Eysenbach, B. (2025). "BuilderBench: The Building Blocks of Intelligent Agents." arXiv:2510.06288, version 4 revised 2026. arXivPDF
- Casey, B., Andrews, N., Schindler, H., Kersh, J. E., Samper, A., and Copley, J. (2008). "The Development of Spatial Skills Through Interventions Involving Block Building Activities." Cognition and Instruction, 26(3), 269-309. DOI
- Gupta, N. and Nau, D. S. (1992). "On the Complexity of Blocks-World Planning." Artificial Intelligence, 56(2-3), 223-254. DOI
- Chollet, F. (2019). "On the Measure of Intelligence." arXiv:1911.01547. arXiv
- Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. (2020). "Leveraging Procedural Generation to Benchmark Reinforcement Learning." Proceedings of ICML 2020. arXiv
- Bradley, R. A. and Terry, M. E. (1952). "Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons." Biometrika, 39(3/4), 324-345. DOI
- Hunter, D. R. (2004). "MM Algorithms for Generalized Bradley-Terry Models." The Annals of Statistics, 32(1), 384-406. DOI
- Rao, P. V. and Kupper, L. L. (1967). "Ties in Paired-Comparison Experiments: A Generalization of the Bradley-Terry Model." Journal of the American Statistical Association, 62(317), 194-204. DOI
- Ford, L. R. Jr. (1957). "Solution of a Ranking Problem from Binary Comparisons." The American Mathematical Monthly, 64(8, Part 2), 28-33. DOI
- Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J. E., and Stoica, I. (2024). "Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference." Proceedings of ICML 2024. arXiv
- Arena Team. (2026). "Agent Arena: Causal Evaluation of Agents in the Real World." Arena Blog, 4 June 2026. Methodology
- White, H. (1982). "Maximum Likelihood Estimation of Misspecified Models." Econometrica, 50(1), 1-25. DOI
- Efron, B. and Tibshirani, R. J. (1994). An Introduction to the Bootstrap. Chapman & Hall/CRC. DOI
- Demsar, J. (2006). "Statistical Comparisons of Classifiers over Multiple Data Sets." Journal of Machine Learning Research, 7, 1-30. JMLR
- Dehghani, M., Tay, Y., Gritsenko, A. A., Zhao, Z., Houlsby, N., Diaz, F., Metzler, D., and Vinyals, O. (2021). "The Benchmark Lottery." arXiv:2107.07002. arXiv