New articles · Material article revisions · Corrections
Content updates
New articles, material changes to published article evidence, scope or conclusions, and important corrections.
新文章 · 文章重要修订 · 更正。这里只记录值得首次阅读或重新阅读的已发布内容变化。
Published changes / 已发布变化
| Date | Type | Update | Why it matters |
|---|---|---|---|
| 2026-08-08 | Material revision | Added the evaluator itself to the evaluation guide's nine checks | All nine checks in the guide asked about the model; none asked about the instrument producing the score. Two paragraphs and one claim were added, drawn from MiraBench (arXiv:2605.29360v1, 28 May 2026) and an independent decision-centric position paper (arXiv:2606.15032v2). MiraBench scores its optimism-bias level with one vision-language model but two different prompts split by model type, and the single model given the more lenient prompt is also the one where agreement with human annotators collapses to 33.3%, against 78.3% to 100% elsewhere; overall reported agreement is 87.8%. The benchmark discloses all of this itself, and an under-detecting judge inflates a reliability score rather than deflating it. The action-consequence check gained a second common false positive: a model that does respond to the action and still shows the task succeeding when the action should have made it fail. The guide also now notes that post-training for task success lowered one model's failure-preservation score from 48.7 to 23.1 while raising task completion from 14.0 to 92.0 - bounded, because that post-trained model is the benchmark authors' own fine-tune and the same comparison does not move at 14B scale. The underlying mechanism is the same one this site already described as distribution bias in the autonomous-driving guide; what is new is that it is measured on the action-outcome axis. No conclusion in the guide was reversed and no check was removed. |
| 2026-08-05 | Material revision | Updated the autonomous-driving guide for Wayve's GAIA-3 and GAIA-4 | The guide's Wayve route row was written for GAIA-2 and was two generations out of date. It now covers the GAIA line through GAIA-4, published 3 August 2026, which adds closed-loop simulation with the driving policy in the loop. Two claims were added. First, both GAIA-3 documents of 2 December 2025 assert that simulation agrees with on-road results, and neither publishes a study, metric or result for that agreement; the figure the press release does report measures the share of generated tests discarded as unusable, which on this site's reading is a different quantity, bounded to those two documents. Second, the GAIA-4 page names three levels for measuring simulator behaviour fidelity, including a frame-by-frame comparison of the trajectory planned in simulation against the trajectory planned on the road, and reports no numeric result for any of the three on that page. Wayve names DriveSafeSim, a UK government-funded project with WMG at the University of Warwick, in the two GAIA-3 documents; the GAIA-4 page does not mention it. No conclusion in the guide was reversed. |
| 2026-07-28 | Article | Published: What do spatial 3D models actually measure? | Added a spatial 3D capability map separating measured geometry outputs from benchmark performance, generated worlds and spatial reasoning, and corrected the World Labs Marble source record. |
| 2026-07-12 | Article | Published: What have world models actually achieved in robotics? | Added a bilingual evidence map separating physical-robot learning, image-goal planning, target-robot finetuning, controlled benchmarks and author-overlap reproductions. |
| 2026-07-10 | Article | Published: How do you evaluate a world model? | Added a bilingual evaluation framework with nine checks grouped into capability, task utility and evidence quality, plus route-specific metric priorities and a reusable checklist. |
| 2026-07-10 | Article + material revision | Published and expanded the autonomous-driving reference guide | Published the bilingual domain guide, then separated traditional simulation from learned world models, added independent RAND and IEEE evidence with page-level locators, and clarified that simulation evidence does not replace a system-level safety case. The core conclusion was narrowed, not reversed. |
| 2026-07-08 | Articles | Published the first three bilingual explainers | Published introductory, boundary and route-comparison explainers for world-model beginners. |