Tencent Hunyuan releases ExplorationBench, an open-source benchmark for AI exploration in verifiable alien worlds
Tencent Hunyuan released ExplorationBench, an open-source benchmark with two verifiable alien worlds, 55 discovery targets, 140 held-out tasks, and results for 10 frontier AI systems. The repository is available on GitHub, with a paper on arXiv and an official website.
China context
- Original name
- 腾讯混元
- Outside China
- Open weights · github.com
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can use the open-source benchmark to evaluate their models' exploration abilities and compare against frontier systems.
- For investors
- Investors can monitor adoption of ExplorationBench as a signal of Tencent Hunyuan's influence in AI research and open-source contributions.
ExplorationBench measures AI systems' ability to explore in verifiable alien worlds where rules are executable and conflict with familiar knowledge. It includes two sandboxes: AlienCode (program synthesis with hidden semantics) and AlienLogic (formal proof with patched inference rules). The benchmark evaluates 10 frontier AI systems over four rounds of autonomous exploration, with leaderboards showing held-out accuracy. Key findings indicate that exploration, not recall or thinking alone, produces required knowledge, and autonomous experiment design outperforms replaying probes.
The benchmark uses deterministic feedback from interpreters or proof checkers, with no LLM judge. Systems start from a flawed manual and fixed worked examples, then probe the world over four rounds. At each milestone, a tool-disabled copy reports believed rules and answers held-out tasks three times. AlienCode has 31 discovery targets and 70 held-out tasks; AlienLogic has 70 held-out theorems, 25 unprovable. Leaderboards show Best@3 and Mean@3 accuracy, with Claude Opus 5 leading AlienCode (89.0% Best@3) and Grok 4.6 leading AlienLogic (83.8% Best@3).
AI developers and researchers outside China gain a new open-source tool to evaluate and improve exploration capabilities in frontier models, potentially shifting benchmark competition toward autonomous scientific discovery. The release of a Tencent benchmark with results for models like Claude, GPT, Gemini, and DeepSeek provides a common reference for comparing exploration performance across labs.
For AI developers, ExplorationBench offers a rigorous, verifiable way to test and differentiate models on exploration tasks, which could influence model selection for research and development. For investors, it highlights Tencent Hunyuan's focus on advanced AI capabilities and open-source contributions, potentially signaling strategic priorities.
Observable next signals include whether other labs adopt ExplorationBench for model evaluation, whether Tencent releases additional worlds or tasks, and whether models improve on the benchmark in subsequent versions. The paper and website may provide further analysis and trajectories.