ASI-Bench: At the Dawn of Artificial Superintelligence

Core Authors: Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen

Preprint: arXiv:2608.17271

Abstract

ASI-Bench is the first benchmark to jointly evaluate AI systems on both innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to measure how far an agent can proceed on its own. Built by over 40 experts at a cost of 31,000+ human hours, it contains 60 project-level research tasks across 11 scientific domains. Across 18 agent–model configurations, performance drops sharply as guidance is removed — from 50.91 with full methodological guidance, to 29.10 when only the method is specified, to 26.62 when the agent must determine the method itself — indicating that current systems remain heavily dependent on human guidance and are still far from end-to-end autonomous research.

See the official arXiv record for the full paper, the complete author list, and citation details.