Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 7 days ago • 203
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 6 days ago • 302
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction Paper • 2609.10715 • Published 8 days ago • 318
SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published 7 days ago • 259
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching Paper • 2609.01404 • Published 16 days ago • 28
Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement Paper • 2609.01481 • Published 16 days ago • 19
The Mechanics of Democratic Dominance: A System Dynamics Paradigm for Dynamic Consent Engineering Paper • 2608.27509 • Published 17 days ago • 7
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence Paper • 2608.31075 • Published 17 days ago • 31
Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities Paper • 2608.28122 • Published 20 days ago • 67
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Paper • 2608.28281 • Published 20 days ago • 106
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 16 days ago • 104
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published 17 days ago • 386
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution Paper • 2608.31106 • Published 17 days ago • 102
UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City Paper • 2608.27456 • Published 21 days ago • 112