Improving Test-Time Scaling with Adaptive Looped Transformers Paper • 2609.35748 • Published 7 days ago • 58
WorldPlay2: Extending Real-Time Interactive World Models in Control and Horizon Paper • 2609.35560 • Published 7 days ago • 31
Learning to Learn from Context: Synthetic Training from Perturbed Public Documents Paper • 2609.33642 • Published 8 days ago • 32
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning Paper • 2609.33781 • Published 8 days ago • 45
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning Paper • 2609.35767 • Published 7 days ago • 49
Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge Paper • 2609.34327 • Published 7 days ago • 40
CompoWorld: Compositional Environment Scaling for General Agents Paper • 2609.33665 • Published 8 days ago • 42
EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks Paper • 2609.28236 • Published 12 days ago • 39
Improving Test-Time Scaling with Adaptive Looped Transformers Paper • 2609.35748 • Published 7 days ago • 58
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 25 days ago • 174
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation Paper • 2608.21500 • Published Aug 21 • 41
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Paper • 2605.15565 • Published May 15 • 16
The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Paper • 2604.02029 • Published Apr 2 • 111
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis Paper • 2603.20278 • Published Mar 17 • 103