FrontierChallenge: Evaluating Scientific Workflow Completion Paper β’ 2608.24979 β’ Published Aug 25 β’ 152
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper β’ 2608.23283 β’ Published Aug 24 β’ 212
TransMamba: Flexibly Switching between Transformer and Mamba Paper β’ 2503.24067 β’ Published Mar 31, 2025 β’ 22 β’ 2
TransMamba: Flexibly Switching between Transformer and Mamba Paper β’ 2503.24067 β’ Published Mar 31, 2025 β’ 22
TransMamba: Flexibly Switching between Transformer and Mamba Paper β’ 2503.24067 β’ Published Mar 31, 2025 β’ 22
Scale-Distribution Decoupling: Enabling Stable and Effective Training of Large Language Models Paper β’ 2502.15499 β’ Published Feb 21, 2025 β’ 15
HMoE: Heterogeneous Mixture of Experts for Language Modeling Paper β’ 2408.10681 β’ Published Aug 20, 2024 β’ 10
Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent Paper β’ 2411.02265 β’ Published Nov 4, 2024 β’ 26
Scaling Laws for Floating Point Quantization Training Paper β’ 2501.02423 β’ Published Jan 5, 2025 β’ 26