SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking Paper • 2609.13141 • Published 14 days ago • 65
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents Paper • 2609.06702 • Published 19 days ago • 26
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks Paper • 2609.11042 • Published 15 days ago • 63
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse Paper • 2603.12201 • Published Mar 12 • 69
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling Paper • 2607.02980 • Published Jul 3 • 85
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference Paper • 2605.25475 • Published May 25
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging Paper • 2506.23266 • Published Jun 29, 2025 • 1
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling Paper • 2607.02980 • Published Jul 3 • 85