SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning Paper • 2608.23493 • Published 21 days ago • 1
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Paper • 2609.02783 • Published 12 days ago • 120
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild Paper • 2608.23181 • Published 21 days ago • 34
ParaTempo: Efficient Parallel Reasoning via Temporal Confidence Paper • 2608.16425 • Published 27 days ago • 40
SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution Paper • 2608.18933 • Published 26 days ago • 13
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe Paper • 2608.13667 • Published Aug 13 • 17
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Paper • 2608.09802 • Published Aug 10 • 136
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Paper • 2607.01211 • Published Jul 1 • 14
How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study Paper • 2607.10856 • Published Jul 12 • 7
Dockerless: Environment-Free Program Verifier for Coding Agents Paper • 2606.28436 • Published Jun 26 • 118
FastContext: Training Efficient Repository Explorer for Coding Agents Paper • 2606.14066 • Published Jun 12 • 96
SWE-Explore: Benchmarking How Coding Agents Explore Repositories Paper • 2606.07297 • Published Jun 5 • 124
Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents Paper • 2602.07900 • Published Feb 8 • 4
DLLM-Searcher: Adapting Diffusion Large Language Model for Search Agents Paper • 2602.07035 • Published Feb 3 • 31
CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding Paper • 2602.01785 • Published Feb 2 • 97