Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation Paper • 2609.04298 • Published 15 days ago
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work? Paper • 2606.07682 • Published Jun 5 • 2
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces Paper • 2604.05172 • Published Apr 6 • 24