ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
Abstract
Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical web development, however, agents may need to infer behavior from working software and implement it in an incomplete application. We introduce ProgramDistill, a benchmark evaluating coding agents on features discovered through interaction with fully functional reference applications. We build ProgramDistill by factorizing applications into features of different granularities, each associated with replayable behaviors executable via its gold patch. Our pipeline, mine-craft-patch, discovers 1,975 replay-verified behaviors across 26 applications and constructs 4,063 tasks without human intervention. Across nine frontier coding agents, GPT-6 Astra and Claude Opus 5 achieve 49.2% and 28.8% success on cumulative workflows in full-application reconstruction. In partial-application reconstruction, success falls from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8. ProgramDistill thus provides a scalable benchmark with controlled difficulty for evaluating and diagnosing coding agents, and a natural basis for future curriculum-based training.
Community
For more information, please check out our:
๐ Blog: https://microsoft.github.io/debug-gym/blog/2026/09/programdistill/
๐ Paper: https://microsoft.github.io/debug-gym/static/papers/ProgramDistill_arxiv.pdf
Weโre working on a public release of the ProgramDistill benchmark, so you'll be able to play with it yourself soon.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- GameXpert-Bench: How Far Are Coding Agents from Expert Game Development? (2026)
- ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders (2026)
- EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants (2026)
- LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation (2026)
- ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions (2026)
- Repo2Skill-Evo: Repository Skills Go Stale in Silence (2026)
- IWC-Bench: Evaluating Web Application Generation from a Software Testing Perspective (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.18805 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper